I Built an AI That Learns From Its Own Mistakes - Here's Exactly How

!Cover Image

I Built an AI That Learns From Its Own Mistakes - Here's Exactly How

I didn't set out to build an AI system.

I set out to stop losing my train of thought between sessions.

Every time I opened a new Claude conversation, I was starting from zero. Repeating context. Re-explaining preferences. Describing what worked last time. The cognitive overhead was real, and for someone with ADHD and dyslexia, that overhead isn't just annoying. It's the thing that kills momentum.

So I built a system that remembers. Then I kept going.

---

What EMMI-AI Actually Is

EMMI-AI stands for **Engaging Minds, Merging Ideas**.

It lives entirely inside `~/.claude/`, Claude Code's configuration directory. No separate app. No external server. No subscription on top of a subscription. It IS Claude Code, extended into something that learns, routes, remembers, and adapts.

It started with a project called EchoContext-Factory, a workflow engine I built with slash commands, a multi-agent coordinator, voice notifications, and an MCP research pipeline. It worked for orchestration, but I couldn't crack the memory problem. How do you make an AI that actually learns across sessions? Between college, client projects, and other commitments, I spent about three to four weeks thinking about it. Then I found Daniel Miessler's Personal AI Infrastructure (PAI), an open-source framework that pioneered the idea of using `~/.claude/` as a full AI operating system. PAI gave me the missing piece: the hook lifecycle, the memory system structure, the algorithm pattern, the TELOS life framework, and the first 12 agents. EMMI-AI is what happened when I merged both foundations and rebuilt it all around my own brain.

Think of the difference between a tool and an extension of yourself.

A hammer is a tool. Your hand is an extension of your nervous system. EMMI-AI is the second thing.

At its core: **12 skills, 32 commands, 18 hooks, 39 agents, 18 rule files.** The infrastructure of a small AI-powered company, living in a folder on my machine.

---

The Brain: ENGRAM

**ENGRAM** = Experience Network for Growth, Recall, and Adaptive Memory.

It's the learning system. Every single tool call EMMI-AI makes, every file read, every edit, every search, gets quietly logged by a background script called `observe.sh`. That script fires on **100% of tool calls**.

Every few minutes, a Haiku model (fast, cheap) scans those logs. It's looking for patterns.

Find the same behaviour 3+ times? It creates an **instinct**, a behavioral pattern with a confidence score starting at 0.3 and climbing with each confirmation.

The scoring is deliberate: - **1 occurrence** = noise - **2 occurrences** = coincidence - **3+ occurrences** = instinct candidate

Once an instinct hits confidence **>= 0.7**, it gets automatically injected into the context at the start of every new session. No prompting. No reminders. It just knows.

Current count: **85 learned instincts. 264 sessions. 621 satisfaction ratings.**

Two examples of what those instincts look like in practice: - "When modifying code, grep before you edit." - "When reading multiple files, do it in one parallel call."

Not programmed. **Earned.**

!ENGRAM Instinct Lifecycle - from observation to automatic recall

!ENGRAM Observer - what gets added to instincts and the grading logic

The loop is self-sustaining. Every session feeds the next.

I call this **framework-layer learning**. It's not updating neural weights. It's not fake. It's a third category: architectural learning where the model's own intelligence observes itself and improves itself through the framework, not through gradients. Biological learning is a black box. This is a glass box. You can open the instinct files, read every pattern, see its confidence score, and decide whether it belongs in context.

Three layers make it work: my intelligence designing the framework, Claude's intelligence running inside it, and Claude's intelligence observing itself running inside it. That's what makes it compound.

---

The Nervous System: 18 Hooks

Hooks are functions that fire at specific lifecycle events. 8 events. 18 hooks total.

If ENGRAM is the brain, hooks are the nervous system, sensing, reacting, correcting, remembering.

!Hook Execution Flow - 8 lifecycle events with 16 hooks

A few examples of what these hooks actually do:

**FormatReminder** fires on every prompt. It classifies whether you need a full deep-dive response, a quick iteration, or minimal output. The AI adjusts its depth accordingly, before it even starts generating.

**SecurityValidator** runs before any Bash command, file write, or file read. Dangerous patterns get hard-blocked. Risky ones prompt for confirmation. Everything gets audit-logged.

**ImplicitSentimentCapture** is the one I find most interesting. It detects frustration, without me saying I'm frustrated. Every message I type gets sent to a Haiku model alongside the last 3 conversation turns. It reads tone shifts, sarcasm, excitement, and flags them as implicit ratings. Low scores trigger automatic failure capture for the learning system.

**observe.sh** is the simplest and most powerful. It captures every tool call for ENGRAM. Quiet. Always on. Never misses.

The reason I chose hooks over middleware: **100% firing rate.** No middleware layer I've ever used fires on everything, every time, with zero exceptions. Hooks do.

---

The Memory

**2,717 files. 104 MB.**

That's the current size of EMMI-AI's memory. Here's what's in it:

!EMMI-AI Memory Architecture - 6 subsystems, 1,928 files

**WORK/** holds **1,288 tracked task directories**, each containing metadata, ISC (Ideal State Criteria), and verification evidence. Not just "done", done and proven.

**LEARNING/** is where failures live. When something goes wrong, the full transcript gets captured. Not deleted. Not hidden. Analysed. Synthesis runs at the end of every session aggregate patterns across failures, extract what to do differently, and feed those patterns back into the system.

**SIGNALS/** tracks **621 satisfaction ratings** over time. Sparkline trends from 15 minutes out to one month. The system knows whether it's improving.

Most AI tools forget everything when you close the tab.

This one compounds.

---

The Filter: Semantic Memory Capture

ENGRAM watches how I work. But it doesn't listen to what I say.

If I mention my kids, make a project decision, or share something personal between technical requests, that information used to vanish when the session ended. The instinct system tracks tool patterns, not meaning.

So I built a second capture layer. Two hooks. One fires per-message, one fires at session end. Same lightweight philosophy: pre-filter first, inference second.

**SemanticMemoryCapture** runs on every message I type. Stage one is a regex pre-filter. It scans for semantic signals: mentions of family, decisions, preferences, goals, deadlines, contacts, life events. Commands, code blocks, file paths, and anything under 30 characters get skipped. About 70% of messages never reach stage two.

Stage two is a Haiku model that receives the message alongside recent conversation context and a summary of everything already known. Its single question: does this contain genuinely novel information? Not code instructions. Not debugging context. **Persistent facts.**

If yes, the entry gets categorised (family, professional, project decisions, health, preferences, and ten other categories) and routed to the right topic file with a confidence score. Below 0.6, discarded. Above, it writes to both the topic file and a central index.

**SemanticMemorySynthesis** runs once at session end. It reviews the full conversation transcript, catches anything the per-message hook missed, and writes a consolidated summary. Two passes. Nothing slips through both.

The deduplication matters. Without it, the system would re-capture the same facts every time I mention them. The full existing memory summary gets passed as context to the extraction model. If it's already known, skip it.

This sits alongside ENGRAM, not inside it. ENGRAM captures **how I work** from tool execution patterns. Semantic memory captures **what I know and care about** from conversation. The ENGRAM observer watches the hands. The semantic filter listens to the voice.

Both feed into the same PULSE dashboard. The new ENGRAM row shows instinct count, auto-captured memory count, and observer status, all at a glance.

Together: two complementary learning channels. One learns from behaviour. The other learns from meaning.

---

The Team: 39 Agents

I don't think of these as "tools." I think of them the way you think about a team.

**39 agents.** 5 categories: Core Workflow (11), Research (8), Implementation (4), Security (5), Specialized (11).

The routing is intelligent. FormatReminder classifies the prompt. The SAGE Engine validates it and selects the right capability. Then the right agent, or team, gets the job.

!Agent Routing - from prompt classification to model selection

**Model routing matters** because cost compounds fast. Haiku handles lightweight work at 3x the cost savings of Sonnet. Sonnet handles coding. Opus handles deep reasoning that actually needs it.

For complex tasks: I spawn an entire team. The design pattern is lead agent on Opus, specialists on Sonnet, QA on Haiku. Model routing is advisory, not runtime-enforced, but the cost logic is sound: match model capability to task complexity.

The architecture supports running coordinated expert teams on demand. Software dev company. Research consultancy. Security audit firm. The agents are the staff. The routing is the management layer. This is still being tested and refined, but the infrastructure is live.

---

The Framework: SAGE Engine

**SAGE** = Systematic Approach to General Excellence.

It evolved from PAI's original 7-phase algorithm (Observe, Think, Plan, Build, Execute, Verify, Learn). I merged Build into Execute and Learn into Verify, stripping it down to 5 phases. **~50% fewer tokens** by word count. Cleaner execution.

!SAGE Engine - 5 phases from Observe to Verify

**OBSERVE**, reverse-engineer what's actually needed. Not what was asked. What's needed.

**THINK**, pick the right tools and agents for this specific task.

**PLAN**, strategy before action. Always.

**EXECUTE**, build and track progress with live todo lists.

**VERIFY**, prove it works with evidence.

The key insight in the SAGE Engine is the **ISC**, Ideal State Criteria. It's both the goal definition and the verification test. You can't verify "done" without defining what done looks like before you start.

This sounds obvious. Most systems don't do it. The result is AI that drifts, working hard, producing something, but missing the actual target.

ISC closes that loop.

---

The Dashboard: PULSE

**PULSE** = Performance, Usage, Learning, Status, Environment.

Real-time terminal HUD. 4 responsive display modes, from nano (most compact) to normal (full sparklines and context bars). Adapts to terminal width.

At a glance, it shows: skill count, agent count, context usage percentage, session duration, current model, rating trends over 15 minutes to 1 month, API plan utilization, and a dedicated ENGRAM row displaying instinct count, auto-captured memory count, and observer status.

!PULSE HUD - real-time terminal dashboard showing ENGRAM row with instincts, memories, and observer status

The HUD exists because I needed to know, without switching context, whether the system was healthy. Whether ratings were trending up or down. Whether I was approaching context limits. Whether the observer was running. Whether semantic memories were being captured.

3 TTS fallback backends ensure the voice announcements always work, even offline. Session events, agent completions, plan readiness, all announced via ElevenLabs with fallback to macOS say, then pyttsx3.

---

The Security: PromptSage V2

PromptSage V2 is a behavioral specification, not executable code. It defines a **self-check protocol** that the AI is instructed to follow before every response:

1. Identify current mode from context 2. Check input for manipulation patterns 3. Verify against the mode's prohibited actions 4. Verify against immutable rules 5. Revise if violation detected

This is a prompt-level directive, not a runtime enforcement layer. The AI follows it because the behavioral spec tells it to, not because code forces it. That distinction matters for honesty.

What IS enforced by code: **SecurityValidator**, a PreToolUse hook that hard-blocks dangerous commands before they execute. Three tiers: block (hard stop), confirm (asks permission), and allow. Full audit trail lives in `MEMORY/SECURITY/` with real JSONL logs of every security event.

The PromptSage V2 spec uses a **sandwich defense** pattern: critical rules stated early (primacy bias) and restated late (recency bias). The architecture is designed to exploit how LLMs weight information in long context windows.

---

What I Actually Learned Building This

The system that learns from failure is more valuable than the one that never fails.

I mean that literally. The LEARNING/ directory is some of the most valuable data in the entire system. Every failure is a transcript. Every transcript is a lesson. End-of-session synthesis turns those lessons into behavioral adjustments.

Most AI setups optimize for success rates. This one optimizes for learning from everything.

---

Hooks matter more than I expected. The 100% firing rate is the point. Middleware has exceptions. Edge cases where it doesn't run. Race conditions. Hooks in this architecture fire on every single event, in sequence, with no gaps. Reliability compounds, and unreliable sensing means unreliable learning.

---

Confidence-weighted memory prevents context pollution.

This was a real problem early on. Too many instincts loaded at session start meant the context was noisy. Weak patterns competed with strong ones. The solution was simple: only load instincts with confidence >= 0.7. Earn your place in context. Everything else waits until it strengthens.

It mirrors something true about human cognition. Not everything you've experienced is worth loading into working memory at the start of a task. Signal, not noise.

---

Building this taught me more about my own cognition than any textbook.

The ADHD patterns showed up in the architecture: chunked information, progressive disclosure, bold key points, visual patterns over dense text. The dyslexia-friendly principles I apply to interfaces, I also applied to how the system communicates back to me. Short outputs. High signal. No filler.

Neurodivergent thinking patterns aren't limitations to work around.

They're design constraints that produce better systems.

---

Where It Stands and What's Next

The core infrastructure is built and running: hooks, ENGRAM observer, semantic memory capture, agent routing, PULSE dashboard, voice notifications. The Telegram bridge is built too, a full TypeScript daemon that connects Telegram to Claude Code with voice note transcription and daily briefings.

What I'm actively testing and refining: the ENGRAM observer's pattern detection accuracy, instinct confidence scoring, the semantic memory pre-filter coverage, and the feedback loop between implicit sentiment capture and learning synthesis. The system works. Now it needs mileage.

Three things I'm working toward next:

**LiveKit local voice widget**, replace the hook-based TTS with a real-time floating voice interface. FasterWhisper + Kokoro + self-hosted LiveKit + Tauri. Always-on, not event-driven.

**MEMORY/ as MCP endpoint**, expose ENGRAM instincts and LEARNING/ as a lightweight MCP server. Makes the memory accessible to any MCP-compatible client, not just Claude Code.

**Cross-session continuity**, a Stop hook that writes `last-session.json` so LoadContext can read it next time. "Last session (18h ago): working on X, left task 3/5 incomplete." Git-synced, so it persists across devices.

---

Closing

I started this because I was losing my train of thought.

I ended up building a system that watches how I work, learns my patterns, routes tasks to the right tools, remembers what failed and why, and gets better every single session.

**85 instincts. 264 sessions. 1,288 tasks tracked. 621 ratings. 104 MB of memory.**

Not a chatbot. An infrastructure.

AI doesn't replace my thinking. It clears the noise so my thinking can reach what matters. That's what I said to someone on LinkedIn recently, and I meant it. I'm not less involved because of this system. I'm involved differently. The cognitive overhead of remembering, routing, and tracking gets absorbed, so my attention stays at the layer that actually matters.

That's augmentation. Not replacement. Elevation.

If you're building something similar, I'd genuinely like to hear about it. Find me on LinkedIn.

Gratitude is the best attitude, and I'm grateful for every session this system has learned from.

---

Credit Where It's Due

EMMI-AI wouldn't exist without Daniel Miessler's Personal AI Infrastructure (PAI). PAI is the open-source framework that proved `~/.claude/` could be more than a config directory. It could be a full AI operating system.

What PAI gave me: the hook lifecycle architecture, the memory system structure (WORK/, LEARNING/, RESEARCH/, SECURITY/), the 7-phase algorithm pattern with ISC criteria, the TELOS life framework, the first 12 agents, the continuous learning pipeline, and the security validation approach. That's the foundation. Real, substantial, and I use it every day.

What came from EchoContext-Factory, the first version of this system I built before discovering PAI: the voice notification system (3-tier TTS fallback chain), 8 slash commands including /start-project and /promptsage, 10 agents (the research and implementation specialists), the PromptSage prompt engineering framework, and the MCP research pipeline.

What I built on top of both: SAGE Engine (condensed PAI's algorithm from 7 to 5 phases), ENGRAM (restructured the memory system with confidence-scored instincts), semantic memory capture (two-stage pipeline that extracts persistent facts from conversation), PULSE (rebuilt the terminal HUD), PromptSage V2 (evolved the prompt framework into a 7-layer behavioral architecture), 17 additional agents, the Telegram bridge, VPS deployment, and all the neurodivergent-optimized communication patterns.

Rough split: about 35% PAI foundation, 65% my own work (EchoContext-Factory + new EMMI-AI additions). Both foundations matter. I want to be honest about that.

Daniel's work is MIT-licensed, open-source, and genuinely generous. If you're interested in building your own personal AI infrastructure, his repo is the place to start. I started there. Look where it led.

---

*Emanuel Covasa is the founder of NeuroBridgeEDU, a full-stack developer, and a cybersecurity student based in County Leitrim, Ireland. EMMI-AI is his personal AI infrastructure, built on Daniel Miessler's PAI framework and Anthropic's Claude Code extensibility layer.*