Claude-Mem: Persistent Context Across Sessions with 3-Layer Progressive Disclosure and 10x Token Savings
An architectural deep-dive into Claude-Mem (thedotmack/claude-mem): how autonomous AI coding agents capture tool observations, diffs, and decisions via 5 lifecycle hooks, index them in a hybrid SQLite FTS5 + Chroma store, and achieve ~10x context token reduction through 3-layer progressive disclosure across Claude Code, OpenClaw, Codex, and Hermes Agent.

One of the most persistent operational hurdles when working with AI coding agents—such as Claude Code, OpenClaw, Codex, or Hermes Agent—is amnesia between sessions.
Every time you type /exit or restart an agent process:
- The model forgets the architectural decisions reached thirty minutes prior.
- It forgets why a specific refactoring was chosen over an alternative.
- It re-runs expensive exploratory commands (
find,grep,cat), burning thousands of input tokens and driving up API costs.
Worse, the naive solution—dumping hundreds of past session transcripts into the system prompt—creates catastrophic context window pollution, degrading reasoning quality and exploding token bills.
Enter Claude-Mem (thedotmack/claude-mem), an open-source autonomous memory engine with over 97,000 GitHub stars. It tackles cross-session continuity not by hoarding raw conversation history, but through an event-driven 5-hook interception pipeline combined with 3-Layer Progressive Disclosure, achieving up to 10x token savings.
Here is an architectural breakdown of how Claude-Mem works under the hood, how its hybrid storage engine indexes reality, and how to deploy it in production agent workflows.
1. Architectural Blueprint: The 5 Lifecycle Hooks Pipeline
Rather than requiring manual note-taking commands, Claude-Mem operates as an invisible, event-driven observer that hooks directly into the agent runtime:
┌────────────────────────────────────────────────────────────────────────┐
│ Agent Execution Lifecycle │
│ [ SessionStart ] ──> [ UserPromptSubmit ] ──> [ PostToolUse ] │
│ │ │
│ [ SessionEnd ] <── [ Stop / Interrupted ] <───────┘ │
└──────────┬─────────────────────────────────────────────┬───────────────┘
│ │
▼ (Intercept Tool Execution & Output) ▼ (Session Lifecycle)
┌────────────────────────────────────────────────────────────────────────┐
│ Local Bun Worker Service │
│ (Lightweight HTTP Daemon on Localhost) │
└──────────────────────────────────┬─────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────────────┐
│ Observation Processing & Summarization │
│ • Strip boilerplate & terminal noise │
│ • Extract code diffs, file paths, and exit codes │
│ • Generate semantic summary via background LLM worker │
└──────────────────────────────────┬─────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────────────┐
│ Hybrid Storage Engine │
│ ┌───────────────────────────────┐ ┌────────────────────────────────┐ │
│ │ SQLite Database with FTS5 │ │ Chroma Vector Database │ │
│ │ (Exact keyword, paths, logs) │ │ (Dense semantic embeddings) │ │
│ └───────────────────────────────┘ └────────────────────────────────┘ │
└────────────────────────────────────────────────────────────────────────┘
The 5 Interception Hooks:
SessionStart: Triggered when the agent process launches. It initializes the local worker connection and injects high-level context from recent project sessions.UserPromptSubmit: Intercepts the user’s intent, enabling the memory layer to pre-query relevant past observations before tool execution begins.PostToolUse: The critical capture hook. Every time the agent runs a shell command, edits a file, or queries an API, Claude-Mem captures:- The tool name and input arguments.
- File diffs (patches applied).
- Execution status and stdout/stderr results.
Stop: Handles voluntary pauses, user interruptions, or safety blocks, capturing the incomplete state so the next session knows where work halted.SessionEnd: Runs final consolidation routines, commits transaction logs, and updates the project’s cumulative timeline.
2. The 3-Layer Progressive Disclosure Strategy (~10x Token Cut)
The breakthrough insight in Claude-Mem is that an agent rarely needs full conversation transcripts from past sessions. Injecting massive text blobs into prompt context exhausts tokens and causes “needle-in-a-haystack” retrieval degradation.
Claude-Mem structures memory retrieval into three discrete, progressive tiers:
[ Step 1: User Asks a Question ]
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Layer 1: Compact Search Index │
│ Returns only matching observation IDs, dates & titles. │
│ Cost: ~50 - 100 Tokens Total │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Layer 2: Chronological Timeline │
│ The agent inspects the sequence of events surrounding IDs. │
│ Cost: ~150 - 300 Tokens Total │
└──────────────────────────────┬──────────────────────────────┘
│
▼ (Only if deep specifics are needed)
┌─────────────────────────────────────────────────────────────┐
│ Layer 3: Batch Fetch Full Observation │
│ Retrieves full code patches and terminal outputs ONLY for │
│ the specific target observation ID. │
│ Cost: ~500 - 1,000 Tokens (Targeted) │
└─────────────────────────────────────────────────────────────┘
Why This Architecture Wins:
- Conventional Approaches: Dump 15,000–30,000 tokens of past chat logs on every turn.
- Claude-Mem: Queries start at 50 tokens. If the compact index provides enough architectural signal, the agent proceeds immediately. If it needs the exact cryptographic salt or regex pattern used two days ago, it selectively requests Layer 3 for that single item.
- Result: A consistent ~90% to 92% reduction in memory overhead, allowing agents to stay within cheap context tiers indefinitely.
3. The Storage Engine: SQLite FTS5 + Chroma Vector Search
Context recovery fails if an agent relies solely on vector similarity (which struggles with exact file paths and error codes) or solely on keyword search (which misses conceptual synonyms).
Claude-Mem resolves this with a Dual-Engine Hybrid Store:
| Component | Technology | Primary Function |
|---|---|---|
| Relational & Exact Search | SQLite with FTS5 | Full-text search over function names, file paths (/src/auth/token.ts), package versions, and error codes. |
| Semantic & Conceptual Search | Chroma Vector DB | Dense vector embeddings capturing semantic intent (e.g., matching “user login optimization” with “JWT validation caching”). |
| Local Daemon Worker | Bun / Node.js Runtime | Background HTTP daemon providing sub-millisecond API responses and zero cold-boot latency. |
| Visual Inspection UI | Local Web Viewer | Real-time web dashboard running on localhost allowing developers to inspect, edit, or delete captured memory entries. |
4. Cross-Harness Interoperability
While originally developed for Claude Code, Claude-Mem has evolved into a universal, multi-harness memory protocol. It integrates natively across:
- Claude Code: Installed as a native plugin via
/plugin install claude-mem. - OpenClaw: Deployed as an ambient memory node for omnichannel bots.
- Hermes Agent (Nous Research): Seamlessly bridges session memory with Hermes’s progressive skills and local vault.
- OpenCode & Codex: Integrates via CLI hooks and MCP protocols.
- Antigravity CLI: Provides cross-workspace persistence during deep architectural refactoring.
5. Practical Guide: Installation and Setup
Step 1: Automated Installation
You can install Claude-Mem via npx with automatic environment detection:
# General automated installation
npx claude-mem install
# Or target a specific agent environment
npx claude-mem install --ide opencode
Step 2: Running with Claude Code
Inside your Claude Code session, enable the plugin:
/plugin install claude-mem
The background worker automatically launches on localhost:3773. You can verify active observation recording in your terminal or navigate to the web viewer:
http://localhost:3773
Step 3: Privacy Controls (<private> Tags)
To prevent sensitive credentials, production tokens, or private customer records from being captured by memory summaries, wrap sensitive blocks in <private> tags:
<private>
STRIPE_SECRET_KEY="sk_live_51Mz..."
DATABASE_PASSWORD="super-secret-production-pw"
</private>
Claude-Mem’s observation pre-processor automatically redacts any content within <private> tags prior to database commitment and summarization.
6. Architectural Summary & Production Implications
| Metric | Claude-Mem Specification |
|---|---|
| Open Source License | MIT License |
| Runtime Architecture | Local Bun / Node.js HTTP worker daemon |
| Storage Topology | SQLite (FTS5 indexed) + Chroma Vector Store |
| Token Efficiency | ~10x savings via 3-Layer Progressive Disclosure |
| Interception Method | 5 Non-blocking event hooks (SessionStart to SessionEnd) |
| Supported Agents | Claude Code, OpenClaw, Hermes, Codex, OpenCode, Antigravity |
Claude-Mem addresses the foundational bottleneck of agentic computing: turning stateless conversational models into stateful, persistent systems engineering partners. By combining deterministic lifecycle hooks with progressive retrieval, it delivers continuous project memory without sacrificing token economics.
Written by Fouad Salkini (فؤاد سلقيني)
General Manager & Tech Lead at Tripnologies and Sync Studios. Systems Architect focusing on AI coding agents, DevOps, and quantitative systems.