Long-term memory plugin for Claude Code. Sessions, rules, and facts that persist across conversations.
Sessions — the agent automatically saves session transcripts and can write compact summaries (goals, decisions, blockers, dead ends). A new session can load a previous one and continue where it left off, with full context.
Session chaining via /clear — /clear is not just a reset. The plugin detects the outgoing session, loads its summary/compact into the next session, and links them. This works as an alternative to the native compact: instead of compressing a bloated context, you start fresh with a clean window while keeping the essential context from the previous session.
Lazy rules (in development) — rules and conventions are not dumped into context all at once. They're loaded dynamically: relevant rules are prefetched based on what you're discussing and injected right before the agent acts. Plan-relevant rules are loaded after planning, before execution.
Facts & knowledge — not limited to sessions and rules. You can ask the agent to remember anything — debugging insights, API quirks, deployment notes, personal preferences — and search for it later.
Plain Markdown — everything is stored as .md files with YAML front-matter. No background process, no database server. You can read, edit, and organize your memory with any text editor, commit it to git, or share across machines. Files are Obsidian-compatible — session transcripts use callout blocks, and cross-references use [[wikilinks]].
Requires: Python 3.10+, Claude Code with plugin support.
# Register the repository as a plugin marketplace (once)
/plugin marketplace add git@github.com:dankinsoid/ai-memory.git
# Install the plugin
/plugin install ai-memoryFor development:
claude --plugin-dir ./plugins/ai-memoryCodex support is installed separately because MCP and hooks are configured via
~/.codex/config.toml and ~/.codex/hooks.json.
bash scripts/install-codex.shThe installer:
- enables
features.codex_hooks = truein~/.codex/config.toml - registers the local MCP server in
~/.codex/config.toml - merges ai-memory hooks into
~/.codex/hooks.json - bakes absolute script paths and
AI_MEMORY_*/OPENAI_API_KEYvalues from~/.claude/settings.jsondirectly into hook commands, so hooks do not depend on shell startup files like~/.zshrc
Restart Codex after running the installer.
OpenAI features are opt-in to avoid silently spending tokens when OPENAI_API_KEY is set globally. Add to your Claude Code settings.json:
{
"env": {
"OPENAI_API_KEY": "sk-...",
"AI_MEMORY_EMBEDDING": "true",
"AI_MEMORY_EMBEDDING_MODEL": "text-embedding-3-small",
"AI_MEMORY_LLM": "true",
"AI_MEMORY_LLM_MODEL": "gpt-4.1-nano",
"AI_MEMORY_LLM_PROVIDER": "openai"
}
}| Variable | Default | Description |
|---|---|---|
AI_MEMORY_EMBEDDING |
false |
Enable vector embeddings |
AI_MEMORY_EMBEDDING_MODEL |
text-embedding-3-small |
OpenAI embedding model |
AI_MEMORY_LLM |
false |
Enable LLM calls (future) |
AI_MEMORY_LLM_MODEL |
gpt-4.1-nano / haiku |
Chat model (default depends on provider) |
AI_MEMORY_LOAD_BUDGET |
20000 |
Chars returned by memory_load_session (compact + facts + transcript tail) |
AI_MEMORY_LLM_PROVIDER |
auto | openai or claude-cli. When omitted, uses openai if OPENAI_API_KEY is set, otherwise falls back to claude-cli |
OPENAI_API_KEY |
— | Required for embeddings and openai LLM provider. claude-cli provider uses CLI auth — no key needed |
Without these, the plugin works fully via tag-based filtering. The claude-cli provider requires no API key — it uses your existing Claude Code CLI authentication.
By default, embeddings are stored in the local SQLite cache. To use an external Qdrant instance instead:
export QDRANT_URL="http://localhost:6333"
export QDRANT_API_KEY="..." # optional, for Qdrant CloudSet AI_MEMORY_DIR to point at a folder inside your Obsidian vault:
export AI_MEMORY_DIR="$HOME/ObsidianVault/ai-memory"A minimal Obsidian plugin for styling session transcripts is included in obsidian-chat-view/. Install via BRAT or copy the folder manually into your vault's .obsidian/plugins/ directory.
| Command | What it does |
|---|---|
/load |
Load any session and resume. Accepts last or free-text search |
/rule |
Save a rule, preference, or convention to memory |
/save |
Write a session compact — a condensed summary of the entire conversation. Not required for /load (transcripts are saved automatically), but useful as a checkpoint when context matters |
The plugin runs a local MCP server (stdio, no daemon) with these tools:
| Tool | Purpose |
|---|---|
memory_session |
Create or update a session summary (title, project, tags, compact notes) |
memory_remember |
Save a fact or rule with tags and title |
memory_search |
Search by tags (AND/OR/exclude), date range, and optional semantic query |
memory_read |
Read full content of a memory by [[wikilink]] ref |
memory_load_session |
Load a previous session for deep recovery |
memory_explore_tags |
List all tags with file counts |
All MCP tool calls are auto-approved — no confirmation prompts.
The plugin uses Claude Code hooks to manage memory without manual intervention.
Start — loads universal rules/facts (up to 5), project-scoped facts and recent sessions (up to 10). On /clear, automatically loads the previous session's summary/compact (see session chaining above).
During — reminds the agent to save session summaries sometimes. Tracks context token usage and prompts /save before the context window fills up (~100K tokens).
Every turn — appends the latest conversation chunk to the session .md file (async). Records git context (branch, commits) in front-matter.
Currently, rules are loaded at session start and when saving a session (matching rules are returned alongside the save confirmation). The agent can also search for rules at any time via memory_search(tags=["rule"]).
Fully automatic lazy loading (async prefetch based on conversation topics, injection before tool calls and after plan finalization) is in development.
Default location: ~/.ai-memory/ (override with AI_MEMORY_DIR). If that path
does not exist but legacy ~/.claude/ai-memory/ does, ai-memory uses the
legacy location automatically.
ai-memory/
├── universal/
│ ├── rules/ # Apply everywhere
│ └── facts/
├── languages/
│ └── <lang>/ # Language-specific
├── projects/
│ └── <project>/
│ ├── rules/ # Project-specific
│ ├── facts/
│ └── sessions/
└── sessions/
└── YYYY-MM-DD/ # Date-organized transcripts
Each file has YAML front-matter (tags, date, title, etc.). Tags are both explicit and derived from directory structure — a file in projects/myapp/rules/ automatically gets project/myapp and rule tags.
A SQLite cache indexes the files for fast tag/date search. Location: ~/Library/Caches/ai-memory/index.db (macOS) or ~/.cache/ai-memory/index.db (Linux). The cache is purely derived from the filesystem — delete it anytime, it rebuilds automatically.
Three tiers:
- Scope:
universal,project/<name>,lang/<name> - Aspect:
testing,architecture,debugging,deployment,performance,security,tooling,workflow,error-handling,api-design,concurrency,ci-cd, ... - Specific: any topic tag (
react,docker,auth, ...)
Rules (tagged rule) get special treatment — loaded at session start and dynamically injected when relevant.
How often semantic search puts the right note in the top 5, measured on a real vault of 2848 notes with 398 queries:
| Retriever | recall@1 | recall@5 | recall@10 | MRR |
|---|---|---|---|---|
Semantic (text-embedding-3-small) |
0.65 | 0.86 | 0.89 | 0.73 |
| BM25 (lexical baseline) | 0.51 | 0.72 | 0.78 | 0.60 |
Method. The benchmark indexes exactly what production indexes — corpus building reuses the same functions as storage.reindex(). Queries are LLM-generated: for each sampled note the model writes what a developer would type weeks later from faded memory, and that note is the ground truth. Full harness and reproduction steps in eval/.
Read the low-leak number, not the headline. Because queries are derived from the notes, some reuse the original wording. Each query carries a leak score — the share of its content words found in the note — and the subset with low leak (n=150) is the honest measure:
| Retriever | recall@5, low-leak queries |
|---|---|
| Semantic | 0.75 |
| BM25 | 0.38 |
The gap is where embeddings earn their cost: on queries phrased by meaning rather than by matching words, lexical search loses half its accuracy while semantic search holds.
Two more findings from the same run:
- Failures are retrieval failures, not ranking ones — recall only moves from 0.86 to 0.89 between k=5 and k=10. A note missing from the top 5 is usually missing entirely.
- The hardcoded
threshold=0.25discards nothing. The lowest cosine score on a correct note was 0.368, versus a median of 0.562.
Caveats. The vault is ~99% session summaries (2832 sessions vs 16 facts), so these numbers describe session retrieval and say little about facts and rules. Ground truth is LLM-generated, so it measures whether search can find the note a query was written from — not user satisfaction.
MIT