Agent Memory
A memory layer beneath the agent: knowledge that connects like a graph, reads like a vault, and matures with use. Runs locally.
An agent that closes its session forgets everything it learned in it. agent-memory keeps it in plain Markdown, shared across your agents.
Two lines, one store
Agent memory has grown along two lines. A retrieval engine finds the right thing but hands the agent an opaque chunk it cannot inspect. A filesystem is legible and free to run, but does not rank, and stops scaling once the tree outgrows a listing.
agent-memory is both in one store: the retrieval engine indexes a filesystem the agent can also just read. Relations live as links inside the memories, a local index ranks them, and every hit resolves to a whole Markdown file on disk.
| Approach | What you get | What you give up |
|---|---|---|
| Retrieval engine | Ranked recall over embeddings or a graph | A store you cannot read or migrate off |
| Filesystem | Markdown the agent reads with ls and grep | Ranking, once the tree grows |
| agent-memory | Ranked recall over files you can read, grep, and commit | Nothing in the read path calls a model or the network |
Retrieve by path, then read by level
Recall does not paste text into your context. It answers with a short list — one-line abstract, file path, anchor, score — and the agent opens each hit only as deep as the task needs: index line, abstract, outline, full file, then the raw messages it was distilled from. Each rung costs an order of magnitude more than the last, and each is a place to stop.
Design commitments
| Three read tracks | Deterministic MEMORY.md injection at session start, BM25 recall with an optional vector index fused in, and the plain directory tree when both miss. |
|---|---|
| Writes are the system's job | Distillation fires at conversation boundaries without holding up the task, and the full trace is copied first — missed by the distiller never means lost by the system. |
| Files are the truth | Every index is a rebuildable cache: delete it, rebuild, lose nothing. Enforced by a test, not promised in a doc. |
| A Manage layer on its own clock | A sleep-time pass consolidates and forgets by value. It may add and update unattended; deletion only ever arrives as a proposal you confirm. |
| No LLM client inside | Judgement is borrowed from the host agent's own CLI, so every write stays visible in your transcript and there is no billing surface. |
Proof it works
Measured on LongMemEval-S with a bounded haystack of 12 sessions per episode, 120 episodes, Claude Haiku 4.5 as host and one calibrated judge, two exam replays per arm.
| Arm | Pooled accuracy |
|---|---|
| agent-memory | 127/240 = 52.9% |
| MemCore | 86/240 = 35.8% |
| No memory | 7/120 = 5.8% |
The bounded haystack makes this a write-strategy study, so the numbers are not comparable to published LongMemEval scores. Across Claude Code, Codex CLI, and Hermes, all 9 ordered writer/reader pairs pass: what one host writes, another finds.
Get started
Requires Python 3.12+ and uv. Install from a checkout:
git clone https://github.com/tigerless-labs/agent-memory.git
cd agent-memory
uv sync --all-packages
export PATH="$PWD/.venv/bin:$PATH"Create the store, then wire it into your agent:
mem init
mem setup --host claude-codeUse --host codex for Codex CLI. Setup adds its hook to the host and leaves the rest of your settings alone: session start injects, session end distils. Agents that speak MCP get the same core calls through mem-mcp; anything that can run a shell command can use the mem CLI directly.
More in the works — built in the open, shipped fast.