Agent Memory

A memory layer beneath the agent: knowledge that connects like a graph, reads like a vault, and matures with use. Runs locally.

An agent that closes its session forgets everything it learned in it. agent-memory keeps it in plain Markdown, shared across your agents.

Two lines, one store

Agent memory has grown along two lines. A retrieval engine finds the right thing but hands the agent an opaque chunk it cannot inspect. A filesystem is legible and free to run, but does not rank, and stops scaling once the tree outgrows a listing.

agent-memory is both in one store: the retrieval engine indexes a filesystem the agent can also just read. Relations live as links inside the memories, a local index ranks them, and every hit resolves to a whole Markdown file on disk.

ApproachWhat you getWhat you give up
Retrieval engineRanked recall over embeddings or a graphA store you cannot read or migrate off
FilesystemMarkdown the agent reads with ls and grepRanking, once the tree grows
agent-memoryRanked recall over files you can read, grep, and commitNothing in the read path calls a model or the network

Retrieve by path, then read by level

Recall does not paste text into your context. It answers with a short list — one-line abstract, file path, anchor, score — and the agent opens each hit only as deep as the task needs: index line, abstract, outline, full file, then the raw messages it was distilled from. Each rung costs an order of magnitude more than the last, and each is a place to stop.

Design commitments

Three read tracksDeterministic MEMORY.md injection at session start, BM25 recall with an optional vector index fused in, and the plain directory tree when both miss.
Writes are the system's jobDistillation fires at conversation boundaries without holding up the task, and the full trace is copied first — missed by the distiller never means lost by the system.
Files are the truthEvery index is a rebuildable cache: delete it, rebuild, lose nothing. Enforced by a test, not promised in a doc.
A Manage layer on its own clockA sleep-time pass consolidates and forgets by value. It may add and update unattended; deletion only ever arrives as a proposal you confirm.
No LLM client insideJudgement is borrowed from the host agent's own CLI, so every write stays visible in your transcript and there is no billing surface.

Proof it works

Measured on LongMemEval-S with a bounded haystack of 12 sessions per episode, 120 episodes, Claude Haiku 4.5 as host and one calibrated judge, two exam replays per arm.

ArmPooled accuracy
agent-memory127/240 = 52.9%
MemCore86/240 = 35.8%
No memory7/120 = 5.8%

The bounded haystack makes this a write-strategy study, so the numbers are not comparable to published LongMemEval scores. Across Claude Code, Codex CLI, and Hermes, all 9 ordered writer/reader pairs pass: what one host writes, another finds.

Get started

Requires Python 3.12+ and uv. Install from a checkout:

git clone https://github.com/tigerless-labs/agent-memory.git
cd agent-memory
uv sync --all-packages
export PATH="$PWD/.venv/bin:$PATH"

Create the store, then wire it into your agent:

mem init
mem setup --host claude-code

Use --host codex for Codex CLI. Setup adds its hook to the host and leaves the rest of your settings alone: session start injects, session end distils. Agents that speak MCP get the same core calls through mem-mcp; anything that can run a shell command can use the mem CLI directly.

More in the works — built in the open, shipped fast.

Explore Tigerless Labs