Best AI Agent Memory Systems 2026: Hindsight, Mem0, Zep
The short answer
Pick by what the agent needs to do with its memory, not by leaderboard. As of September 2026: Mem0 for the fastest, most widely integrated memory layer; Hindsight for agents that must learn and consolidate across sessions, with the only independently reproduced LongMemEval results; Zep / Graphiti for temporal reasoning over facts that change; Letta if you want a stateful-agent runtime rather than a library; cognee for open-source knowledge-graph memory with Python, TypeScript and Rust SDKs. All five are open source; the hosted tiers range from usage-based (Hindsight, cognee) to $20/month (Letta Pro) to about $125/month (Zep Flex).
The ranking
| Rank | System | Core idea | License | Hosted price (Sep 2026) | Best for |
|---|---|---|---|---|---|
| 1 | Mem0 | Extract facts from conversations, add/search API, optional graph memory | Open source | Free tier, usage above | Most production chatbots and assistants |
| 2 | Hindsight (Vectorize) | Facts → observations → mental models; retain / recall / reflect | MIT | Usage-based, free credits, no per-seat | Agents that learn over time; always-on agents |
| 3 | Zep / Graphiti | Temporal knowledge graph with fact-validity windows | Open source (Graphiti) | Flex ~$125/mo reported | Time-sensitive facts, CRM-style memory |
| 4 | Letta (ex-MemGPT) | Agent runtime with self-editing in-context vs archival memory | Open source | Pro $20/mo | Teams that want the whole agent, not a layer |
| 5 | cognee | Knowledge-graph memory pipelines; Python/TS/Rust SDKs, MCP | Apache 2.0 | Free, then $2.50 per 1M tokens | Graph-heavy, multi-language stacks |
1. Mem0 — the default
Mem0 is the most adopted memory layer, at roughly 62,500 GitHub stars in September 2026, and the one most other tools compare themselves against. The model is simple: pass conversation turns to add, and it extracts and stores facts per user, agent or session; call search at inference time to retrieve them. It reports 92.5% on LoCoMo, 94.4% on LongMemEval and 64.1% on BEAM-1M on its own research page (updated September 22, 2026). Graph memory is optional. It is the right answer when the requirement is “the assistant should remember what the user said three weeks ago” and nothing more exotic. Weakness: the numbers are self-reported, and consolidation is shallower than Hindsight’s — facts accumulate rather than resolve into beliefs.
2. Hindsight — memory that learns
Hindsight’s difference is the two layers above facts. Observations are deduplicated beliefs consolidated in the background with supporting quotes and a proof count; new evidence strengthens or weakens them instead of overwriting. Mental models are standing answers to a question about a memory bank that Hindsight rewrites as it learns and that the agent reads as a plain database read at boot. The third operation, reflect, reasons over memories rather than retrieving them. It runs on Postgres + pgvector via Docker, embedded Python, Helm or Cloud; supports 25+ LLM providers including fully local ones; and can use existing ChatGPT, Claude, Cursor or Copilot subscriptions as the LLM with no API key. Its LongMemEval results are the only ones in this list independently reproduced (Virginia Tech’s Sanghani Center, The Washington Post). It gained about 4,500 GitHub stars on September 27-28, 2026 as persistent agents (Meta Muse, Microsoft Autopilot, OpenAI “o”) became the story. Detail in what is Hindsight.
3. Zep / Graphiti — when time matters
Zep wraps Graphiti, an open-source temporal knowledge graph that records when each fact became true and when it stopped being true. That makes it the strongest option for memory where facts get superseded — a customer’s plan, an employee’s role, a project’s status — and for questions like “what was true in March.” Zep’s own LoCoMo claim is 94.7%, but a third-party run put it at 75.1%, and its consistently cited LongMemEval figure is 71.2% under a GPT-4o judge. The hosted Flex plan is reported at about $125/month; self-hosting Graphiti is free but you own a graph database.
4. Letta — the runtime, not the layer
Letta (formerly MemGPT) is an agent framework in which the agent itself manages what stays in context versus what goes to archival storage, using memory-management tools. It positions around 74% on LoCoMo under an agent-autonomy framing. If you want to adopt an opinionated stateful-agent server with a REST API and development environment, Letta is the most complete; if you already have an agent and want to add memory to it, it is the heaviest option here. A $20/month Pro cloud tier shipped in 2026.
5. cognee — graph memory, many languages
cognee (topoteretes/cognee) builds knowledge-graph memory pipelines from documents and conversations, with Python, TypeScript and Rust SDKs and an MCP server. It is Apache 2.0, with a free cloud tier and $2.50 per 1M tokens above it. It trended alongside Hindsight in late September 2026. Choose it when your memory is document-and-entity heavy rather than conversation heavy, or when you need a non-Python client.
How to choose
- Chatbot that should remember users: Mem0. Done in an afternoon.
- Agent that runs unattended and must improve: Hindsight. Mental models at boot are the feature you will notice.
- Facts that expire or change: Zep / Graphiti.
- Greenfield stateful agent, no existing framework: Letta.
- Graph-shaped knowledge, Rust or TypeScript backend: cognee.
- Any of them: run LoCoMo or LongMemEval yourself on your own logs. Vendor scores do not share a stack, and the gaps between claims (Zep 94.7% vs 75.1%) are bigger than the gaps between products.
For the conceptual split between agent memory, chat memory and RAG, see agent memory vs chat memory vs RAG. For isolation, injection and retention risks, see how to secure AI agent memory.
Last verified: September 28, 2026. Hosted prices from vendor pages and September 2026 reporting; benchmark figures are vendor-reported unless stated.