What Is Hindsight? Vectorize's Agent Memory That Learns 2026
The short answer
Hindsight is Vectorize’s MIT-licensed agent memory system whose pitch is “agents that learn, not just remember.” It stores memories in isolated banks as world facts, experiences, observations and mental models; exposes retain, recall and reflect; runs on PostgreSQL + pgvector (server, embedded, Helm or Cloud); and claims state-of-the-art LongMemEval accuracy that has been independently reproduced. It surged on GitHub in the last week of September 2026 — roughly +4,500 stars on September 27-28 on top of ~35,000 — because the always-on agent wave (Meta Muse, Microsoft Autopilot, OpenAI “o”) made persistent, consolidating memory the obvious missing primitive.
The four memory types
Most memory layers store one thing: extracted facts. Hindsight stores four, and the top two are the reason it exists.
| Type | What it is | Example |
|---|---|---|
| World facts | Facts about the world | ”The stove gets hot” |
| Experiences | The agent’s own experiences | ”I touched the stove and it hurt” |
| Observations | Consolidated, evidence-backed beliefs formed from many memories | ”This customer churns when onboarding stalls past day 3” (with quotes + proof count) |
| Mental models | Standing answers to a question about a bank, rewritten in the background | ”What are this user’s preferences?” → a maintained page |
When you retain something, an LLM extracts facts, temporal data, entities and relationships, normalizes them into canonical entities and time series, and indexes them with sparse and dense vectors. In the background, related facts are consolidated into observations — deduplicated beliefs that keep their supporting quotes and a proof count, and that get strengthened, weakened or extended when new evidence arrives instead of silently overwritten. Mental models go one step further: you define a question once, Hindsight writes the answer, stores it, and rewrites it as the bank learns. Reading one is a database read — no retrieval, no LLM call — so an agent can boot with a page of settled knowledge. Knowledge pages are mental models organized like a wiki and projectable onto disk as Markdown.
The three operations
from hindsight_client import Hindsight
client = Hindsight(base_url="http://localhost:8888")
client.retain(bank_id="my-bank", content="Alice loves hiking in Yosemite")
client.recall(bank_id="my-bank", query="What does Alice like?")
client.reflect(bank_id="my-bank", query="What should I know about Alice?")
- retain stores and extracts.
- recall retrieves across all memory types, merges with reciprocal rank fusion, reranks with a cross-encoder and trims to a token budget.
- reflect reasons over memories to form new connections or answer a question that needs thinking rather than lookup — a project-manager agent reflecting on risks, a sales agent on why some outreach got replies.
Banks are strictly isolated (“one brain per user, agent or project,” no cross-bank leakage) and carry disposition traits — skepticism, literalism, empathy — that shape how reflect reasons.
How you run it
- Docker, one line:
docker run … ghcr.io/vectorize-io/hindsight:latestwithHINDSIGHT_API_LLM_API_KEY; API on :8888, UI on :9999. - Embedded Python, no server:
pip install hindsight-api. - Helm for Kubernetes; Oracle AI Database as an alternative store for enterprise; air-gapped and Windows setups documented.
- LLM providers: 25+ — OpenAI, Anthropic, Gemini, Groq, Bedrock, Vertex, MiniMax, DeepSeek, Meta, fully local via Ollama / LM Studio / llama.cpp, any OpenAI-compatible endpoint, and LiteLLM gateways. Existing ChatGPT Plus/Pro, Claude Pro/Max, Cursor and GitHub Copilot subscriptions work as the LLM with no API key.
- Clients: Python, Node/TypeScript, Go, CLI.
wrap_openai()/wrap_anthropic()add memory to an existing SDK call in two lines. - Coding agents: one package gives Claude Code, Cursor and similar a per-repo bank built from git history and past sessions, plus knowledge pages on architecture and conventions. There is also an MCP server exposing retain/recall/reflect as tools.
- Hindsight Cloud: usage-based, free credits to start, no per-seat fee, 99.9% SLA.
Benchmarks, with the caveat
Hindsight’s README calls it “the most accurate agent memory system ever tested” on LongMemEval, and says its numbers were independently reproduced by Virginia Tech’s Sanghani Center and The Washington Post, while competitors’ scores are self-reported. Vectorize’s own comparison page reports 94.6% against Supermemory’s self-reported 81.6%. Mem0 reports 94.4% on LongMemEval on its own research page (updated September 22, 2026). None of these runs share a model stack, judge or retrieval config, so treat the leaderboard as a starting point. Live per-model accuracy, latency and cost are published at benchmarks.hindsight.vectorize.io.
When Hindsight is the right choice
- Agents that run across sessions and need to get better — AI employees, support agents, research agents, personal assistants — especially now that Muse, Autopilot and “o”-style agents make cross-session memory a product requirement.
- Teams frustrated with RAG latency who want precomputed answers (mental models) at boot.
- Anyone who needs strict per-user isolation with metadata filtering for compliance.
It is overkill for a chatbot that only needs last-week’s conversation. For the wider field — Mem0, Zep, Letta, cognee — see best AI agent memory systems 2026, and for the security side, how to secure AI agent memory.
Last verified: September 28, 2026.