TL;DR

Hindsight is an open-source agent memory server from Vectorize.io. You push text into it (retain), pull relevant memories back out (recall), and ask it to reason over everything it knows (reflect). Underneath is PostgreSQL with pgvector, an LLM that extracts facts, entities and timestamps on the way in, and a four-way retrieval pipeline (vector, BM25 keyword, entity graph, time range) on the way out. The pitch is in the tagline: memory that learns — consolidating raw facts into “observations” and standing “mental models” instead of piling up chat snippets.

As of October 5, 2026 the repository has 45,676 stars (10,623 of them added in the past week, which put it near the top of GitHub’s weekly trending list), 5,921 forks, 143 open issues, and its latest release is v0.10.2 (September 29, 2026). License: MIT.

  • Benchmarks (the project’s numbers): the team’s arXiv paper reports 91.4% on LongMemEval and up to 89.61% on LoCoMo, and says a 20B open model with Hindsight goes from 39% to 83.6% over a full-context baseline. The README says independent researchers reproduced the results.
  • Integrations: 60+ listed — Claude Code, Codex, Cursor, LangGraph, CrewAI, Pydantic AI, Vercel AI SDK, n8n and more — plus a built-in MCP endpoint per memory bank.
  • LLM providers: 25+, including local ones (Ollama, LM Studio, llama.cpp) and a none mode that runs it as a plain chunk store with no LLM at all.
  • Our lab result: the README’s one-line Python install (pip install hindsight-all -U) failed after 226 seconds with “No space left on device” on a standard GitHub-hosted runner. Details in the hands-on section below.

New to the project? Our what-is explainer covers the concepts in brief; this review goes further into setup, the install footprint, published benchmarks and alternatives.

Who should look at it: teams building agents that run for weeks or months against the same users or projects, and who are willing to operate a Postgres-backed service for that. Who should skip it: anyone who just needs “remember the last few turns.”

What problem Hindsight is trying to solve

Most agent memory today works like this: after each conversation, an LLM pulls out “salient” snippets, they get embedded, and the next session retrieves the top-k nearest snippets into the prompt. That is fine for “the user’s name is Alice.” It falls apart in three places that the Hindsight paper names directly:

  1. Evidence and inference blur together. A stored snippet might be something the user said, or something the agent guessed. Retrieval cannot tell them apart.
  2. Long horizons get messy. “Alice works at Google” from January and “Alice joined Anthropic” from June are both close to the query “where does Alice work?” Similarity search does not know which is current.
  3. No explanation. When an agent acts on a memory, you usually cannot trace which piece of evidence drove the decision.

Hindsight’s answer is to give memory structure before it hits the vector index, and to keep a background process that reconciles what has been stored.

How it works

Four kinds of memory

Every memory bank (one isolated “brain” per user, agent or project) separates what it stores into:

TypeWhat it holdsExample from the README
World factsThings true about the world”The stove gets hot”
ExperiencesWhat the agent itself did and saw”I touched the stove and it really hurt”
ObservationsConsolidated beliefs built from many facts, each with quoted evidence and a proof countFormed in the background
Mental modelsStanding answers to questions you define once”What are this user’s preferences?”

The observation layer is the interesting part. According to the docs, when new evidence arrives an existing observation is refined — strengthened, weakened or extended — rather than silently overwritten. That is the direct fix for the “January Alice vs. June Alice” problem.

Mental models are pre-computed answers. Reading one is a database read with no retrieval and no LLM call, so an agent can boot with a page of settled knowledge. “Knowledge pages” are the same idea exposed as a wiki of markdown documents the bank writes about itself.

Three operations

from hindsight_client import Hindsight

client = Hindsight(base_url="http://localhost:8888")

# Retain: store information (an LLM extracts facts, entities, time)
client.retain(
    bank_id="my-bank",
    content="Alice got promoted to senior engineer",
    context="career update",
    timestamp="2025-06-15T10:00:00Z",
)

# Recall: four retrieval strategies in parallel, fused and reranked
client.recall(bank_id="my-bank", query="What happened in June?")

# Reflect: reason over the whole bank, not just look things up
client.reflect(bank_id="my-bank", query="What should I know about Alice?")

Recall runs semantic (vector), keyword (BM25), graph (entity/temporal/causal links) and temporal (time range) retrieval in parallel, merges them with reciprocal rank fusion, reranks with a cross-encoder, then trims to a token budget. Hybrid retrieval is not new, but shipping all four plus a reranker as the default is more than most memory libraries do out of the box.

Reflect is the expensive one: an LLM reasons over the bank to answer questions like “why did some outreach emails get replies and others didn’t?” Banks can carry “disposition traits” (skepticism, literalism, empathy) that shape how reflect reasons.

Quick start

Docker (the README’s first option)

export OPENAI_API_KEY=sk-xxx

docker run -it --pull always --name hindsight --restart unless-stopped \
  -p 8888:8888 -p 9999:9999 \
  -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
  -v hindsight-data:/home/hindsight/.pg0 \
  ghcr.io/vectorize-io/hindsight:latest

That gives you the API on port 8888 and a web UI on port 9999. The image bundles an embedded PostgreSQL (pg0), so you do not need a separate database to try it. The default model is gpt-5-mini; switch with HINDSIGHT_API_LLM_PROVIDER and HINDSIGHT_API_LLM_MODEL. Notably, existing subscriptions work as providers too: claude-code, openai-codex, cursor and github-copilot need no separate API key, per the README.

Embedded Python (no server)

import os
from hindsight import HindsightServer, HindsightClient

with HindsightServer(
    llm_provider="openai",
    llm_model="gpt-5-mini",
    llm_api_key=os.environ["OPENAI_API_KEY"],
) as server:
    client = HindsightClient(base_url=server.url)
    client.retain(bank_id="my-bank", content="Alice works at Google")
    results = client.recall(bank_id="my-bank", query="Where does Alice work?")

This is the pip install hindsight-all path — the one we tried in the lab.

The two-line LLM wrapper

The lowest-effort integration swaps your OpenAI or Anthropic client for a wrapped one. Memories are recalled before each call and the conversation is retained after it:

from openai import OpenAI
from hindsight_litellm import wrap_openai

client = wrap_openai(
    OpenAI(),
    bank_id="user-123",
    hindsight_api_url="http://localhost:8888",
)

response = client.chat.completions.create(
    model="gpt-5-mini",
    messages=[{"role": "user", "content": "What do you know about me?"}],
)

One gotcha from the README’s own comment: the wrapper defaults to Hindsight Cloud unless you pass hindsight_api_url. If you are self-hosting for privacy, set it explicitly.

Coding agents and MCP

npx @vectorize-io/hindsight-coding-agents install all

This wires a per-repo memory bank, built from git history and past sessions, into Claude Code, Codex CLI, Cursor CLI, Copilot CLI, opencode and several others. Separately, every server exposes an MCP endpoint per bank at http://localhost:8888/mcp/{bank_id}/, so any MCP client gets retain, recall and reflect as tools.

Running without any LLM

Buried in the configuration reference is HINDSIGHT_API_LLM_PROVIDER=none. In that mode, retain stores raw chunks (no fact extraction), recall still works through embeddings and keyword search, and reflect, consolidation and mental-model refresh are disabled — the project’s own test suite checks that reflect returns an error. It turns Hindsight into a hybrid-search chunk store, which is useful for kicking the tyres without spending tokens, but it is not the product the benchmarks describe.

Hands-on test (October 5, 2026)

We tried to run the README’s embedded-Python path in a clean python:3.12-bookworm container on a GitHub-hosted runner (x86_64, 2 CPUs, 8 GB RAM, no GPU), against the repository at commit 2105a458afc4. The plan was four steps: install with the README’s command, print versions, create a non-root user (the embedded Postgres will not run as root), then start an embedded server with llm_provider="none" — so no API key was needed — retain three short facts and recall them.

It stopped at step one. pip install hindsight-all -U ran for 226 seconds and exited with:

ERROR: Could not install packages due to an OSError: [Errno 28] No space left on device

The other three steps were skipped. Verdict: failed, 0 of 4 steps passed. The full record, with every command and its output, is in the box at the top of this page.

What this does and does not tell you:

  • It is about install footprint, not about memory quality. We never got to retain or recall, so this review says nothing first-hand about retrieval accuracy or speed.
  • The likely cause is the dependency tree. hindsight-all 0.10.2 depends on hindsight-api-slim[all], and per the project’s pyproject.toml the all extra pulls in local-ml (PyTorch, sentence-transformers, transformers, flashrank) plus local-onnx — that is, local embedding and reranking models. On Linux x86_64, a default PyTorch install is large. We did not measure the final size, so treat this as an inference from the dependency list, not a measurement.
  • The project’s numbers do not tell you any of this. The benchmark leaderboard reports accuracy, latency and cost per model. Disk footprint and cold-install time are not on it, and they matter on small CI runners, cheap VPSes and serverless builders.
  • Practical takeaway: if you are tight on disk, use the Docker image (dependencies are pre-built into layers you pull once) or pip install hindsight-api against an external Postgres, and budget several gigabytes either way. The README recommends hindsight-all-slim only for Intel Macs; whether it suits small Linux boxes is something we have not tested.

We did not re-run with a different install method: the README’s command is what most readers will paste first, and this is what happened when we did.

Benchmarks (as published by the project)

Every number in this section comes from Vectorize, its paper, or its benchmark site — not from us.

  • LongMemEval: 91.4% with a larger backbone model (arXiv:2512.12818).
  • LoCoMo: up to 89.61%, versus 75.78% for what the paper calls the strongest prior open system.
  • Same-model lift: with an open-source 20B model, overall accuracy rises from 39% to 83.6% compared to stuffing the full conversation into context, and the paper says it beats full-context GPT-4o.
  • Reproduction: the README states the results were independently reproduced by Virginia Tech’s Sanghani Center for AI and Data Analytics and The Washington Post, and that competitors’ scores in its comparison chart are self-reported by those vendors.
  • Live leaderboard: benchmarks.hindsight.vectorize.io publishes per-model accuracy, latency and cost for different LLM, embedding and reranker choices.

Two cautions. First, LongMemEval and LoCoMo measure conversational recall — “what did the user say three sessions ago?” — which is narrower than “did the agent get better at its job?” Second, the memory-benchmark space is crowded with vendor-run comparisons, and several competing systems publish their own LongMemEval scores under their own settings. Differences of a point or two between systems run by different teams under different settings are not meaningful.

Community reactions

Hindsight has been around since late 2025 — the repository was created October 30, 2025, and the launch post on r/Rag in December headlined the 91.4% LongMemEval score. The discussion since then has a consistent shape:

  • The architecture gets respect; the operational weight gets hesitation. A May 2026 r/AI_Agents post titled “Goldfish brains: Why my 5-agent setup forgets everything — I tested Hindsight, here’s why I’m waiting” says the design (per-agent banks, Postgres backend, embeddings) “is sound” but that the author held off adopting it for their pipeline.
  • People compare footprints. In an r/hermesagent thread comparing memory providers, Hindsight is listed as an external Postgres instance (184 MB in that user’s setup) that “survives all clients dying,” next to single-SQLite-file alternatives. That is the trade-off in one line: durability and a real database versus one more service to run.
  • Self-hosting questions dominate. In an r/AI_Agents “Supermemory vs Hindsight” thread, one reply focuses on operations — pgvector load, metadata tables, summarization on the write path — rather than features.
  • There is a counter-argument to agent memory itself. A recurring critique this year is that similarity search tells you two snippets are close, not which one is correct or current, and that agents are better served by maintained documentation. Hindsight’s observations and knowledge pages are, in effect, an attempt to meet that critique halfway: memory that maintains its own documents.

Honest limitations

  1. Heavy install. Our lab hit a disk-space failure with the README’s pip command on a standard 2-CPU runner. Plan for Docker or an external Postgres plus hindsight-api.
  2. LLM cost on every write. In normal mode, retain calls an LLM to extract facts, and consolidation and mental-model refreshes run more LLM calls in the background. The docs offer knobs (per-operation models, a MIN_REFRESH_INTERVAL to batch refreshes, flex service tiers) — the fact that those knobs exist tells you the cost is real at volume.
  3. It is a service, not a library. Postgres with pgvector (or Oracle 23ai), migrations, a worker, an admin CLI for “stuck operations.” The production feature list (Prometheus metrics, webhooks, per-tenant config) is good, but it is also a list of things to operate.
  4. Overkill for simple flows. The README itself says Hindsight “may be overkill” for simple n8n-style workflows.
  5. Cloud defaults in places. The LLM wrapper points at Hindsight Cloud unless told otherwise; read integration defaults before sending user data anywhere.
  6. Benchmarks are conversational. Strong LongMemEval numbers do not prove an agent learns a task better over time, which is the bigger claim in the tagline.
  7. Fast-moving. Three releases in September alone (v0.10.0 on Sept 14, v0.10.1 on Sept 21, v0.10.2 on Sept 29), with 143 open issues. Pin versions.

When to use Hindsight — and what to pick instead

Use it if you run long-lived agents (support, sales, project management, coding assistants) against the same people or repositories for months; you want memories that reconcile contradictions over time rather than accumulate them; and you already run Postgres or are happy to.

Look elsewhere if:

  • You want a lighter, managed-first API → Supermemory.
  • You want to model users as “peers” with a theory-of-mind angle → Honcho.
  • You want full control over a document-to-graph pipeline → Cognee.
  • You want everything local in a single process with no database server → MemPalace or a SQLite-based option.
  • You only need the last few turns → no memory system at all; a rolling summary is enough.

FAQ

Is Hindsight free and open source?

Yes. The code is MIT-licensed and fully self-hostable. Vectorize also sells Hindsight Cloud, a managed version with usage-based billing, free starting credits, a dashboard and a 99.9% uptime SLA.

Can Hindsight run without an OpenAI key?

Yes. It supports 25+ providers, including fully local Ollama, LM Studio and llama.cpp, plus claude-code, openai-codex, cursor and github-copilot subscriptions without a separate key. With HINDSIGHT_API_LLM_PROVIDER=none it runs with no LLM at all, but then only stores chunks and searches them — no fact extraction, reflect or consolidation.

What database does Hindsight need?

PostgreSQL with pgvector, or Oracle AI Database 23ai. The Docker image and the hindsight-all Python package include an embedded Postgres (pg0) for local use; that embedded database refuses to run as root.

How is Hindsight different from RAG?

RAG retrieves chunks of documents by similarity. Hindsight extracts facts, entities and timestamps on write, keeps world facts separate from the agent’s own experiences, retrieves with four strategies at once, and consolidates facts into observations and mental models in the background. The project’s “RAG vs Memory” doc page covers its own view of the difference.

Does Hindsight work with Claude Code and other coding agents?

Yes. npx @vectorize-io/hindsight-coding-agents install all wires a per-repository memory bank into Claude Code, Codex CLI, Cursor CLI, Copilot CLI, opencode and others, built from git history and past sessions. Any MCP client can also connect to http://localhost:8888/mcp/{bank_id}/.

Did you run Hindsight yourselves?

We tried. On October 5, 2026 the README’s pip install hindsight-all -U failed with “No space left on device” after 226 seconds on a 2-CPU, 8 GB GitHub-hosted runner, so we did not reach retain or recall. All accuracy figures in this review are the project’s.

Verdict

Hindsight is the most ambitious open-source take on agent memory right now: structured memory types, hybrid retrieval with reranking, background consolidation that refines beliefs instead of overwriting them, and a strong — and partly independently reproduced — benchmark story. The integration list is long enough that wiring it into an existing agent is a small job.

The cost is weight. It is a Postgres-backed service with an LLM in the write path, and the quick-start Python install was too big for a standard CI runner in our lab. If you are building agents meant to live for months, it deserves a serious trial — start with the Docker image, not pip, and measure the token bill from retain and consolidation before you scale it.

Sources