Which AI Agent Frameworks Support Long-Term Memory? 2026
The short answer
LangGraph, Mastra, Letta, CrewAI and Google’s ADK all have long-term memory built in; the OpenAI Agents SDK has session history only. They differ in who decides what to remember: in LangGraph and Mastra your code (or a background process) writes memories; in Letta the agent edits its own memory blocks with tools; in CrewAI an LLM infers scope and importance when saving. Facts below are from each framework’s documentation, read October 8, 2026.
The comparison
| Framework | Language | Short-term memory | Long-term memory | Who writes memories | Pick it when |
|---|---|---|---|---|---|
| LangGraph | Python, TypeScript | Thread state persisted by a checkpointer | Stores with custom namespaces, shared across threads | Your graph nodes | You want explicit control and durable, resumable workflows |
| Mastra | TypeScript | Message history per thread | Working memory, semantic recall, Observational Memory | Mastra (background agents) + config | You build in TypeScript and want memory that compacts itself |
| Letta | Python, TypeScript SDKs; server | In-context messages | Core memory blocks pinned in context + full persisted history | The agent, via memory tools | Long-lived agents that should learn and edit their own memory |
| CrewAI | Python | Inside the unified Memory | Unified Memory with scopes, categories and importance | LLM-inferred on save | Multi-agent crews that share project knowledge |
| Google ADK | Python, Java and more | Agent Platform Sessions | Agent Platform Memory Bank | Memory Bank service | You run on Google Cloud |
| OpenAI Agents SDK | Python, TypeScript | Sessions (e.g. SQLiteSession) | Not built in — add a memory layer | — | You want OpenAI’s runtime and will bring memory |
Framework notes
LangGraph. The cleanest mental model. Short-term memory is thread-scoped: the conversation lives in the graph’s state, saved to a database by a checkpointer so a thread can resume at any time. Long-term memory lives in stores, keyed by custom namespaces (for example a user ID) rather than thread ID, so it can be recalled in any thread. Nothing is remembered unless a node writes it, which makes behaviour predictable and testable.
Mastra. Memory is layered and configurable. Message history keeps recent turns; working memory stores persistent structured data such as names, preferences and goals; semantic recall retrieves past messages by meaning; and Observational Memory — the option Mastra’s docs now recommend — uses background agents to maintain a dense observation log that replaces raw history as it grows, keeping the context window small. Multi-user threads let several users share one thread, and memory processors filter or trim when the total would exceed the model’s limit.
Letta. Built around stateful agents. Every message, tool call and reasoning step is persisted in a database, so nothing is lost when it leaves the context window. Important “core” memories are memory blocks pinned into the system prompt, and the agent rewrites them with memory tools. Blocks can be attached to several agents at once, which is how teams share state between agents.
CrewAI. Replaced its separate short-term, long-term, entity and external memory types with one Memory class. On save an LLM infers scope, categories and importance; recall uses a composite score of semantic similarity, recency and importance, with tunable weights such as a recency half-life. It works standalone, with crews, with single agents or inside Flows.
Google ADK. Google’s 2026 rename turned Vertex AI Agent Engine Sessions and Memory Bank into Gemini Enterprise Agent Platform Sessions and Memory Bank; ADK agents use them for session state and long-term memory when deployed on Google Cloud.
OpenAI Agents SDK. Sessions automatically store conversation history across runs so you don’t pass message lists by hand, and can be swapped for OpenAI’s server-managed continuation. That is conversation memory, not extracted long-term facts — for those, call a memory layer from a tool.
How to choose
- Python, production workflows, want to see every write: LangGraph.
- TypeScript product team: Mastra.
- Agents that run for months and should learn: Letta.
- Several cooperating agents sharing knowledge: CrewAI or Letta shared blocks.
- Multi-framework estate or strict tenant isolation: a dedicated layer — see the best AI agent memory systems and agent memory for multi-user chatbots.
Whichever you choose, scope every memory by user and tenant ID and enforce the filter in code, and give memories an expiry: stale memories cause more wrong answers than missing ones. For the difference between memory and retrieval, see agent memory vs chat memory vs RAG.
Last verified: October 8, 2026, against each framework’s documentation.