An AI agent should not carry every past conversation into every new prompt. That makes the prompt grow as history accumulates, increasing latency and cost, while still leaving the agent with a retrieval problem: it must find and correctly interpret the information that matters now. Better memory is not unlimited storage. It is a system for deciding what to retain, how to update it, and when to bring it back.
Why not include the entire conversation history?
Putting the full conversation into each new prompt is the simplest continuity strategy: the model can refer directly to what was said earlier. But every additional exchange lengthens the prompt. Redis AI Research describes the resulting trade-off as longer prompts, slower responses, and greater expense as history grows.
External memory changes the flow. The system processes earlier interactions into a store, then retrieves selected material for a later task. This can keep the active prompt smaller, but it replaces the full-history problem with several linked jobs: ingest information, retain or update it, retrieve it when relevant, and interpret it in the new context. A stored fact that is never retrieved—or is retrieved at the wrong moment—does not provide useful continuity.
What can go wrong when memory is compressed or retrieved?
Extracted facts can omit the detail a later task needs
A system that converts conversations into compact facts can consolidate information across sessions and represent changes. But the compression is selective. If an exact name, date, number, qualification, or piece of wording was not extracted, it may not be available in the fact store when a later question depends on it.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors#1 Best Overall
Similarity search can find related text but miss the relationship
Raw excerpts preserve what was actually said, but the system still has to retrieve the right passage. A later request may use different wording, or depend on when something happened, why an action was taken, or how several steps relate. The AMA-Bench authors argue that agent trajectories contain states, actions, observations, and tool outputs, and report that systems relying heavily on lossy similarity-based retrieval can miss causal or objective information.
Old information can conflict with a change
Suppose a person first says they plan to visit one city, then later changes the destination. A memory system needs to distinguish the current plan from the earlier statement, rather than treating both as equally current or retrieving whichever is more textually similar. The same issue applies to preferences, project decisions, and constraints. Keeping a dated source or other provenance alongside a summary can help a system establish what a memory came from and whether it has been superseded.
Rank #2
What kinds of agent memory are available?
| Approach | What it retains or organizes | Main design pressure |
|---|---|---|
| Raw-text storage or indexing | Conversation passages, which can preserve exact wording and detail | Retrieval must locate the relevant passage, including when the later query is phrased differently |
| Extracted facts | Compact statements distilled from interactions, potentially including updates | Details not captured during extraction may not be recoverable from the fact store |
| Structured or graph-like memory | Information organized into explicit entities or relationships | The structure and update rules must fit the task; an organized store still has to retrieve the right information |
| Hierarchical memory systems | Multiple layers or components coordinating storage, updating, retrieval, and response generation | More coordination can create additional design choices; the architecture is not a universal guarantee of better results |
| Hybrid facts plus raw excerpts | Compact extracted facts alongside source passages | Can make both summaries and exact evidence available, but results depend on the implementation and evaluation setup |
These are design families, not a ranking. The useful choice depends on what an agent must remember and the cost of getting it wrong. For an assistant that needs exact figures or wording, preserving source excerpts may matter. For recurring preferences or a plan that changes over time, a system also needs a way to consolidate and update facts.
What do the reported benchmark results actually show?
Published results offer evidence about particular tasks and configurations, not a single league table for agent memory. The studies below use different benchmarks, so their figures should not be compared as though they measured the same thing.
| Work and evaluation | Reported result | What the result is evidence for |
|---|---|---|
| SimpleMem authors, LoCoMo, 2026 | 26.4% average F1 improvement | The authors’ result for SimpleMem on LoCoMo; it is not a universal improvement for memory systems. |
| SimpleMem authors, inference-time token consumption, 2026 | Up to 30× lower | An experimental claim from the same paper. “Up to” matters; it is not a guaranteed reduction in every deployment. |
| Redis AI Research, LongMemEval Small, 2026 | 86.1% task-averaged accuracy for a configuration combining raw excerpts and extracted facts | Redis’s reported evaluation on the Small split, described as 500 questions across multi-session chat histories. It supports that configuration in that evaluation, not a controlled conclusion about every production setting. |
| AMA-Agent authors, AMA-Bench, 2026 | 57.22% accuracy and an 11.16 percentage-point lead over the strongest baseline | The PMLR record’s reported result on AMA-Bench, which focuses on realistic agent trajectories. |
| Microsoft Research, Memora, standard long-conversation benchmarks, 2026 | Up to 98% fewer context tokens than full-history prompting | Microsoft Research’s project-blog claim for those benchmarks. It should not be generalized to every agent or memory workload. |
The Redis evaluation is notable because it combines extracted facts with raw excerpts: the facts offer a compact representation, while excerpts retain access to source detail. That is a plausible hybrid pattern, not proof that every agent should use the same design. Likewise, SimpleMem’s LoCoMo and token-consumption findings, AMA-Agent’s AMA-Bench accuracy, and Microsoft’s Memora token claim answer different evaluation questions.
How should a team choose what to retain?
There is no standardized universal score for memory quality. A practical review should make the trade-offs explicit across the following dimensions:
Rank #4
- Recall and fidelity: Can the system recover the detail the task needs, including names, dates, numbers, and qualifications?
- Change over time: Can it represent a revised preference, plan, or decision without presenting stale information as current?
- Retrieval quality: Can it find relevant information when the query uses different wording or depends on causal, temporal, or multi-step relationships?
- Cost and latency: What processing is required when information is written, and what must happen each time the agent reads or queries memory?
- Transparency and control: Can people inspect, correct, approve, or remove information, and understand why it influenced a response?
These dimensions help expose where a system is making a trade rather than hiding it behind a single memory-size or benchmark number. For example, compact facts may reduce what has to be searched at answer time, while source excerpts may provide stronger access to exact evidence. Which matters more depends on the task and its failure costs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What should a practical memory pipeline do?
For builders, separate the work that happens when information enters memory from the work that happens when a new request arrives. This makes it easier to test whether a failure came from extraction, updating, retrieval, or interpretation.
Recommended Free Tools
Best Value
- Ingest with provenance. Store the source or enough traceable evidence to check where an extracted fact came from. Preserve raw excerpts when exact wording or detail could matter later.
- Represent updates explicitly. Record when a statement was made and provide a way to mark a later statement as a correction or replacement. Do not assume that two conflicting memories can be resolved by similarity alone.
- Retrieve for the current task. Select relevant facts and excerpts for the request rather than treating all stored material as automatically useful. Test queries that are phrased differently from the original conversation and that depend on time or cause.
- Interpret with the source in view. Distinguish a remembered statement from a current fact or commitment. If evidence is ambiguous or contradictory, the response should not silently present one version as certain.
- Make memory inspectable. Give people a way to see, correct, or remove remembered information and, where appropriate, approve how it has been interpreted.
- Evaluate each stage. Check whether important information was captured, whether changes were represented, whether the right material was retrieved, and whether the final answer used it correctly. Measure latency and token use in the same task setup as answer quality.
Why does user control belong in the design?
A research poster on user perceptions of AI memory uses questions such as “Does it save everything?”, “What does the AI take in?”, and “Why did it bring that up?” as examples of concerns raised in the study—not as evidence that every user asks those questions. The poster reports that participants evaluated memory partly through how prior information was recalled and interpreted, and points to interest in transparency and the ability to see, edit, or approve interpretations.
That concern is practical as well as personal. When a response draws on an earlier conversation, people need a way to understand which remembered information shaped it and to correct a mistaken or outdated interpretation. Memory quality therefore includes more than whether information was stored: it also includes whether its use can be understood and challenged.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




