The strongest agent-memory test pairs an assertion about what the system stored or retrieved with an assertion about what the agent later did because of it. Recall alone can pass even when memory is stale, out of scope, or ignored during tool use.
What memory decay looks like in an agent
Decay is not limited to a system forgetting a fact. A memory layer can lose important details during compression or maintenance, keep a superseded value active, merge claims that belong to different projects, retrieve a correct fact but apply it incorrectly, or answer confidently when it has no supporting evidence. The AgingBench paper record discusses degradation mechanisms and diagnostic probes; MELT’s lifecycle dimensions cover correction, contradiction, scope, maintenance, provenance, and abstention.
As an Amazon Associate I earn from qualifying purchases.
Test the full path: a fact enters memory, changes or survives as intended, is retrieved in the right context, and affects a later decision or state change. Exact memory wording is usually less useful to assert than whether the required meaning, scope, and evidence remain intact.
Assertions to add to a memory test suite
1. Verify the write, not just the conversation
Give the agent a decision-relevant fact, then inspect the normalized memory representation. Assert that it preserves the essential fact and any scope or source needed to use it safely. Avoid matching a particular sentence unless exact wording is part of the memory layer’s contract. MELT treats write quality and provenance as evaluation dimensions (MELT documentation).
#1 Best Overall
2. Separate correction from historical recall
Store an initial value, then provide an explicit correction. A current-time query should return the corrected value. If the product is expected to retain history, an as-of query should still return the earlier value for the earlier time. These are separate assertions: the system should not treat an obsolete value as current, but it also should not erase history that the test requires. MELT distinguishes correction from temporal recall (MELT documentation).
3. Test real conflicts without inventing them
Supply two incompatible claims with the same scope and no explicit correction. Assert that the agent preserves the conflict or qualifies its answer rather than silently combining the claims. Then vary the project, user, or time. A difference that is valid in its own context should not be treated as a contradiction. MELT identifies contradiction and conflict precision as distinct evaluation dimensions (MELT documentation).
Rank #2
4. Run maintenance between writing and checking
Place consolidation or another maintenance operation between the memory write and the later query. Assert that durable preferences or identity facts remain available, and that explicitly expired or revoked information is not presented as current truth. Set the expiration policy in the fixture: there is no universal decay interval established by the cited evaluations. MELT includes maintenance, decay, and core memory among its test dimensions (MELT documentation).
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →5. Check project and user isolation
Write similar facts into two projects, users, or workspaces, then query each scope independently. Assert that each response uses only the appropriate memory unless sharing was explicitly enabled. Similarity is not authorization: a fact that resembles the current task can still belong to the wrong context. Project scope is one of MELT’s lifecycle dimensions (MELT documentation).
6. Preserve provenance and require abstention
When the agent answers from stored information, check that the relevant source identity and scope survived updates and retrieval. For a query unsupported by memory, assert an abstention or a clear statement that the information is unknown—not a plausible-sounding invention. MELT lists provenance and abstention as evaluation dimensions (MELT documentation).
7. Prove the memory changes a later action
Across interrupted sessions, establish a preference or task state, then give the agent a tool-using task where that information should affect tool choice or arguments. Assert both the selected action and its parameters, then verify the resulting state where it is observable. A separate recall question is insufficient: the agent might retrieve the right fact but fail to use it. Mem2ActBench targets proactive memory use for tool selection and parameter grounding; MemoryArena tests interdependent multi-session tasks where prior experience should guide later actions.
Rank #4
8. Assert external state transitions deterministically
For tools that alter records or other external state, verify the final state directly and assert any required procedural steps. Do not use a fluent completion message as proof that the action happened. STATE-Bench describes pre-populated task environments with deterministic state assertions, a useful model for this kind of check.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Use paired counterfactuals to diagnose failures
For a downstream task, run comparable cases with the relevant memory present, corrected, missing, or stored under another scope. Treat this as a diagnostic design, not a standardized protocol. If the outcome stays the same when the relevant fact changes, the agent may be ignoring memory. If it changes when only an irrelevant or out-of-scope fact changes, retrieval or isolation may be too broad. AgingBench describes paired counterfactual probes and temporal dependency graphs for diagnosing write, retrieval, and utilization stages (paper record).
Keep the task and other inputs fixed as closely as possible, and compare the action and final state—not only the answer text. A useful failure trace records which memory was written, which was retrieved, its scope and provenance, the chosen tool and arguments, and the observed state transition. Those checkpoints help distinguish a write failure from a retrieval failure or a failure to act on retrieved information.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Why recall-only scores can miss the problem
Recall questions test whether an answer can be retrieved, but agent memory also has to support decisions across sessions and tool interactions. MemoryArena’s 2026 paper says existing evaluations often assess memorization and action in isolation; its interdependent tasks connect experience from one session to later decisions, and it reports that systems near saturation on LoCoMo perform poorly in its agentic setting (MemoryArena paper record). AMA-Bench argues that realistic agent memory includes trajectories of states, actions, observations, and tool outputs—not only dialogue—and identifies missed causal or objective information and lossy similarity-based retrieval as problems (AMA-Bench paper record). Mem2ActBench focuses on applying memory in tool execution rather than merely recalling isolated facts (Mem2ActBench paper record).
What the available test suites emphasize
| Suite | Emphasis | Scale or design stated in its source |
|---|---|---|
| MemoryArena | Interdependent tasks across sessions; prior experience should influence later actions. | The cited paper record describes the benchmark design; no task count is stated here. |
| AMA-Bench | Long-horizon agent memory that includes state, action, observation, and tool-output trajectories. | The cited paper record describes its focus; no task count is stated here. |
| Mem2ActBench | Using long-term memory for tool selection and parameter grounding. | Its 2026 construction comprised 2,029 synthesized sessions, averaging 12 user–assistant–tool turns; it included 400 tool-use tasks, 91.3% of which human evaluation judged strongly memory-dependent. These are benchmark construction and evaluation figures, not production score targets. |
| STATE-Bench | Tasks in pre-populated environments with deterministic assertions about state. | Microsoft Open Source’s 2026-05-19 announcement describes 450 tasks across customer support, travel, and shopping; this is the announced release scale, not a universal coverage requirement. |
| MELT | Lifecycle dimensions including correction, contradiction, scope, maintenance, provenance, and abstention. | The project documentation describes these evaluation dimensions; no task count is stated here. |
| AgingBench | Agent aging and diagnostic probes, including paired counterfactuals and temporal dependency graphs. | The 2026 paper record reports about 400 runs across seven scenarios and 14 models, spanning 8–200 sessions. This describes study scale, not a benchmark score or a prediction that all systems age alike. |
No single cited source establishes a universally complete assertion suite. When choosing or adapting a suite, check whether it covers active use as well as recall, multiple sessions, tool calls and visible state changes, correction versus contradiction, temporal queries, scope isolation, maintenance, provenance, and abstention. Reproducible tasks, baselines, seeds, and scoring also matter when comparing results.
Recommended Free Tools
Quick Recap
Turn the assertions into a regression gate
- Define the contract. Specify what counts as a durable fact, an explicit correction, an unresolved conflict, an expired item, and a permitted sharing scope.
- Build a multi-session fixture. Include an initial memory event, later updates or maintenance, and a downstream task that depends on the expected memory state.
- Assert state and behavior separately. Check memory content, time and scope where relevant; then check the agent’s response, tool selection, parameters, and external state as appropriate.
- Add counterfactual cases. Change or remove only the fact under test to reveal whether it affects the dependent action and whether unrelated memories improperly affect it.
- Keep failure evidence. Record the inputs, retrieved evidence, action, and resulting state so a failure can be localized to writing, maintenance, retrieval, scope, or use.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




