October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk6 min

Building a Temporal Memory Graph for Agents with Hindsight

Hindsight combines four kinds of structured agent memory with vector, keyword, graph, and temporal retrieval. Here is how its architecture works and what its reported scores can—and cannot—tell you.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Hindsight is an agent-memory architecture that organizes information into four logical networks and combines vector search, keyword matching, graph traversal, and temporal filtering. Its design aims to help an agent retrieve not only semantically related material, but also relevant entities, relationships, and information that changes over time. The published results are promising, but they are benchmark findings reported by Hindsight’s authors—not a guarantee of performance for every agent or workload.

What Hindsight means by a temporal memory graph

Hindsight treats memory as structured information an agent can query and reason over, rather than only a collection of selected past conversation snippets. Its authors describe a temporal, entity-aware layer that incrementally turns conversational streams into a queryable memory bank. The architecture separates that memory into four logical networks:

Network What it represents in Hindsight Why the separation matters
World Facts about the world Separates claims treated as facts from the agent’s personal history or beliefs.
Experience The agent’s experiences Preserves what the agent has done or encountered, rather than blending it with general world knowledge.
Observation Synthesized summaries about entities Provides a consolidated view of an entity from the information retained about it.
Opinion Evolving beliefs Allows the agent’s beliefs to be represented distinctly from world facts and updated over time.

This four-part model is Hindsight’s design, not a universal standard for agent memory. It is useful to think of the networks as different kinds of records that may relate to the same entity: a person’s stated preference might be an observation, an agent’s prior interaction with that person an experience, and the agent’s current interpretation an opinion. The distinction helps avoid treating an inference or belief as if it were an unqualified fact.

Time matters because stored information can become stale or be superseded. A temporal memory system needs to preserve enough context to distinguish what was true at one point from what is believed or known now. Hindsight presents its entity-aware and temporal design as a way to maintain that context and update information traceably. The sources describe the architectural intent, but do not establish one universal schema or configuration for every application.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How retain, recall, and reflect work

Hindsight groups its workflow into three operations: retain, recall, and reflect. Together, they cover storing information, retrieving it, and reasoning over what has been stored.

Retain: ingest and organize information

Retain handles ingestion. The system takes conversational information and builds it into structured memory rather than leaving it only as raw dialogue. The four-network model provides the categories Hindsight uses to distinguish world facts, agent experiences, synthesized entity summaries, and evolving beliefs.

Recall: retrieve relevant memory

Recall retrieves information from memory. The ACL 2026 demonstration paper describes a pipeline combining vector search, keyword matching, graph traversal, and temporal filtering, backed by PostgreSQL with pgvector. These methods address different retrieval needs: vector search can help surface semantically related material, keyword matching can find explicit terms, graph traversal can follow entity relationships, and temporal filtering can narrow results by time.

Reflect: reason and update

Reflect reasons over the stored memory. In the authors’ description, a reflection layer can produce answers and update information in a traceable way. That is distinct from simply returning a retrieved passage: reflection is intended to use the memory bank to form or revise a response while retaining a connection to the supporting information.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How Hindsight differs from vector search and temporal knowledge graphs

A vector database can be part of a memory system, but vector similarity alone answers a narrower question: which stored items are semantically close to this query? Hindsight’s stated retrieval pipeline adds keyword, graph, and temporal operations, while its four-network model distinguishes types of information. That makes the comparison less about whether vectors are present and more about what other structure and operations the system provides.

Approach Memory representation Retrieval or temporal handling described in the sources What to verify for a real comparison
Hindsight Four logical networks for world facts, agent experiences, synthesized entity summaries, and evolving beliefs. Its ACL paper describes vector search, keyword matching, graph traversal, and temporal filtering. Model and prompt, storage and deployment setup, latency, cost, tuning effort, and the evaluation task.
Vector retrieval alone Not specified here as a single product or schema; the phrase describes an approach based on semantic similarity. Vector similarity can retrieve semantically related items. Temporal and relationship handling depend on the broader system. Whether it supports entity relationships, time-aware updates, evidence traceability, and the same workload and evaluation conditions.
Zep’s Graphiti The Zep authors describe a temporally aware knowledge graph combining conversational information with structured business data. The Zep preprint says Graphiti retains historical relationships. Whether model configuration, prompts, datasets, scoring, deployment, latency, and cost match the Hindsight setup.

Hindsight’s paper also discusses MemGPT, Zep, and Mem0 in describing its feature set. That comparison should be understood in the context of the paper, not as proof that Hindsight is the only system with these capabilities; implementations and alternatives can change. The sources identify Graphiti as a relevant temporal-graph comparison, but the reported scores below are not directly interchangeable with Hindsight’s results.

What Hindsight’s benchmark scores show—and what they do not

The Hindsight authors report results on LongMemEval and LoCoMo under different model configurations. The figures are useful evidence about those evaluations, but they are not a universal ranking: the model, prompts, baseline, split, and scoring procedure affect what a score means.

Source and evaluation Reported result Configuration and qualification
Hindsight authors’ 2025 preprint, LongMemEval 83.6% Reported with an open-source 20B model. The authors compare it with 39% for their full-context baseline using the same backbone.
Hindsight authors’ 2025 preprint, LongMemEval 91.4% Reported with a larger-backbone configuration; the preprint’s headline figure should be read in its stated evaluation setup.
Hindsight authors’ 2025 preprint, LoCoMo 89.61% Reported with the stronger configuration described in the preprint. The authors compare it with 75.78% for the strongest prior open system in their evaluation context.
Association for Computational Linguistics (ACL), 2026, LongMemEval 83.6% Reported with a 20B open-source model; the ACL abstract also reports 91.4% with Gemini-3 Pro.
Association for Computational Linguistics (ACL), 2026, LoCoMo 83.2% Reported with a 20B open-source model.
Zep authors’ 2025 preprint, DMR 94.8% versus 93.4% The Zep authors report these figures in their own evaluation context. They also describe LongMemEval improvements against their stated baselines; these results should not be compared directly with Hindsight’s without aligned evaluation conditions.

The 91.4% LongMemEval result appears in both the 2025 preprint and the 2026 ACL publication, with the ACL abstract specifying Gemini-3 Pro. The figures should not be collapsed into a claim that a system will achieve the same accuracy on an agent’s production workload.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In a March 2026 commentary, the Hindsight team argues that LongMemEval and LoCoMo remain useful but may not distinguish memory architectures well when large-context models can fit evaluation material. The team also says these datasets emphasize chatbot-style conversational recall more than multi-step agent tasks. That is the project team’s assessment of benchmark limits, not an independent audit of every memory system.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to judge a memory benchmark fairly

Before using a published score to choose an architecture, check whether the test matches the system and task you care about. The Hindsight team’s benchmark commentary emphasizes that accuracy is only one production concern; speed, cost, and usability also matter. For a meaningful comparison, establish:

  • Model and prompt: Which model, answer-generation prompt, and judge prompt were used? Prompt and model choices can materially change measured accuracy.
  • Baseline definition: What does the baseline include, and does it use the same backbone and context available to the memory-enabled system?
  • Evaluation protocol: Which benchmark split and scoring procedure produced the result?
  • Operational cost: What were the latency and inference costs, not just the final accuracy?
  • Setup burden: How much configuration, data preparation, and tuning were required?
  • Task fit: Does the test resemble conversational recall, or the multi-step autonomous workflow the agent will actually perform?

Hindsight’s scores are evidence for the specific configurations its authors evaluated. A team evaluating an agent should also test representative changes over time, conflicting information, entity relationships, and the consequences of retrieving or updating the wrong fact.

Can you run Hindsight locally?

The ACL 2026 publication describes Hindsight as open source under the MIT license and says it is distributed as a Python package named hindsight-all and as a Docker image. That supports local use as a software deployment option, but the publication alone does not establish current package commands, dependencies, model support, or the steps needed for a particular environment. Consult the project’s current documentation before choosing a setup; those details can change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The project README positions Hindsight for conversational agents and autonomous task-oriented agents, particularly agents that should adapt to feedback and build capability across complex tasks. That is the project’s intended-use framing, not independent evidence that every deployment will achieve those outcomes. The ACL paper also reports production use at Fortune 500 enterprises; it does not, in the cited description, identify customers or provide deployment details.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.