No single database wins for AI agents. The right choice depends on three separate jobs: storing what the agent must remember across sessions, retrieving knowledge it does not hold in its prompt, and keeping execution state intact through interruptions. Each job has different read, write, and correctness requirements, so the strongest setup is often a composition of capabilities, either inside one multi-model database or across several systems. Product features and SDK backend support change over time, so confirm current behavior in the documentation linked below before you commit.
Separate the three storage questions
Questions such as “Where do AI agents store memory?” bundle three different problems. Pulling them apart is the first step in choosing storage.
As an Amazon Associate I earn from qualifying purchases.
- What must persist as memory? Recent conversation, active task context, and durable facts such as user preferences.
- How does the agent retrieve knowledge? Documents, records, or entities that are too large or too numerous to place in the prompt.
- What execution state must survive an interruption? Task status, checkpoints, tool outcomes, and records that need exact updates.
Memory is not one data type
MongoDB’s agent documentation separates short-term session context from long-term memory. Short-term memory is typically recent turns kept under a session identifier. Long-term memory is selected information extracted from interactions and stored for later sessions. The two have different lifecycles, so they usually need different retention and cleanup rules. See MongoDB’s documentation on AI agents for the pattern.
Knowledge retrieval depends on the search method
Vector search finds content by semantic similarity, which helps when a user’s wording differs from the source text. Full-text search matches terms, which matters for product codes, error strings, and names. Hybrid search combines both. MongoDB documents an agent choosing among vector, full-text, and hybrid retrieval tools depending on the task.
#1 Best Overall
Execution state needs exact, ordered writes
An agent that books a flight, files a ticket, or runs a multi-step analysis must know which step completed, what a tool returned, and what to do next after a crash or timeout. These are exact, ordered updates. A retrieval index can return a relevant note about a task, but it cannot guarantee that the task’s status was written once and in the right order. That requirement calls for transactional writes and ordering, whatever memory and retrieval layers sit alongside.
Is a vector database enough for an AI agent?
Usually not. A vector database is enough when the agent answers questions from a document collection, keeps no state between requests, and needs no exact lookups. Once the agent keeps conversations, updates records, or resumes multi-step work, four gaps appear:
- Exact state: task status and tool outcomes need the transactional, ordered writes described above.
- Session history: transcripts must be read back in sequence, which similarity ranking does not provide.
- Exact terms: identifiers such as invoice numbers, SKUs, and error codes often need keyword matching.
- Multi-hop questions: chains of relationships need traversal, which a similarity index does not perform.
The usual fix is to pair the vector index with a store for state, rather than asking the index to do both jobs.
Compare the storage patterns
| Pattern | Best fit | Read and write pattern | Correctness and recovery |
|---|---|---|---|
| Relational (PostgreSQL) | Defined records, transactional state, joins already in the application | Exact lookup and update, joins, SQL filters | Transactional writes; vector, graph, and full-text extensions run in the same engine |
| Key-value or session store (Redis, or Dapr-configured state stores) | Keyed session state; shared low-latency access across workers | Keyed get and set; ordered history depends on how you model it | Not stated in the SDK guidance; depends on the chosen backend and its deployment |
| Vector and hybrid retrieval (MongoDB, dedicated vector stores) | Semantic similarity, plus keyword or hybrid search for identifiers | Similarity search, full-text search, hybrid search, metadata filters | Not a substitute for exact, transactional state |
| Graph (Neo4j) | Questions about how several people, events, or records connect | Relationship traversal and multi-hop queries | Not stated in Neo4j’s architecture guidance; check the product’s transaction and recovery documentation |
| Files and SQLite | Local prototypes, single-user assistants, small memory profiles | Simple reads and writes from one application | File-backed SQLite persists conversations; the SDK lists no shared-access guarantees for it |
| Extract-and-update memory service | Production deployments where several agents share memory and token cost matters | Extract facts, add, update, merge, or delete them, summarize asynchronously, then retrieve | Depends on the underlying stores; extraction quality needs evaluation |
Relational databases
Use a relational store when agent state and business records already have defined structures, when transactions matter, or when joins are already part of the application. Microsoft’s Azure HorizonDB documentation for AI agents describes PostgreSQL with pgvector, Apache AGE, and full-text search as options for agent workloads. Treat that as a description of a Microsoft product, not an independent comparison. Having vector, graph, and full-text features in one engine reduces the number of systems you run, but feature availability does not show that the setup meets your scale or query requirements.
Key-value and session stores
The OpenAI Agents SDK lists Redis sessions for shared memory across workers and services, described as suitable for low-latency distributed deployments. It also lists Dapr sessions, which let teams switch the configured state-store backend while keeping agent code stable. These are SDK guidance points, documented in the OpenAI Agents SDK sessions page. Durability, consistency, and failover depend on how the backend is deployed, so verify them against your own configuration rather than assuming them from the SDK listing.
Vector and hybrid retrieval
MongoDB’s documentation describes vector, full-text, and hybrid retrieval as tools an agent can choose among depending on the task. A dedicated vector database may fit when similarity retrieval dominates and graph relationships are limited. Which option fits depends on filtering, update behavior, scale, and the results of your own evaluation. Keyword search earns its place when users look for exact codes, error strings, or names, where exact terms matter more than meaning.
Rank #3
Graph databases
Choose a graph when the agent must follow relationships among people, events, entities, or records, especially when the question asks how several links connect. A graph makes those relationships explicit and traversable. A relational model can represent the same relationships through joins, and a vector store can retrieve similar content with supported filters, so the graph’s advantage is strongest when multi-hop traversal is the core query. The Neo4j graph memory architecture guidance advises evaluating actual queries and operating requirements rather than adopting a graph by default. It is a weaker choice when the application mostly updates keyed state or runs similarity search with few relational hops.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Files and SQLite
Microsoft’s memory patterns guidance describes a structured relational profile or a small Markdown file as transparent, cheap, and auditable, and sufficient in many semantic-memory cases. The OpenAI Agents SDK lists in-memory SQLite for temporary conversations and file-backed SQLite for persistent conversations. This is a sound starting point for a local prototype or a single-user assistant. Move to a shared service once concurrency, availability, or access boundaries require it.
Extract-and-update memory services
A separate memory layer can extract candidate facts from conversations, decide whether to add, update, merge, or delete them, summarize interactions asynchronously, and serve retrieval through vector search, optionally augmented by a graph. Microsoft presents this pattern for production deployments where several agents share memory and cost matters. Its trade-offs are the cost of operating another service and the work of evaluating extraction quality, because an incorrect extraction is stored as if it were a fact.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Example compositions
The following are illustrative designs, not tested configurations.
- Local prototype or single-user assistant: file-backed SQLite for conversation history, plus a small structured profile or Markdown file for durable facts. A shared service is unnecessary until concurrency or access boundaries change.
- Multi-worker product with shared session state: a shared session store such as Redis for live session state, a transactional database for business records, and a vector index for retrieved content. Add an extraction service once several agents need the same long-term memory.
- Support agent over governed enterprise content with relationship questions: retrieve from source systems through permission-aware search rather than copying content into agent memory. Add a graph only if questions that chain accounts, tickets, and products across several hops are frequent and central to the product.
A decision sequence
- List what must survive a restart. Transcript, checkpoint, task state, source records, extracted facts, or a combination.
- List the operations. Exact keyed access, transactional writes, ordered history, keyword search, semantic similarity, or relationship traversal.
- Start with the fewest systems that meet your correctness and retrieval requirements. A multi-model database can reduce integration work because memory, retrieval, and state share one system. MongoDB’s documentation describes its database as supporting several search methods for agentic RAG and storing short- and long-term agent memory in the same database; that is the vendor’s own description of its capabilities. Separate systems are justified when a specialized capability is central to the product, such as graph traversal or shared low-latency sessions, and you accept the cost of keeping data consistent across them.
- Define governance before persisting user facts or indexing enterprise content. See the governance section below.
- Benchmark candidates on your own workload. See the testing checklist below.
Governance and deletion shape the architecture
Microsoft’s memory patterns guidance describes retrieval from governed enterprise systems as a way to keep source data fresh, reduce leakage, and make deletion tractable. It also notes that a permission-aware index and retrieval quality remain requirements. Settle these questions early:
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →- Scope: which user, agent, and tenant can read and write each memory store.
- Retention: how long each memory type lives. Session transcripts and extracted long-term facts often need different schedules.
- Correction: how a wrong fact is edited, and whether summaries or cached answers built from it are refreshed.
- Deletion: what must be removed on request, including derived indexes and embeddings, not just the source row.
- Audit: which memory changes are logged, and who can read the log.
What the performance evidence does and does not show
Do not infer a winner from the storage category alone. The evidence cited here is limited in the following ways:
- Neo4j’s guidance states that it does not provide a reproducible PostgreSQL-versus-Neo4j benchmark for the workloads it describes, and it implies no measured latency, storage estimate, or universal asymptotic comparison.
- None of the vendor or project documentation cited in this article offers an independent cross-database benchmark.
- Microsoft’s memory guidance gives two token-cost figures: summarization produces “Roughly a 43% token reduction while retaining most of the context,” and fact extraction costs “Around 2K tokens per query in published benchmarks.” The guidance does not name the original benchmark publisher. Treat these as approximate figures about prompt cost in an architecture document, not as verified database comparisons or evidence that one storage option is faster.
Test a workload before you commit
Build a representative test set, then compare candidates on equivalent results, latency, and resource use. Record the following for each candidate:
Quick Recap
- Schema, indexes, and representative data volume
- Vector dimensions and the embedding model used to produce them
- The exact queries the agent will run, including metadata filters and keyword lookups
- Concurrency level, including concurrent writes to the same session or record
- Cache state, measured cold and warm
- Latency and resource use at equal result quality
- Stale-information cases and restart or recovery scenarios where they matter to the product
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




