To give a LlamaIndex agent memory that lasts across sessions without mixing users’ data, keep recent chat in the framework’s short-term queue, send extracted facts to a store keyed by an end-user ID that your application derives from an authenticated session, and treat that ID as the isolation boundary. MemorySync’s service filters reads, searches and deletes by the identifiers it receives. It cannot tell whether the caller is entitled to those identifiers. That check belongs in your application.
Short-term context and durable memory are separate layers
LlamaIndex’s Memory object holds two layers. The first is a first-in, first-out queue of ChatMessage objects that carries the recent conversation. When the queue exceeds its configured boundary, messages are archived and flushed to memory blocks, which process them into longer-term context. At retrieval time the framework merges the short-term and long-term layers. LlamaIndex’s developer documentation, “Memory in LlamaIndex,” puts it this way:
The
Memoryclass in LlamaIndex is used to store and retrieve both short-term and long-term memory.
| Layer | What it holds | How long it lasts | How it reaches the model |
|---|---|---|---|
| Short-term queue | Recent ChatMessage objects |
Until the queue exceeds its configured boundary | Part of the active conversation context |
| Memory blocks | Messages flushed from the queue, processed by the block | Longer term; the overview does not state how long block storage is kept | Merged with short-term memory at retrieval |
LlamaIndex documents three built-in block types: static memory, fact extraction, and vector memory. Each block also has a priority that determines what is kept when memory exceeds the token budget. The section on token pressure below covers that behavior in more detail.
Recommended Free Tools
#1 Best Overall
- EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
Decide the tenant key before writing any code
Most isolation failures in this kind of system are identity failures, not storage failures. Settle the identity chain first:
- Authenticate the caller at your API boundary using whatever your stack already uses, such as a session cookie or a bearer token.
- Resolve the authenticated principal to a stable, opaque
user_idstored in your own database. Do not use an email address, and do not accept the value from the request body. - Authorize the principal for the specific memory operation requested, such as reading, writing, or deleting memories for that user.
- Only after those checks, construct the MemorySync memory object with that user ID and a conversation ID.
MemorySync describes three scope identifiers. Their roles differ, and the table below separates them.
| Identifier | Role according to MemorySync | Who should set it | Notes |
|---|---|---|---|
| Project | Tenant boundary for the deployment; the developer FAQ says project boundaries are enforced | Your application configuration | Keep one project per boundary you need to keep separate |
| End user | Required for API-key calls; reads, searches and deletes are filtered by it | Your server, from the authenticated principal | Use an opaque, stable identifier |
| Session | Optional; groups stored facts by conversation thread | Your server, from the conversation record | Grouping only; it is not a user-isolation mechanism |
MemorySync’s integration guide treats the user ID as required and the session ID as a way to group facts by thread. The developer FAQ describes project and end-user identifiers as tenant coordinates and session as optional context. Both descriptions lead to the same design: the end-user ID carries the isolation, and the session ID carries only conversational grouping.
A common mistake is letting the client choose the scope. If a request body can name a user_id, a logged-in user can ask for another user’s facts. The service’s filter will then match the ID it was given and return that user’s data, because it cannot know the caller was not entitled to it.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Choose an integration surface by control flow
MemorySync documents four integration surfaces. They differ mainly in who decides when memory is read or written.
| Surface | Control flow | Typical use | Who decides when memory is touched |
|---|---|---|---|
MemorySyncMemory |
A subclass of LlamaIndex Memory, passed to the agent’s memory parameter |
Chat agents that need cross-session recall with minimal wiring | The framework, on each run |
MemorySyncMemoryBlock |
A composable block inside a custom Memory |
Custom memory stacks where you control block order and composition | Your composition code |
MemorySyncRetriever |
A BaseRetriever for retrieval query engines and retriever tools |
Retrieval-augmented question answering over stored facts | Your query pipeline |
| Explicit memory tools | A tool factory exposing add, search, list, update and delete | Agents that should decide when to read or change memories | The model, within the tools you expose |
MemorySyncMemory: the ready-made memory object
This is the simplest path when your agent is a chat agent. According to the integration guide, user messages are sent for fact extraction on the asynchronous aput path, and recall is inserted through the framework’s memory-block template. The guide also says the short-term buffer and the standard memory options remain available. Choose this surface when you want memory to behave like the rest of the chat history without custom code.
Rank #2
- Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
- 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
- AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
- Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
- Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.
MemorySyncMemoryBlock: composing a custom memory
Use the block when you already build a LlamaIndex Memory with several blocks and want MemorySync’s facts as one of them. You control ordering and priority, which matters for the token-budget behavior described later.
MemorySyncRetriever: retrieval, not conversation
The retriever fits a question-answering path where stored facts are one source among several. The integration guide describes retriever errors separately from an empty result, which matters for the failure handling later in this article. The guide does not describe whether the retriever can write, so treat it as a read path in your design unless you verify otherwise in the current guide.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteExplicit memory tools: the agent decides
The tool factory exposes add, search, list, update and delete operations. This gives the model direct control, which is useful when the user asks the agent to remember or forget something. It also widens the permission surface. The read-only mode and how to constrain it are covered in their own section below.
A minimal integration
MemorySync’s integration guide lists these requirements:
- Package:
llamaindex-memorysync, version 1.1.0 as shown in the guide - LlamaIndex core:
llama-index-core0.13 or later - Python: 3.10 or later
Package indexes change, so check the current release before pinning. The guide’s indexed content was reviewed on 2026-10-01, and the version figures are the ones it reports at that time.
pip install "llamaindex-memorysync==1.1.0" "llama-index-core>=0.13"
The guide’s example follows the shape below. It assumes agent is an existing LlamaIndex agent and that authenticated_user_id and conversation_id come from your server, not from the client. Take the import path from the integration guide, because it is not reproduced here.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Rank #3
- EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
memory = MemorySyncMemory.from_defaults(
user_id=authenticated_user_id, # derived after application authorization
session_id=conversation_id,
)
response = await agent.run(user_message, memory=memory)
Run this in a staging environment with two test accounts before relying on it. Confirm that account A’s recall never returns account B’s facts, even when the test deliberately sends account B’s ID from account A’s session.
Limiting the agent to read-only memory
The explicit tool factory has a read_only=True mode, which the guide describes as returning search and list operations only. Use it when the agent’s job is to answer from what is already known, such as a support assistant that can say what it remembers about a customer’s account but should not edit those records.
Mutation-capable tools deserve more care. Delete in particular is a permission-sensitive operation. A practical configuration looks like this:
- Expose add and update only where the user has asked the agent to change a stored fact.
- Expose delete only behind an application-level confirmation step, so the model cannot remove a memory on its own initiative.
- Check the authenticated user’s authorization inside your tool wrapper, not only in the prompt.
Isolation: what the service enforces and what your code must enforce
MemorySync’s developer FAQ says that API-key calls must include an end-user ID, and that reads, searches and deletes are filtered by user, project and environment. It describes project boundaries as enforced. It also says the application decides which end user a request is for. Those statements divide the responsibility cleanly:
- The service’s scope filters are a defense at the data-access layer. They prevent a correctly scoped request from reaching another project or user’s records.
- Your application is responsible for binding each authenticated principal to the correct scope. The service has no independent way to verify that binding.
In practice, a tenant leak through this design is an authorization bug in your code. The storage layer will faithfully serve whatever scope it is given.
Retrieved memory is untrusted input
MemorySync’s tenant operations documentation advises treating retrieved memory text and metadata as untrusted data, not as system instructions. This matters most when recalled facts are inserted into the prompt or returned from a retriever. A fact that a user once typed can contain text that looks like an instruction. Place recalled memories in a clearly labeled context section, and never let memory content change tool permissions, scopes, or identifiers.
Rank #4
Privacy and data-handling claims to check
MemorySync’s developer FAQ makes several data-handling claims. They come from the vendor’s documentation and have not been independently audited in the material reviewed for this article:
- Encryption at rest, with a per-end-user scheme as the FAQ describes it.
- HTTPS-only transit.
- Memory text is sent to a model provider for fact extraction and for embeddings.
The last point deserves the most attention. Any end-user memory your agent stores is sent to a third-party model provider as part of the service’s normal operation. Before deploying, review MemorySync’s current contract and retention settings, its subprocessor list, and the regulatory requirements that apply to your users and region.
Failure handling and degradation
The integration guide documents how each surface behaves when something goes wrong. Your production design should decide, per operation, whether the conversation can continue without that memory.
| Operation | Documented behavior | Suggested policy |
|---|---|---|
| Short-term buffer update | Happens first, before external persistence | Treat as always available; the conversation continues from it |
| External persistence | Errors can be routed through an error handler | Log, alert, and decide whether to retry; the guide does not describe automatic retry |
| Recall through the memory block | A failed recall can omit the memory block while the conversation continues | Allow degradation for personalization; record each omission |
| Retriever call | Errors are reported separately from an empty result | Handle an error as a failure and an empty result as no matching memory |
| Update or delete through tools | Not stated in the guide’s overview | Fail closed, tell the user the change did not happen, and never assume it succeeded |
Monitor memory failures as their own metric. A rising rate of omitted recalls can look like a quality problem in the agent when it is really a connectivity or authorization problem.
Token pressure: two different mechanisms
Two behaviors are easy to confuse. LlamaIndex’s documented model uses block priority: when memory exceeds the token budget, priority decides which blocks are retained. MemorySync’s MemorySyncMemoryBlock is described as performing partial truncation under token pressure. That is a product-specific behavior of the block, not part of LlamaIndex’s priority model. The guide’s summary does not say which facts are cut first, so do not assume that the most recent or most important memories survive.
Set the token budget deliberately and check the result in staging, using a conversation long enough to trigger truncation. Confirm that the facts you need most are still present in the assembled context.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




