October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk8 min

Building Multi-Tenant Memory Layers for AI Agents in Python with LlamaIndex and MemorySync

How to give a LlamaIndex agent durable memory across sessions without mixing users' data: separate short-term chat from long-term facts, derive the tenant ID from authenticated identity, and choose the right MemorySync integration surface.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To give a LlamaIndex agent memory that lasts across sessions without mixing users’ data, keep recent chat in the framework’s short-term queue, send extracted facts to a store keyed by an end-user ID that your application derives from an authenticated session, and treat that ID as the isolation boundary. MemorySync’s service filters reads, searches and deletes by the identifiers it receives. It cannot tell whether the caller is entitled to those identifiers. That check belongs in your application.

Short-term context and durable memory are separate layers

LlamaIndex’s Memory object holds two layers. The first is a first-in, first-out queue of ChatMessage objects that carries the recent conversation. When the queue exceeds its configured boundary, messages are archived and flushed to memory blocks, which process them into longer-term context. At retrieval time the framework merges the short-term and long-term layers. LlamaIndex’s developer documentation, “Memory in LlamaIndex,” puts it this way:

The Memory class in LlamaIndex is used to store and retrieve both short-term and long-term memory.

Layer What it holds How long it lasts How it reaches the model
Short-term queue Recent ChatMessage objects Until the queue exceeds its configured boundary Part of the active conversation context
Memory blocks Messages flushed from the queue, processed by the block Longer term; the overview does not state how long block storage is kept Merged with short-term memory at retrieval

LlamaIndex documents three built-in block types: static memory, fact extraction, and vector memory. Each block also has a priority that determines what is kept when memory exceeds the token budget. The section on token pressure below covers that behavior in more detail.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
GMKtec AI Mini PC Ryzen Al Max+ 395 (up to 5.1GHz) Mini Gaming Computers
  • EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.

Decide the tenant key before writing any code

Most isolation failures in this kind of system are identity failures, not storage failures. Settle the identity chain first:

  1. Authenticate the caller at your API boundary using whatever your stack already uses, such as a session cookie or a bearer token.
  2. Resolve the authenticated principal to a stable, opaque user_id stored in your own database. Do not use an email address, and do not accept the value from the request body.
  3. Authorize the principal for the specific memory operation requested, such as reading, writing, or deleting memories for that user.
  4. Only after those checks, construct the MemorySync memory object with that user ID and a conversation ID.

MemorySync describes three scope identifiers. Their roles differ, and the table below separates them.

Identifier Role according to MemorySync Who should set it Notes
Project Tenant boundary for the deployment; the developer FAQ says project boundaries are enforced Your application configuration Keep one project per boundary you need to keep separate
End user Required for API-key calls; reads, searches and deletes are filtered by it Your server, from the authenticated principal Use an opaque, stable identifier
Session Optional; groups stored facts by conversation thread Your server, from the conversation record Grouping only; it is not a user-isolation mechanism

MemorySync’s integration guide treats the user ID as required and the session ID as a way to group facts by thread. The developer FAQ describes project and end-user identifiers as tenant coordinates and session as optional context. Both descriptions lead to the same design: the end-user ID carries the isolation, and the session ID carries only conversational grouping.

A common mistake is letting the client choose the scope. If a request body can name a user_id, a logged-in user can ask for another user’s facts. The service’s filter will then match the ID it was given and return that user’s data, because it cannot know the caller was not entitled to it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an integration surface by control flow

MemorySync documents four integration surfaces. They differ mainly in who decides when memory is read or written.

Surface Control flow Typical use Who decides when memory is touched
MemorySyncMemory A subclass of LlamaIndex Memory, passed to the agent’s memory parameter Chat agents that need cross-session recall with minimal wiring The framework, on each run
MemorySyncMemoryBlock A composable block inside a custom Memory Custom memory stacks where you control block order and composition Your composition code
MemorySyncRetriever A BaseRetriever for retrieval query engines and retriever tools Retrieval-augmented question answering over stored facts Your query pipeline
Explicit memory tools A tool factory exposing add, search, list, update and delete Agents that should decide when to read or change memories The model, within the tools you expose

MemorySyncMemory: the ready-made memory object

This is the simplest path when your agent is a chat agent. According to the integration guide, user messages are sent for fact extraction on the asynchronous aput path, and recall is inserted through the framework’s memory-block template. The guide also says the short-term buffer and the standard memory options remain available. Choose this surface when you want memory to behave like the rest of the chat history without custom code.

Rank #2
AMD Ryzen™ AI Halo - Personal AI Desktop Computer - Developer Platform - Linux OS
  • Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
  • 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
  • AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
  • Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
  • Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.

MemorySyncMemoryBlock: composing a custom memory

Use the block when you already build a LlamaIndex Memory with several blocks and want MemorySync’s facts as one of them. You control ordering and priority, which matters for the token-budget behavior described later.

MemorySyncRetriever: retrieval, not conversation

The retriever fits a question-answering path where stored facts are one source among several. The integration guide describes retriever errors separately from an empty result, which matters for the failure handling later in this article. The guide does not describe whether the retriever can write, so treat it as a read path in your design unless you verify otherwise in the current guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Explicit memory tools: the agent decides

The tool factory exposes add, search, list, update and delete operations. This gives the model direct control, which is useful when the user asks the agent to remember or forget something. It also widens the permission surface. The read-only mode and how to constrain it are covered in their own section below.

A minimal integration

MemorySync’s integration guide lists these requirements:

  • Package: llamaindex-memorysync, version 1.1.0 as shown in the guide
  • LlamaIndex core: llama-index-core 0.13 or later
  • Python: 3.10 or later

Package indexes change, so check the current release before pinning. The guide’s indexed content was reviewed on 2026-10-01, and the version figures are the ones it reports at that time.

pip install "llamaindex-memorysync==1.1.0" "llama-index-core>=0.13"

The guide’s example follows the shape below. It assumes agent is an existing LlamaIndex agent and that authenticated_user_id and conversation_id come from your server, not from the client. Take the import path from the integration guide, because it is not reproduced here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD
  • EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
memory = MemorySyncMemory.from_defaults(
    user_id=authenticated_user_id,  # derived after application authorization
    session_id=conversation_id,
)
response = await agent.run(user_message, memory=memory)

Run this in a staging environment with two test accounts before relying on it. Confirm that account A’s recall never returns account B’s facts, even when the test deliberately sends account B’s ID from account A’s session.

Limiting the agent to read-only memory

The explicit tool factory has a read_only=True mode, which the guide describes as returning search and list operations only. Use it when the agent’s job is to answer from what is already known, such as a support assistant that can say what it remembers about a customer’s account but should not edit those records.

Mutation-capable tools deserve more care. Delete in particular is a permission-sensitive operation. A practical configuration looks like this:

  • Expose add and update only where the user has asked the agent to change a stored fact.
  • Expose delete only behind an application-level confirmation step, so the model cannot remove a memory on its own initiative.
  • Check the authenticated user’s authorization inside your tool wrapper, not only in the prompt.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Isolation: what the service enforces and what your code must enforce

MemorySync’s developer FAQ says that API-key calls must include an end-user ID, and that reads, searches and deletes are filtered by user, project and environment. It describes project boundaries as enforced. It also says the application decides which end user a request is for. Those statements divide the responsibility cleanly:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The service’s scope filters are a defense at the data-access layer. They prevent a correctly scoped request from reaching another project or user’s records.
  • Your application is responsible for binding each authenticated principal to the correct scope. The service has no independent way to verify that binding.

In practice, a tenant leak through this design is an authorization bug in your code. The storage layer will faithfully serve whatever scope it is given.

Retrieved memory is untrusted input

MemorySync’s tenant operations documentation advises treating retrieved memory text and metadata as untrusted data, not as system instructions. This matters most when recalled facts are inserted into the prompt or returned from a retriever. A fact that a user once typed can contain text that looks like an instruction. Place recalled memories in a clearly labeled context section, and never let memory content change tool permissions, scopes, or identifiers.

Privacy and data-handling claims to check

MemorySync’s developer FAQ makes several data-handling claims. They come from the vendor’s documentation and have not been independently audited in the material reviewed for this article:

  • Encryption at rest, with a per-end-user scheme as the FAQ describes it.
  • HTTPS-only transit.
  • Memory text is sent to a model provider for fact extraction and for embeddings.

The last point deserves the most attention. Any end-user memory your agent stores is sent to a third-party model provider as part of the service’s normal operation. Before deploying, review MemorySync’s current contract and retention settings, its subprocessor list, and the regulatory requirements that apply to your users and region.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Failure handling and degradation

The integration guide documents how each surface behaves when something goes wrong. Your production design should decide, per operation, whether the conversation can continue without that memory.

Operation Documented behavior Suggested policy
Short-term buffer update Happens first, before external persistence Treat as always available; the conversation continues from it
External persistence Errors can be routed through an error handler Log, alert, and decide whether to retry; the guide does not describe automatic retry
Recall through the memory block A failed recall can omit the memory block while the conversation continues Allow degradation for personalization; record each omission
Retriever call Errors are reported separately from an empty result Handle an error as a failure and an empty result as no matching memory
Update or delete through tools Not stated in the guide’s overview Fail closed, tell the user the change did not happen, and never assume it succeeded

Monitor memory failures as their own metric. A rising rate of omitted recalls can look like a quality problem in the agent when it is really a connectivity or authorization problem.

Token pressure: two different mechanisms

Two behaviors are easy to confuse. LlamaIndex’s documented model uses block priority: when memory exceeds the token budget, priority decides which blocks are retained. MemorySync’s MemorySyncMemoryBlock is described as performing partial truncation under token pressure. That is a product-specific behavior of the block, not part of LlamaIndex’s priority model. The guide’s summary does not say which facts are cut first, so do not assume that the most recent or most important memories survive.

Set the token budget deliberately and check the result in staging, using a conversation long enough to trigger truncation. Confirm that the facts you need most are still present in the assembled context.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.