Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk5 min

Why Local LLMs Use More Memory as Context Grows: The KV Cache Explained

The KV cache stores attention data for earlier tokens so a local LLM can reuse it while generating. As retained context grows, cache memory usually grows too—but weights, compute buffers, and runtime choices also shape the memory reading.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Local LLMs use more memory as a chat grows because they retain attention data for earlier tokens in a structure called the KV cache. The cache helps the model generate the next token without recalculating all earlier attention state. Its size usually grows with the number of tokens kept in context, while model weights take up a separate, mostly fixed allocation.

“RAM” can mean system RAM, GPU VRAM, or unified memory. Which one rises depends on where your inference runtime places the model, cache, and working buffers.

What the KV cache does

When a model generates text one token at a time, its attention layers process each new token in relation to earlier tokens. Those layers produce key (K) and value (V) data. The runtime keeps this data for prior positions so it can reuse it on the next generation step instead of recomputing it. Hugging Face explains this reuse in its KV cache documentation.

Think of each retained token as adding another slice of K and V data across the model’s cache-bearing attention layers. That makes the cache a speed-memory tradeoff: retaining prior state costs memory, but avoids repeating work. The model’s weights do not need to grow for the cache to grow.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
CORSAIR Vengeance LPX DDR4 RAM 32GB (2x16GB) Up to 3200MHz CL16-20-20-38 1.35V Intel XMP AMD EXPO Computer Memory – Black (CMK32GX4M2E3200C16)
  • Disclaimer: Maximum Speed requires overclocking/PC BIOS adjustments. Maximum speed and performance depend on system components, including motherboard and CPU
  • Hand-sorted memory chips ensure high performance with generous overclocking headroom
  • VENGEANCE LPX is optimized for wide compatibility with the latest Intel and AMD DDR4 motherboards
  • A low-profile height of just 34mm ensures that VENGEANCE LPX even fits in most small-form-factor builds
  • A solid aluminum heatspreader efficiently dissipates heat from each module so that they consistently run at high clock speeds

Why longer chats increase memory

In a standard full-attention model, cache storage grows roughly linearly with the number of retained tokens. Both the prompt and the generated continuation count: if the runtime keeps them in the active context, each occupies a position whose attention state may be cached. Hugging Face’s Transformers v4.56.0 cache guide describes cache tensors with a sequence-length dimension that advances as tokens are processed.

The configured context maximum is not necessarily the amount of memory already occupied. Some implementations grow the cache as tokens arrive; others reserve capacity ahead of time. Sliding-window attention can also stop retaining older positions after its window is full. These behaviors depend on the model architecture and runtime, so a long context setting alone does not tell you the exact allocation.

Estimate KV cache memory per token

For a conventional cache, use this first-order estimate:

KV cache bytes ≈ B × T × 2 × L × Hkv × D × S

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Crucial 32GB DDR5 RAM Kit (2x16GB), 5600MHz (or 5200MHz or 4800MHz) Laptop Memory 262-Pin SODIMM, Compatible with Intel Core and AMD Ryzen 7000, Black - CT2K16G56C46S5
  • Boosts System Performance: 32GB DDR5 RAM laptop memory kit (2x16GB) that operates at 5600MHz, 5200MHz, or 4800MHz to improve multitasking and system responsiveness for smoother performance
  • Accelerated gaming performance: Every millisecond gained in fast-paced gameplay counts—power through heavy workloads and benefit from versatile downclocking and higher frame rates
  • Optimized DDR5 compatibility: Best for 12th Gen Intel Core and AMD Ryzen 7000 Series processors — Intel XMP 3.0 and AMD EXPO also supported on the same RAM module
  • Trusted Micron Quality: Backed by 42 years of memory expertise, this DDR5 RAM is rigorously tested at both component and module levels, ensuring top performance and reliability
  • ECC Type = Non-ECC, Form Factor = SODIMM, Pin Count = 262-Pin, PC Speed = PC5-44800, Voltage = 1.1V, Rank And Configuration = 1Rx8
  • B: number of concurrent sequences or batch items.
  • T: retained tokens per sequence.
  • 2: one set of keys and one set of values.
  • L: attention layers that retain cache.
  • Hkv: key/value heads per layer.
  • D: head dimension.
  • S: bytes per cached value; FP16 or BF16 commonly uses two bytes per value.

Use KV heads, not automatically the model’s total query-head count. Grouped-query and multi-query attention use fewer KV heads than query heads and can therefore reduce cache size. The formula is an estimate: quantization metadata, tensor layout, hybrid attention, and runtime allocation details can change the actual figure. Hugging Face discusses cache shape and sliding-window behavior in its cache guide.

To make a forecast, identify the model’s cache-bearing layer count, KV-head count, head dimension, intended retained-token count, cache element type, and simultaneous sequence count. Then allow additional memory for weights, compute buffers, the operating system, and implementation overhead. Without a named model and configuration, a single claim such as “X GB for Y tokens” is not reliable.

What else is using memory?

A rising memory meter does not show the KV cache alone. A llama.cpp maintainer’s allocation breakdown distinguishes model weights, KV buffer, output buffer, and compute buffers; it is a conceptual guide, not a universal report format for every version or backend. See the llama.cpp discussion.

  • Model weights: memory used by the loaded or memory-mapped model; driven mainly by model size and weight representation.
  • KV cache: attention state for retained tokens; affected by context, attention architecture, cache type, and concurrency.
  • Compute buffers: temporary inference workspace. In llama.cpp, batch settings and Flash Attention can affect this allocation.
  • Output and runtime buffers: additional structures whose size and reporting vary by backend.

Why memory use varies between models and runtimes

Attention architecture

The number of cache-bearing layers, KV heads, and head dimension all affect the per-token footprint. Full-attention layers generally retain state across the active context; sliding-window layers can cap retained positions at their window size. Models with hybrid attention may therefore have a different cache profile from a model using full attention throughout.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Corsair Vengeance RGB RS DDR5 16GB (2 x 8GB) Up to 6000MHz AMD Intel RAM
  • Disclaimer: Maximum Speed requires overclocking/PC BIOS adjustments. Maximum speed and performance depend on system components, including motherboard and CPU
  • AMD EXPO & Intel XMP 3.0 Compatible Only: Dual memory profiles allow you to easily select optimized settings for your platform, whether you’re running an AMD or Intel processor
  • Dynamic RGB Lighting: Individually addressable RGB lighting delivers vibrant effects through a sleek, understated panoramic diffuser
  • Onboard Voltage Regulation: Onboard voltage regulation for reliable power at high frequencies
  • Maximum Bandwidth and Tight Response Times: Optimized for peak performance on the latest AMD and Intel DDR5 motherboards

Cache type and precision

Lower-precision cache values can reduce storage, but the resulting speed or quality tradeoff depends on the model and implementation. llama.cpp’s rolling server documentation lists separate K and V cache type options, including floating-point and quantized types; check the documentation matching your installed version for supported values and behavior: llama.cpp server documentation.

Batching and simultaneous chats

More active sequences mean more context state to maintain, although a runtime may pool cache storage or allocate it per slot. llama.cpp documents unified KV and per-slot context controls in its server documentation. Batch configuration can also affect compute-buffer needs, so concurrency may raise memory beyond the cache alone.

Allocation strategy and offloading

A runtime may reserve cache capacity for a configured maximum, grow it dynamically, or combine strategies. Offloading cache or model state between GPU and host memory shifts pressure between VRAM and system RAM and may affect performance. Exact behavior is runtime- and configuration-specific; consult its documentation rather than assuming that an option moves only the cache. Transformers describes different cache strategies and their behavior in its cache guide.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to diagnose a rising memory reading

  1. Identify the memory pool. Check whether the reading is system RAM, GPU VRAM, or unified memory, and note which process or runtime is using it.
  2. Compare stages. Record memory after model load, after prompt ingestion, and during generation. A largely fixed increase at load points toward weights; growth with prompt processing or generation is consistent with cache or working-buffer allocations.
  3. Inspect runtime logs. If available, use allocation logs to distinguish weights, KV cache, and compute buffers. The labels and categories may differ between runtime versions and backends.
  4. Check the actual configuration. Compare context capacity, cache types, batch or slot count, and offload settings with the documentation for your runtime and version.

Ways to reduce memory pressure

  • Use a shorter context if the task does not need the full conversation history.
  • Reduce simultaneous sequences or server slots when concurrency is not essential.
  • Check supported cache precision options for your runtime and model; verify the actual memory and generation behavior rather than assuming a fixed benefit.
  • Check for sliding-window attention in the model and whether your runtime supports its behavior.
  • Consider cache offloading only if shifting pressure to another memory pool suits your hardware and performance needs.

These controls affect different parts of the memory and performance tradeoff. The llama.cpp documentation covers context, cache types, offloading, and server cache controls; Hugging Face documents cache strategies in its KV cache guide. Defaults and options can change, so use documentation for the version you run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.