October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk4 min

Estimate LLM VRAM in JavaScript: Weights, KV Cache and Headroom

A short JavaScript calculator estimates LLM weights and KV-cache memory per GPU, while making runtime headroom and its assumptions explicit.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To estimate whether an LLM will fit on your GPU, add three things: model-weight storage, KV-cache storage for the context and concurrent sequences you plan to run, and a deliberate allowance for runtime allocations. The JavaScript calculator below turns those inputs into a rough per-GPU estimate. It is a screening tool, not a promise of peak runtime usage.

What the estimate includes

Model weights are only the starting point. NVIDIA describes weights and the KV cache as the two main contributors to GPU memory use; serving runtimes also allocate memory for activations, communication and workspace buffers, CUDA graphs, I/O tensors, and other needs. The amount varies with the model, runtime, and workload, so there is no universal headroom percentage.

  • Weights: parameter count multiplied by effective bytes per stored weight.
  • KV cache: storage for keys and values across the cached tokens and concurrent sequences.
  • Runtime headroom: an explicit allowance for allocations not captured by the first two calculations.

For tensor-parallel inference across multiple GPUs, the code estimates weights per GPU by dividing across the configured GPU count. That is a rough allocation model, not a guarantee of how a runtime shards weights or uses memory. It reports the KV cache and total estimated memory per GPU using an even-share assumption; actual cache placement depends on the runtime.

Use the model’s actual dimensions

For a common transformer, approximate KV-cache bytes per token as 2 × layers × KV heads × head dimension × bytes per cache value. The factor of two accounts for keys and values. Multiply that by the number of sequences being cached and their retained token count.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • 0dB technology lets you enjoy light gaming in relative silence
  • Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
  • Dual ball fan bearings last up to twice as long as sleeve bearing designs

Use the model’s KV-head count, not automatically its query-head count. In grouped-query attention, those counts differ. NVIDIA’s broader formula uses hidden size in architectures where the combined attention-head dimensions equal it, but that shortcut is not universal. Check the model configuration and make the dimensions and cache representation match the target runtime.

Calculate an estimate in JavaScript

Set the inputs to match the model and workload. The cache token count should cover all tokens retained at the point you want to size for, typically prompt plus generated tokens—not just the prompt. Headroom is an explicit modeling assumption in bytes; replace the example allowance with one appropriate to your runtime and validate it in an actual deployment.

Rank #2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070 Ti
  • Integrated with 16GB GDDR7 256bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system
const parameters = 7e9;
const weightBytes = 2;             // effective bytes per stored weight
const tensorParallelGpuCount = 1;
const layers = 32;
const kvHeads = 32;
const headDim = 128;
const kvBytesPerValue = 2;         // e.g. half-precision cache
const sequences = 1;
const cachedTokensPerSequence = 4096;
const runtimeHeadroomBytes = 2 * 1024 ** 3; // chosen assumption

const weightsPerGpu = parameters * weightBytes / tensorParallelGpuCount;
const kvBytes = sequences * cachedTokensPerSequence * 2 * layers *
  kvHeads * headDim * kvBytesPerValue;
const estimatedPerGpuBytes = weightsPerGpu + kvBytes + runtimeHeadroomBytes;
const GiB = 1024 ** 3;

console.log({
  weightsPerGpuGiB: weightsPerGpu / GiB,
  kvCacheGiB: kvBytes / GiB,
  estimatedPerGpuGiB: estimatedPerGpuBytes / GiB,
  estimatedClusterGiB: estimatedPerGpuBytes * tensorParallelGpuCount / GiB
});

This uses 7 billion parameters, 32 layers, 32 KV heads, a 128-dimension head, half-precision weights and cache, one sequence, and 4,096 cached tokens as illustrative inputs. The 2 GiB runtime allowance is a chosen assumption, not a generally applicable recommendation. The result is an estimate in GiB (bytes divided by 1024³); it does not include extra safety margin beyond the allowance you enter.

Choose weight bytes carefully

NVIDIA NIM’s heuristic assigns 2 bytes per parameter to BF16 and FP16, 1 byte to FP8, and 0.5 byte to INT4 and NVFP4. Use these as planning values, not exact file sizes: quantized formats and implementation overhead can change actual storage. For tensor parallelism, the heuristic divides the weight estimate by the number of GPUs used.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5060
  • Integrated with 8GB GDDR7 128bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

As a sense check, NVIDIA’s Technical Blog estimates roughly 14 GB of FP16 weights for a 7-billion-parameter model. That is a weights-only example, not a total-memory requirement. Its Llama 2 7B example also estimates roughly 2 GB of KV cache at batch 1 and sequence length 4,096 in half precision, under that article’s architecture assumptions. Neither figure is a universal constant.

Understand what changes the result

  • Lower weight precision reduces the weight-storage estimate; the actual quantized format and runtime determine the real footprint.
  • Longer retained context increases KV-cache use in proportion to cached tokens, all else equal.
  • More concurrent sequences increases the cache estimate in proportion to sequence count, if each retains the same number of tokens.
  • More GPUs for tensor parallelism can reduce the simple per-GPU weight estimate, but does not by itself establish actual sharding or total usable capacity.
  • Runtime memory settings can control cache sizing or reserve GPU memory for other uses, so configuration can make observed allocation differ from the formula.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Check fit against the deployment, not just the arithmetic

Before downloading or deploying a model, compare the estimate with usable VRAM on each GPU and check the intended runtime, precision, tensor-parallel layout, context, and concurrency together. The remaining capacity after weights and cache must accommodate runtime allocations. If it does not, consider reducing context or concurrency, using a lower-weight-storage precision, or changing the GPU layout, then recalculate. Confirm the choice with the selected serving stack: activations, workspaces, graph capture, I/O tensors, cache-block allocation, and workload shape can all shift peak usage.

Rank #4
Sale
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
  • Powered by Radeon RX 9070 XT
  • WINDFORCE Cooling System
  • Hawk Fan
  • Server-grade Thermal Conductive Gel
  • RGB Lighting

VRAM capacity alone does not rank GPUs for a deployment; this calculation does not measure speed or runtime compatibility. NVIDIA’s 24 GB RTX 4090 example in its NIM guide is illustrative, not a general recommendation for every model or workload.

Quick Recap

Bestseller No. 1
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$529.00
Bestseller No. 2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070 Ti; Integrated with 16GB GDDR7 256bit memory interface
$1,162.49
SaleBestseller No. 3
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5060; Integrated with 8GB GDDR7 128bit memory interface
$459.99
SaleBestseller No. 4
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
Powered by Radeon RX 9070 XT; WINDFORCE Cooling System; Hawk Fan; Server-grade Thermal Conductive Gel
$814.99
SaleBestseller No. 5
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$829.00
Best Value
Sale
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • Phase-change GPU thermal pad helps ensure optimal heat transfer, lowering GPU temperatures for enhanced performance and reliability
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • Dual-ball fan bearings last up to twice as long as standard conventional sleeve bearings designs
  • 0dB technology lets you enjoy light gaming in relative silence

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.