Uniform INT8 can reduce recurrent-state memory, but it should not be assumed to preserve accuracy. Because a quantized recurrent state feeds into later updates, errors may persist or propagate during decoding. Two recent arXiv preprints report task-dependent accuracy costs and explore selective precision as an alternative. Their results make a case for testing INT8 against the actual model and serving workload—not for rejecting it categorically.
Why recurrent-state quantization is different
Hybrid language models can combine softmax-attention layers, whose key-value cache grows with prior tokens, with linear-attention components such as Gated DeltaNet (GDN) or Kimi Delta Attention (KDA). These components summarize history in fixed-size recurrent states. Those states can still use substantial memory at high concurrency, and they are read and updated during decoding. Compressing them can reduce storage and memory traffic, but it also changes data used in future updates.
As an Amazon Associate I earn from qualifying purchases.
That repeated use is the central risk. A quantized value is not necessarily a one-time approximation: it becomes part of the next state, which may carry its error forward. The DAMP authors describe this mechanism in their arXiv preprint: “Quantization error therefore enters subsequent updates and propagates through the recurrence, as formalized in Section 3.3.” The extent to which error persists depends on the update dynamics; DAMP analyzes learned decay and delta-rule updates, while STEPQuant examines temporal persistence and differences in how state rows affect outputs.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsThis does not mean uniform INT8 always fails. It means that a result from a single task, architecture, or quantizer is not enough to establish that it is safe for every deployment.
#1 Best Overall
What recent experiments say about uniform INT8
DAMP: accuracy depends on the task
DAMP (Decay-Aware Mixed-Precision Recurrent-State Quantization) is a post-training method for GDN and KDA states. In its calibration process, it ranks risky key channels using quantization error and decay-based error retention, then keeps selected channels in FP16 while storing the remainder in INT8 with stochastic rounding (SR). Its main configuration uses 16 selected key channels per head in FP16; the authors report an effective 9.9 bits per state value.
In its v2 report, DAMP evaluates Qwen3.6-35B, Kimi-Linear-48B, and Kimi-K3 on reasoning and code-generation benchmarks. The authors report that uniform INT8 and FP8 degrade complex-reasoning accuracy in their experiments, while tested INT4 and NVFP4 configurations cause severe degradation. But the results vary substantially by benchmark: INT8+SR is within 0.1 percentage points of FP32 on GPQA-Diamond and MMLU-Pro for Qwen3.6-35B, while accuracy drops by more than 20 percentage points on AIME 2026 and LiveCodeBench-v6. These are results for the paper’s particular models, benchmarks, and configurations—not a general prediction for other tasks.
Rank #2
At its 9.9-bit configuration, DAMP reports average accuracy close to FP32 across its three evaluated checkpoints. Its RULER long-context evaluation covers 4K to 128K tokens; on Qwen3.6-35B and Kimi-Linear-48B, the maximum absolute difference between DAMP and FP32 accuracy was 0.04 and 0.02 percentage points, respectively. Those figures describe that benchmark and those models, not every long-context workload.
STEPQuant: precision can be allocated selectively
STEPQuant (When and Where Errors Matter in Delta-Rule Recurrent State Quantization) allocates precision according to error magnitude and memory lifetime. It fits key-row and value-column scales using state distributions and estimates how key-row error affects output error. The authors evaluate Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct. They report that their nominal 6-bit setting closely matches FP32-state accuracy on the tested benchmarks, and that their 4-bit configuration outperforms uniform INT8 in their experiments.
Rank #3
That is evidence that the label “INT8” alone does not determine the accuracy-memory trade-off: how precision is assigned matters. It is not proof that STEPQuant, DAMP, or any particular bit width will outperform uniform INT8 on an untested architecture or workload.
Memory and speed gains are implementation-specific
DAMP reports a 69.1% reduction in recurrent-state storage at its 9.9-bit configuration, up to 2.59× speedup for the recurrent-state update kernel, and up to 19.0% lower full-model time per output token (TPOT), relative to FP32-state inference. These are author-reported results in SGLang. In the paper’s batch-size-256 decoding results, TPOT fell 19.0% on Qwen3.6, 14.5% on Kimi-Linear, and 7.3% on Kimi-K3; the authors suggest inter-device communication may contribute to the smaller Kimi-K3 reduction. In a multi-turn Kimi-K3 setting, DAMP also reduced mean time to first token by 20.7% versus FP32 and 14.5% versus BF16. Do not assume those gains will transfer unchanged to different hardware, batch sizes, or serving stacks. DAMP bibliographic record
STEPQuant reports more than 5× recurrent-state compression at nominal 6 bits and up to 68.7% lower total serving memory when integrated into SGLang with optimized GPU kernels. One Qwen serving measurement used 28.609 MiB per request for packed pages versus 144 MiB for FP32, a 5.03× reduction. These figures come from STEPQuant’s own implementation and configuration; they should not be treated as a direct head-to-head comparison with DAMP’s results. STEPQuant bibliographic record
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Kernel speedup and end-to-end latency are different measurements. A faster state update does not guarantee an equal reduction in TPOT, because other work—including communication and the rest of the model—also contributes to serving time.
Best Value
How to decide for a deployment
Treat uniform INT8 as a candidate configuration to validate, not an automatic setting. Compare it with the precision alternatives your model and serving stack support, using the same workload and quality criteria.
- Test the tasks you actually serve. Include representative reasoning, coding, and long-context evaluations where relevant. DAMP’s benchmark results show why an average or a single favorable score can conceal a large task-specific drop.
- Measure total state memory, not just nominal bit width. Include packed codes, scales, precision maps or pivots, and any retained state or cache data. Compare the memory that the implementation actually allocates.
- Measure both update-kernel and end-to-end performance. Record recurrent-update latency and full-model TPOT under deployment-relevant batch size, concurrency, context length, and generation length. If tensor parallelism or communication is involved, test that setup too.
- Match the method to the state geometry. GDN, KDA, and Delta-rule state structures should not be treated as interchangeable. Confirm that the method and kernels support the architecture you deploy.
- Include calibration and operational overhead. Selective methods can require calibration, precision layouts, and compatible quantized update kernels. Count that work, along with software integration, against the measured memory or latency benefit.
- Keep the comparison reproducible. Record model checkpoint, quantizer configuration, benchmark versions, serving software, hardware, and workload settings. Results without those conditions are difficult to interpret or reproduce.
What the evidence does—and does not—establish
DAMP and STEPQuant are recent research preprints, not independent production-wide evaluations. Their experiments support caution about choosing uniform INT8 by habit and show that selective precision is a plausible alternative in the studied settings. They do not establish that INT8 is categorically unsuitable, that one method is best across architectures, or that the reported gains will recur on a different serving stack.
The scope also matters: these results concern recurrent states in particular linear-attention and Delta-rule architectures. They should not automatically be applied to ordinary transformer KV caches, all recurrent neural networks, every quantizer, or every production configuration. Read STEPQuant’s experimental report alongside the DAMP experimental report for the methods and conditions behind the reported measurements.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




