a new KV-cache eviction mechanism using single-hop, drift-free rotation for up to 4.6x speedup than continuous re-rotation, no accuracy cost.
-
Updated
Sep 19, 2026 - Jupyter Notebook
a new KV-cache eviction mechanism using single-hop, drift-free rotation for up to 4.6x speedup than continuous re-rotation, no accuracy cost.
[NeurIPS 2025] NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
Intent-aware KV execution prototype for agentic long-context inference: semantic block selection, dynamic scoring, KV quantization modeling, speculative prefetch simulation, CPU references, and future Triton/CUDA kernels.
Transformer attention, worked through numerically; from self-attention and RoPE to KV cache, MQA, and GQA.
MeritKV: Proactive KV cache admission for Large Language Models. Reduces prefill latency, energy consumption, and extends effective cache capacity by selectively caching tokens based on long-term utility.
0ns zero-copy, autograd-free hybrid guide layer (Pre-Transformer Packet Rectifier). Uses viscous Burgers' & Vorticity under FNG V3 to pre-rectify high-order skewness & stream clean tensor manifolds straight into LLM Attention.
To associate your repository with the kv-cache-optimization topic, visit your repo's landing page and select "manage topics."