Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
-
Updated
Sep 11, 2026 - C++
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
[Neurips 2025] R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
Inference-native Tokenmaxxing Agent Harness for Loop Engineering
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
PiKV: KV Cache Management System for Mixture of Experts [Efficient ML System]
Recomputation-free, context-independent KV caching for large language models
(ACL2025 oral) SCOPE: Optimizing KV Cache Compression in Long-context Generation
🔥 [ICML'26] ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
Span Queries: What if we had a way to plan and optimize GenAI like we do for SQL?
A TurboQuant implementation with Llama.cpp for AMD with Vulkan runtime
KV Cache with PagedAttention vs PagedAttention + TurboQuant - experiments across token sizes comparing memory, latency, and accuracy.
RestoreKV: Recovering full-cache behavior under aggressive query-agnostic KV cache eviction.
To associate your repository with the kvcache topic, visit your repo's landing page and select "manage topics."