KV cache / paged attention
16 papers in this thread, across 6 domains.
Progress0 of 13
- 2026HERALD: High-Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrievalno summary yetcs lg2606.21633Berkeley0 citesJun 19, 2026cs lg
- 2026Epiphany-Aware KV Cache Eviction Without the Attention Matrixno summary yetcs lg2606.26472CMU0 citesJun 25, 2026cs lg
- 2026OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache QuantizationInference Optimization2605.17757TogetherMay 18, 2026~114 minInference Optimization
- 2026KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compressionno summary yetcs cl2607.01237CMU0 citesMay 1, 2026cs cl
- 2026HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache SharingArchitecture2602.03560Xiaomi MiMoFeb 3, 2026score 9~105 minArchitecture
- 2025Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision BoostLLM Systems2511.18643Together AINov 23, 2025~108 minLLM Systems
- 2025The Pitfalls of KV Cache CompressionInference Optimization2510.00231Sep 30, 2025~125 minInference Optimization
- 2025Inference-Time Hyper-Scaling with KV Cache CompressionInference Optimization2506.05345NVIDIAJun 5, 2025score 9~102 minInference Optimization
- 2025Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel DecodingInference Optimization2505.22618May 28, 2025score 10~107 minInference Optimization
- 2024ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM InferenceServing2410.21465ByteDance SeedOct 28, 2024score 9~114 minServing
- 2024GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache CompressionArchitecture2407.12077Jul 16, 2024score 9~124 minArchitecture
- 2024A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache CompressionInference Optimization2406.11430Jun 17, 2024score 8~93 minInference Optimization
- 2024InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache ManagementInference Optimization2406.19707Jun 28, 2024~107 minInference Optimization
- 2024Layer-Condensed KV Cache for Efficient Inference of Large Language ModelsArchitecture2405.10637May 17, 2024score 9~96 minArchitecture
- 2024ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase PartitionServing2402.15220Feb 23, 2024score 10~108 minServing
- 2023Efficient Memory Management for Large Language Model Serving with PagedAttentionServing2309.06180Sep 12, 2023score 10~122 minServing