State-space / Mamba / linear attention
35 papers in this thread, across 9 domains.
Progress0 of 31
- 2026Gated DeltaNet-2: Decoupling Erase and Write in Linear AttentionArchitecture2605.22791NVIDIAMay 21, 2026~109 minArchitecture
- 2026Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningArchitecture2604.12374NVIDIAApr 14, 2026score 10~121 minArchitecture
- 2026Mamba-3: Improved Sequence Modeling using State Space PrinciplesArchitecture2603.15569Together AIMar 16, 2026~121 minArchitecture
- 2026MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context ModelingArchitecture2602.11761Feb 12, 2026~104 minArchitecture
- 2026SLA2: Sparse-Linear Attention with Learnable Routing and QATDiffusion2602.12675Feb 13, 2026~115 minDiffusion
- 2026Test-Time Training with KV Binding Is Secretly Linear AttentionLLM Systems2602.21204NVIDIAFeb 24, 2026score 9~131 minLLM Systems
- 2026MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-HeadArchitecture2601.07832DAGroup-PKUJan 12, 2026score 9~105 minArchitecture
- 2025Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningArchitecture2512.20848NVIDIADec 23, 2025~114 minArchitecture
- 2025Higher-order Linear AttentionArchitecture2510.27258Oct 31, 2025~112 minArchitecture
- 2025SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear AttentionArchitecture2509.24006Tsinghua UniversitySep 28, 2025score 10~110 minArchitecture
- 2025NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning ModelReasoning2508.14444NVIDIAAug 20, 2025score 10~119 minReasoning
- 2025A Systematic Analysis of Hybrid Linear AttentionArchitecture2507.06457Jul 8, 2025score 9~92 minArchitecture
- 2025ZeCO: Zero Communication Overhead Sequence Parallelism for Linear AttentionDistributed Training2507.01004Jul 1, 2025score 9~111 minDistributed Training
- 2025Log-Linear AttentionArchitecture2506.04761Jun 5, 2025~122 minArchitecture
- 2025RADLADS: Rapid Attention Distillation to Linear Attention Decoders at ScaleTraining Methods2505.03005May 5, 2025score 9~133 minTraining Methods
- 2025Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM PruningTraining Methods2504.11409Apr 15, 2025score 9~124 minTraining Methods
- 2025RWKV-7 "Goose" with Expressive Dynamic State EvolutionArchitecture2503.14456Mar 18, 2025score 9~123 minArchitecture
- 2024Stuffed Mamba: Oversized States Lead to the Inability to ForgetTraining Methods2410.07145Oct 9, 2024score 9~127 minTraining Methods
- 2024Jamba-1.5: Hybrid Transformer-Mamba Models at ScaleArchitecture2408.12570Aug 22, 2024~102 minArchitecture
- 2024GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache CompressionArchitecture2407.12077Jul 16, 2024score 9~124 minArchitecture
- 2024Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingArchitecture2406.07522Jun 11, 2024score 9~115 minArchitecture
- 2024Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityArchitecture2405.21060Amazon ScienceMay 31, 2024score 10~123 minArchitecture
- 2024Eagle and Finch: RWKV with Matrix-Valued States and Dynamic RecurrenceArchitecture2404.05892Apr 8, 2024score 9~133 minArchitecture
- 2024Jamba: A Hybrid Transformer-Mamba Language ModelArchitecture2403.19887AI21Mar 28, 2024~130 minArchitecture
- 2024Simple linear attention language models balance the recall-throughput tradeoffArchitecture2402.18668Together AIFeb 28, 2024score 9~112 minArchitecture
- 2024The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax MimicryLLM Systems2402.04347Feb 6, 2024score 9~113 minLLM Systems
- 2024MambaByte: Token-free Selective State Space ModelArchitecture2401.13660Jan 24, 2024score 9~116 minArchitecture
- 2024MoE-Mamba: Efficient Selective State Space Models with Mixture of ExpertsMixture of Experts2401.04081Jan 8, 2024score 9~97 minMixture of Experts
- 2023Codestral MambaArchitecture2312.00752MistralDec 1, 2023~109 minArchitecture
- 2023Read MoreArchitecture2305.13048EleutherAIMay 22, 2023score 9~106 minArchitecture
- 2022Hungry Hungry Hippos: Towards Language Modeling with State Space ModelsArchitecture2212.14052Together AIDec 28, 2022~118 minArchitecture
- 2021Structured Denoising Diffusion Models in Discrete State-Spacesno summary yetcs lg2107.03006Google Research82 citesJul 7, 2021cs lg
- 2019Deep Physiological State Space Model for Clinical Forecastingno summary yetcs lg1912.01762Google Research1 citesDec 4, 2019cs lg
- 2019Explicit Explore-Exploit Algorithms in Continuous State Spacesno summary yetcs lg1911.00617Microsoft Research14 citesNov 1, 2019cs lg
- 2019Continuous Value Iteration (CVI) Reinforcement Learning and Imaginary Experience Replay (IER) for learning multi-goal, continuous action and state space controllersno summary yetcs ai1908.10255Sony5 citesAug 27, 2019cs ai