89 papers
Mixture of Experts
0/89MoE routing, sparsity, and expert architectures.
Progress0 of 89
2026
24- MaySlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-trainingarchitecture2605.08738May 9, 2026~106 min
- MayEMO: Pretraining Mixture of Experts for Emergent Modularitymoe2605.06663Allen Institute for AIMay 7, 2026~100 min
- MayUniPool: A Globally Shared Expert Pool for Mixture-of-Expertsmoe2605.06665CUHKMay 7, 2026~113 min
- MayDECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devicesmoe2605.10933Tsinghua NLP GroupMay 11, 2026~105 min
- MayBEAM: Binary Expert Activation Masking for Dynamic Routing in MoEmoe2605.14438alibaba-incMay 14, 2026~112 min
- AprNemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoningarchitecture2604.12374NVIDIAApr 14, 2026score 10~121 min
- AprCoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generationdiffusion2604.19636alibaba-incApr 21, 2026~120 min
- AprQwen3.5-Omni Technical Reportmultimodal2604.15804Apr 17, 2026~123 min
- AprHY-Embodied-0.5: Embodied Foundation Models for Real-World Agentsuncategorized2604.07430Tencent HunyuanApr 8, 2026score 9~127 min
- MarTimer-S1: A Billion-Scale Time Series Foundation Model with Serial Scalingllm-systems2603.04791ByteDanceMar 5, 2026score 3~108 min
- MarBeyond Language Modeling: An Exploration of Multimodal Pretrainingmultimodal2603.03276AI at MetaMar 3, 2026score 9~119 min
- MarLongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learningreasoning2603.21065Meituan LongCatMar 22, 2026score 9~110 min
- FebOmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scalemoe2602.05711Feb 5, 2026~115 min
- FebStep 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parametersmoe2602.10604StepFunFeb 11, 2026score 8~135 min
- FebArcee Trinity Large Technical Reportmoe2602.17004Feb 19, 2026~114 min
- FebERNIE 5.0 Technical Reportmultimodal2602.04705Feb 4, 2026~138 min
- FebSPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learningpretraining2602.02472ByteDance SeedFeb 2, 2026score 2~112 min
- FebQwen3-Coder-Next Technical Reporttraining-methods2603.00729QwenFeb 28, 2026score 9~123 min
- JanLongCat-Flash-Thinking-2601 Technical Reportagents2601.16725LongCatJan 23, 2026score 9~101 min
- JanScaling Embeddings Outperforms Scaling Experts in Language Modelsarchitecture2601.21204Meituan LongCatJan 29, 2026score 8~115 min
- JanConceptMoE: Adaptive Token-to-Concept Compression for Implicit Compute Allocationarchitecture2601.21420ByteDance SeedJan 29, 2026score 3~103 min
- JanMoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUsdistributed-training2601.05296Jan 8, 2026~113 min
- JanK-EXAONE Technical Reportmoe2601.01739LG EXAONEJan 5, 2026~112 min
- JanTAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Expertsmoe2601.08881Tencent HunyuanJan 12, 2026score 7~101 min
2025
29- DecNemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoningarchitecture2512.20848NVIDIADec 23, 2025~114 min
- DecNVIDIA Nemotron 3: Efficient and Open Intelligencearchitecture2512.20856NVIDIADec 24, 2025~74 min
- DecLLaDA2.0: Scaling Up Diffusion Language Models to 100Bdiffusion2512.15745Dec 10, 2025~122 min
- DecUnderstanding and Harnessing Sparsity in Unified Multimodal Modelsmoe2512.02351ByteDance SeedDec 2, 2025score 9~121 min
- DecSonicMoE: Accelerating MoE with IO and Tile-aware Optimizationsmoe2512.14080Dec 16, 2025~105 min
- DecCoupling Experts and Routers in Mixture-of-Experts via an Auxiliary Lossmoe2512.23447Dec 29, 2025~111 min
- DecBridging Your Imagination with Audio-Video Generation via a Unified Directormultimodal2512.23222ByteDanceDec 29, 2025score 6~102 min
- DecStabilizing Reinforcement Learning with LLMs: Formulation and Practicesrl-training2512.01374Dec 1, 2025~126 min
- NovQwen3-VL Technical Reportmultimodal2511.21631Nov 26, 2025~107 min
- NovSoft Adaptive Policy Optimizationrl-training2511.20347QwenNov 25, 2025score 9~108 min
- NovFP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Errortraining-methods2511.02302Nov 4, 2025~122 min
- OctEvery Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundationmoe2510.22115inclusionAIOct 25, 2025score 10~127 min
- OctLongCat-Flash-Omni Technical Reportmultimodal2511.00279Meituan LongCatOct 31, 2025score 9~117 min
- OctRDMA Point-to-Point Communication for LLM Systemsserving2510.27656Oct 31, 2025~117 min
- SepHunyuanImage 3.0 Technical Reportmoe2509.23951Tencent HunyuanSep 28, 2025score 9~121 min
- SepSAIL-VL2 Technical Reportmultimodal2509.14033Sep 17, 2025~130 min
- SepQwen3-Omni Technical Reportmultimodal2509.17765Sep 22, 2025~129 min
- AugGLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Modelsllm-systems2508.06471Zhipu / GLMAug 8, 2025~138 min
- AugIntern-S1: A Scientific Multimodal Foundation Modelmultimodal2508.15763InternLM / Shanghai AI LabAug 21, 2025~122 min
- JulFlexOlmo: Open Language Models for Flexible Data Usemoe2507.07024Jul 9, 2025~105 min
- JulGSPO: Towards Scalable Reinforcement Learning for Language Modelsrl-training2507.18071Qwen / Alibaba CloudJul 24, 2025~94 min
- JunMiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attentionreasoning2506.13585DeepSeekJun 16, 2025score 9~132 min
- MayQwen3 Technical Reportllm-systems2505.09388Qwen / Alibaba CloudMay 14, 2025~128 min
- MayInsights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architecturesmoe2505.09343NVIDIAMay 14, 2025score 10~111 min
- MayModel Merging in Pre-training of Large Language Modelspretraining2505.12082May 17, 2025~98 min
- AprKimi-VL Technical Reportmultimodal2504.07491NVIDIAApr 10, 2025~121 min
- AprMegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelismserving2504.02263Apr 3, 2025~102 min
- MarPhi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAsmultimodal2503.01743Microsoft ResearchMar 3, 2025score 10~111 min
- JanMiniMax-01: Scaling Foundation Models with Lightning Attentionarchitecture2501.08313MiniMaxJan 14, 2025~123 min
2024
22- DecQwen2.5 Technical Reportpretraining2412.15115Qwen / Alibaba CloudDec 19, 2024~109 min
- NovMixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Modelsarchitecture2411.04996Nov 7, 2024score 10~118 min
- NovHunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencentmoe2411.02265Tencent HunyuanNov 4, 2024score 9~110 min
- NovMH-MoE: Multi-Head Mixture-of-Expertsmoe2411.16205Nov 25, 2024score 9~107 min
- SepConfigurable Foundation Models: Building LLMs from a Modular Perspectivearchitecture2409.02877OpenBMBSep 4, 2024score 9~129 min
- SepOLMoE: Open Mixture-of-Experts Language Modelsmoe2409.02060Allen Institute for AISep 3, 2024~96 min
- SepMM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-Tuningmultimodal2409.20566AppleSep 30, 2024~110 min
- AugJamba-1.5: Hybrid Transformer-Mamba Models at Scalearchitecture2408.12570Aug 22, 2024~102 min
- AugLayerwise Recurrent Router for Mixture-of-Expertsmoe2408.06793Aug 13, 2024score 8~110 min
- AugAUXILIARY-LOSS-FREE LOAD BALANCING STRATEGY FOR MIXTURE-OF-EXPERTStraining-methods2408.15664Aug 28, 2024~97 min
- JulQwen2 Technical Reportllm-systems2407.10671Qwen / Alibaba CloudJul 15, 2024~114 min
- JulLet the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Modelsmoe2407.01906DeepSeekJul 2, 2024score 9~107 min
- JunDeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligencepretraining2406.11931DeepSeekJun 17, 2024score 10~11 min
- AprJetMoE: Reaching Llama2 Performance with 0.1M Dollarsmoe2404.07413Apr 11, 2024score 9~79 min
- AprMiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategiestraining-methods2404.06395OpenBMBApr 9, 2024score 9~139 min
- MarJamba: A Hybrid Transformer-Mamba Language Modelarchitecture2403.19887AI21Mar 28, 2024~130 min
- MarGemini 1.5: Unlocking multimodal understanding across millions of tokens of contextcontext-optimization2403.05530Mar 8, 2024score 10~133 min
- MarBranch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLMmoe2403.07816Mar 12, 2024score 9~107 min
- MarMini-Gemini: Mining the Potential of Multi-modality Vision Language Modelsmultimodal2403.18814Mar 27, 2024score 9~109 min
- JanMoE-Mamba: Efficient Selective State Space Models with Mixture of Expertsmoe2401.04081Jan 8, 2024score 9~97 min
- JanMixtral of Expertsmoe2401.04088Jan 8, 2024~103 min
- JanDeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Modelsmoe2401.06066DeepSeekJan 11, 2024score 10~107 min
2023
6- DecSwitchHead: Accelerating Transformers with Mixture-of-Experts Attentionmoe2312.07987Dec 13, 2023score 9~109 min
- NovMemory Augmented Language Models through Mixture of Word Expertsarchitecture2311.10768Nov 15, 2023score 9~93 min
- OctQMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Modelsmoe2310.16795Oct 25, 2023score 9~111 min
- AugFrom Sparse to Soft Mixture of Expertsarchitecture2308.00951DeepMindAug 2, 2023score 9~109 min
- MayRead Morearchitecture2305.13048EleutherAIMay 22, 2023score 9~106 min
- MayBrainformers: Trading Simplicity for Efficiencyarchitecture2306.00008May 29, 2023score 9~88 min