Mixture of Experts
37 papers in this thread, across 10 domains.
Progress0 of 31
- 2026Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Modelsno summary yetcs dc2607.01844Salesforce0 citesJul 2, 2026cs dc
- 2026SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-trainingArchitecture2605.08738May 9, 2026~106 minArchitecture
- 2026EMO: Pretraining Mixture of Experts for Emergent ModularityMixture of Experts2605.06663Allen Institute for AIMay 7, 2026~100 minMixture of Experts
- 2026UniPool: A Globally Shared Expert Pool for Mixture-of-ExpertsMixture of Experts2605.06665CUHKMay 7, 2026~113 minMixture of Experts
- 2026DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side DevicesMixture of Experts2605.10933Tsinghua NLP GroupMay 11, 2026~105 minMixture of Experts
- 2026BEAM: Binary Expert Activation Masking for Dynamic Routing in MoEMixture of Experts2605.14438alibaba-incMay 14, 2026~112 minMixture of Experts
- 2026Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningArchitecture2604.12374NVIDIAApr 14, 2026score 10~121 minArchitecture
- 2026OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at ScaleMixture of Experts2602.05711Feb 5, 2026~115 minMixture of Experts
- 2026ConceptMoE: Adaptive Token-to-Concept Compression for Implicit Compute AllocationArchitecture2601.21420ByteDance SeedJan 29, 2026score 3~103 minArchitecture
- 2026MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUsDistributed Training2601.05296Jan 8, 2026~113 minDistributed Training
- 2026TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-ExpertsMixture of Experts2601.08881Tencent HunyuanJan 12, 2026score 7~101 minMixture of Experts
- 2025Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningArchitecture2512.20848NVIDIADec 23, 2025~114 minArchitecture
- 2025SonicMoE: Accelerating MoE with IO and Tile-aware OptimizationsMixture of Experts2512.14080Dec 16, 2025~105 minMixture of Experts
- 2025Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary LossMixture of Experts2512.23447Dec 29, 2025~111 minMixture of Experts
- 2025OlmoEarth: Stable Latent Image Modeling for Multimodal Earth ObservationPretraining2511.13655Ai2Nov 17, 2025score 6~130 minPretraining
- 2025FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization ErrorTraining Methods2511.02302Nov 4, 2025~122 minTraining Methods
- 2025MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert ParallelismServing2504.02263Apr 3, 2025~102 minServing
- 2024Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by TencentMixture of Experts2411.02265Tencent HunyuanNov 4, 2024score 9~110 minMixture of Experts
- 2024MH-MoE: Multi-Head Mixture-of-ExpertsMixture of Experts2411.16205Nov 25, 2024score 9~107 minMixture of Experts
- 2024OLMoE: Open Mixture-of-Experts Language ModelsMixture of Experts2409.02060Allen Institute for AISep 3, 2024~96 minMixture of Experts
- 2024Layerwise Recurrent Router for Mixture-of-ExpertsMixture of Experts2408.06793Aug 13, 2024score 8~110 minMixture of Experts
- 2024AUXILIARY-LOSS-FREE LOAD BALANCING STRATEGY FOR MIXTURE-OF-EXPERTSTraining Methods2408.15664Aug 28, 2024~97 minTraining Methods
- 2024JetMoE: Reaching Llama2 Performance with 0.1M DollarsMixture of Experts2404.07413Apr 11, 2024score 9~79 minMixture of Experts
- 2024Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLMMixture of Experts2403.07816Mar 12, 2024score 9~107 minMixture of Experts
- 2024MoE-Mamba: Efficient Selective State Space Models with Mixture of ExpertsMixture of Experts2401.04081Jan 8, 2024score 9~97 minMixture of Experts
- 2024DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsMixture of Experts2401.06066DeepSeekJan 11, 2024score 10~107 minMixture of Experts
- 2023SwitchHead: Accelerating Transformers with Mixture-of-Experts AttentionMixture of Experts2312.07987Dec 13, 2023score 9~109 minMixture of Experts
- 2023HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Expertsno summary yetcs lg2312.07035Salesforce0 citesDec 12, 2023cs lg
- 2023QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter ModelsMixture of Experts2310.16795Oct 25, 2023score 9~111 minMixture of Experts
- 2023COMET: Learning Cardinality Constrained Mixture of Experts with Trees and Local Searchno summary yetcs lg2306.02824Google Research2 citesJun 5, 2023cs lg
- 2022Tutel: Adaptive Mixture-of-Experts at ScaleMixture of Experts2206.03382Jun 7, 2022~114 minMixture of Experts
- 2021GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsMixture of Experts2112.06905Dec 13, 2021~114 minMixture of Experts
- 2021Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inferenceno summary yetcs cl2110.03742Google Research0 citesSep 24, 2021cs cl
- 2021DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learningno summary yetcs lg2106.03760Google Research16 citesJun 7, 2021cs lg
- 2021Scaling Vision with Sparse Mixture of Expertsno summary yetcs cv2106.05974Amazon29 citesJun 10, 2021cs cv
- 2017DeepETA: How Uber Predicts Arrival Times Using Deep LearningMixture of Experts1701.06538Uber268 citesJan 23, 2017~112 minMixture of Experts
- 2013Learning Factored Representations in a Deep Mixture of ExpertsArchitecture1312.4314137 citesDec 16, 2013~116 minArchitecture