Distillation
100 papers in this thread, across 23 domains.
Progress0 of 39
- 2026Procedural Memory Distillation: Online Reflection for Self-Improving Language Modelsno summary yetcs ai2607.01480Salesforce0 citesJul 1, 2026cs ai
- 2026Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agentsno summary yetcs lg2606.12634Amazon0 citesJun 10, 2026cs lg
- 2026Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillationno summary yetcs lg2606.13657Alibaba0 citesJun 11, 2026cs lg
- 2026RT-VLA: Real-Time Vision-Language-Action Models via Knowledge Distillationno summary yetcs cv2606.14010CMU0 citesJun 12, 2026cs cv
- 2026HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon AgentsAgents2605.17873KAIST AIMay 18, 2026~112 minAgents
- 2026It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMsAlignment2605.20258KAIST AIMay 18, 2026~122 minAlignment
- 2026SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-trainingArchitecture2605.08738May 9, 2026~106 minArchitecture
- 2026Learning from Language Feedback via Variational Policy DistillationReasoning2605.15113Salesforce AI ResearchMay 14, 2026~108 minReasoning
- 2026Continuous-Time Distribution Matching for Few-Step Diffusion DistillationTraining Methods2605.06376alibaba-incMay 7, 2026~110 minTraining Methods
- 2026Adaptive Teacher Exposure for Self-Distillation in LLM ReasoningTraining Methods2605.11458ByteDanceMay 12, 2026~124 minTraining Methods
- 2026Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy DistillationTraining Methods2605.11739Tencent HunyuanMay 12, 2026~113 minTraining Methods
- 2026D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Modelsno summary yetcs cv2605.05204Tongyi-MAIMay 6, 2026cs cv
- 2026A Survey of On-Policy Distillation for Large Language ModelsTraining Methods2604.00626Apr 1, 2026~136 minTraining Methods
- 2026Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense SupervisionTraining Methods2604.12002Princeton UniversityApr 13, 2026score 9~122 minTraining Methods
- 2026Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy DistillationTraining Methods2604.13010NVIDIAApr 14, 2026~132 minTraining Methods
- 2026Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and RecipeTraining Methods2604.13016OpenBMBApr 14, 2026score 10~136 minTraining Methods
- 2026The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy DistillationTraining Methods2604.16830Salesforce AI ResearchApr 18, 2026~115 minTraining Methods
- 2026TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous AgentsTraining Methods2604.24005TongyiLabApr 27, 2026~112 minTraining Methods
- 2026Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?Reasoning2603.24472Microsoft ResearchMar 25, 2026score 9~109 minReasoning
- 2026Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy DistillationTraining Methods2603.19220NVIDIAMar 19, 2026~94 minTraining Methods
- 2026ArcFlow: Unleashing 2-Step Text-to-Image Generation via High-Precision Non-Linear Flow DistillationDiffusion2602.09014Fudan UniversityFeb 9, 2026score 8~99 minDiffusion
- 2026SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-TuningInference Optimization2602.13515Feb 13, 2026~108 minInference Optimization
- 2026Learning beyond Teacher: Generalized On-Policy Distillation with Reward ExtrapolationTraining Methods2602.12125Tencent HunyuanFeb 12, 2026score 9~91 minTraining Methods
- 2026jina-embeddings-v5-text: Task-Targeted Embedding DistillationTraining Methods2602.15547Feb 17, 2026~106 minTraining Methods
- 2026Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long ContextsArchitecture2601.22156OpenBMBJan 29, 2026score 9~127 minArchitecture
- 2026Transition Matching Distillation for Fast Video GenerationInference Optimization2601.09881NVIDIAJan 14, 2026score 9~115 minInference Optimization
- 2026Distribution-Aligned Sequence Distillation for Superior Long-CoT ReasoningReasoning2601.09088Alibaba Cloud Apsara Lab Jan 14, 2026score 6~88 minReasoning
- 2025Few-Step Distillation for Text-to-Image Generation: A Practical GuideDiffusion2512.13006Alibaba DAMODec 15, 2025score 6~86 minDiffusion
- 20254D-RGPT: Toward Region-level 4D Understanding via Perceptual DistillationMultimodal2512.17012NVIDIADec 18, 2025score 6~109 minMultimodal
- 2025Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode SupervisionTraining Methods2512.15489NVIDIADec 17, 2025score 8~103 minTraining Methods
- 2025Black-Box On-Policy Distillation of Large Language ModelsTraining Methods2511.10643Nov 13, 2025~100 minTraining Methods
- 2025BitNet DistillationTraining Methods2510.13998Microsoft ResearchOct 15, 2025score 9~127 minTraining Methods
- 2025RADLADS: Rapid Attention Distillation to Linear Attention Decoders at ScaleTraining Methods2505.03005May 5, 2025score 9~133 minTraining Methods
- 2025Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM PruningTraining Methods2504.11409Apr 15, 2025score 9~124 minTraining Methods
- 2024Pre-training Distillation for Large Language Models: A Design Space ExplorationPretraining2410.16215Oct 21, 2024score 10~124 minPretraining
- 2024Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillationno summary yetcs cv2409.02555Baidu23 citesSep 4, 2024cs cv
- 2024LLM Pruning and Distillation in Practice: The Minitron ApproachTraining Methods2408.11796Aug 21, 2024~110 minTraining Methods
- 2024Compact Language Models via Pruning and Knowledge DistillationPretraining2407.14679Jul 19, 2024~140 minPretraining
- 2024SuperPADL: Scaling Language-Directed Physics-Based Control with Progressive Supervised Distillationno summary yetcs lg2407.10481NVIDIA2 citesJul 15, 2024cs lg
- 2024LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt CompressionContext Optimization2403.12968Microsoft ResearchMar 19, 2024score 8~113 minContext Optimization
- 2023DDistill-SR: Reparameterized Dynamic Distillation Network for Lightweight Image Super-Resolutionno summary yeteess iv2312.14551Baidu34 citesDec 22, 2023eess iv
- 2023RLCD: Reinforcement Learning from Contrast Distillation for Language Model AlignmentRL Training2307.12950Jul 24, 2023score 9~85 minRL Training
- 2023On-Policy Distillation of Language Models: Learning from Self-Generated MistakesTraining Methods2306.13649DeepMindJun 23, 2023score 10~116 minTraining Methods
- 2022UM4: Unified Multilingual Multiple Teacher-Student Model for Zero-Resource Neural Machine Translationno summary yetcs cl2207.04900Tencent11 citesJul 11, 2022cs cl
- 2022Non-Local Latent Relation Distillation for Self-Adaptive 3D Human Pose Estimationno summary yetcs cv2204.01971Google Research3 citesApr 5, 2022cs cv
- 2022Spatial Likelihood Voting with Self-Knowledge Distillation for Weakly Supervised Object Detectionno summary yetcs cv2204.06899Alibaba4 citesApr 14, 2022cs cv
- 2021Unsupervised Representation Learning Meets Pseudo-Label Supervised Self-Distillation: A New Approach to Rare Disease Classificationno summary yetcs cv2110.04558Tencent12 citesOct 9, 2021cs cv
- 2021Improving Neural Ranking via Lossless Knowledge Distillationno summary yetcs ir2109.15285Google Research1 citesSep 30, 2021cs ir
- 2021Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inferenceno summary yetcs cl2110.03742Google Research0 citesSep 24, 2021cs cl
- 2021ALBEF: Align Before FuseMultimodal2107.07651Jul 16, 2021~118 minMultimodal
- 2021ROD: Reception-aware Online Distillation for Sparse Graphsno summary yetcs lg2107.11789Alibaba22 citesJul 25, 2021cs lg
- 2021Dataset Distillation with Infinitely Wide Convolutional Networksno summary yetcs lg2107.13034Google Research22 citesJul 27, 2021cs lg
- 2021ERNIE-Tiny : A Progressive Distillation Framework for Pretrained Transformer Compressionno summary yetcs cl2106.02241Baidu3 citesJun 4, 2021cs cl
- 2021MergeDistill: Merging Pre-trained Language Models using Distillationno summary yetcs cl2106.02834Google Research1 citesJun 5, 2021cs cl
- 2021Does Knowledge Distillation Really Work?no summary yetcs lg2106.05945Google Research11 citesJun 10, 2021cs lg
- 2021Teacher's pet: understanding and mitigating biases in distillationno summary yetcs lg2106.10494Google Research5 citesJun 19, 2021cs lg
- 2021Local-Global Knowledge Distillation in Heterogeneous Federated Learning with Non-IID Datano summary yetcs lg2107.00051Google Research36 citesJun 30, 2021cs lg
- 2021Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matchingno summary yetcs cv2105.08252Baidu1 citesMay 18, 2021cs cv
- 2021Selective Knowledge Distillation for Neural Machine Translationno summary yetcs cl2105.12967Tencent2 citesMay 27, 2021cs cl
- 2020Training data-efficient image transformers & distillation through attentionVision2012.12877Meta AI / FAIRDec 23, 2020~106 minVision
- 2020Efficient Knowledge Distillation for RNN-Transducer Modelsno summary yeteess as2011.06110Google Research3 citesNov 11, 2020eess as
- 2020Neighbourhood Distillation: On the benefits of non end-to-end distillationno summary yetcs lg2010.01189Google Research0 citesOct 2, 2020cs lg
- 2020DiPair: Fast and Accurate Distillation for Trillion-Scale Text Matching and Pair Modelingno summary yetcs cl2010.03099Google Research23 citesOct 7, 2020cs cl
- 2020Unsupervised Distillation of Syntactic Information from Contextualized Word Representationsno summary yetcs cl2010.05265AllenAI1 citesOct 11, 2020cs cl
- 2020Anti-Distillation: Improving reproducibility of deep networksno summary yetcs lg2010.09923Google Research6 citesOct 19, 2020cs lg
- 2020A Joint Learning Approach based on Self-Distillation for Keyphrase Extraction from Scientific Documentsno summary yetcs cl2010.11980Adobe2 citesOct 22, 2020cs cl
- 2020Improving Streaming Automatic Speech Recognition With Non-Streaming Model Distillation On Unsupervised Datano summary yetcs sd2010.12096Google Research1 citesOct 22, 2020cs sd
- 2020Iterative Graph Self-Distillationno summary yetcs lg2010.12609Salesforce10 citesOct 23, 2020cs lg
- 2020Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillationno summary yetcs cv2007.01951Tencent8 citesJul 3, 2020cs cv
- 2020Interpretable Foreground Object Search As Knowledge Distillationno summary yetcs cv2007.09867Alibaba1 citesJul 20, 2020cs cv
- 2020DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion ParametersTraining Methods2006.05525Jun 9, 2020~126 minTraining Methods
- 2020Why distillation helps: a statistical perspectiveno summary yetcs lg2005.10419Google Research16 citesMay 21, 2020cs lg
- 2020Syntactic Structure Distillation Pretraining For Bidirectional Encodersno summary yetcs cl2005.13482DeepMind5 citesMay 27, 2020cs cl
- 2020Transferring Inductive Biases through Knowledge Distillationno summary yetcs lg2006.00555Google Research13 citesMay 31, 2020cs lg
- 2020XtremeDistil: Multi-stage Distillation for Massive Multilingual Modelsno summary yetcs cl2004.05686Microsoft Research4 citesApr 12, 2020cs cl
- 2020PoseNet3D: Learning Temporally Consistent 3D Human Pose via Knowledge Distillationno summary yetcs cv2003.03473Amazon2 citesMar 7, 2020cs cv
- 2020Collaborative Distillation for Ultra-Resolution Universal Style Transferno summary yetcs cv2003.08436Adobe7 citesMar 18, 2020cs cv
- 2020Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation from a Blackbox Modelno summary yetcs cv2003.13960Google Research4 citesMar 31, 2020cs cv
- 2020MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersDistributed Training2002.10957Microsoft ResearchFeb 25, 2020~97 minDistributed Training
- 2020Understanding and Improving Knowledge Distillationno summary yetcs lg2002.03532Google Research90 citesFeb 10, 2020cs lg
- 2020Improving Face Recognition from Hard Samples via Distribution Distillation Lossno summary yetcs cv2002.03662Tencent3 citesFeb 10, 2020cs cv
- 2020Self-Distillation Amplifies Regularization in Hilbert Spaceno summary yetcs lg2002.05715Google Research98 citesFeb 13, 2020cs lg
- 2019MKD: a Multi-Task Knowledge Distillation Approach for Pretrained Language Modelsno summary yetcs cl1911.03588Salesforce19 citesNov 9, 2019cs cl
- 2019Few Shot Network Compression via Cross Distillationno summary yetcs lg1911.09450Tencent17 citesNov 21, 2019cs lg
- 2019Contrastive Representation Distillationno summary yetcs lg1910.10699Google Research64 citesOct 23, 2019cs lg
- 2019Patient Knowledge Distillation for BERT Model Compressionno summary yetcs cl1908.09355Microsoft Research85 citesAug 25, 2019cs cl
- 2019Privileged Features Distillation at Taobao Recommendationsno summary yetcs ir1907.05171Alibaba1 citesJul 11, 2019cs ir
- 2019Scalable Syntax-Aware Language Models Using Knowledge Distillationno summary yetcs cl1906.06438Google Research3 citesJun 14, 2019cs cl
- 2019Distilling Policy Distillationno summary yetcs lg1902.02186DeepMind39 citesFeb 6, 2019cs lg
- 2019Improved Knowledge Distillation via Teacher Assistantno summary yetcs lg1902.03393DeepMind116 citesFeb 9, 2019cs lg
- 2019DDFlow: Learning Optical Flow with Unlabeled Data Distillationno summary yetcs cv1902.09145Tencent17 citesFeb 25, 2019cs cv
- 2019Multilingual Neural Machine Translation with Knowledge Distillationno summary yetcs cl1902.10461Microsoft Research129 citesFeb 27, 2019cs cl
- 2018Spatial Knowledge Distillation to aid Visual Reasoningno summary yetcs cv1812.03631Adobe2 citesDec 10, 2018cs cv
- 2018Optimal Completion Distillation for Sequence Learningno summary yetcs lg1810.01398Google Research5 citesOct 2, 2018cs lg
- 2018Exploration by Random Network Distillationno summary yetcs lg1810.12894OpenAI259 citesOct 30, 2018cs lg
- 2018Large scale distributed neural network training through online distillationno summary yetcs lg1804.03235Google Research152 citesApr 9, 2018cs lg
- 2018Model compression via distillation and quantizationno summary yetcs ne1802.05668Google Research262 citesFeb 15, 2018cs ne
- 2017Graph Distillation for Action Detection with Privileged Modalitiesno summary yetcs cv1712.00108Google Research3 citesNov 30, 2017cs cv
- 2017Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillationno summary yetstat ml1710.06169Microsoft Research24 citesOct 17, 2017stat ml
- 2015Policy Distillationno summary yetcs lg1511.06295DeepMind87 citesNov 19, 2015cs lg