166 papers
Alignment
0/166RLHF, DPO, preference learning, and safety tuning.
Progress0 of 166
2026
29- MayQwen-Scope: Turning Sparse Features into Development Tools for Large Language Modelsalignment2605.11887Qwen / Alibaba CloudMay 12, 2026~138 min
- MayIt Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMsalignment2605.20258KAIST AIMay 18, 2026~122 min
- MayConditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignmentalignment2605.20834May 20, 2026~103 min
- MayLearning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillationtraining-methods2605.11739Tencent HunyuanMay 12, 2026~113 min
- MaySee What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understandingtraining-methods2605.18018TongyiLabMay 18, 2026~118 min
- AprBeyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Modelsevaluation2604.02315Salesforce AI ResearchApr 2, 2026score 7~112 min
- AprModeling Multiple Support Strategies within a Single Turn for Emotional Support Conversationsrl-training2604.17972Qwen DianJinApr 20, 2026~118 min
- AprRethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipetraining-methods2604.13016OpenBMBApr 14, 2026score 10~136 min
- AprLeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectoriestraining-methods2604.15311ByteDance SeedApr 16, 2026score 8~110 min
- AprThe Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillationtraining-methods2604.16830Salesforce AI ResearchApr 18, 2026~115 min
- MarCharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Productionalignment2603.01973Meta LlamaMar 2, 2026score 9~140 min
- MarLearning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Usealignment2603.03205Microsoft ResearchMar 3, 2026score 8~111 min
- MarReasoning Models Struggle to Control their Chains of Thoughtalignment2603.05706OpenAIMar 5, 2026score 9~99 min
- MarRubricBench: Aligning Model-Generated Rubrics with Human Standardsevaluation2603.01562Tencent HunyuanMar 2, 2026score 9~105 min
- MarHow Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularitiesevaluation2603.02578alibaba-incMar 3, 2026score 6~107 min
- MarAI Can Learn Scientific Tasterl-training2603.14473OpenMOSSMar 15, 2026~105 min
- MarBenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMssafety2603.16557LG AI ResearchMar 17, 2026score 8~101 min
- FebWhy Steering Works: Toward a Unified View of Language Model Parameter Dynamicsalignment2602.02343alibabaFeb 2, 2026score 4~107 min
- FebVLS: Steering Pretrained Robot Policies via Vision-Language Modelsalignment2602.03973Ai2Feb 3, 2026score 2~100 min
- FebOutcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Modelsalignment2602.04649QwenFeb 4, 2026score 9~109 min
- FebP-GenRM: Personalized Generative Reward Model with Test-time User-based Scalingalignment2602.12116Tongyi-ConvAIFeb 12, 2026score 9~115 min
- FebOn the Optimal Reasoning Length for RL-Trained Language Modelsreasoning2602.09591DeepSeekFeb 10, 2026~137 min
- FebGood SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learningrl-training2602.01058Feb 1, 2026~103 min
- FebThe Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societiessafety2602.09877AnthropicFeb 10, 2026score 3~121 min
- FebPhyCritic: Multimodal Critic Models for Physical AItraining-methods2602.11124NVIDIAFeb 11, 2026score 6~105 min
- JanLost in the Noise: How Reasoning Models Fail with Contextual Distractorsreasoning2601.07226KAIST AIJan 12, 2026score 3~106 min
- JanGDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimizationrl-training2601.05242Jan 8, 2026~109 min
- JanBuilding Production-Ready Probes For Geminisafety2601.11516Jan 16, 2026score 3~145 min
- JanDenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignmenttraining-methods2601.20218TongyiLabJan 28, 2026score 6~100 min
2025
33- DecAligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Modelsalignment2512.04981KAIST AIDec 4, 2025score 4~119 min
- DecFantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Processreasoning2512.23988Dec 30, 2025~119 min
- DecTaxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Modelssafety2512.05339RobloxDec 5, 2025score 5~96 min
- NovDynamic Reflections: Probing Video Representations with Text Alignmentalignment2511.02767DeepMindNov 4, 2025score 9~115 min
- NovVideo Generation Models Are Good Latent Reward Modelsalignment2511.21541Tencent HunyuanNov 26, 2025score 8~104 min
- NovWUSH: Near-Optimal Adaptive Transforms for LLM Quantizationalignment2512.00956 IST Austria Distributed Algorithms and Systems LabNov 30, 2025score 9~117 min
- NovThe Path Not Taken: RLVR Provably Learns Off the Principalsrl-training2511.08567AI at MetaNov 11, 2025score 9~110 min
- NovBlack-Box On-Policy Distillation of Large Language Modelstraining-methods2511.10643Nov 13, 2025~100 min
- OctBeyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Modelsalignment2510.21978AI at MetaOct 24, 2025score 10~112 min
- OctVibe Checker: Aligning Code Evaluation with Human Preferenceevaluation2510.07315Oct 8, 2025~114 min
- OctBeyond Correctness: Evaluating Subjective Writing Preferences Across Culturesevaluation2510.14616ByteDance SeedOct 16, 2025score 8~104 min
- OctLarge Reasoning Models Learn Better Alignment from Flawed Thinkingsafety2510.00938Oct 1, 2025~96 min
- OctDistractor Injection Attacks on Large Reasoning Models: Characterization and Defensesafety2510.16259Amazon ScienceOct 17, 2025score 6~101 min
- OctAny-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depthsafety2510.18081ByteDance SeedOct 20, 2025score 8~133 min
- SepRLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewardsalignment2509.21319NVIDIASep 25, 2025score 9~105 min
- SepMultiplayer Nash Preference Optimizationalignment2509.23102Sep 27, 2025~110 min
- SepInverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?evaluation2509.04292ByteDance SeedSep 4, 2025score 6~97 min
- SepCARE: Cognitive-reasoning Augmented Reinforcement for Emotional Support Conversationtraining-methods2510.05122Qwen DianJinSep 30, 2025score 1~100 min
- AugDuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimizationtraining-methods2508.14460ByteDance SeedAug 20, 2025score 9~112 min
- AugHermes 4 Technical Reporttraining-methods2508.18255Aug 25, 2025~110 min
- JulA Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligenceagents2507.21046Jul 28, 2025score 10~120 min
- JunWhen Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Frameworkalignment2506.16411Together AIJun 19, 2025~132 min
- MayBPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenizationalignment2505.24689May 30, 2025~124 min
- AprReinforcement Learning from Human Feedbackrl-training2504.12501Apr 16, 2025~84 min
- MarOn the Acquisition of Shared Grammatical Representations in Bilingual Language Modelsalignment2503.03962Mar 5, 2025score 4~111 min
- JanInternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Modelalignment2501.12368InternLM / Shanghai AI LabJan 21, 2025score 9~95 min
- JanMONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hackingalignment2501.13011DeepMindJan 22, 2025~107 min
- JanSparse Autoencoders Trained on the Same Data Learn Different Featuresalignment2501.16615Jan 28, 2025~112 min
- JanPartially Rewriting a Transformer in Natural Languagealignment2501.18838Jan 31, 2025~101 min
- JanKimi k1.5: Scaling Reinforcement Learning with LLMsreasoning2501.12599Jan 22, 2025~107 min
- JanREINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalizationrl-training2501.03262Jan 4, 2025score 9~107 min
- JanOpen Problems in Mechanistic Interpretabilitysafety2501.16496Jan 27, 2025score 6~110 min
- JanThe Lessons of Developing Process Reward Models in Mathematical Reasoningtraining-methods2501.07301Qwen / Alibaba CloudJan 13, 2025score 10~115 min
2024
49- DecOpenAI o1 System Cardsafety2412.16720OpenAIDec 21, 2024score 9~170 min
- DecLearnLM: Improving Gemini for Learningtraining-methods2412.16429Dec 21, 2024score 7~111 min
- NovSample-Efficient Alignment for LLMsalignment2411.01493Nov 3, 2024score 8~124 min
- NovBenchmarking Distributional Alignment of Large Language Modelsevaluation2411.05403Nov 8, 2024~118 min
- NovTülu 3: Pushing Frontiers in Open Language Model Post-Trainingtraining-methods2411.15124Allen Institute for AINov 22, 2024~132 min
- OctTLDR: Token-Level Detective Reward Model for Large Vision Language Modelsalignment2410.04734Meta LlamaOct 7, 2024score 8~119 min
- OctSelf-Boosting Large Language Models with Synthetic Preference Dataalignment2410.06961Oct 9, 2024score 9~118 min
- OctControllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirementsalignment2410.08968Oct 11, 2024score 8~113 min
- OctMA-RLHF: Reinforcement Learning from Human Feedback with Macro Actionsrl-training2410.02743BAIDUOct 3, 2024score 8~117 min
- OctGPT-4o System Cardsafety2410.21276OpenAIOct 25, 2024score 9~137 min
- OctMerge to Learn: Efficiently Adding Skills to Language Models with Model Mergingtraining-methods2410.12937Oct 16, 2024~116 min
- OctUFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Functiontraining-methods2410.21438Oct 28, 2024~111 min
- SepTowards a Unified View of Preference Learning for Large Language Models: A Surveyalignment2409.02795Sep 4, 2024score 10~110 min
- SepLanguage Models Learn to Mislead Humans via RLHFalignment2409.12822Sep 19, 2024score 8~98 min
- SepInstruction Following without Instruction Tuningalignment2409.14254Sep 21, 2024score 9~115 min
- SepIs Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysisalignment2409.20059Sep 30, 2024score 8~107 min
- SepLaw of the Weakest Link: Cross Capabilities of Large Language Modelsevaluation2409.19951Sep 30, 2024score 9~114 min
- AugHermes 3 Technical Reportalignment2408.11857Aug 15, 2024~105 min
- AugUnleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Modelsdata2408.02085Aug 4, 2024score 10~182 min
- JulComposable Interventions for Language Modelsalignment2407.06483Jul 9, 2024~101 min
- JulQwen2-Audio Technical Reportmultimodal2407.10759Qwen / Alibaba CloudJul 15, 2024score 6~104 min
- JulOn scalable oversight with weak LLMs judging strong LLMssafety2407.04622Jul 5, 2024~137 min
- JunAligning Language Models with Demonstrated Feedbackalignment2406.00888Jun 2, 2024score 8~102 min
- JunBootstrapping Language Models with DPO Implicit Rewardsalignment2406.09760Jun 14, 2024score 9~116 min
- JunChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Toolsllm-systems2406.12793Zhipu / GLMJun 18, 2024score 9~108 min
- JunWPO: Enhancing RLHF with Weighted Preference Optimizationrl-training2406.11827Zoom AIJun 17, 2024score 9~112 min
- MaySelf-Play Preference Optimization for Language Model Alignmentalignment2405.00675May 1, 2024score 9~96 min
- MayUnderstanding the performance gap between online and offline alignment algorithmsalignment2405.08448May 14, 2024score 9~107 min
- MayOpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Frameworkalignment2405.11143May 20, 2024score 9~99 min
- MayOffline Regularised Reinforcement Learning for Large Language Models Alignmentrl-training2405.19107May 29, 2024score 8~97 min
- MayRLHF Workflow: From Reward Modeling to Online RLHFtraining-methods2405.07863SalesforceMay 13, 2024score 10~96 min
- AprPhi-3 Technical Report: A Highly Capable Language Model Locally on Your Phonepretraining2404.14219Microsoft ResearchApr 22, 2024score 9~114 min
- AprReFT: Representation Finetuning for Language Modelstraining-methods2404.03592Apr 4, 2024~102 min
- AprLearn Your Reference Model for Real Good Alignmenttraining-methods2404.09656Apr 15, 2024score 9~83 min
- MarLong-form factuality in large language modelsalignment2403.18802DeepMindMar 27, 2024score 8~110 min
- MarChatbot Arena: An Open Platform for Evaluating LLMs by Human Preferenceevaluation2403.04132Mar 7, 2024~118 min
- MarRewardBench: Evaluating Reward Models for Language Modelingevaluation2403.13787InternLM / Shanghai AI LabMar 20, 2024~124 min
- MarInternLM2 Technical Reporttraining-methods2403.17297InternLM / Shanghai AI LabMar 26, 2024score 9~92 min
- MarAtP*: An efficient and scalable method for localizing LLM behaviour to componentsuncategorized2403.00745DeepMind0 citesMar 1, 2024score 3~123 min
- FebKTO: Model Alignment as Prospect Theoretic Optimizationalignment2402.01306OpenBMBFeb 2, 2024~117 min
- FebODIN: Disentangled Reward Mitigates Hacking in RLHFalignment2402.07319Feb 11, 2024score 9~112 min
- FebSuppressing Pink Elephants with Direct Principle Feedbackalignment2402.07896EleutherAIFeb 12, 2024score 9~102 min
- FebExperts Don't Cheat: Learning What You Don't Know By Predicting Pairsalignment2402.08733DeepMindFeb 13, 2024~109 min
- FebDirect Language Model Alignment from Online AI Feedbackrl-training2402.04792Feb 7, 2024score 9~100 min
- JanSelf-Play Fine-Tuning Converts Weak Language Models to Strong Language Modelsalignment2401.01335Jan 2, 2024score 9~128 min
- JanSecrets of RLHF in Large Language Models Part II: Reward Modelingalignment2401.06080Jan 11, 2024score 9~130 min
- JanSelf-Rewarding Language Modelsalignment2401.10020Jan 18, 2024score 9~116 min
- JanTuning Language Models by Proxytraining-methods2401.08565Jan 16, 2024~104 min
- JanWARM: On the Benefits of Weight Averaged Reward Modelstraining-methods2401.12187Jan 22, 2024score 9~112 min
2023
40- DecNash Learning from Human Feedbackalignment2312.00886Dec 1, 2023score 9~110 min
- DecThe Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learningalignment2312.01552Dec 4, 2023score 9~116 min
- DecAlignment for Honestyalignment2312.07000Dec 12, 2023score 9~124 min
- DecHelping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hackingalignment2312.09244Dec 14, 2023score 9~116 min
- DecWeak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervisionalignment2312.09390Dec 14, 2023score 9~115 min
- DecChallenges with unsupervised LLM knowledge discoveryalignment2312.10029DeepMindDec 15, 2023score 5~131 min
- NovReport of the 1st Workshop on Generative AI and Lawalignment2311.06477DeepMindNov 11, 2023~147 min
- NovLevels of AGI for Operationalizing Progress on the Path to AGIevaluation2311.02462DeepMindNov 4, 2023score 5~122 min
- NovFine-tuning Language Models for Factualitytraining-methods2311.08401Nov 14, 2023score 9~104 min
- NovCamels in a Changing Climate: Enhancing LM Adaptation with Tulu 2training-methods2311.10702Allen Institute for AINov 17, 2023score 10~106 min
- OctA General Theoretical Paradigm to Understand Learning from Human Preferencesalignment2310.12036DeepMindOct 18, 2023~103 min
- OctSafe RLHF: Safe Reinforcement Learning from Human Feedbackalignment2310.12773Oct 19, 2023score 9~114 min
- OctControlled Decoding from Language Modelsalignment2310.17022Oct 25, 2023score 9~110 min
- OctLarge Language Models Cannot Self-Correct Reasoning Yetreasoning2310.01798DeepMindOct 3, 2023score 8~111 min
- OctContrastive Prefence Learning: Learning from Human Feedback without RLrl-training2310.13639Oct 20, 2023score 9~109 min
- SepRLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackalignment2309.00267Sep 1, 2023score 9~119 min
- SepRead Morealignment2309.08600EleutherAISep 15, 2023score 2~103 min
- SepStabilizing RLHF through Advantage Model and Selective Rehearsalalignment2309.10202Sep 18, 2023score 9~119 min
- SepQwen Technical Reportalignment2309.16609Stability AISep 28, 2023score 9~119 min
- SepScaling Catalog Attribute Extraction with Multi-modal LLMsprompting2309.03409InstacartSep 7, 2023~116 min
- SepStatistical Rejection Sampling Improves Preference Optimizationrl-training2309.06657Sep 13, 2023score 9~114 min
- SepEfficient RLHF: Reducing the Memory Usage of PPOtraining-methods2309.00754Sep 1, 2023score 9~102 min
- AugSIMPLE SYNTHETIC DATA REDUCES SYCOPHANCY IN LARGE LANGUAGE MODELSalignment2308.03958Aug 7, 2023~107 min
- AugSelf-Alignment with Instruction Backtranslationalignment2308.06259Aug 11, 2023score 9~109 min
- AugDeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scalesllm-systems2308.01320Aug 2, 2023score 9~96 min
- JulLlama 2: Open Foundation and Fine-Tuned Chat Modelsalignment2307.09288NVIDIAJul 18, 2023~120 min
- JulOpen Problems and Fundamental Limitations of Reinforcement Learning from Human Feedbackalignment2307.15217Jul 27, 2023score 9~121 min
- JulSecrets of RLHF in Large Language Models: Part I: PPOrl-training2307.04964Jul 11, 2023~131 min
- JulRLCD: Reinforcement Learning from Contrast Distillation for Language Model Alignmentrl-training2307.12950Jul 24, 2023score 9~85 min
- JunRead Morealignment2306.03819EleutherAIJun 6, 2023score 2~103 min
- JunPreference Ranking Optimization for Human Alignmentalignment2306.17492Jun 30, 2023score 9~118 min
- JunMistral 7Bevaluation2306.05685MistralJun 9, 2023~109 min
- MaySLiC-HF: Sequence Likelihood Calibration with Human Feedbackalignment2305.10425May 17, 2023score 9~108 min
- MayLIMA: Less Is More for Alignmentalignment2305.11206May 18, 2023~96 min
- MayLet's Verify Step by Stepreasoning2305.20050May 31, 2023~124 min
- MayRead Moresafety2305.16367EleutherAIMay 25, 2023score 2~123 min
- MayDirect Preference Optimization: Your Language Model is Secretly a Reward Modeltraining-methods2305.18290Stability AIMay 29, 2023~122 min
- AprImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generationalignment2304.05977Zhipu / GLMApr 12, 2023~113 min
- MarReclaiming the Digital Commons: A Public Data Trust for Training Dataalignment2303.09001EleutherAIMar 16, 2023~115 min
- MarGPT-4 Technical Reportpretraining2303.08774Mar 15, 2023~118 min
2022
5- DecConstitutional AI: Harmlessness from AI Feedbackalignment2212.08073Dec 15, 2022~123 min
- DecBLOOM+1: Adding Language Support to BLOOM for Zero-Shot Promptingalignment2212.09535EleutherAIDec 19, 2022~121 min
- OctEleutherAI: Going Beyond "Open Science" to "Science in the Open"alignment2210.06413EleutherAIOct 12, 2022~97 min
- AprTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedbackalignment2204.05862Apr 12, 2022~132 min
- MarTraining language models to follow instructions with human feedbackalignment2203.02155Berkeley BAIRMar 4, 2022~125 min