RLHF / preference learning
105 papers in this thread, across 25 domains.
Progress0 of 69
- 2026Offline Preference-Based Trajectory Evaluationno summary yetcs lg2606.17541CMU0 citesJun 16, 2026cs lg
- 2026Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimizationno summary yetcs ai2606.26899Princeton0 citesJun 25, 2026cs ai
- 2026Beyond expert users: agents should help users construct preferences, not just elicit themno summary yetcs ai2606.30863Stanford0 citesJun 29, 2026cs ai
- 2026How Human Feedback Shapes AI-generated Community Notesno summary yetcs cy2606.30905UW0 citesJun 29, 2026cs cy
- 2026Freeform Preference Learning for Robotic Manipulationno summary yetcs ro2606.32027Stanford0 citesJun 30, 2026cs ro
- 2026Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable AlignmentAlignment2605.20834May 20, 2026~103 minAlignment
- 2026DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and VerificationTraining Methods2605.09269Tencent HunyuanMay 10, 2026~113 minTraining Methods
- 2026Beyond Length Scaling: Synergizing Breadth and Depth for Generative Reward ModelsReasoning2603.01571Tencent HunyuanMar 2, 2026score 9~108 minReasoning
- 2026BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMsSafety2603.16557LG AI ResearchMar 17, 2026score 8~101 minSafety
- 2026Visual-ERM: Reward Modeling for Visual EquivalenceVision2603.13224InternLM / Shanghai AI LabMar 13, 2026score 9~127 minVision
- 2026Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward ModelsAlignment2602.04649QwenFeb 4, 2026score 9~109 minAlignment
- 2026P-GenRM: Personalized Generative Reward Model with Test-time User-based ScalingAlignment2602.12116Tongyi-ConvAIFeb 12, 2026score 9~115 minAlignment
- 2026GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL OptimizationRL Training2601.05242Jan 8, 2026~109 minRL Training
- 2025ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual ReasoningRL Training2512.05111Intern Large ModelsDec 4, 2025score 8~102 minRL Training
- 2025Video Generation Models Are Good Latent Reward ModelsAlignment2511.21541Tencent HunyuanNov 26, 2025score 8~104 minAlignment
- 2025DocReward: A Document Reward Model for Structuring and StylizingData2510.11391Microsoft ResearchOct 13, 2025score 6~95 minData
- 2025Vibe Checker: Aligning Code Evaluation with Human PreferenceEvaluation2510.07315Oct 8, 2025~114 minEvaluation
- 2025Beyond Correctness: Evaluating Subjective Writing Preferences Across CulturesEvaluation2510.14616ByteDance SeedOct 16, 2025score 8~104 minEvaluation
- 2025Fine-Grained GRPO for Precise Preference Alignment in Flow ModelsTraining Methods2510.01982IXCLab@Shanghai AI LabOct 2, 2025score 8~101 minTraining Methods
- 2025RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable RewardsAlignment2509.21319NVIDIASep 25, 2025score 9~105 minAlignment
- 2025Multiplayer Nash Preference OptimizationAlignment2509.23102Sep 27, 2025~110 minAlignment
- 2025Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language ModelsAgents2508.15202Qwen DianJinAug 21, 2025score 9~111 minAgents
- 2025DuPO: Enabling Reliable LLM Self-Verification via Dual Preference OptimizationTraining Methods2508.14460ByteDance SeedAug 20, 2025score 9~112 minTraining Methods
- 2025Reinforcement Learning from Human FeedbackRL Training2504.12501Apr 16, 2025~84 minRL Training
- 2025Inference-Time Scaling for Generalist Reward ModelingScaling Laws2504.02495Apr 3, 2025~101 minScaling Laws
- 2025InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward ModelAlignment2501.12368InternLM / Shanghai AI LabJan 21, 2025score 9~95 minAlignment
- 2025The Lessons of Developing Process Reward Models in Mathematical ReasoningTraining Methods2501.07301Qwen / Alibaba CloudJan 13, 2025score 10~115 minTraining Methods
- 2024TLDR: Token-Level Detective Reward Model for Large Vision Language ModelsAlignment2410.04734Meta LlamaOct 7, 2024score 8~119 minAlignment
- 2024Self-Boosting Large Language Models with Synthetic Preference DataAlignment2410.06961Oct 9, 2024score 9~118 minAlignment
- 2024MA-RLHF: Reinforcement Learning from Human Feedback with Macro ActionsRL Training2410.02743BAIDUOct 3, 2024score 8~117 minRL Training
- 2024Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language ModelsRL Training2410.18252Oct 23, 2024score 9~117 minRL Training
- 2024Towards a Unified View of Preference Learning for Large Language Models: A SurveyAlignment2409.02795Sep 4, 2024score 10~110 minAlignment
- 2024Language Models Learn to Mislead Humans via RLHFAlignment2409.12822Sep 19, 2024score 8~98 minAlignment
- 2024Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical AnalysisAlignment2409.20059Sep 30, 2024score 8~107 minAlignment
- 2024Minimizing Live Experiments in Recommender Systems: User Simulation to Evaluate Preference Elicitation Policiesno summary yetcs ir2409.17436Google Research5 citesSep 26, 2024cs ir
- 2024Bootstrapping Language Models with DPO Implicit RewardsAlignment2406.09760Jun 14, 2024score 9~116 minAlignment
- 2024From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder PipelineEvaluation2406.11939Jun 17, 2024score 9~105 minEvaluation
- 2024WPO: Enhancing RLHF with Weighted Preference OptimizationRL Training2406.11827Zoom AIJun 17, 2024score 9~112 minRL Training
- 2024Self-Play Preference Optimization for Language Model AlignmentAlignment2405.00675May 1, 2024score 9~96 minAlignment
- 2024OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF FrameworkAlignment2405.11143May 20, 2024score 9~99 minAlignment
- 2024RLHF Workflow: From Reward Modeling to Online RLHFTraining Methods2405.07863SalesforceMay 13, 2024score 10~96 minTraining Methods
- 2024Advancing LLM Reasoning Generalists with Preference TreesReasoning2404.02078OpenBMBApr 2, 2024score 9~115 minReasoning
- 2024Iterative Reasoning Preference OptimizationReasoning2404.19733Apr 30, 2024score 9~101 minReasoning
- 2024Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceEvaluation2403.04132Mar 7, 2024~118 minEvaluation
- 2024RewardBench: Evaluating Reward Models for Language ModelingEvaluation2403.13787InternLM / Shanghai AI LabMar 20, 2024~124 minEvaluation
- 2024Parameter Efficient Reinforcement Learning from Human FeedbackTraining Methods2403.10704Mar 15, 2024score 8~106 minTraining Methods
- 2024ODIN: Disentangled Reward Mitigates Hacking in RLHFAlignment2402.07319Feb 11, 2024score 9~112 minAlignment
- 2024Learning to Learn Faster from Human Feedback with Language Model Predictive ControlTraining Methods2402.11450DeepMindFeb 18, 2024score 9~104 minTraining Methods
- 2024Secrets of RLHF in Large Language Models Part II: Reward ModelingAlignment2401.06080Jan 11, 2024score 9~130 minAlignment
- 2024WARM: On the Benefits of Weight Averaged Reward ModelsTraining Methods2401.12187Jan 22, 2024score 9~112 minTraining Methods
- 2024Two-pass Endpoint Detection for Speech Recognitionno summary yeteess as2401.08916Amazon0 citesJan 17, 2024eess as
- 2023Nash Learning from Human FeedbackAlignment2312.00886Dec 1, 2023score 9~110 minAlignment
- 2023Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward HackingAlignment2312.09244Dec 14, 2023score 9~116 minAlignment
- 2023Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2Training Methods2311.10702Allen Institute for AINov 17, 2023score 10~106 minTraining Methods
- 2023A General Theoretical Paradigm to Understand Learning from Human PreferencesAlignment2310.12036DeepMindOct 18, 2023~103 minAlignment
- 2023Safe RLHF: Safe Reinforcement Learning from Human FeedbackAlignment2310.12773Oct 19, 2023score 9~114 minAlignment
- 2023Contrastive Prefence Learning: Learning from Human Feedback without RLRL Training2310.13639Oct 20, 2023score 9~109 minRL Training
- 2023RLAIF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackAlignment2309.00267Sep 1, 2023score 9~119 minAlignment
- 2023Stabilizing RLHF through Advantage Model and Selective RehearsalAlignment2309.10202Sep 18, 2023score 9~119 minAlignment
- 2023Statistical Rejection Sampling Improves Preference OptimizationRL Training2309.06657Sep 13, 2023score 9~114 minRL Training
- 2023Efficient RLHF: Reducing the Memory Usage of PPOTraining Methods2309.00754Sep 1, 2023score 9~102 minTraining Methods
- 2023DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All ScalesLLM Systems2308.01320Aug 2, 2023score 9~96 minLLM Systems
- 2023printf: Preference Modeling Based on User Reviews with Item Images and Textual Information via Graph Learningno summary yetcs ir2308.09943Amazon1 citesAug 19, 2023cs ir
- 2023Open Problems and Fundamental Limitations of Reinforcement Learning from Human FeedbackAlignment2307.15217Jul 27, 2023score 9~121 minAlignment
- 2023Secrets of RLHF in Large Language Models: Part I: PPORL Training2307.04964Jul 11, 2023~131 minRL Training
- 2023Preference Ranking Optimization for Human AlignmentAlignment2306.17492Jun 30, 2023score 9~118 minAlignment
- 2023Rank-heterogeneous Preference Models for School Choiceno summary yetstat ap2306.01801Amazon2 citesJun 1, 2023stat ap
- 2023SLiC-HF: Sequence Likelihood Calibration with Human FeedbackAlignment2305.10425May 17, 2023score 9~108 minAlignment
- 2023Direct Preference Optimization: Your Language Model is Secretly a Reward ModelTraining Methods2305.18290Stability AIMay 29, 2023~122 minTraining Methods
- 2023ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationAlignment2304.05977Zhipu / GLMApr 12, 2023~113 minAlignment
- 2023Robust Preference-Guided Denoising for Graph based Social Recommendationno summary yetcs ir2303.08346Tencent70 citesMar 15, 2023cs ir
- 2022Robust Preference Learning for Storytelling via Contrastive Reinforcement LearningTraining Methods2210.07792EleutherAIOct 14, 2022~121 minTraining Methods
- 2022RESUS: Warm-Up Cold Users via Meta-Learning Residual User Preferences in CTR Predictionno summary yetcs ir2210.16080Tencent8 citesOct 28, 2022cs ir
- 2022Training a Helpful and Harmless Assistant with Reinforcement Learning from Human FeedbackAlignment2204.05862Apr 12, 2022~132 minAlignment
- 2022Human Preferences as Dueling Banditsno summary yetcs ir2204.10362Microsoft Research7 citesApr 21, 2022cs ir
- 2022Training language models to follow instructions with human feedbackAlignment2203.02155Berkeley BAIRMar 4, 2022~125 minAlignment
- 2022Dual Preference Distribution Learning for Item Recommendationno summary yetcs ir2201.09490Alibaba4 citesJan 24, 2022cs ir
- 2021WebGPT: Browser-Assisted Question-Answering with Human FeedbackAgents2112.09332Dec 17, 2021~112 minAgents
- 2021Modelling of Bi-directional Spatio-Temporal Dependence and Users' Dynamic Preferences for Missing POI Check-in Identificationno summary yetcs lg2112.15285Baidu13 citesDec 31, 2021cs lg
- 2021Recursively Summarizing Books with Human FeedbackTraining Methods2109.1086266 citesSep 22, 2021~99 minTraining Methods
- 2021PR-Net: Preference Reasoning for Personalized Video Highlight Detectionno summary yetcs cv2109.01799Tencent0 citesSep 4, 2021cs cv
- 2021Understanding WeChat User Preferences and "Wow" Diffusionno summary yetcs si2103.02930Microsoft Research27 citesMar 4, 2021cs si
- 2020Fairness Preferences, Actual and Hypothetical: A Study of Crowdworker Incentivesno summary yetcs ai2012.04216Google Research0 citesDec 8, 2020cs ai
- 2020Offline Reinforcement Learning from Human Feedback in Real-World Sequence-to-Sequence Tasksno summary yetcs cl2011.02511Google Research0 citesNov 4, 2020cs cl
- 2020Incorporating Stylistic Lexical Preferences in Generative Language Modelsno summary yetcs cl2010.11553Adobe0 citesOct 22, 2020cs cl
- 2020Learning to summarize from human feedbackRL Training2009.01325Sep 2, 2020~94 minRL Training
- 2020Learning Personalized Risk Preferences for Recommendationno summary yetcs ir2007.02478Alibaba19 citesJul 6, 2020cs ir
- 2020Preferences Single-Peaked on a Tree: Multiwinner Elections and Structural Resultsno summary yetcs gt2007.06549Google Research6 citesJul 13, 2020cs gt
- 2020Learning-to-Rank with Partitioned Preference: Fast Estimation for the Plackett-Luce Modelno summary yetcs lg2006.05067Google Research4 citesJun 9, 2020cs lg
- 2020DyHGCN: A Dynamic Heterogeneous Graph Convolutional Network to Learn Users' Dynamic Preferences for Information Diffusion Predictionno summary yetcs si2006.05169Alibaba6 citesJun 9, 2020cs si
- 2020Eliciting User Preferences for Personalized Explanations for Video Summariesno summary yetcs hc2005.00465Google Research2 citesMay 1, 2020cs hc
- 2020Social diversity and social preferences in mixed-motive reinforcement learningno summary yetcs ma2002.02325DeepMind14 citesFeb 6, 2020cs ma
- 2020RL agents Implicitly Learning Human Preferencesno summary yetcs ai2002.06137Google Research0 citesFeb 14, 2020cs ai
- 2019Gradient-based Optimization for Bayesian Preference Elicitationno summary yetcs lg1911.09153Google Research2 citesNov 20, 2019cs lg
- 2019Reinforcing an Image Caption Generator Using Off-Line Human Feedbackno summary yetcs cv1911.09753Google Research3 citesNov 21, 2019cs cv
- 2019Controlling Text Generation with Plug and Play Language ModelsRL Training1909.08593UberSep 18, 2019~120 minRL Training
- 2019Intelligent social bots uncover the link between user preference and diversity of news consumptionno summary yetcs si1907.02703Tencent35 citesJul 5, 2019cs si
- 2018How Many Pairwise Preferences Do We Need to Rank A Graph Consistently?no summary yetcs lg1811.02161Google Research2 citesNov 6, 2018cs lg
- 2018Reward learning from human preferences and demonstrations in Atarino summary yetcs lg1811.06521Baidu39 citesNov 15, 2018cs lg
- 2018Hybrid-MST: A Hybrid Active Sampling Strategy for Pairwise Preference Aggregationno summary yetcs lg1810.08851Tencent22 citesOct 20, 2018cs lg
- 2018Preference-based Online Learning with Dueling Bandits: A Surveyno summary yetcs lg1807.11398Google Research24 citesJul 30, 2018cs lg
- 2018Task Transfer by Preference-Based Cost Learningno summary yetcs lg1805.04686Tencent1 citesMay 12, 2018cs lg
- 2017Reinforcement Learning for Bandit Neural Machine Translation with Simulated Human Feedbackno summary yetcs cl1707.07402Microsoft Research27 citesJul 24, 2017cs cl
- 2017Deep reinforcement learning from human preferencesRL Training1706.03741Jun 12, 2017~104 minRL Training
- 2017Preference-driven Similarity Joinno summary yetcs db1706.04266Google Research1 citesJun 13, 2017cs db