107 papers
Safety
0/106Jailbreak resistance, alignment, and red-teaming.
Progress0 of 106
RelatedAlignment
2026
17- MayIt Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMsalignment2605.20258KAIST AIMay 18, 2026~122 min
- AprProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluationevaluation2604.23099DeepMindApr 25, 2026~137 min
- AprRethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capabilityreasoning2604.06628AI45ResearchApr 8, 2026score 9~124 min
- AprThe Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillationtraining-methods2604.16830Salesforce AI ResearchApr 18, 2026~115 min
- MarLLM2Vec-Gen: Generative Embeddings from Large Language Modelsagents2603.10913McGill NLP GroupMar 11, 2026~94 min
- MarT-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Searchagents2603.22341KAIST AIMar 21, 2026score 9~112 min
- MarLearning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Usealignment2603.03205Microsoft ResearchMar 3, 2026score 8~111 min
- MarReasoning Models Struggle to Control their Chains of Thoughtalignment2603.05706OpenAIMar 5, 2026score 9~99 min
- MarHow Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularitiesevaluation2603.02578alibaba-incMar 3, 2026score 6~107 min
- MarBenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMssafety2603.16557LG AI ResearchMar 17, 2026score 8~101 min
- MarSNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detectionsafety2603.20686KAIST AIMar 21, 2026score 3~120 min
- FebWhy Steering Works: Toward a Unified View of Language Model Parameter Dynamicsalignment2602.02343alibabaFeb 2, 2026score 4~107 min
- FebPrivasis: Synthesizing the Largest "Public" Private Dataset from Scratchdata2602.03183NVIDIAFeb 3, 2026score 2~109 min
- FebThe Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societiessafety2602.09877AnthropicFeb 10, 2026score 3~121 min
- JanA Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5safety2601.10527Jan 15, 2026score 4~118 min
- JanBuilding Production-Ready Probes For Geminisafety2601.11516Jan 16, 2026score 3~145 min
- JanShaping capabilities with token-level data filteringsafety2601.21571AnthropicJan 29, 2026score 3~110 min
2025
32- DecAligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Modelsalignment2512.04981KAIST AIDec 4, 2025score 4~119 min
- DecThe FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factualityevaluation2512.10791DeepmindDec 11, 2025score 9~127 min
- DecTaxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Modelssafety2512.05339RobloxDec 5, 2025score 5~96 min
- DecEvaluating Gemini Robotics Policies in a Veo World Simulatorvision2512.10675DeepmindDec 11, 2025score 4~115 min
- OctMagentic Marketplace: An Open-Source Environment for Studying Agentic Marketsagents2510.25779Microsoft ResearchOct 27, 2025score 8~114 min
- OctLarge Reasoning Models Learn Better Alignment from Flawed Thinkingsafety2510.00938Oct 1, 2025~96 min
- OctQwen3Guard Technical Reportsafety2510.14276Qwen / Alibaba CloudOct 16, 2025score 8~136 min
- OctDistractor Injection Attacks on Large Reasoning Models: Characterization and Defensesafety2510.16259Amazon ScienceOct 17, 2025score 6~101 min
- OctAny-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depthsafety2510.18081ByteDance SeedOct 20, 2025score 8~133 min
- OctA Pragmatic View of AI Personhoodsafety2510.26396DeepMindOct 30, 2025~115 min
- SepWhy Language Models Hallucinatealignment2509.04664Sep 4, 2025~118 min
- SepReviewScore: Misinformed Peer Review Detection with Large Language Modelsevaluation2509.21679KAIST AISep 25, 2025score 3~108 min
- SepThe Pitfalls of KV Cache Compressioninference-optimization2510.00231Sep 30, 2025~125 min
- SepApertus: Democratizing Open and Compliant LLMs for Global Language Environmentspretraining2509.14233Swiss AI InitiativeSep 17, 2025score 9~107 min
- SepTruthRL: Incentivizing Truthful LLMs via Reinforcement Learningtraining-methods2509.25760Sep 30, 2025~100 min
- AugDeep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMspretraining2508.06601Aug 8, 2025score 6~110 min
- JulThe Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigationsafety2507.05578Jul 8, 2025score 9~120 min
- JulDoes More Inference-Time Compute Really Help Robustness?safety2507.15974DeepSeekJul 21, 2025~111 min
- MayWhen AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Researchevaluation2505.11855May 17, 2025score 8~108 min
- MayLessons from Defending Gemini Against Indirect Prompt Injectionssafety2505.14534DeepMindMay 20, 2025score 5~116 min
- AprThe AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Searchagents2504.08066Apr 10, 2025score 9~93 min
- AprThe Leaderboard Illusionevaluation2504.20879Apr 29, 2025~124 min
- AprDeepSeek-R1 Thoughtology: Let's think about LLM Reasoningreasoning2504.07128DeepSeekApr 2, 2025score 10~120 min
- AprMechanistic Anomaly Detection for "Quirky" Language Modelssafety2504.08812Apr 9, 2025~121 min
- MarGemini Robotics: Bringing AI into the Physical Worldmultimodal2503.20020Mar 25, 2025score 9~134 min
- FebSlowing Learning by Erasing Simple Featuressafety2502.02820Feb 5, 2025~131 min
- FebMaximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flipssafety2502.07408NVIDIAFeb 11, 2025~127 min
- JanTowards Best Practices for Open Datasets for LLM Trainingdata2501.08365Jan 14, 2025~104 min
- JanTrading Inference-Time Compute for Adversarial Robustnessreasoning2501.18841OpenAIJan 31, 2025score 10~141 min
- JanOpen Problems in Mechanistic Interpretabilitysafety2501.16496Jan 27, 2025score 6~110 min
- JanEarly External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluationsafety2501.17749OpenAIJan 29, 2025score 6~108 min
- Jano3-mini vs DeepSeek-R1: Which One is Safer?safety2501.18438OpenAIJan 30, 2025score 3~96 min
2024
24- DecMachine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Researchsafety2412.06966DeepMindDec 9, 2024~132 min
- DecOpenAI o1 System Cardsafety2412.16720OpenAIDec 21, 2024score 9~170 min
- NovSteering Language Model Refusal with Sparse Autoencoderssafety2411.11296Nov 18, 2024~107 min
- OctControllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirementsalignment2410.08968Oct 11, 2024score 8~113 min
- OctUndesirable Memorization in Large Language Models: A Surveysafety2410.02650Oct 3, 2024~133 min
- OctGPT-4o System Cardsafety2410.21276OpenAIOct 25, 2024score 9~137 min
- OctMerge to Learn: Efficiently Adding Skills to Language Models with Model Mergingtraining-methods2410.12937Oct 16, 2024~116 min
- SepLanguage Models Learn to Mislead Humans via RLHFalignment2409.12822Sep 19, 2024score 8~98 min
- AugRecent Surge in Public Interest in Transportation: Sentiment Analysis of Baidu Apollo Go Using Weibo Datacontext-optimization2408.10088Aug 19, 2024score 1~117 min
- JulSpectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scalepretraining2407.12327Jul 17, 2024score 9~98 min
- JulPhi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cyclesafety2407.13833Microsoft ResearchJul 18, 2024score 9~110 min
- JulThe Llama 3 Herd of Modelstraining-methods2407.21783Jul 31, 2024~129 min
- AprCapabilities of Gemini Models in Medicinemultimodal2404.18416Apr 29, 2024score 9~125 min
- AprPhi-3 Technical Report: A Highly Capable Language Model Locally on Your Phonepretraining2404.14219Microsoft ResearchApr 22, 2024score 9~114 min
- AprDoes Transformer Interpretability Transfer to RNNs?safety2404.05971Apr 9, 2024~107 min
- MarLong-form factuality in large language modelsalignment2403.18802DeepMindMar 27, 2024score 8~110 min
- MarGemma: Open Models Based on Gemini Research and Technologyllm-systems2403.08295Google ResearchMar 13, 2024score 9~105 min
- MarEvaluating Frontier Models for Dangerous Capabilitiessafety2403.13793DeepMindMar 20, 2024score 8~140 min
- MarFew-Shot Recalibration of Language Modelssafety2403.18286DeepMindMar 27, 2024~111 min
- MarRecourse for reclamation: Chatting with generative language modelsno summary yetcs-hc2403.144671 citesMar 21, 2024score 2
- FebSuppressing Pink Elephants with Direct Principle Feedbackalignment2402.07896EleutherAIFeb 12, 2024score 9~102 min
- FebDo Membership Inference Attacks Work on Large Language Models?safety2402.07841Feb 12, 2024~125 min
- FebSimulacra as Conscious Exoticasafety2402.12422DeepMindFeb 19, 2024~136 min
- JanFrom GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalitiesevaluation2401.15071Jan 26, 2024score 8~96 min
2023
21- DecAlignment for Honestyalignment2312.07000Dec 12, 2023score 9~124 min
- DecHelping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hackingalignment2312.09244Dec 14, 2023score 9~116 min
- NovLevels of AGI for Operationalizing Progress on the Path to AGIevaluation2311.02462DeepMindNov 4, 2023score 5~122 min
- NovSystem 2 Attention (is something you might need too)prompting2311.11829Nov 20, 2023score 9~121 min
- NovScalable AI Safety via Doubly-Efficient Debatesafety2311.14125DeepMindNov 23, 2023~117 min
- OctSafe RLHF: Safe Reinforcement Learning from Human Feedbackalignment2310.12773Oct 19, 2023score 9~114 min
- OctRepresentation Engineering: A Top-Down Approach to AI Transparencysafety2310.01405EleutherAIOct 2, 2023~134 min
- OctLinear Representations of Sentiment in Large Language Modelssafety2310.15154EleutherAIOct 23, 2023~131 min
- OctDetecting Pretraining Data from Large Language Modelssafety2310.16789Oct 25, 2023score 9~111 min
- SepM3DSYNTH: A DATASET OF MEDICAL 3D IMAGES WITH AI-GENERATED LOCAL MANIPULATIONSdata2309.07973Sep 14, 2023~120 min
- SepDoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Modelsinference-optimization2309.03883Sep 7, 2023score 9~113 min
- SepLMSYS-CHAT-1M: A LARGE-SCALE REAL-WORLD LLM CONVERSATION DATASETsafety2309.11998Sep 21, 2023~126 min
- AugSIMPLE SYNTHETIC DATA REDUCES SYCOPHANCY IN LARGE LANGUAGE MODELSalignment2308.03958Aug 7, 2023~107 min
- JulLlama 2: Open Foundation and Fine-Tuned Chat Modelsalignment2307.09288NVIDIAJul 18, 2023~120 min
- JulOpen Problems and Fundamental Limitations of Reinforcement Learning from Human Feedbackalignment2307.15217Jul 27, 2023score 9~121 min
- JulA Survey on Evaluation of Large Language Modelsevaluation2307.03109Jul 6, 2023~108 min
- JulHow is ChatGPT's behavior changing over time?evaluation2307.09009Jul 18, 2023score 9~107 min
- JunRead Morealignment2306.03819EleutherAIJun 6, 2023score 2~103 min
- MayPaLM 2 Technical Reportpretraining2305.10403May 17, 2023~108 min
- MayRead Moresafety2305.16367EleutherAIMay 25, 2023score 2~123 min
- MayHow Language Model Hallucinations Can Snowballuncategorized2305.13534May 22, 2023score 9~113 min
2022
3- JunBIG-bench: Beyond the Imitation Game Benchmarkscaling-laws2206.04615MistralJun 9, 2022~99 min
- AprTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedbackalignment2204.05862Apr 12, 2022~132 min
- FebRed Teaming Language Models with Language Modelsalignment2202.03286Feb 7, 2022~119 min
2021
4- DecGopherscaling-laws2112.11446Dec 8, 2021~142 min
- SepTruthfulQA: Measuring How Models Mimic Human Falsehoodssafety2109.07958MistralSep 8, 2021~117 min
- JulEvaluating Large Language Models Trained on Codecode2107.03374MistralJul 7, 2021~126 min
- AprTowards Measuring Fairness in AI: the Casual Conversations Datasetsafety2104.02821NVIDIAApr 6, 2021~111 min