2026
84- MayWildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluationagents2605.10912Intern Large ModelsMay 11, 2026~116 min
- MayKernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernelscode2605.04956Tsinghua UniversityMay 6, 2026~132 min
- MayCovering Human Action Space for Computer Use: Data Synthesis and Benchmarkdata2605.12501MicrosoftMay 12, 2026~108 min
- MayPlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Modelsdata2605.20873Tencent HunyuanMay 20, 2026~120 min
- MaySoohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMsevaluation2605.09063EleutherAIMay 9, 2026~115 min
- MayMEME: Multi-entity & Evolving Memory Evaluationevaluation2605.12477KAIST AIMay 12, 2026~130 min
- MayMSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generationevaluation2605.20183TongyiLabMay 19, 2026~122 min
- MayFrom Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skillsevaluation2605.23899Microsoft ResearchMay 22, 2026~120 min
- MayRethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?retrieval2605.10848May 11, 2026~110 min
- MayMeasuring Maximum Activations in Open Large Language Modelsserving2605.15572BAIDUMay 15, 2026~113 min
- MayDeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verificationtraining-methods2605.09269Tencent HunyuanMay 10, 2026~113 min
- AprCORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discoveryagents2604.01658Massachusetts Institute of TechnologyApr 2, 2026score 9~131 min
- AprSynthetic Computers at Scale for Long-Horizon Productivity Simulationagents2604.28181MicrosoftApr 30, 2026~128 min
- AprBeyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Modelsevaluation2604.02315Salesforce AI ResearchApr 2, 2026score 7~112 min
- AprRoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policiesevaluation2604.09860NVIDIAApr 10, 2026~118 min
- AprDiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domainevaluation2604.10425meituanApr 12, 2026score 3~113 min
- AprOccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Modelsevaluation2604.10866QwenApr 13, 2026score 8~104 min
- AprDo Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understandingevaluation2604.11177VideoDBApr 13, 2026score 8~119 min
- AprLARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignmentevaluation2604.11689LongCatApr 13, 2026~111 min
- AprMind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMsevaluation2604.16054Microsoft ResearchApr 17, 2026~137 min
- AprAJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluationevaluation2604.18240LongCatApr 20, 2026~107 min
- AprProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluationevaluation2604.23099DeepMindApr 25, 2026~137 min
- AprGeneral365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasksreasoning2604.11778LongCatApr 13, 2026score 9~125 min
- AprRAGEN-2: Reasoning Collapse in Agentic RLtraining-methods2604.06268MLL LabApr 7, 2026score 9~118 min
- MarT-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Searchagents2603.22341KAIST AIMar 21, 2026score 9~112 min
- MarVision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verificationagents2603.26648Z.aiMar 27, 2026score 9~132 min
- MarCharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Productionalignment2603.01973Meta LlamaMar 2, 2026score 9~140 min
- MarReasoning Models Struggle to Control their Chains of Thoughtalignment2603.05706OpenAIMar 5, 2026score 9~99 min
- MarRubricBench: Aligning Model-Generated Rubrics with Human Standardsevaluation2603.01562Tencent HunyuanMar 2, 2026score 9~105 min
- MarHow Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularitiesevaluation2603.02578alibaba-incMar 3, 2026score 6~107 min
- MarCan Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streamsevaluation2603.07392KAIST AIMar 8, 2026score 9~117 min
- MarMA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agentsevaluation2603.09827KAIST AIMar 10, 2026score 3~119 min
- MarStrategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collectionsevaluation2603.12180SnowflakeMar 12, 2026score 9~134 min
- MarMMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videosevaluation2603.14145NVIDIAMar 14, 2026score 9~133 min
- MarECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretationevaluation2603.14326KAIST AIMar 15, 2026score 9~120 min
- MarOmni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Modelsevaluation2603.22212alibaba-incMar 23, 2026score 6~118 min
- MarEgo2Web: A Web Agent Benchmark Grounded in Egocentric Videosevaluation2603.22529DeepmindMar 23, 2026score 8~120 min
- MarFinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocolevaluation2603.24943Qwen DianJinMar 26, 2026score 3~111 min
- MarBizGenEval: A Systematic Benchmark for Commercial Visual Content Generationevaluation2603.25732Microsoft ResearchMar 26, 2026score 3~121 min
- MarConsistency Amplifies: How Behavioral Variance Shapes Agent Accuracyevaluation2603.25764SnowflakeMar 26, 2026score 9~98 min
- MarRealChart2Code: Advancing Chart-to-Code Generation with Real Data and Multi-Task Evaluationevaluation2603.25804QwenMar 26, 2026score 4~89 min
- MarViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?evaluation2603.25823meituanMar 26, 2026score 9~108 min
- MarTAPS: Task Aware Proposal Distributions for Speculative Samplinginference-optimization2603.27027Image and Video Understanding LabMar 27, 2026score 9~125 min
- MarQianfan-OCR: A Unified End-to-End Model for Document Intelligencemultimodal2603.13398BAIDUMar 11, 2026score 9~134 min
- MarUnderstanding Reasoning in LLMs through Strategic Information Allocation under Uncertaintyreasoning2603.15500Microsoft ResearchMar 16, 2026score 9~160 min
- MarAI Can Learn Scientific Tasterl-training2603.14473OpenMOSSMar 15, 2026~105 min
- MarBenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMssafety2603.16557LG AI ResearchMar 17, 2026score 8~101 min
- MarVisual-ERM: Reward Modeling for Visual Equivalencevision2603.13224InternLM / Shanghai AI LabMar 13, 2026score 9~127 min
- FebTRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenariosagents2602.01675meituanFeb 2, 2026score 3~122 min
- FebSWE-Universe: Scale Real-World Verifiable Environments to Millionsagents2602.02361QwenFeb 2, 2026score 4~109 min
- FebWebWorld: A Large-Scale World Model for Web Agent Trainingagents2602.14721Qwen / Alibaba CloudFeb 16, 2026score 9~101 min
- FebMobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenariosagents2602.22638alibaba-incFeb 26, 2026score 3~104 min
- FebOutcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Modelsalignment2602.04649QwenFeb 4, 2026score 9~109 min
- FebP-GenRM: Personalized Generative Reward Model with Test-time User-based Scalingalignment2602.12116Tongyi-ConvAIFeb 12, 2026score 9~115 min
- FebBABE: Biology Arena BEnchmarkevaluation2602.05857ByteDance SeedFeb 5, 2026score 9~111 min
- FebHow2Everything: Mining the Web for How-To Procedures to Evaluate and Improve LLMsevaluation2602.08808Ai2Feb 9, 2026score 3~131 min
- FebSPEED-Bench: A Unified and Diverse Benchmark for Speculative Decodinginference-optimization2604.09557NVIDIAFeb 10, 2026score 9~89 min
- FebAdapting Vision-Language Models for E-commerce Understanding at Scalemultimodal2602.11733Feb 12, 2026~133 min
- FebThinking with Drafting: Optical Decompression via Logical Reconstructionreasoning2602.11731ByteDanceFeb 12, 2026score 3~102 min
- FebScaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgmentsscaling-laws2602.23234AppleFeb 26, 2026~104 min
- FebPhyCritic: Multimodal Critic Models for Physical AItraining-methods2602.11124NVIDIAFeb 11, 2026score 6~105 min
- FebTOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Roboticsuncategorized2602.19313Ai2Feb 22, 2026score 3~97 min
- JanThinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalizationagents2601.05432alibaba-incJan 8, 2026score 3~115 min
- JanOver-Searching in Search-Augmented Large Language Modelsagents2601.05503AppleJan 9, 2026score 3~131 min
- JanDeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraintsagents2601.18137QwenJan 26, 2026score 2~101 min
- JanMEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineeringagents2601.22859ernie-researchJan 30, 2026score 5~120 min
- JanAACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Contextcode2601.19494AoneJan 27, 2026score 3~111 min
- JanMotion Attribution for Video Generationdata2601.08828NVIDIAJan 13, 2026score 2~115 min
- JanEvasionBench: A Large-Scale Benchmark for Detecting Managerial Evasion in Earnings Call Q&Aevaluation2601.09142Jan 14, 2026score 2~119 min
- JanEverything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Modelsevaluation2601.20354alibaba-incJan 28, 2026score 2~100 min
- JanRetrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilitiesevaluation2601.21937ByteDance SeedJan 29, 2026score 2~117 min
- JanWorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Modelsevaluation2602.02537Moonshot AIJan 28, 2026score 5~108 min
- JanK-EXAONE Technical Reportmoe2601.01739LG EXAONEJan 5, 2026~112 min
- JanQwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Rankingmultimodal2601.04720Qwen / Alibaba CloudJan 8, 2026~114 min
- JanRethinking Video Generation Model for the Embodied Worldmultimodal2601.15282ByteDance SeedJan 21, 2026score 2~104 min
- JanLost in the Noise: How Reasoning Models Fail with Contextual Distractorsreasoning2601.07226KAIST AIJan 12, 2026score 3~106 min
- JanVisual Generation Unlocks Human-Like Reasoning through Multimodal World Modelsreasoning2601.19834ByteDance SeedJan 27, 2026score 3~111 min
- JanArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Rankingrl-training2601.06487Alibaba DAMOJan 10, 2026score 9~120 min
- JanHow AI Impacts Skill Formationrl-training2601.20245AnthropicJan 28, 2026score 1~107 min
- JanA Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5safety2601.10527Jan 15, 2026score 4~118 min
- JanBuilding Production-Ready Probes For Geminisafety2601.11516Jan 16, 2026score 3~145 min
- JanHow Do Large Language Models Learn Concepts During Continual Pre-Training?training-methods2601.03570Jan 7, 2026score 4~121 min
- JanChaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewardstraining-methods2601.06021Z.aiJan 9, 2026score 9~112 min
- JanQwen3-ASR Technical Reporttraining-methods2601.21337Qwen / Alibaba CloudJan 29, 2026score 2~125 min
2025
108- DecTowards a Science of Scaling Agent Systemsagents2512.08296Dec 9, 2025~118 min
- DecSeed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experienceagents2512.17260ByteDance SeedDec 19, 2025score 8~129 min
- DecEgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editingdiffusion2512.06065SnapDec 5, 2025score 2~105 min
- DecDAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycleevaluation2512.04324ByteDance SeedDec 3, 2025score 8~111 min
- DecEcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerceevaluation2512.08868TongyiLabDec 9, 2025score 7~95 min
- DecThe FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factualityevaluation2512.10791DeepmindDec 11, 2025score 9~127 min
- DecMMGR: Multi-Modal Generative Reasoningevaluation2512.14691Dec 16, 2025~123 min
- DecProbing Scientific General Intelligence of LLMs with Scientist-Aligned Workflowsevaluation2512.16969Dec 18, 2025score 9~138 min
- DecMobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environmentsevaluation2512.19432TongyiLabDec 22, 2025score 6~122 min
- DecLLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamicsevaluation2512.21010ByteDance SeedDec 24, 2025score 9~109 min
- Dec4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillationmultimodal2512.17012NVIDIADec 18, 2025score 6~109 min
- DecSpatialTree: How Spatial Abilities Branch Out in MLLMsmultimodal2512.20617ByteDance SeedDec 23, 2025score 8~112 min
- DecARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoningrl-training2512.05111Intern Large ModelsDec 4, 2025score 8~102 min
- DecTaxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Modelssafety2512.05339RobloxDec 5, 2025score 5~96 min
- DecFaithLens: Detecting and Explaining Faithfulness Hallucinationtraining-methods2512.20182Tsinghua NLP GroupDec 23, 2025score 6~102 min
- DecEvaluating Gemini Robotics Policies in a Veo World Simulatorvision2512.10675DeepmindDec 11, 2025score 4~115 min
- NovThe Collaboration Gapagents2511.02687Microsoft ResearchNov 4, 2025score 6~123 min
- NovWhat Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversityagents2511.15593Nov 19, 2025~125 min
- NovGeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalizationagents2511.15705Tencent HunyuanNov 19, 2025score 7~96 min
- NovFara-7B: An Efficient Agentic Model for Computer Useagents2511.19663Microsoft ResearchNov 24, 2025score 6~124 min
- NovDynamic Reflections: Probing Video Representations with Text Alignmentalignment2511.02767DeepMindNov 4, 2025score 9~115 min
- NovCodeClash: Benchmarking Goal-Oriented Software Engineeringcode2511.00839Nov 2, 2025~100 min
- NovMMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generationdiffusion2511.09611ByteDanceNov 12, 2025score 9~109 min
- NovWhen Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thoughtevaluation2511.02779ByteDance SeedNov 4, 2025score 8~114 min
- NovDiscoX: Benchmarking Discourse-Level Translation task in Expert Domainsevaluation2511.10984ByteDance SeedNov 14, 2025score 5~111 min
- NovInferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulationinference-optimization2511.20714DAMO AcademyNov 25, 2025score 9~100 min
- NovMME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacitymultimodal2511.03146ByteDance SeedNov 5, 2025score 9~107 min
- NovDeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoningreasoning2511.22570Nov 27, 2025~110 min
- NovTimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learningrl-training2511.05489ByteDanceNov 7, 2025score 9~106 min
- NovReinforcement Learning Improves Traversal of Hierarchical Knowledge in LLMstraining-methods2511.05933Meta AINov 8, 2025score 9~116 min
- OctMagentic Marketplace: An Open-Source Environment for Studying Agentic Marketsagents2510.25779Microsoft ResearchOct 27, 2025score 8~114 min
- OctDocReward: A Document Reward Model for Structuring and Stylizingdata2510.11391Microsoft ResearchOct 13, 2025score 6~95 min
- OctAgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesisdata2510.24695TongyiLabOct 28, 2025score 9~113 min
- OctVibe Checker: Aligning Code Evaluation with Human Preferenceevaluation2510.07315Oct 8, 2025~114 min
- OctBeyond Correctness: Evaluating Subjective Writing Preferences Across Culturesevaluation2510.14616ByteDance SeedOct 16, 2025score 8~104 min
- OctProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judgeevaluation2510.18941NVIDIAOct 21, 2025score 9~101 min
- OctOSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agentsevaluation2510.24563TongyiLabOct 28, 2025score 5~111 min
- OctSTAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligenceevaluation2510.24693Intern Large ModelsOct 28, 2025score 6~114 min
- OctAMO-Bench: Large Language Models Still Struggle in High School Math Competitionsevaluation2510.26768LongCatOct 30, 2025score 9~107 min
- OctWhen to Ensemble: Identifying Token-Level Points for Stable and Fast LLM Ensemblingllm-systems2510.15346KAIST AIOct 17, 2025score 5~120 min
- OctGenerative Universal Verifier as Multimodal Meta-Reasonermultimodal2510.13804ByteDance SeedOct 15, 2025score 8~106 min
- OctR-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?reasoning2510.08189OpenAIOct 9, 2025score 8~107 min
- OctTracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoningreasoning2510.10494Microsoft ResearchOct 12, 2025score 9~120 min
- OctQwen3Guard Technical Reportsafety2510.14276Qwen / Alibaba CloudOct 16, 2025score 8~136 min
- OctE2Rank: Your Text Embedding can Also be an Effective and Efficient Listwise Rerankertraining-methods2510.22733Alibaba DAMOOct 26, 2025score 9~108 min
- OctDSI-Bench: A Benchmark for Dynamic Spatial Intelligencevision2510.18873alibaba-incOct 21, 2025score 8~102 min
- OctTowards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculumvision2510.27571Alibaba DAMOOct 31, 2025score 8~135 min
- SepTowards General Agentic Intelligence via Environment Scalingagents2509.13311Sep 16, 2025~114 min
- SepWhy Language Models Hallucinatealignment2509.04664Sep 4, 2025~118 min
- SepInverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?evaluation2509.04292ByteDance SeedSep 4, 2025score 6~97 min
- SepReviewScore: Misinformed Peer Review Detection with Large Language Modelsevaluation2509.21679KAIST AISep 25, 2025score 3~108 min
- SepVitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applicationsevaluation2509.26490LongCatSep 30, 2025score 6~136 min
- SepThe Pitfalls of KV Cache Compressioninference-optimization2510.00231Sep 30, 2025~125 min
- SepHunyuan-MT Technical Reportllm-systems2509.05209Tencent HunyuanSep 5, 2025score 3~99 min
- AugEvaluating, Synthesizing, and Enhancing for Customer Support Conversationdata2508.04423Qwen DianJinAug 6, 2025score 3~117 min
- AugLiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queriesevaluation2508.15760Zoom AIAug 21, 2025score 7~104 min
- AugReportBench: Evaluating Deep Research Agents via Academic Survey Tasksevaluation2508.15804ByteDanceAug 14, 2025score 6~112 min
- AugIs Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lensreasoning2508.01191Aug 2, 2025~140 min
- AugOn the Theoretical Limitations of Embedding-Based Retrievalretrieval2508.21038Aug 28, 2025~102 min
- AugREINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translationuncategorized2508.04946Roblox CorporationAug 7, 2025score 2~114 min
- AugFutureX: An Advanced Live Benchmark for LLM Agents in Future Predictionuncategorized2508.11987ByteDance SeedAug 16, 2025score 8~116 min
- JulA Systematic Analysis of Hybrid Linear Attentionarchitecture2507.06457Jul 8, 2025score 9~92 min
- JulSWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?code2507.12415Jul 16, 2025score 10~126 min
- JulEvaluating Morphological Alignment of Tokenizers in 70 Languagesevaluation2507.06378Jul 8, 2025~112 min
- JulEvaluating SAE interpretability without explanationsevaluation2507.08473Jul 11, 2025~108 min
- JulThe Invisible Leash: Why RLVR May or May Not Escape Its Originevaluation2507.14843Jul 20, 2025score 10~137 min
- JulTraceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodologymultimodal2507.07999OpenAIJul 10, 2025score 8~99 min
- JulVoxtralmultimodal2507.13264MistralJul 17, 2025~111 min
- JulThe Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigationsafety2507.05578Jul 8, 2025score 9~120 min
- JunThe Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvementscode2506.22419Jun 27, 2025score 10~118 min
- JunInfini-gram mini: Exact n-gram Search at the Internet Scale with FM-Indexdata2506.12229Jun 13, 2025score 9~123 min
- JunM$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Datasetevaluation2506.02510Qwen DianJinJun 3, 2025score 6~122 min
- JunLongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMsinference-optimization2506.14429Jun 17, 2025score 9~112 min
- JunQwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Modelspretraining2506.05176Qwen / Alibaba CloudJun 5, 2025score 10~82 min
- JunPerformance Prediction for Large Systems via Text-to-Text Regressionserving2506.21718DeepMindJun 26, 2025score 3~109 min
- MayWhen AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Researchevaluation2505.11855May 17, 2025score 8~108 min
- MayEfficientLLM: Efficiency in Large Language Modelsevaluation2505.13840May 20, 2025score 10~104 min
- MayAn Empirical Study of Qwen3 Quantizationlow-precision2505.02214May 4, 2025score 9~89 min
- MayREASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewardsreasoning2505.24760May 30, 2025score 9~109 min
- MayLessons from Defending Gemini Against Indirect Prompt Injectionssafety2505.14534DeepMindMay 20, 2025score 5~116 min
- AprPaperBench: Evaluating AI's Ability to Replicate AI Researchagents2504.01848Apr 2, 2025score 10~102 min
- AprMulti-SWE-bench: A Multilingual Benchmark for Issue Resolvingevaluation2504.02605ByteDance SeedApr 3, 2025score 9~99 min
- AprDeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?evaluation2504.08120OpenAIApr 10, 2025score 8~97 min
- AprThe Leaderboard Illusionevaluation2504.20879Apr 29, 2025~124 min
- AprInternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Modelsmultimodal2504.10479Apr 14, 2025score 10~101 min
- AprKimi-Audio Technical Reportmultimodal2504.18425Moonshot AIApr 25, 2025score 9~104 min
- AprDeepSeek-R1 Thoughtology: Let's think about LLM Reasoningreasoning2504.07128DeepSeekApr 2, 2025score 10~120 min
- AprDoes Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?rl-training2504.13837Apr 18, 2025score 9~107 min
- AprInference-Time Scaling for Generalist Reward Modelingscaling-laws2504.02495Apr 3, 2025~101 min
- AprOLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokenstraining-methods2504.07096Apr 9, 2025score 9~92 min
- MarPokéChamp: an Expert-level Minimax Language Agentagents2503.04094Mar 6, 2025score 4~130 min
- MarWhy Do Multi-Agent LLM Systems Fail?agents2503.13657Mar 17, 2025~133 min
- MarUVE: Are MLLMs Unified Evaluators for AI-Generated Videos?evaluation2503.09949ByteDance SeedMar 13, 2025score 6~111 min
- MarRankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluationevaluation2503.19092Mar 24, 2025~123 min
- MarQuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?evaluation2503.22674DeepMindMar 28, 2025~121 min
- MarGemini Embedding: Generalizable Embeddings from Geminillm-systems2503.07891DeepMindMar 10, 2025~99 min
- MarA Comprehensive Survey on Long Context Language Modelingllm-systems2503.17407Mar 20, 2025~121 min
- FebKernelBench: Can LLMs Write Efficient GPU Kernels?evaluation2502.10517Feb 14, 2025~112 min
- FebSuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplinesevaluation2502.14739ByteDance SeedFeb 20, 2025score 9~86 min
- FebV2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Modelsllm-systems2502.09980NVIDIAFeb 14, 2025score 2~119 min
- FebGold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2reasoning2502.03544Feb 5, 2025score 10~127 min
- JanInternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Modelalignment2501.12368InternLM / Shanghai AI LabJan 21, 2025score 9~95 min
- JanSFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingalignment2501.17161Jan 28, 2025score 10~103 min
- JanHumanity's Last Examevaluation2501.14249AnthropicJan 24, 2025score 10~122 min
- JanPeople who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textevaluation2501.15654Jan 26, 2025score 10~110 min
- JanEarly External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluationsafety2501.17749OpenAIJan 29, 2025score 6~108 min
- Jano3-mini vs DeepSeek-R1: Which One is Safer?safety2501.18438OpenAIJan 30, 2025score 3~96 min
- JanThe Lessons of Developing Process Reward Models in Mathematical Reasoningtraining-methods2501.07301Qwen / Alibaba CloudJan 13, 2025score 10~115 min
2024
56- DecLAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotationsdata2412.08580Dec 11, 2024score 6~88 min
- DecEXAONE 3.5: Series of Large Language Models for Real-world Use Casespretraining2412.04862LG EXAONEDec 6, 2024~97 min
- DecQwen2.5 Technical Reportpretraining2412.15115Qwen / Alibaba CloudDec 19, 2024~109 min
- DecDensing Law of LLMsscaling-laws2412.04315OpenBMBDec 5, 2024score 8~113 min
- NovBenchmarking Distributional Alignment of Large Language Modelsevaluation2411.05403Nov 8, 2024~118 min
- NovDo Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?evaluation2411.16679DeepMindNov 25, 2024~124 min
- Nov"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantizationinference-optimization2411.02355DropboxNov 4, 2024score 10~118 min
- OctTLDR: Token-Level Detective Reward Model for Large Vision Language Modelsalignment2410.04734Meta LlamaOct 7, 2024score 8~119 min
- OctNot All LLM Reasoners Are Created Equalreasoning2410.01748Oct 2, 2024~109 min
- OctA Comparative Study on Reasoning Patterns of OpenAI's o1 Modelreasoning2410.13639Oct 17, 2024score 9~110 min
- OctGPT-4o System Cardsafety2410.21276OpenAIOct 25, 2024score 9~137 min
- SepLanguage Models Learn to Mislead Humans via RLHFalignment2409.12822Sep 19, 2024score 8~98 min
- SepIs Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysisalignment2409.20059Sep 30, 2024score 8~107 min
- SepA Case Study of Web App Coding with OpenAI Reasoning Modelscode2409.13773Sep 19, 2024score 8~108 min
- SepMMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmarkevaluation2409.02813Sep 4, 2024score 9~99 min
- SepMichelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queriesevaluation2409.12640DeepMindSep 19, 2024~114 min
- SepLaw of the Weakest Link: Cross Capabilities of Large Language Modelsevaluation2409.19951Sep 30, 2024score 9~114 min
- SepAttention Heads of Large Language Models: A Surveyreasoning2409.03752Sep 5, 2024~127 min
- SepCoffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Coderl-training2409.19715Sep 29, 2024score 8~111 min
- SepA Controlled Study on Long Context Extension and Generalization in LLMstraining-methods2409.12181Sep 18, 2024~146 min
- AugRecent Surge in Public Interest in Transportation: Sentiment Analysis of Baidu Apollo Go Using Weibo Datacontext-optimization2408.10088Aug 19, 2024score 1~117 min
- AugUnleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Modelsdata2408.02085Aug 4, 2024score 10~182 min
- AugSelf-Taught Evaluatorsevaluation2408.02666Aug 5, 2024score 9~103 min
- AugEXAONE 3.0 7.8B Instruction Tuned Language Modelllm-systems2408.03541LG EXAONEAug 7, 2024score 9~114 min
- JulLMMs-Eval: Reality Check on the Evaluation of Large Multimodal Modelsevaluation2407.12772Jul 17, 2024score 9~133 min
- JulKiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Modelsevaluation2407.17773DeepMindJul 25, 2024~112 min
- JulSpectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scalepretraining2407.12327Jul 17, 2024score 9~98 min
- JulOn scalable oversight with weak LLMs judging strong LLMssafety2407.04622Jul 5, 2024~137 min
- JulLarge Language Monkeys: Scaling Inference Compute with Repeated Samplingscaling-laws2407.21787Jul 31, 2024~130 min
- JulGENERALIZATION V.S. MEMORIZATION: TRACING LANGUAGE MODELS’ CAPABILITIES BACK TO PRETRAINING DATAtraining-methods2407.14985Jul 20, 2024~131 min
- JunDataComp-LM: In Search of the Next Generation of Training Sets for Language Modelsdata2406.11794AppleJun 17, 2024score 10~93 min
- JunMMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmarkevaluation2406.01574MistralJun 3, 2024score 10~94 min
- JunVideoPhy: Evaluating Physical Commonsense for Video Generationevaluation2406.03520Jun 5, 2024~105 min
- JunGenAI Arena: An Open Evaluation Platform for Generative Modelsevaluation2406.04485Jun 6, 2024score 9~98 min
- JunThe BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Modelsevaluation2406.05761Jun 9, 2024~107 min
- JunOLMES: A Standard for Language Model Evaluationsevaluation2406.08446Allen Institute for AIJun 12, 2024~111 min
- JunFrom Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipelineevaluation2406.11939Jun 17, 2024score 9~105 min
- AprRULER: What's the Real Context Size of Your Long-Context Language Models?evaluation2404.06654Qwen / Alibaba CloudApr 9, 2024score 10~109 min
- AprReplacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Modelsevaluation2404.18796Apr 29, 2024score 9~103 min
- AprLLM2Vec: Large Language Models Are Secretly Powerful Text Encodersllm-systems2404.05961Apr 9, 2024~105 min
- AprCapabilities of Gemini Models in Medicinemultimodal2404.18416Apr 29, 2024score 9~125 min
- MarLong-form factuality in large language modelsalignment2403.18802DeepMindMar 27, 2024score 8~110 min
- MarGemini 1.5: Unlocking multimodal understanding across millions of tokens of contextcontext-optimization2403.05530Mar 8, 2024score 10~133 min
- MarChatbot Arena: An Open Platform for Evaluating LLMs by Human Preferenceevaluation2403.04132Mar 7, 2024~118 min
- MarCLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Koreanevaluation2403.06412Microsoft ResearchMar 11, 2024~111 min
- MarRewardBench: Evaluating Reward Models for Language Modelingevaluation2403.13787InternLM / Shanghai AI LabMar 20, 2024~124 min
- MarGemma: Open Models Based on Gemini Research and Technologyllm-systems2403.08295Google ResearchMar 13, 2024score 9~105 min
- MarEvaluating Frontier Models for Dangerous Capabilitiessafety2403.13793DeepMindMar 20, 2024score 8~140 min
- MarAtP*: An efficient and scalable method for localizing LLM behaviour to componentsuncategorized2403.00745DeepMind0 citesMar 1, 2024score 3~123 min
- FebOLMo: Accelerating the Science of Language Modelsagents2402.00838Allen Institute for AIFeb 1, 2024score 10~113 min
- FebPrior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Banditsevaluation2402.05878DeepMind0 citesFeb 8, 2024~114 min
- FebOmniPred: Language Models as Universal Regressorspretraining2402.14547DeepMindFeb 22, 2024score 5~101 min
- FebPremise Order Matters in Reasoning with Large Language Modelsreasoning2402.08939DeepMindFeb 14, 2024~116 min
- FebDo Membership Inference Attacks Work on Large Language Models?safety2402.07841Feb 12, 2024~125 min
- JanInfini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokensdata2401.17377Jan 30, 2024score 10~114 min
- JanFrom GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalitiesevaluation2401.15071Jan 26, 2024score 8~96 min
2023
28- DecGenerative agent-based modeling with actions grounded in physical, social, or digital space using Concordiaagents2312.03664DeepMindDec 6, 2023score 6~96 min
- DecAlignment for Honestyalignment2312.07000Dec 12, 2023score 9~124 min
- DecChallenges with unsupervised LLM knowledge discoveryalignment2312.10029DeepMindDec 15, 2023score 5~131 min
- DecGemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Modelsevaluation2312.17661Dec 29, 2023score 3~127 min
- DecGemini vs GPT-4V: A Preliminary Comparison and Combination of Vision-Language Models Through Qualitative Casesmultimodal2312.15011Dec 22, 2023score 5~118 min
- DecA Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertisevision2312.12436Dec 19, 2023score 9~92 min
- NovLevels of AGI for Operationalizing Progress on the Path to AGIevaluation2311.02462DeepMindNov 4, 2023score 5~122 min
- NovInstruction-Following Evaluation for Large Language Modelsevaluation2311.07911Google ResearchNov 14, 2023score 9~89 min
- NovGPQA: A Graduate-Level Google-Proof Q&A Benchmarkevaluation2311.12022MistralNov 20, 2023~119 min
- NovMMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmarkevaluation2311.1650201-AINov 27, 2023~97 min
- NovFine-tuning Language Models for Factualitytraining-methods2311.08401Nov 14, 2023score 9~104 min
- NovRoboVQA: Multimodal Long-Horizon Reasoning for Roboticsuncategorized2311.00899DeepMindNov 1, 2023score 5~140 min
- OctAssessing Large Language Models on Climate Informationevaluation2310.02932DeepMindOct 4, 2023~118 min
- OctDetecting Pretraining Data from Large Language Modelssafety2310.16789Oct 25, 2023score 9~111 min
- SepM3DSYNTH: A DATASET OF MEDICAL 3D IMAGES WITH AI-GENERATED LOCAL MANIPULATIONSdata2309.07973Sep 14, 2023~120 min
- SepLMSYS-CHAT-1M: A LARGE-SCALE REAL-WORLD LLM CONVERSATION DATASETsafety2309.11998Sep 21, 2023~126 min
- AugStudying Large Language Model Generalization with Influence Functionstraining-methods2308.03296Aug 7, 2023score 8~102 min
- JulLost in the Middle: How Language Models Use Long Contextscontext-optimization2307.03172Together AIJul 6, 2023score 9~101 min
- JulA Survey on Evaluation of Large Language Modelsevaluation2307.03109Jul 6, 2023~108 min
- JulHow is ChatGPT's behavior changing over time?evaluation2307.09009Jul 18, 2023score 9~107 min
- JunMistral 7Bevaluation2306.05685MistralJun 9, 2023~109 min
- JunOrca: Progressive Learning from Complex Explanation Traces of GPT-4reasoning2306.02707Stability AIJun 5, 2023score 9~91 min
- JunEstimating the Causal Effect of Early ArXiving on Paper Acceptancesafety2306.13891Jun 24, 2023~116 min
- MayStarCoder: may the source be with you!code2305.06161Stability AIMay 9, 2023score 8~124 min
- MayCheaply Evaluating Inference Efficiency Metrics for Autoregressive Transformer APIsserving2305.02440May 3, 2023score 9~111 min
- MayThe False Promise of Imitating Proprietary LLMstraining-methods2305.15717May 25, 2023score 9~122 min
- MayHow Language Model Hallucinations Can Snowballuncategorized2305.13534May 22, 2023score 9~113 min
- MarSparks of Artificial General Intelligence: Early experiments with GPT-4evaluation2303.12712Mar 22, 2023~117 min
2022
11- DecRobust Speech Recognition via Large-Scale Weak Supervisionpretraining2212.04356OpenAIDec 6, 2022~119 min
- DecHyDE: Precise Zero-Shot Dense Retrieval without Relevance Labelsretrieval2212.10496DoorDashDec 20, 2022~94 min
- NovHolistic Evaluation of Language Modelsevaluation2211.09110Stanford NLPNov 16, 2022~115 min
- OctThe Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMsevaluation2210.14986EleutherAIOct 26, 2022~100 min
- AugAtlas: Few-shot Learning with Retrieval Augmented Language Modelsretrieval2208.03299Aug 5, 2022~124 min
- JunBIG-bench: Beyond the Imitation Game Benchmarkscaling-laws2206.04615MistralJun 9, 2022~99 min
- JunEmergent Abilities of Large Language Modelsscaling-laws2206.07682Jun 15, 2022~128 min
- MayFew-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learningalignment2205.05638May 11, 2022~115 min
- MayFLEURS: Few-shot Learning Evaluation of Universal Representations of Speechevaluation2205.12446NVIDIAMay 25, 2022~94 min
- AprGPT-NeoX-20B: An Open-Source Autoregressive Language Modelllm-systems2204.06745Stability AIApr 14, 2022~111 min
- FebRed Teaming Language Models with Language Modelsalignment2202.03286Feb 7, 2022~119 min
2021
9- DecGopherscaling-laws2112.11446Dec 8, 2021~142 min
- OctGSM8K: A Benchmark for Grade School Math Word Problemsreasoning2110.14168MistralOct 27, 2021~113 min
- SepTruthfulQA: Measuring How Models Mimic Human Falsehoodssafety2109.07958MistralSep 8, 2021~117 min
- SepRecursively Summarizing Books with Human Feedbacktraining-methods2109.1086266 citesSep 22, 2021~99 min
- JulEvaluating Large Language Models Trained on Codecode2107.03374MistralJul 7, 2021~126 min
- MarMATH: Measuring Mathematical Problem Solving with the MATH Datasetevaluation2103.03874MistralMar 5, 2021~102 min
- FebProbing Classifiers: Promises, Shortcomings, and Advancesevaluation2102.12452Feb 24, 2021~128 min
- FebCalibrate Before Use: Improving Few-Shot Performance of Language Modelsprompting2102.09690Feb 19, 2021~102 min
- FebMeasuring and Improving Consistency in Pretrained Language Modelsuncategorized2102.0101732 citesFeb 1, 2021~101 min
2020
22019
5- OctMLQA: Evaluating Cross-lingual Extractive Question Answeringevaluation1910.07475Microsoft ResearchOct 16, 2019~121 min
- SepLanguage Models as Knowledge Bases?evaluation1909.01066Sep 3, 2019~114 min
- JunLatent Retrieval for Weakly Supervised Open Domain Question Answeringretrieval1906.0030067 citesJun 1, 2019~102 min
- MaySuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systemsevaluation1905.00537991 citesMay 2, 2019~114 min
- MayHellaSwag: Can a Machine Really Finish Your Sentence?evaluation1905.07830MistralMay 19, 2019~119 min
2018
7- OctModel Cards for Model Reportingalignment1810.03993Meta AI / FAIROct 5, 2018~118 min
- OctHow Powerful Are Graph Neural Networks?architecture1810.00826Oct 1, 2018~119 min
- SepXNLI: Evaluating Cross-lingual Sentence Representationsevaluation1809.05053Microsoft ResearchSep 13, 2018~96 min
- AugSWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inferenceevaluation1808.05326103 citesAug 16, 2018~123 min
- JunSQuAD 2.0: The Stanford Question Answering Datasetevaluation1806.03822Jun 11, 2018~97 min
- AprGLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understandingevaluation1804.07461555 citesApr 20, 2018~108 min
- MarThink you have Solved Question Answering? Try ARC, the AI2 Reasoning Challengeevaluation1803.05457Allen Institute for AIMar 14, 2018~118 min