306 papers
Agents
0/304Agentic systems, tool use, and autonomous reasoning.
Progress0 of 304
2026
127- MayMolmoAct2: Action Reasoning Models for Real-world Deploymentagents2605.02881Allen Institute for AIMay 4, 2026~118 min
- MayWildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluationagents2605.10912Intern Large ModelsMay 11, 2026~116 min
- MayToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agentsagents2605.12481TongyiLabMay 12, 2026~122 min
- MayPREPING: Building Agent Memory without Tasksagents2605.13880KAIST AIMay 11, 2026~94 min
- MayOrchard: An Open-Source Agentic Modeling Frameworkagents2605.15040Microsoft ResearchMay 14, 2026~130 min
- MayHINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agentsagents2605.17873KAIST AIMay 18, 2026~112 min
- MaySkillOpt: Executive Strategy for Self-Evolving Agent Skillsagents2605.23904Microsoft ResearchMay 22, 2026~125 min
- MayCUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agentsagents2605.25624May 25, 2026~120 min
- Mayoptimize_anything: A Universal API for Optimizing any Text Parameterarchitecture2605.19633May 19, 2026~111 min
- MayContext Training with Active Information Seekingcontext-optimization2605.13050DeepmindMay 13, 2026~123 min
- MayCovering Human Action Space for Computer Use: Data Synthesis and Benchmarkdata2605.12501MicrosoftMay 12, 2026~108 min
- MayMEME: Multi-entity & Evolving Memory Evaluationevaluation2605.12477KAIST AIMay 12, 2026~130 min
- MayFrom Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skillsevaluation2605.23899Microsoft ResearchMay 22, 2026~120 min
- MayOpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agentsmultimodal2605.05185Tencent HunyuanMay 6, 2026~104 min
- MayMemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Modelsmultimodal2605.14906NVIDIAMay 14, 2026~125 min
- MayHeavySkill: Heavy Thinking as the Inner Skill in Agentic Harnessreasoning2605.02396LongCatMay 4, 2026~97 min
- MayLLMs Improving LLMs: Agentic Discovery for Test-Time Scalingreasoning2605.08083GoogleMay 8, 2026~135 min
- MayRethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?retrieval2605.10848May 11, 2026~110 min
- MayAEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learningrl-training2605.00425BAIDUMay 1, 2026~125 min
- MayDebiased Model-based Representations for Sample-efficient Continuous Controlrl-training2605.11711Tencent HunyuanMay 12, 2026~113 min
- AprCORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discoveryagents2604.01658Massachusetts Institute of TechnologyApr 2, 2026score 9~131 min
- AprGPA: Learning GUI Process Automation from Demonstrationsagents2604.01676Salesforce AI ResearchApr 2, 2026score 9~114 min
- AprMemory Transfer Learning: How Memories are Transferred Across Domains in Coding Agentsagents2604.14004KAIST AIApr 15, 2026score 8~113 min
- AprMM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generationagents2604.15309Microsoft ResearchApr 16, 2026score 3~94 min
- AprAgent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligenceagents2604.18292ByteDance SeedApr 20, 2026~108 min
- AprToward Scalable Terminal Task Synthesis via Skill Graphsagents2604.25727Tencent HunyuanApr 28, 2026~130 min
- AprSynthetic Computers at Scale for Long-Horizon Productivity Simulationagents2604.28181MicrosoftApr 30, 2026~128 min
- AprRoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policiesevaluation2604.09860NVIDIAApr 10, 2026~118 min
- AprOccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Modelsevaluation2604.10866QwenApr 13, 2026score 8~104 min
- AprLARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignmentevaluation2604.11689LongCatApr 13, 2026~111 min
- AprAJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluationevaluation2604.18240LongCatApr 20, 2026~107 min
- AprMolmoWeb: Open Visual Web Agent and Open Data for the Open Webllm-systems2604.08516Apr 9, 2026~100 min
- AprGLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agentsmultimodal2604.26752Z.aiApr 29, 2026~127 min
- AprProcess Reward Agents for Steering Knowledge-Intensive Reasoningreasoning2604.09482ETH ZurichApr 10, 2026score 9~100 min
- AprTowards Autonomous Mechanistic Reasoning in Virtual Cellsreasoning2604.11661KAIST AIApr 13, 2026score 3~115 min
- AprSandMLE: A Multi-Agent Framework that Generates Synthetic Sandbox Environments for Training Machine Learning Engineering Agentstraining-methods2604.04872AI at MetaApr 6, 2026score 9~97 min
- AprTCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agentstraining-methods2604.24005TongyiLabApr 27, 2026~112 min
- AprHY-Embodied-0.5: Embodied Foundation Models for Real-World Agentsuncategorized2604.07430Tencent HunyuanApr 8, 2026score 9~127 min
- AprActionParty: Multi-Subject Action Binding in Generative Video Gamesvision2604.02330Snap ResearchApr 2, 2026score 9~102 min
- MarProact-VL: A Proactive VideoLLM for Real-Time AI Companionsagents2603.03447Microsoft ResearchMar 3, 2026score 9~94 min
- MarTRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemasagents2603.16448meituanMar 17, 2026score 9~98 min
- MarProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agentsagents2603.18815NVIDIAMar 19, 2026~99 min
- MarA Subgoal-driven Framework for Improving Long-Horizon LLM Agentsagents2603.19685DeepmindMar 20, 2026score 10~135 min
- MarLumosX: Relate Any Identities with Their Attributes for Personalized Video Generationagents2603.20192Alibaba DAMOMar 20, 2026score 4~109 min
- MarT-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Searchagents2603.22341KAIST AIMar 21, 2026score 9~112 min
- MarVision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verificationagents2603.26648Z.aiMar 27, 2026score 9~132 min
- MarUnify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesisagents2603.29620Tencent HunyuanMar 31, 2026score 8~107 min
- MarLearning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Usealignment2603.03205Microsoft ResearchMar 3, 2026score 8~111 min
- MarPOLCA: Stochastic Generative Optimization with LLMcode2603.14769DeepmindMar 16, 2026score 9~113 min
- MarCan Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streamsevaluation2603.07392KAIST AIMar 8, 2026score 9~117 min
- MarMA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agentsevaluation2603.09827KAIST AIMar 10, 2026score 3~119 min
- MarStrategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collectionsevaluation2603.12180SnowflakeMar 12, 2026score 9~134 min
- MarEgo2Web: A Web Agent Benchmark Grounded in Egocentric Videosevaluation2603.22529DeepmindMar 23, 2026score 8~120 min
- MarFinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocolevaluation2603.24943Qwen DianJinMar 26, 2026score 3~111 min
- MarConsistency Amplifies: How Behavioral Variance Shapes Agent Accuracyevaluation2603.25764SnowflakeMar 26, 2026score 9~98 min
- MarArtLLM: Generating Articulated Assets via 3D LLMmultimodal2603.01142Tencent HunyuanMar 1, 2026score 2~105 min
- MarIntern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scalemultimodal2603.25040InternLM / Shanghai AI LabMar 26, 2026~113 min
- MarUnderstanding by Reconstruction: Reversing the Software Development Process for LLM Pretrainingpretraining2603.11103ByteDance SeedMar 11, 2026score 9~111 min
- MarLearning to Retrieve from Agent Trajectoriesretrieval2604.04949RUC-GSAI-IIRLabMar 30, 2026score 9~118 min
- MarHeterogeneous Agent Collaborative Reinforcement Learningrl-training2603.02604ByteDanceMar 3, 2026score 9~112 min
- MarCode-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Modelsrl-training2603.10098DeepmindMar 10, 2026score 9~109 min
- MarComplementary Reinforcement Learningrl-training2603.17621alibaba-incMar 18, 2026score 9~129 min
- MarPivotRL: High Accuracy Agentic Post-Training at Low Compute Costrl-training2603.21383NVIDIAMar 22, 2026~123 min
- MarScaling Agentic Capabilities, Not Context: Efficient Reinforcement Finetuning for Large Toolspacestraining-methods2603.06713Microsoft ResearchMar 5, 2026score 9~112 min
- MarMeta-Reinforcement Learning with Self-Reflection for Agentic Searchtraining-methods2603.11327Ai2Mar 11, 2026~97 min
- MarOnline Experiential Learning for Language Modelstraining-methods2603.16856Microsoft ResearchMar 17, 2026score 9~121 min
- MarNemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillationtraining-methods2603.19220NVIDIAMar 19, 2026~94 min
- MarUnderstanding the Challenges in Iterative Generative Optimization with LLMstraining-methods2603.23994DeepmindMar 25, 2026score 9~94 min
- MarMolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulationuncategorized2603.16861Ai2Mar 17, 2026score 9~120 min
- MarSIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLMuncategorized2603.23386ByteDance SeedMar 24, 2026score 3~116 min
- MarGrounding World Simulation Models in a Real-World Metropolisvision2603.15583NAVER AI LabMar 16, 2026score 9~120 min
- FebTRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenariosagents2602.01675meituanFeb 2, 2026score 3~122 min
- FebSWE-Universe: Scale Real-World Verifiable Environments to Millionsagents2602.02361QwenFeb 2, 2026score 4~109 min
- FebSEAD: Self-Evolving Agent for Multi-Turn Service Dialogueagents2602.03548meituanFeb 3, 2026score 5~110 min
- FebProAct: Agentic Lookahead in Interactive Environmentsagents2602.05327Tencent HunyuanFeb 5, 2026score 8~95 min
- FebAgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Researchagents2602.06540OpenBMBFeb 6, 2026score 9~98 min
- FebScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Trainingagents2602.06820LongCatFeb 6, 2026score 9~103 min
- FebAgent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learningagents2602.10090SnowflakeFeb 10, 2026score 9~111 min
- FebSAGE: Scalable Agentic 3D Scene Generation for Embodied AIagents2602.10116NVIDIAFeb 10, 2026score 2~115 min
- FebSingle-minus gluon tree amplitudes are nonzeroagents2602.12176OpenAIFeb 12, 2026score 1~121 min
- FebWebWorld: A Large-Scale World Model for Web Agent Trainingagents2602.14721Qwen / Alibaba CloudFeb 16, 2026score 9~101 min
- FebRynnBrain: Open Embodied Foundation Modelsagents2602.14979Alibaba DAMOFeb 13, 2026score 9~112 min
- FebMobile-Agent-v3.5: Multi-platform Fundamental GUI Agentsagents2602.16855TongyiLabFeb 15, 2026score 2~131 min
- FebDecoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking Systemagents2602.18640Feb 20, 2026~114 min
- FebMobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenariosagents2602.22638alibaba-incFeb 26, 2026score 3~104 min
- FebDr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generationscode2602.05885Feb 5, 2026~111 min
- FebOn Data Engineering for Scaling LLM Terminal Capabilitiesdata2602.21193NVIDIAFeb 24, 2026score 3~102 min
- FebK-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Modelllm-systems2602.19128Feb 22, 2026~109 min
- FebStep 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parametersmoe2602.10604StepFunFeb 11, 2026score 8~135 min
- FebKimi K2.5: Visual Agentic Intelligencemultimodal2602.02276Moonshot / KimiFeb 2, 2026score 8~145 min
- FebVideoWorld 2: Learning Transferable Knowledge from Real-world Videospretraining2602.10102ByteDance SeedFeb 10, 2026score 2~107 min
- FebMolHIT: Advancing Molecular-Graph Generation with Hierarchical Discrete Diffusion Modelspretraining2602.17602KAIST AIFeb 19, 2026score 2~123 min
- FebAccelerating Scientific Research with Gemini: Case Studies and Common Techniquesreasoning2602.03837Feb 3, 2026score 2~123 min
- FebSearch-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaborationrl-training2602.03647Feb 3, 2026~113 min
- FebThe Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societiessafety2602.09877AnthropicFeb 10, 2026score 3~121 min
- FebDualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inferenceserving2602.21548DeepSeekFeb 25, 2026score 9~109 min
- FebReinforcement World Model Learning for LLM-based Agentstraining-methods2602.05842Microsoft ResearchFeb 5, 2026score 3~88 min
- FebSeeUPO: Sequence-Level Agentic-RL with Convergence Guaranteestraining-methods2602.06554Tongyi-MAIFeb 6, 2026score 9~111 min
- FebABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learningtraining-methods2602.11236Alibaba AMAP CV LabFeb 11, 2026score 2~111 min
- FebNanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Actstraining-methods2602.13367Feb 13, 2026~100 min
- FebWorld Guidance: World Modeling in Condition Space for Action Generationtraining-methods2602.22010ByteDance SeedFeb 25, 2026score 2~102 min
- FebCUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generationtraining-methods2602.24286ByteDance SeedFeb 27, 2026score 9~110 min
- FebQwen3-Coder-Next Technical Reporttraining-methods2603.00729QwenFeb 28, 2026score 9~123 min
- FebDreamDojo: A Generalist Robot World Model from Large-Scale Human Videosuncategorized2602.06949NVIDIAFeb 6, 2026score 2~117 min
- FebGLM-5: from Vibe Coding to Agentic Engineeringuncategorized2602.15763Zhipu / GLMFeb 17, 2026~148 min
- FebMedXIAOHE: A Comprehensive Recipe for Building Medical MLLMsvision2602.12705ByteDanceFeb 13, 2026score 7~121 min
- JanNitroGen: An Open Foundation Model for Generalist Gaming Agentsagents2601.02427NVIDIAJan 4, 2026score 2~125 min
- JanThinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalizationagents2601.05432alibaba-incJan 8, 2026score 3~115 min
- JanExpSeek: Self-Triggered Experience Seeking for Web Agentsagents2601.08605TongyiLabJan 13, 2026score 5~99 min
- JanVLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memoryagents2601.08665ByteDance SeedJan 13, 2026score 3~124 min
- JanFast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planningagents2601.09708NVIDIAJan 14, 2026score 3~113 min
- JanUnlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Textagents2601.10355LongCatJan 15, 2026score 9~111 min
- JanAgentic Reasoning for Large Language Modelsagents2601.12538University of Illinois at Urbana-ChampaignJan 18, 2026score 9~116 min
- JanFantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigationagents2601.13976Alibaba AMAP CV LabJan 20, 2026score 3~126 min
- JanEvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experienceagents2601.15876meituanJan 22, 2026score 9~120 min
- JanComputer Environments Elicit General Agentic Intelligence in LLMsagents2601.16206Microsoft ResearchJan 22, 2026score 5~94 min
- JanLongCat-Flash-Thinking-2601 Technical Reportagents2601.16725LongCatJan 23, 2026score 9~101 min
- JanThe Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generationagents2601.17737Tencent HunyuanJan 25, 2026score 2~118 min
- JanDeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraintsagents2601.18137QwenJan 26, 2026score 2~101 min
- JanMEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineeringagents2601.22859ernie-researchJan 30, 2026score 5~120 min
- JanRigMo: Unifying Rig and Motion Learning for Generative Animationarchitecture2601.06378SnapJan 10, 2026score 1~115 min
- JanSWE-Pruner: Self-Adaptive Context Pruning for Coding Agentscode2601.16746ByteDanceJan 23, 2026score 8~126 min
- JanEvasionBench: A Large-Scale Benchmark for Detecting Managerial Evasion in Earnings Call Q&Aevaluation2601.09142Jan 14, 2026score 2~119 min
- JanLost in the Noise: How Reasoning Models Fail with Contextual Distractorsreasoning2601.07226KAIST AIJan 12, 2026score 3~106 min
- JanArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Rankingrl-training2601.06487Alibaba DAMOJan 10, 2026score 9~120 min
- JanHow AI Impacts Skill Formationrl-training2601.20245AnthropicJan 28, 2026score 1~107 min
- JanMegaFlow: Large-Scale Distributed Orchestration System for the Agentic Eraserving2601.07526QwenJan 12, 2026score 6~109 min
2025
116- DecSIMA 2: A Generalist Embodied Agent for Virtual Worldsagents2512.04797DeepmindDec 4, 2025score 9~116 min
- DecTowards a Science of Scaling Agent Systemsagents2512.08296Dec 9, 2025~118 min
- DecLong-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solvingagents2512.10739Dec 11, 2025~115 min
- DecMemory in the Age of AI Agentsagents2512.13564Dec 15, 2025score 9~101 min
- DecSAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learningagents2512.13874Ai2Dec 15, 2025score 9~119 min
- DecSeed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experienceagents2512.17260ByteDance SeedDec 19, 2025score 8~129 min
- DecMAI-UI Technical Report: Real-World Centric Foundation GUI Agentsagents2512.22047TongyiLabDec 26, 2025score 4~124 min
- DecKernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Metaagents2512.23236Dec 29, 2025~128 min
- DecGR-Dexter Technical Reportagents2512.24210ByteDance SeedDec 30, 2025score 2~103 min
- DecLet It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystemagents2512.24873alibaba-incDec 31, 2025score 9~130 min
- DecNemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoningarchitecture2512.20848NVIDIADec 23, 2025~114 min
- DecNVIDIA Nemotron 3: Efficient and Open Intelligencearchitecture2512.20856NVIDIADec 24, 2025~74 min
- DecWeb World Modelsarchitecture2512.23676Dec 29, 2025~112 min
- DecConfucius Code Agent: Scalable Agent Scaffolding for Real-World Codebasescode2512.10398Meta ResearchDec 11, 2025score 9~112 min
- DecDAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycleevaluation2512.04324ByteDance SeedDec 3, 2025score 8~111 min
- DecEcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerceevaluation2512.08868TongyiLabDec 9, 2025score 7~95 min
- DecProbing Scientific General Intelligence of LLMs with Scientist-Aligned Workflowsevaluation2512.16969Dec 18, 2025score 9~138 min
- DecMobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environmentsevaluation2512.19432TongyiLabDec 22, 2025score 6~122 min
- DecLLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamicsevaluation2512.21010ByteDance SeedDec 24, 2025score 9~109 min
- DecDEER: Draft with Diffusion, Verify with Autoregressive Modelsinference-optimization2512.15176Dec 17, 2025score 10~125 min
- DecDeepSeek-V3.2: Pushing the Frontier of Open Large Language Modelsllm-systems2512.02556DeepSeekDec 2, 2025~127 min
- DecARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoningrl-training2512.05111Intern Large ModelsDec 4, 2025score 8~102 min
- DecCUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learningtraining-methods2512.02551NVIDIADec 2, 2025~96 min
- DecSpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RLtraining-methods2512.04069NVIDIADec 3, 2025score 7~118 min
- DecQwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Managementtraining-methods2512.12967Dec 15, 2025~110 min
- DecNested Browser-Use Learning for Agentic Information Seekinguncategorized2512.23647TongyiLabDec 29, 2025score 9~115 min
- NovThe Collaboration Gapagents2511.02687Microsoft ResearchNov 4, 2025score 6~123 min
- NovScaling Agent Learning via Experience Synthesisagents2511.03773Nov 5, 2025~120 min
- NovRobot Learning from a Physical World Modelagents2511.07416DeepmindNov 10, 2025score 2~105 min
- NovLumine: An Open Recipe for Building Generalist Agents in 3D Open Worldsagents2511.08892ByteDance SeedNov 12, 2025score 6~115 min
- NovPart-X-MLLM: Part-aware 3D Multimodal Large Language Modelagents2511.13647Tencent HunyuanNov 17, 2025score 9~107 min
- NovWhat Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversityagents2511.15593Nov 19, 2025~125 min
- NovGeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalizationagents2511.15705Tencent HunyuanNov 19, 2025score 7~96 min
- NovFara-7B: An Efficient Agentic Model for Computer Useagents2511.19663Microsoft ResearchNov 24, 2025score 6~124 min
- NovCodeClash: Benchmarking Goal-Oriented Software Engineeringcode2511.00839Nov 2, 2025~100 min
- NovAccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimizationcode2511.15915Stanford UniversityNov 19, 2025~115 min
- NovToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestrationtraining-methods2511.21689NVIDIANov 26, 2025score 9~106 min
- NovRynnVLA-002: A Unified Vision-Language-Action and World Modeluncategorized2511.17502DAMO AcademyNov 21, 2025score 6~109 min
- NovDeep Research: A Systematic Surveyuncategorized2512.02038Nov 24, 2025~129 min
- OctLearning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasksagents2510.08002KnowledgeXLab@Shanghai AI LabOct 9, 2025score 9~110 min
- OctAgent Learning via Early Experienceagents2510.08558Oct 9, 2025~108 min
- OctDemystifying Reinforcement Learning in Agentic Reasoningagents2510.11701Oct 13, 2025score 9~127 min
- OctScaling Long-Horizon LLM Agent via Context-Foldingagents2510.11967ByteDance SeedOct 13, 2025score 8~108 min
- OctExploring Conditions for Diffusion models in Robotic Controlagents2510.15510NAVER AI LabOct 17, 2025score 2~102 min
- OctGame-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agentsagents2510.23691ByteDanceOct 27, 2025score 8~124 min
- OctWebLeaper: Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seekingagents2510.24697TongyiLabOct 28, 2025score 9~124 min
- OctAgentFold: Long-Horizon Web Agents with Proactive Context Managementagents2510.24699TongyiLabOct 28, 2025score 9~101 min
- OctTongyi DeepResearch Technical Reportagents2510.24701Oct 28, 2025~115 min
- OctMagentic Marketplace: An Open-Source Environment for Studying Agentic Marketsagents2510.25779Microsoft ResearchOct 27, 2025score 8~114 min
- OctVLA-0: Building State-of-the-Art VLAs with Zero Modificationarchitecture2510.13054NVIDIAOct 15, 2025score 6~100 min
- OctCode Aesthetics with Agentic Reward Feedbackcode2510.23272Microsoft ResearchOct 27, 2025score 9~110 min
- OctAgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesisdata2510.24695TongyiLabOct 28, 2025score 9~113 min
- OctOSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agentsevaluation2510.24563TongyiLabOct 28, 2025score 5~111 min
- OctAgentic Context Engineering: Evolving Contexts for Self-Improving Language Modelsllm-systems2510.04618Oct 6, 2025score 9~107 min
- OctUI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoningreasoning2510.20286TongyiLabOct 23, 2025score 6~107 min
- OctRepurposing Synthetic Data for Fine-grained Search Agent Supervisionrl-training2510.24694TongyiLabOct 28, 2025score 9~105 min
- OctParallelMuse: Agentic Parallel Thinking for Deep Information Seekingtraining-methods2510.24698TongyiLabOct 28, 2025score 9~120 min
- OctFrom Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priorsuncategorized2510.17439ByteDance SeedOct 20, 2025score 9~122 min
- OctLayerComposer: Multi-Human Personalized Generation via Layered Canvasvision2510.20820Snap ResearchOct 23, 2025score 3~109 min
- SepUI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learningagents2509.02544ByteDance SeedSep 2, 2025score 6~126 min
- SepThe Landscape of Agentic Reinforcement Learning for LLMs: A Surveyagents2509.02547Sep 2, 2025~127 min
- SepAstra: A Multi-Agent System for GPU Kernel Performance Optimizationagents2509.07506Sep 9, 2025~102 min
- SepTowards General Agentic Intelligence via Environment Scalingagents2509.13311Sep 16, 2025~114 min
- SepLIMI: Less is More for Agencyagents2509.17567Sep 22, 2025~104 min
- SepScaling Generalist Data-Analytic Agentsagents2509.25084QwenSep 29, 2025score 9~103 min
- SepMultiplayer Nash Preference Optimizationalignment2509.23102Sep 27, 2025~110 min
- SepRPG: A Repository Planning Graph for Unified and Scalable Codebase Generationcode2509.16198Sep 19, 2025score 10~108 min
- SepHunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generationdiffusion2509.12815Sep 16, 2025score 3~119 min
- SepVitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applicationsevaluation2509.26490LongCatSep 30, 2025score 6~136 min
- SepScaling Agents via Continual Pre-trainingtraining-methods2509.13310Sep 16, 2025~117 min
- SepSocratic-Zero : Bootstrapping Reasoning via Data-Free Agent Co-evolutiontraining-methods2509.24726alibaba-incSep 29, 2025score 9~105 min
- SepV2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughtsvision2509.18053NVIDIASep 22, 2025score 3~97 min
- AugFin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Modelsagents2508.15202Qwen DianJinAug 21, 2025score 9~111 min
- AugMemento: Fine-tuning LLM Agents without Fine-tuning LLMsagents2508.16153Aug 22, 2025score 9~114 min
- AugUniversal Deep Research: Bring Your Own Model and Strategyagents2509.00244Aug 29, 2025~92 min
- AugEvaluating, Synthesizing, and Enhancing for Customer Support Conversationdata2508.04423Qwen DianJinAug 6, 2025score 3~117 min
- AugOpen Data Synthesis For Deep Researchdata2509.00375Aug 30, 2025~139 min
- AugLiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queriesevaluation2508.15760Zoom AIAug 21, 2025score 7~104 min
- AugReportBench: Evaluating Deep Research Agents via Academic Survey Tasksevaluation2508.15804ByteDanceAug 14, 2025score 6~112 min
- AugGLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Modelsllm-systems2508.06471Zhipu / GLMAug 8, 2025~138 min
- AugWebWatcher: Breaking New Frontier of Vision-Language Deep Research Agentmultimodal2508.05748Alibaba DAMOAug 7, 2025score 9~109 min
- AugSeeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memorymultimodal2508.09736ByteDance SeedAug 13, 2025score 8~110 min
- AugSemantic IDs for Joint Generative Search and Recommendationpretraining2508.10478DoorDashAug 14, 2025score 3~123 min
- AugMolmoAct: Action Reasoning Models that can Reason in Spaceuncategorized2508.07917Aug 11, 2025~111 min
- AugFutureX: An Advanced Live Benchmark for LLM Agents in Future Predictionuncategorized2508.11987ByteDance SeedAug 16, 2025score 8~116 min
- JulThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planningagents2507.16815NVIDIAJul 22, 2025score 9~110 min
- JulA Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligenceagents2507.21046Jul 28, 2025score 10~120 min
- JulWebShaper: Agentically Data Synthesizing via Information-Seeking Formalizationdata2507.15061Alibaba DAMOJul 20, 2025score 9~109 min
- JulGemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilitiesllm-systems2507.06261Jul 7, 2025score 10~104 min
- JulGLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learningmultimodal2507.01006Zhipu / GLMJul 1, 2025score 10~109 min
- JulA Survey of Context Engineering for Large Language Modelsprompting2507.13334Jul 17, 2025~88 min
- JulEXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modesreasoning2507.11407LG EXAONEJul 15, 2025score 10~110 min
- JulScaling RL to Long Videosrl-training2507.07966Jul 10, 2025~98 min
- JulAgentic Reinforced Policy Optimizationrl-training2507.19849Jul 26, 2025~95 min
- JulCUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learningtraining-methods2507.14111Jul 18, 2025score 10~138 min
- JulHunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixelsvision2507.21809Tencent HunyuanJul 29, 2025score 3~105 min
- JunLanguage Modeling by Language Modelsarchitecture2506.20249Jun 25, 2025~92 min
- JunThe Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvementscode2506.22419Jun 27, 2025score 10~118 min
- JunMagistralreasoning2506.10910MistralJun 12, 2025~105 min
- JunGUI-Actor: Coordinate-Free Visual Grounding for GUI Agentsvision2506.03143Microsoft ResearchJun 3, 2025score 9~98 min
- MayHunyuan-Game: Industrial-grade Intelligent Game Creation Modelagents2505.14135May 20, 2025score 4~195 min
- MayMiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoderarchitecture2505.07916May 12, 2025score 2~112 min
- MayMMaDA: Multimodal Large Diffusion Language Modelsmultimodal2505.15809May 21, 2025~136 min
- MayG1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learningtraining-methods2505.13426Moonshot AIMay 19, 2025score 9~102 min
- AprPaperBench: Evaluating AI's Ability to Replicate AI Researchagents2504.01848Apr 2, 2025score 10~102 min
- AprThe AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Searchagents2504.08066Apr 10, 2025score 9~93 min
- AprKimi-VL Technical Reportmultimodal2504.07491NVIDIAApr 10, 2025~121 min
- MarPokéChamp: an Expert-level Minimax Language Agentagents2503.04094Mar 6, 2025score 4~130 min
- MarWhy Do Multi-Agent LLM Systems Fail?agents2503.13657Mar 17, 2025~133 min
- MarLarge Language Model Agent: A Survey on Methodology, Applications and Challengesagents2503.21460Mar 27, 2025score 1~89 min
- MarGemini Robotics: Bringing AI into the Physical Worldmultimodal2503.20020Mar 25, 2025score 9~134 min
- FebV2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Modelsllm-systems2502.09980NVIDIAFeb 14, 2025score 2~119 min
- FebGold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2reasoning2502.03544Feb 5, 2025score 10~127 min
- JanEvolving Deeper LLM Thinkingagents2501.09891DeepMindJan 17, 2025~119 min
- JanDeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learningrl-training2501.12948DeepSeekJan 22, 2025score 10~129 min
- JanTransformer-Squared: Self-adaptive LLMstraining-methods2501.06252Jan 9, 2025score 9~115 min
2024
25- DecInternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactionsmultimodal2412.09596InternLM / Shanghai AI LabDec 12, 2024score 6~112 min
- NovSparsing Law: Towards Large Language Models with Greater Activation Sparsityagents2411.02335Nov 4, 2024score 9~108 min
- OctLanguage Model Embeddings Can Be Sufficient for Bayesian Optimizationagents2410.10190DeepMindOct 14, 2024~125 min
- OctREAD MOREserving2410.20399Together AIOct 27, 2024~108 min
- OctMerge to Learn: Efficiently Adding Skills to Language Models with Model Mergingtraining-methods2410.12937Oct 16, 2024~116 min
- SepCan LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchersagents2409.04109Sep 6, 2024score 10~118 min
- AugDiversity Empowers Intelligence: Integrating Expertise of Software Engineering Agentsagents2408.07060SalesforceAug 13, 2024score 10~108 min
- AugAchieving Human Level Competitive Robot Table Tennisarchitecture2408.03906DeepMindAug 7, 2024score 2~120 min
- AugThe Mamba in the Llama: Distilling and Accelerating Hybrid Modelsarchitecture2408.15237Aug 27, 2024~116 min
- JulMindSearch: Mimicking Human Minds Elicits Deep AI Searcheragents2407.20183Jul 29, 2024score 9~93 min
- JulHuman-inspired Episodic Memory for Infinite Context LLMsinference-optimization2407.09450Jul 12, 2024score 9~127 min
- JulOn scalable oversight with weak LLMs judging strong LLMssafety2407.04622Jul 5, 2024~137 min
- JunMixture-of-Agents Enhances Large Language Model Capabilitiesagents2406.04692Jun 7, 2024~91 min
- JunRVT-2: Learning Precise Manipulation from Few Demonstrationsagents2406.08545NVIDIAJun 12, 2024score 2~107 min
- JunScaling Synthetic Data Creation with 1,000,000,000 Personasdata2406.20094Jun 28, 2024~88 min
- FebApproximating the Core via Iterative Coalition Samplingagents2402.03928Feb 6, 2024~100 min
- FebPIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMsagents2402.07872DeepMindFeb 12, 2024score 8~119 min
- FebA Human-Inspired Reading Agent with Gist Memory of Very Long Contextsagents2402.09727DeepMindFeb 15, 2024score 9~103 min
- FebGenie: Generative Interactive Environmentspretraining2402.15391DeepMindFeb 23, 2024score 9~111 min
- FebLearning to Learn Faster from Human Feedback with Language Model Predictive Controltraining-methods2402.11450DeepMindFeb 18, 2024score 9~104 min
- FebLAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editingno summary yetcs-hc2402.1029456 citesFeb 15, 2024score 5
- JanAutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agentsagents2401.12963DeepMindJan 23, 2024score 3~112 min
- JanSteering Language Models with Game-Theoretic Solversagents2402.01704DeepMindJan 24, 2024~116 min
- JanEAGLE: Speculative Sampling Requires Rethinking Feature Uncertaintyinference-optimization2401.15077Jan 26, 2024~133 min
- JanMeta-Prompting: Enhancing Language Models with Task-Agnostic Scaffoldingprompting2401.12954SlackJan 23, 2024~95 min
2023
14- DecGenerative agent-based modeling with actions grounded in physical, social, or digital space using Concordiaagents2312.03664DeepMindDec 6, 2023score 6~96 min
- DecReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agentagents2312.10003Dec 15, 2023score 9~107 min
- OctLearning Interactive Real-World Simulatorsagents2310.06114DeepMindOct 9, 2023~112 min
- SepRevisiting Energy Based Models as Policies: Ranking Noise Contrastive Estimation and Interpolating Energy Modelsagents2309.05803DeepMindSep 11, 2023~103 min
- SepRead Morealignment2309.08600EleutherAISep 15, 2023score 2~103 min
- SepQwen Technical Reportalignment2309.16609Stability AISep 28, 2023score 9~119 min
- AugRetroformer: Retrospective Large Language Agents with Policy Gradient Optimizationrl-training2308.02151Aug 4, 2023score 9~111 min
- JulSecrets of RLHF in Large Language Models: Part I: PPOrl-training2307.04964Jul 11, 2023~131 min
- MayUnlocking the Power of Representations in Long-term Novelty-based Explorationagents2305.01521DeepMind0 citesMay 2, 2023~114 min
- MayRead Moresafety2305.16367EleutherAIMay 25, 2023score 2~123 min
- MayAugmenting Autotelic Agents with Large Language Modelsno summary yetcs-ai2305.124874 citesMay 21, 2023score 3
- AprGenerative Agents: Interactive Simulacra of Human Behavioragents2304.03442Apr 7, 2023~119 min
- MarReflexion: Language Agents with Verbal Reinforcement Learningagents2303.11366Mar 20, 2023~108 min
- FebToolformer: Language Models Can Teach Themselves to Use Toolstraining-methods2302.04761Feb 9, 2023~129 min
2022
4- OctReAct: Synergizing Reasoning and Acting in Language Modelsagents2210.03629Qwen / Alibaba CloudOct 6, 2022~104 min
- AugMoCapAct: A Multi-Task Dataset for Simulated Humanoid Controlagents2208.07363Aug 15, 2022~122 min
- AprDo As I Can, Not As I Say: Grounding Language in Robotic Affordancesuncategorized2204.01691Apr 4, 2022~124 min
- MarLEVEN: A Large-Scale Chinese Legal Event Detection Datasetagents2203.08556Mar 16, 2022~117 min
2021
22020
12019
3- NovMuZero (Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model)rl-training1911.082651,088 citesNov 19, 2019~123 min
- AprHOList: An Environment for Machine Learning of Higher-Order Theorem Provingagents1904.0324120 citesApr 5, 2019~116 min
- JanNatural Questions: A Benchmark for Question Answering Researchrl-training1901.04473Jan 12, 2019~127 min
2018
22017
3- NovInverse Reward Designalignment1711.02827Nov 8, 2017~111 min
- OctRainbow: Combining Improvements in Deep Reinforcement Learningrl-training1710.02298428 citesOct 6, 2017~111 min
- JunFood Discovery with Uber Eats: Using Graph Learning to Power Recommendationsarchitecture1706.02216UberJun 7, 2017~114 min