Tool use & function calling
14 papers in this thread, across 7 domains.
Progress0 of 11
- 2026Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agentsno summary yetcs lg2606.12634Amazon0 citesJun 10, 2026cs lg
- 2026When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agentsno summary yetcs cl2606.23937Amazon0 citesJun 22, 2026cs cl
- 2026Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool UseAlignment2603.03205Microsoft ResearchMar 3, 2026score 8~111 minAlignment
- 2026FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context ProtocolEvaluation2603.24943Qwen DianJinMar 26, 2026score 3~111 minEvaluation
- 2026In-Context Reinforcement Learning for Tool Use in Large Language ModelsTraining Methods2603.08068National University of SingaporeMar 9, 2026~114 minTraining Methods
- 2026ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent TrainingAgents2602.06820LongCatFeb 6, 2026score 9~103 minAgents
- 2026D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool UseTraining Methods2602.02160alibaba-incFeb 2, 2026score 9~101 minTraining Methods
- 2026Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from TextAgents2601.10355LongCatJan 15, 2026score 9~111 minAgents
- 2025MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented EnvironmentsEvaluation2512.19432TongyiLabDec 22, 2025score 6~122 minEvaluation
- 2025ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual ReasoningRL Training2512.05111Intern Large ModelsDec 4, 2025score 8~102 minRL Training
- 2025OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use AgentsEvaluation2510.24563TongyiLabOct 28, 2025score 5~111 minEvaluation
- 2025LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging QueriesEvaluation2508.15760Zoom AIAug 21, 2025score 7~104 minEvaluation
- 2023Toolformer: Language Models Can Teach Themselves to Use ToolsTraining Methods2302.04761Feb 9, 2023~129 minTraining Methods
- 2019Emergent Tool Use From Multi-Agent Autocurriculano summary yetcs lg1909.07528OpenAI337 citesSep 17, 2019cs lg