2026
31- Mayoptimize_anything: A Universal API for Optimizing any Text Parameterarchitecture2605.19633May 19, 2026~111 min
- MayKernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernelscode2605.04956Tsinghua UniversityMay 6, 2026~132 min
- AprMemory Transfer Learning: How Memories are Transferred Across Domains in Coding Agentsagents2604.14004KAIST AIApr 15, 2026score 8~113 min
- AprMM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generationagents2604.15309Microsoft ResearchApr 16, 2026score 3~94 min
- AprInCoder-32B-Thinking: Industrial Code World Model for Thinkingcode2604.03144Apr 3, 2026score 9~100 min
- AprHierarchical SVG Tokenization: Learning Compact Visual Programs for Scalable Vector Graphics Modelingcode2604.05072Tencent HunyuanApr 6, 2026score 3~120 min
- AprLightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillationtraining-methods2604.13010NVIDIAApr 14, 2026~132 min
- MarPOLCA: Stochastic Generative Optimization with LLMcode2603.14769DeepmindMar 16, 2026score 9~113 min
- MarKernel-Smith: A Unified Recipe for Evolutionary Kernel Optimizationcode2603.28342NVIDIAMar 30, 2026~139 min
- MarScaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problemsdata2603.07779Microsoft ResearchMar 8, 2026score 9~118 min
- MarRealChart2Code: Advancing Chart-to-Code Generation with Real Data and Multi-Task Evaluationevaluation2603.25804QwenMar 26, 2026score 4~89 min
- MarCodePercept: Code-Grounded Visual STEM Perception for MLLMsmultimodal2603.10757QwenMar 11, 2026score 9~103 min
- MarBreaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Modelsrl-training2603.07777Microsoft ResearchMar 8, 2026score 9~119 min
- MarCode-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Modelsrl-training2603.10098DeepmindMar 10, 2026score 9~109 min
- FebSWE-Universe: Scale Real-World Verifiable Environments to Millionsagents2602.02361QwenFeb 2, 2026score 4~109 min
- FebDr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generationscode2602.05885Feb 5, 2026~111 min
- FebK-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Modelllm-systems2602.19128Feb 22, 2026~109 min
- FebAccelerating Scientific Research with Gemini: Case Studies and Common Techniquesreasoning2602.03837Feb 3, 2026score 2~123 min
- FebThinking with Drafting: Optical Decompression via Logical Reconstructionreasoning2602.11731ByteDanceFeb 12, 2026score 3~102 min
- FebNanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Actstraining-methods2602.13367Feb 13, 2026~100 min
- FebCUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generationtraining-methods2602.24286ByteDance SeedFeb 27, 2026score 9~110 min
- FebQwen3-Coder-Next Technical Reporttraining-methods2603.00729QwenFeb 28, 2026score 9~123 min
- FebGLM-5: from Vibe Coding to Agentic Engineeringuncategorized2602.15763Zhipu / GLMFeb 17, 2026~148 min
- JanMEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineeringagents2601.22859ernie-researchJan 30, 2026score 5~120 min
- JanSAMTok: Representing Any Mask with Two Wordsarchitecture2601.16093ByteDanceJan 22, 2026score 9~122 min
- JanX-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Testscode2601.06953Microsoft ResearchJan 11, 2026score 8~97 min
- Jan"TODO: Fix the Mess Gemini Created": Towards Understanding GenAI-Induced Self-Admitted Technical Debtcode2601.07786Jan 12, 2026score 2~97 min
- JanStable-DiffCoder: Pushing the Frontier of Code Diffusion Large Language Modelcode2601.15892ByteDance SeedJan 22, 2026score 8~142 min
- JanSWE-Pruner: Self-Adaptive Context Pruning for Coding Agentscode2601.16746ByteDanceJan 23, 2026score 8~126 min
- JanAACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Contextcode2601.19494AoneJan 27, 2026score 3~111 min
- JanHow AI Impacts Skill Formationrl-training2601.20245AnthropicJan 28, 2026score 1~107 min
2025
22- DecKernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Metaagents2512.23236Dec 29, 2025~128 min
- DecConfucius Code Agent: Scalable Agent Scaffolding for Real-World Codebasescode2512.10398Meta ResearchDec 11, 2025score 9~112 min
- DecDataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AIdata2512.16676Peking UniversityDec 18, 2025score 9~98 min
- DecDAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycleevaluation2512.04324ByteDance SeedDec 3, 2025score 8~111 min
- NovCodeClash: Benchmarking Goal-Oriented Software Engineeringcode2511.00839Nov 2, 2025~100 min
- NovAccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimizationcode2511.15915Stanford UniversityNov 19, 2025~115 min
- NovDRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generationtraining-methods2511.06307OpenAINov 9, 2025score 8~114 min
- OctLongCodeZip: Compress Long Context for Code Language Modelscode2510.00446Oct 1, 2025~119 min
- OctCode Aesthetics with Agentic Reward Feedbackcode2510.23272Microsoft ResearchOct 27, 2025score 9~110 min
- OctJanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligencecode2510.23538InternLM / Shanghai AI LabOct 27, 2025score 6~97 min
- OctVibe Checker: Aligning Code Evaluation with Human Preferenceevaluation2510.07315Oct 8, 2025~114 min
- SepAstra: A Multi-Agent System for GPU Kernel Performance Optimizationagents2509.07506Sep 9, 2025~102 min
- SepRPG: A Repository Planning Graph for Unified and Scalable Codebase Generationcode2509.16198Sep 19, 2025score 10~108 min
- AugSeed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inferencecode2508.02193Aug 4, 2025score 9~115 min
- AugGLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Modelsllm-systems2508.06471Zhipu / GLMAug 8, 2025~138 min
- JulSWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?code2507.12415Jul 16, 2025score 10~126 min
- JulCUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learningtraining-methods2507.14111Jul 18, 2025score 10~138 min
- MayCASS: Nvidia to AMD Transpilation with Data, Models, and Benchmarkcode2505.16968NVIDIAMay 22, 2025score 6~112 min
- AprMulti-SWE-bench: A Multilingual Benchmark for Issue Resolvingevaluation2504.02605ByteDance SeedApr 3, 2025score 9~99 min
- FebKernelBench: Can LLMs Write Efficient GPU Kernels?evaluation2502.10517Feb 14, 2025~112 min
- FebCompetitive Programming with Large Reasoning Modelsrl-training2502.06807OpenAIFeb 3, 2025~127 min
- FebSWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolutiontraining-methods2502.18449DeepSeekFeb 25, 2025score 9~116 min
2024
12- NovOPENCODER: THE OPEN COOKBOOK FOR TOP-TIER CODE LARGE LANGUAGE MODELScode2411.04905Nov 7, 2024~110 min
- SepQwen2.5-Coder Technical Reportcode2409.12186Qwen / Alibaba CloudSep 18, 2024~110 min
- SepA Case Study of Web App Coding with OpenAI Reasoning Modelscode2409.13773Sep 19, 2024score 8~108 min
- SepCoffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Coderl-training2409.19715Sep 29, 2024score 8~111 min
- AugDiversity Empowers Intelligence: Integrating Expertise of Software Engineering Agentsagents2408.07060SalesforceAug 13, 2024score 10~108 min
- JulScaling Granite Code Models to 128K Contexttraining-methods2407.13739IBM ResearchJul 18, 2024score 8~113 min
- JunDataComp-LM: In Search of the Next Generation of Training Sets for Language Modelsdata2406.11794AppleJun 17, 2024score 10~93 min
- JunDeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligencepretraining2406.11931DeepSeekJun 17, 2024score 10~11 min
- MarView publicationcode2403.13839AppleMar 14, 2024~98 min
- FebStarCoder 2 and The Stack v2: The Next Generationcode2402.19173RobloxFeb 29, 2024score 9~111 min
- JanE3x: $\mathrm{E}(3)$-Equivariant Deep Learning Made Easycode2401.07595DeepMindJan 15, 2024~86 min
- JanDeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligencecode2401.14196DeepSeekJan 25, 2024score 10~96 min
2023
9- DecMagicoder: Source Code Is All You Needcode2312.02120Dec 4, 2023score 9~82 min
- DecLLM360: Towards Fully Transparent Open-Source LLMspretraining2312.06550Dec 11, 2023score 9~104 min
- DecBeyond Human Data: Scaling Self-Training for Problem-Solving with Language Modelstraining-methods2312.06585Dec 11, 2023score 10~116 min
- NovCamels in a Changing Climate: Enhancing LM Adaptation with Tulu 2training-methods2311.10702Allen Institute for AINov 17, 2023score 10~106 min
- SepQwen Technical Reportalignment2309.16609Stability AISep 28, 2023score 9~119 min
- SepLarge Language Models for Compiler Optimizationcode2309.07062Sep 11, 2023~112 min
- JulHow is ChatGPT's behavior changing over time?evaluation2307.09009Jul 18, 2023score 9~107 min
- JunTextbooks Are All You Needdata2306.11644Microsoft ResearchJun 20, 2023~112 min
- MayStarCoder: may the source be with you!code2305.06161Stability AIMay 9, 2023score 8~124 min