213 papers
LLM Systems
0/210End-to-end systems for serving and orchestrating LLMs.
Progress0 of 210
RelatedAgents
2026
19- MayUniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsificationinference-optimization2605.06221TencentMay 7, 2026~123 min
- MayContinuous Latent Diffusion Language Modelllm-systems2605.06548ByteDance SeedMay 7, 2026~125 min
- AprLeveraging Verifier-Based Reinforcement Learning in Image Editingllm-systems2604.27505ByteDance SeedApr 30, 2026~121 min
- AprA Survey of On-Policy Distillation for Large Language Modelstraining-methods2604.00626Apr 1, 2026~136 min
- MarTimer-S1: A Billion-Scale Time Series Foundation Model with Serial Scalingllm-systems2603.04791ByteDanceMar 5, 2026score 3~108 min
- FebSemantic Search At LinkedInllm-systems2602.07309LinkedInFeb 7, 2026~139 min
- FebWhen to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoningllm-systems2602.10560ByteDance SeedFeb 11, 2026score 8~103 min
- FebK-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Modelllm-systems2602.19128Feb 22, 2026~109 min
- FebUntied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunkingllm-systems2602.21196TogetherFeb 24, 2026score 9~105 min
- FebTest-Time Training with KV Binding Is Secretly Linear Attentionllm-systems2602.21204NVIDIAFeb 24, 2026score 9~131 min
- FebERNIE 5.0 Technical Reportmultimodal2602.04705Feb 4, 2026~138 min
- FebNanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Actstraining-methods2602.13367Feb 13, 2026~100 min
- FebGLM-5: from Vibe Coding to Agentic Engineeringuncategorized2602.15763Zhipu / GLMFeb 17, 2026~148 min
- JanOver-Searching in Search-Augmented Large Language Modelsagents2601.05503AppleJan 9, 2026score 3~131 min
- JanComputer Environments Elicit General Agentic Intelligence in LLMsagents2601.16206Microsoft ResearchJan 22, 2026score 5~94 min
- Jan"TODO: Fix the Mess Gemini Created": Towards Understanding GenAI-Induced Self-Admitted Technical Debtcode2601.07786Jan 12, 2026score 2~97 min
- JanK-EXAONE Technical Reportmoe2601.01739LG EXAONEJan 5, 2026~112 min
- JanMegaFlow: Large-Scale Distributed Orchestration System for the Agentic Eraserving2601.07526QwenJan 12, 2026score 6~109 min
- JanTranslateGemma Technical Reporttraining-methods2601.09012Google ResearchJan 13, 2026~107 min
2025
77- DecMemory in the Age of AI Agentsagents2512.13564Dec 15, 2025score 9~101 min
- DecWeb World Modelsarchitecture2512.23676Dec 29, 2025~112 min
- DecConfucius Code Agent: Scalable Agent Scaffolding for Real-World Codebasescode2512.10398Meta ResearchDec 11, 2025score 9~112 min
- DecDataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AIdata2512.16676Peking UniversityDec 18, 2025score 9~98 min
- DecTurboDiffusion: Accelerating Video Diffusion Models by 100-200 Timesinference-optimization2512.16093University of California, BerkeleyDec 18, 2025score 10~99 min
- DecDeepSeek-V3.2: Pushing the Frontier of Open Large Language Modelsllm-systems2512.02556DeepSeekDec 2, 2025~127 min
- DecBolmo: Byteifying the Next Generation of Language Modelsllm-systems2512.15586Allen Institute for AIDec 17, 2025~120 min
- DecInsight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Languagemultimodal2512.11251Dec 12, 2025~108 min
- DecOlmo 3pretraining2512.13961Allen Institute for AIDec 15, 2025~105 min
- DecEvaluating Parameter Efficient Methods for RLVRrl-training2512.23165DeepSeekDec 29, 2025~105 min
- NovLlama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasksllm-systems2511.07025NVIDIANov 10, 2025score 9~99 min
- NovKitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boostllm-systems2511.18643Together AINov 23, 2025~108 min
- NovToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestrationtraining-methods2511.21689NVIDIANov 26, 2025score 9~106 min
- NovDeep Research: A Systematic Surveyuncategorized2512.02038Nov 24, 2025~129 min
- OctEvery Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoningarchitecture2510.19338Ant GroupOct 22, 2025score 10~141 min
- OctAgentic Context Engineering: Evolving Contexts for Self-Improving Language Modelsllm-systems2510.04618Oct 6, 2025score 9~107 min
- OctTC-LoRA: Temporally Modulated Conditional LoRA for Adaptive Diffusion Controlllm-systems2510.09561NVIDIAOct 10, 2025score 3~86 min
- OctWhen to Ensemble: Identifying Token-Level Points for Stable and Fast LLM Ensemblingllm-systems2510.15346KAIST AIOct 17, 2025score 5~120 min
- OctDeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Searchllm-systems2510.12801AppleOct 14, 2025~14 min
- OctThe Art of Scaling Reinforcement Learning Compute for LLMsrl-training2510.13786Oct 15, 2025~139 min
- OctRDMA Point-to-Point Communication for LLM Systemsserving2510.27656Oct 31, 2025~117 min
- OctEfficient Long-context Language Model Training by Core Attention Disaggregationtraining-methods2510.18121Oct 20, 2025score 10~105 min
- SepInfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptationarchitecture2509.24663OpenBMBSep 29, 2025score 9~116 min
- SepRPG: A Repository Planning Graph for Unified and Scalable Codebase Generationcode2509.16198Sep 19, 2025score 10~108 min
- SepVitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applicationsevaluation2509.26490LongCatSep 30, 2025score 6~136 min
- SepHunyuan-MT Technical Reportllm-systems2509.05209Tencent HunyuanSep 5, 2025score 3~99 min
- SepLongLive: Real-time Interactive Long Video Generationllm-systems2509.22622NVIDIASep 26, 2025score 6~111 min
- SepPixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Imagesllm-systems2509.25185Sep 29, 2025~105 min
- SepApertus: Democratizing Open and Compliant LLMs for Global Language Environmentspretraining2509.14233Swiss AI InitiativeSep 17, 2025score 9~107 min
- SepA Survey of Reinforcement Learning for Large Reasoning Modelsreasoning2509.08827Sep 10, 2025~123 min
- SepSharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharingtraining-methods2509.08721DeepSeekSep 10, 2025score 10~108 min
- SepBenefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspectivetraining-methods2509.22613Microsoft ResearchSep 26, 2025score 9~95 min
- AugUniversal Deep Research: Bring Your Own Model and Strategyagents2509.00244Aug 29, 2025~92 min
- AugA Survey on Diffusion Language Modelsllm-systems2508.10875Aug 14, 2025~126 min
- AugInternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiencymultimodal2508.18265Aug 25, 2025score 10~107 min
- AugSemantic IDs for Joint Generative Search and Recommendationpretraining2508.10478DoorDashAug 14, 2025score 3~123 min
- AugNVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Modelreasoning2508.14444NVIDIAAug 20, 2025score 10~119 min
- AugTaming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inferenceserving2508.19559Aug 27, 2025score 9~115 min
- AugMolmoAct: Action Reasoning Models that can Reason in Spaceuncategorized2508.07917Aug 11, 2025~111 min
- JulDecoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generationarchitecture2507.06607Microsoft ResearchJul 9, 2025score 9~109 min
- JulFalcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performancearchitecture2507.22448Jul 30, 2025score 10~132 min
- JulGemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilitiesllm-systems2507.06261Jul 7, 2025score 10~104 min
- JulThe Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithmllm-systems2507.18553 IST Austria Distributed Algorithms and Systems LabJul 24, 2025score 9~149 min
- JulA Survey of Context Engineering for Large Language Modelsprompting2507.13334Jul 17, 2025~88 min
- JulEXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modesreasoning2507.11407LG EXAONEJul 15, 2025score 10~110 min
- JulAXLearn: Modular Large Model Training on Heterogeneous Infrastructuretraining-methods2507.05411AppleJul 7, 2025score 8~94 min
- JunMiniCPM4: Ultra-Efficient LLMs on End Devicesllm-systems2506.07900OpenBMBJun 9, 2025score 9~122 min
- JunGenRecal: Generation after Recalibration from Large to Small Vision-Language Modelsllm-systems2506.15681NVIDIAJun 18, 2025score 9~102 min
- JunMiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attentionreasoning2506.13585DeepSeekJun 16, 2025score 9~132 min
- JunBeyond the Buzz: A Pragmatic Take on Inference Disaggregationserving2506.05508Jun 5, 2025~100 min
- JunPerformance Prediction for Large Systems via Text-to-Text Regressionserving2506.21718DeepMindJun 26, 2025score 3~109 min
- MayEfficientLLM: Efficiency in Large Language Modelsevaluation2505.13840May 20, 2025score 10~104 min
- MayQwen3 Technical Reportllm-systems2505.09388Qwen / Alibaba CloudMay 14, 2025~128 min
- MayInsights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architecturesmoe2505.09343NVIDIAMay 14, 2025score 10~111 min
- MayPrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applicationsserving2505.07203May 12, 2025~105 min
- AprSleep-time Compute: Beyond Inference Scaling at Test-timeinference-optimization2504.13171Apr 17, 2025~127 min
- AprBitNet b1.58 2B4T Technical Reportllm-systems2504.12285Microsoft ResearchApr 16, 2025score 9~106 min
- AprKimi-VL Technical Reportmultimodal2504.07491NVIDIAApr 10, 2025~121 min
- AprInternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Modelsmultimodal2504.10479Apr 14, 2025score 10~101 min
- AprReinforcement Learning from Human Feedbackrl-training2504.12501Apr 16, 2025~84 min
- AprDoes Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?rl-training2504.13837Apr 18, 2025score 9~107 min
- AprInference-Time Scaling for Generalist Reward Modelingscaling-laws2504.02495Apr 3, 2025~101 min
- MarLarge Language Model Agent: A Survey on Methodology, Applications and Challengesagents2503.21460Mar 27, 2025score 1~89 min
- MarHolistically Evaluating the Environmental Impact of Creating Language Modelsevaluation2503.05804Mar 3, 2025~121 min
- MarRankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluationevaluation2503.19092Mar 24, 2025~123 min
- MarGemini Embedding: Generalizable Embeddings from Geminillm-systems2503.07891DeepMindMar 10, 2025~99 min
- MarA Comprehensive Survey on Long Context Language Modelingllm-systems2503.17407Mar 20, 2025~121 min
- MarQwen2.5-Omni Technical Reportmultimodal2503.20215Qwen / Alibaba CloudMar 26, 2025~106 min
- FebKernelBench: Can LLMs Write Efficient GPU Kernels?evaluation2502.10517Feb 14, 2025~112 min
- FebTransMLA: Multi-Head Latent Attention Is All You Needinference-optimization2502.07864DeepSeekFeb 11, 2025score 10~101 min
- FebV2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Modelsllm-systems2502.09980NVIDIAFeb 14, 2025score 2~119 min
- FebCompetitive Programming with Large Reasoning Modelsrl-training2502.06807OpenAIFeb 3, 2025~127 min
- JanEvolving Deeper LLM Thinkingagents2501.09891DeepMindJan 17, 2025~119 min
- JanMiniMax-01: Scaling Foundation Models with Lightning Attentionarchitecture2501.08313MiniMaxJan 14, 2025~123 min
- JanHunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generationllm-systems2501.12202Tencent HunyuanJan 21, 2025score 6~119 min
- JanSa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videosmultimodal2501.04001ByteDance SeedJan 7, 2025score 8~109 min
- Jans1: Simple test-time scalingreasoning2501.19393OpenAIJan 31, 2025~112 min
2024
59- DecFLEX ATTENTION: A PROGRAMMING MODEL FOR GENERATING OPTIMIZED ATTENTION KERNELSarchitecture2412.05496Dec 7, 2024~116 min
- DecInternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactionsmultimodal2412.09596InternLM / Shanghai AI LabDec 12, 2024score 6~112 min
- DecEXAONE 3.5: Series of Large Language Models for Real-world Use Casespretraining2412.04862LG EXAONEDec 6, 2024~97 min
- DecQwen2.5 Technical Reportpretraining2412.15115Qwen / Alibaba CloudDec 19, 2024~109 min
- NovMARCONI: PREFIX CACHING FOR THE ERA OF HYBRID LLMSserving2411.19379Nov 28, 2024~99 min
- NovMARS: Unleashing the Power of Variance Reduction for Training Large Modelstraining-methods2411.10438Nov 15, 2024score 9~118 min
- NovTülu 3: Pushing Frontiers in Open Language Model Post-Trainingtraining-methods2411.15124Allen Institute for AINov 22, 2024~132 min
- OctRevisiting Reliability in Large-Scale Machine Learning Research Clustersdistributed-training2410.21680Oct 29, 2024~112 min
- OctEoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximationllm-systems2410.21271NVIDIAOct 28, 2024score 6~113 min
- OctMulti-Field Adaptive Retrievalretrieval2410.20056Oct 26, 2024~108 min
- OctOrca: A Distributed Serving System for Transformer-Based Generative Modelsserving2410.17840Oct 23, 2024~108 min
- OctLiger Kernel: Efficient Triton Kernels for LLM Trainingtraining-methods2410.10989Oct 14, 2024~124 min
- SepRetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrievalinference-optimization2409.10516Sep 16, 2024score 8~105 min
- SepEuroLLM: Multilingual Language Models for Europellm-systems2409.16235Sep 24, 2024score 9~103 min
- SepMaskLLM: Learnable Semi-Structured Sparsity for Large Language Modelsllm-systems2409.17481Sep 26, 2024score 9~118 min
- SepMolmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Modelsmultimodal2409.17146Allen Institute for AISep 25, 2024score 9~121 min
- SepAttention Heads of Large Language Models: A Surveyreasoning2409.03752Sep 5, 2024~127 min
- AugHermes 3 Technical Reportalignment2408.11857Aug 15, 2024~105 min
- AugEXAONE 3.0 7.8B Instruction Tuned Language Modelllm-systems2408.03541LG EXAONEAug 7, 2024score 9~114 min
- AugLLM Pruning and Distillation in Practice: The Minitron Approachtraining-methods2408.11796Aug 21, 2024~110 min
- AugMulti-Layer Transformers Gradient Can be Approximated in Almost Linear Timetraining-methods2408.13233Aug 23, 2024score 9~100 min
- AugBuilding and better understanding vision-language models: insights and future directionsvision2408.12637Aug 22, 2024score 10~96 min
- JulMindSearch: Mimicking Human Minds Elicits Deep AI Searcheragents2407.20183Jul 29, 2024score 9~93 min
- JulLearning to (Learn at Test Time): RNNs with Expressive Hidden Statesarchitecture2407.04620Jul 5, 2024~108 min
- JulQwen2 Technical Reportllm-systems2407.10671Qwen / Alibaba CloudJul 15, 2024~114 min
- JulGemma 2: Improving Open Language Models at a Practical Sizepretraining2408.00118Google ResearchJul 31, 2024~106 min
- JunMixture-of-Agents Enhances Large Language Model Capabilitiesagents2406.04692Jun 7, 2024~91 min
- JunInfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Managementinference-optimization2406.19707Jun 28, 2024~107 min
- JunChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Toolsllm-systems2406.12793Zhipu / GLMJun 18, 2024score 9~108 min
- JunLlumnix: Dynamic Scheduling for Large Language Model Servingserving2406.03243Jun 5, 2024~113 min
- JunPowerInfer-2: Fast Large Language Model Inference on a Smartphoneserving2406.06282Jun 10, 2024score 8~108 min
- JunMooncake: A KVCache-centric Disaggregated Architecture for LLM Servingserving2407.00079Jun 24, 2024~127 min
- MayNearest Neighbor Speculative Decoding for LLM Generation and Attributionllm-systems2405.19325May 29, 2024score 9~103 min
- MayParrot: Efficient Serving of LLM-based Applications with Semantic Variableserving2405.19888May 30, 2024score 9~98 min
- MayRLHF Workflow: From Reward Modeling to Online RLHFtraining-methods2405.07863SalesforceMay 13, 2024score 10~96 min
- AprReplacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Modelsevaluation2404.18796Apr 29, 2024score 9~103 min
- AprLLM2Vec: Large Language Models Are Secretly Powerful Text Encodersllm-systems2404.05961Apr 9, 2024~105 min
- AprPhi-3 Technical Report: A Highly Capable Language Model Locally on Your Phonepretraining2404.14219Microsoft ResearchApr 22, 2024score 9~114 min
- MarQuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsinference-optimization2404.00456Mar 30, 2024~106 min
- MarGemma: Open Models Based on Gemini Research and Technologyllm-systems2403.08295Google ResearchMar 13, 2024score 9~105 min
- MarTaming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serveserving2403.02310Mar 4, 2024~86 min
- MarLoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPUtraining-methods2403.06504Mar 11, 2024score 9~120 min
- FebOLMo: Accelerating the Science of Language Modelsagents2402.00838Allen Institute for AIFeb 1, 2024score 10~113 min
- FebMultilingual E5 Text Embeddings: A Technical Reportcontext-optimization2402.05672Feb 8, 2024~118 min
- FebThe Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicryllm-systems2402.04347Feb 6, 2024score 9~113 min
- FebLiRank: Industrial Large Scale Ranking Models at LinkedInllm-systems2402.06859Feb 10, 2024score 3~133 min
- FebWorld Model on Million-Length Video And Language With Blockwise RingAttentionllm-systems2402.08268Feb 13, 2024score 9~106 min
- FebNemotron-4 15B Technical Reportpretraining2402.16819Feb 26, 2024~106 min
- FebThe Era of 1-bit LLMs: All Large Language Models are in 1.58 Bitspretraining2402.17764DropboxFeb 27, 2024~118 min
- FebHydragen: High-Throughput LLM Inference with Shared Prefixesserving2402.05099Feb 7, 2024score 9~105 min
- FebChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partitionserving2402.15220Feb 23, 2024score 10~108 min
- FebMegaScale: Scaling Large Language Model Training to More Than 10,000 GPUstraining-methods2402.15627Feb 23, 2024score 9~116 min
- FebScaling Up LLM Reviews for Google Ads Content Moderationno summary yetcs-ir2402.1459010 citesFeb 7, 2024score 8
- JanDolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Researchdata2402.00159Allen Institute for AIJan 31, 2024score 10~137 min
- JanSoaring from 4K to 400K: Extending LLM's Context with Activation Beaconllm-systems2401.03462Jan 7, 2024score 9~109 min
- JanInference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloadsllm-systems2401.11181Jan 20, 2024~115 min
- JanThe Case for Co-Designing Model Architectures with Hardwarellm-systems2401.14489Jan 25, 2024~103 min
- JanDistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Servingserving2401.09670Jan 18, 2024~102 min
- JanGenerative Expressive Robot Behaviors using Large Language Modelsno summary yetcs-ro2401.1467349 citesJan 26, 2024score 3
2023
35- DecSGLang: Efficient Execution of Structured Language Model Programsagents2312.07104Dec 12, 2023~101 min
- NovFlashDecoding++: Faster Large Language Model Inference on GPUsinference-optimization2311.01282Nov 2, 2023score 9~129 min
- NovRelax: Composable Abstractions for End-to-End Dynamic Machine Learningserving2311.02103Nov 1, 2023~107 min
- NovS-LORA: SERVING THOUSANDS OF CONCURRENT LORA ADAPTERSserving2311.03285Nov 6, 2023~106 min
- NovCamels in a Changing Climate: Enhancing LM Adaptation with Tulu 2training-methods2311.10702Allen Institute for AINov 17, 2023score 10~106 min
- OctRepelling Random Walksllm-systems2310.04854DeepMindOct 7, 2023~115 min
- OctMistral 7Bpretraining2310.06825MistralOct 10, 2023~97 min
- OctDetecting Pretraining Data from Large Language Modelssafety2310.16789Oct 25, 2023score 9~111 min
- OctPUNICA: MULTI-TENANT LORA SERVINGserving2310.18547Oct 28, 2023~100 min
- SepRLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackalignment2309.00267Sep 1, 2023score 9~119 min
- SepQwen Technical Reportalignment2309.16609Stability AISep 28, 2023score 9~119 min
- SepEfficient Memory Management for Large Language Model Serving with PagedAttentionserving2309.06180Sep 12, 2023score 10~122 min
- SepLongLoRA: Efficient Fine-tuning of Long-Context Large Language Modelstraining-methods2309.12307Sep 21, 2023score 9~93 min
- AugAccelerating LLM Inference with Staged Speculative Decodinginference-optimization2308.04623Aug 8, 2023~104 min
- AugOmniQuant: Omnidirectionally Calibrated Quantization for Large Language Modelsinference-optimization2308.13137Aug 25, 2023score 9~100 min
- AugDeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scalesllm-systems2308.01320Aug 2, 2023score 9~96 min
- AugSARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefillsllm-systems2308.16369Aug 31, 2023~105 min
- AugFlexible Isosurface Extraction for Gradient-Based Mesh Optimizationno summary yetcs-gr2308.0537173 citesAug 10, 2023score 2
- JulLlama 2: Open Foundation and Fine-Tuned Chat Modelsalignment2307.09288NVIDIAJul 18, 2023~120 min
- JulLongNet: Scaling Transformers to 1,000,000,000 Tokensarchitecture2307.02486Jul 5, 2023~123 min
- JulHow is ChatGPT's behavior changing over time?evaluation2307.09009Jul 18, 2023score 9~107 min
- JulFocused Transformer: Contrastive Training for Context Scalingllm-systems2307.03170Jul 6, 2023score 9~116 min
- JulScaling TransNormer to 175 Billion Parametersllm-systems2307.14995Jul 27, 2023score 9~114 min
- JulSkeleton-of-Thought: Large Language Models Can Do Parallel Decodingllm-systems2307.15337Jul 28, 2023score 9~96 min
- JunBlock-State Transformerarchitecture2306.09539Jun 15, 2023score 9~108 min
- JunRead Moreinference-optimization2306.17806EleutherAIJun 30, 2023~115 min
- JunAugmenting Language Models with Long-Term Memoryllm-systems2306.07174Jun 12, 2023score 9~107 min
- JunBinary and Ternary Natural Language Generationlow-precision2306.01841Jun 2, 2023score 9~115 min
- MayScaling Catalog Attribute Extraction with Multi-modal LLMsllm-systems2305.05176InstacartMay 9, 2023score 8~97 min
- MayPaLM 2 Technical Reportpretraining2305.10403May 17, 2023~108 min
- MaySpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verificationserving2305.09781May 16, 2023~113 min
- AprGenerative Agents: Interactive Simulacra of Human Behavioragents2304.03442Apr 7, 2023~119 min
- AprInstruction Tuning with GPT-4training-methods2304.03277Allen Institute for AIApr 6, 2023~113 min
- MarREAD MOREinference-optimization2303.06865Together AIMar 13, 2023~115 min
- MarLLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attentiontraining-methods2303.16199Mar 28, 2023~100 min
2022
6- NovHyperTuning: Toward Adapting Large Language Models without Back-propagationllm-systems2211.12485EleutherAINov 22, 2022~111 min
- NovRead Moretraining-methods2211.05100EleutherAINov 9, 2022~119 min
- OctScaling Instruction-Finetuned Language Modelsalignment2210.11416SalesforceOct 20, 2022~105 min
- MayOPT: Open Pre-trained Transformer Language Modelspretraining2205.01068Meta AI / FAIRMay 2, 2022~125 min
- AprGPT-NeoX-20B: An Open-Source Autoregressive Language Modelllm-systems2204.06745Stability AIApr 14, 2022~111 min
- FebRethinking the Role of Demonstrations: What Makes In-Context Learning Work?prompting2202.12837Feb 25, 2022~119 min
2021
6- OctColossal-AI: A Unified Deep Learning System For Large-Scale Parallel Trainingdistributed-training2110.14883Oct 28, 2021~114 min
- JunLoRA: Low-Rank Adaptation of Large Language Modelstraining-methods2106.09685Jun 17, 2021~91 min
- JunBitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-modelstraining-methods2106.10199Jun 18, 2021~102 min
- AprEfficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LMdistributed-training2104.04473Apr 9, 2021~101 min
- FebZero-Shot Text-to-Image Generationarchitecture2102.120921,134 citesFeb 24, 2021~111 min
- FebTeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Modelsdistributed-training2102.07988Feb 16, 2021~109 min
2020
22019
22018
12017
22016
3- SepGoogle's Neural Machine Translation System: Bridging the Gap between Human and Machine Translationarchitecture1609.08144Sep 26, 2016~114 min
- MayTensorFlow: A system for large-scale machine learningllm-systems1605.086958,823 citesMay 27, 2016~115 min
- MarTensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systemsdistributed-training1603.04467Mar 14, 2016~107 min