2026
43- MayCovering Human Action Space for Computer Use: Data Synthesis and Benchmarkdata2605.12501MicrosoftMay 12, 2026~108 min
- MayPlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Modelsdata2605.20873Tencent HunyuanMay 20, 2026~120 min
- MayLens: Rethinking Training Efficiency for Foundational Text-to-Image Modelsdiffusion2605.21573Microsoft ResearchMay 20, 2026~117 min
- AprToward Scalable Terminal Task Synthesis via Skill Graphsagents2604.25727Tencent HunyuanApr 28, 2026~130 min
- AprSynthetic Computers at Scale for Long-Horizon Productivity Simulationagents2604.28181MicrosoftApr 30, 2026~128 min
- AprMolmoWeb: Open Visual Web Agent and Open Data for the Open Webllm-systems2604.08516Apr 9, 2026~100 min
- AprAudio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Musicmultimodal2604.10905NVIDIAApr 13, 2026score 9~106 min
- AprTowards Autonomous Mechanistic Reasoning in Virtual Cellsreasoning2604.11661KAIST AIApr 13, 2026score 3~115 min
- AprDual-View Training for Instruction-Following Information Retrievalretrieval2604.18845SnowflakeApr 20, 2026~95 min
- AprWildDet3D: Scaling Promptable 3D Detection in the Wildvision2604.08626Allen Institute for AIApr 9, 2026score 3~124 min
- MarScaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problemsdata2603.07779Microsoft ResearchMar 8, 2026score 9~118 min
- MarSommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Modelsdata2603.25750KAIST AIMar 20, 2026score 6~111 min
- MarTimer-S1: A Billion-Scale Time Series Foundation Model with Serial Scalingllm-systems2603.04791ByteDanceMar 5, 2026score 3~108 min
- MarUnderstanding by Reconstruction: Reversing the Software Development Process for LLM Pretrainingpretraining2603.11103ByteDance SeedMar 11, 2026score 9~111 min
- MardaVinci-LLM:Towards the Science of Pretrainingpretraining2603.27164SII-GAIRMar 28, 2026score 9~110 min
- MarHopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoningreasoning2603.17024QwenMar 17, 2026score 9~135 min
- MarHiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Imagesvision2603.02210ByteDanceMar 2, 2026score 3~97 min
- FebSEAD: Self-Evolving Agent for Multi-Turn Service Dialogueagents2602.03548meituanFeb 3, 2026score 5~110 min
- FebScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Trainingagents2602.06820LongCatFeb 6, 2026score 9~103 min
- FebAgent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learningagents2602.10090SnowflakeFeb 10, 2026score 9~111 min
- FebSAGE: Scalable Agentic 3D Scene Generation for Embodied AIagents2602.10116NVIDIAFeb 10, 2026score 2~115 min
- FebWebWorld: A Large-Scale World Model for Web Agent Trainingagents2602.14721Qwen / Alibaba CloudFeb 16, 2026score 9~101 min
- FebPrivasis: Synthesizing the Largest "Public" Private Dataset from Scratchdata2602.03183NVIDIAFeb 3, 2026score 2~109 min
- FebHY3D-Bench: Generation of 3D Assetsdata2602.03907Tencent HunyuanFeb 3, 2026score 2~113 min
- FebOPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iterationdata2602.05400QwenFeb 5, 2026score 7~121 min
- FebData Science and Technology Towards AGI Part I: Tiered Data Managementdata2602.09003OpenBMBFeb 9, 2026score 8~115 min
- FebComposition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Modelsdata2602.12036Tencent HunyuanFeb 12, 2026score 5~100 min
- FebRoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learningdata2602.18742KAIST AIFeb 21, 2026score 2~101 min
- FebOn Data Engineering for Scaling LLM Terminal Capabilitiesdata2602.21193NVIDIAFeb 24, 2026score 3~102 min
- FebHow2Everything: Mining the Web for How-To Procedures to Evaluate and Improve LLMsevaluation2602.08808Ai2Feb 9, 2026score 3~131 min
- FebPrincipled Synthetic Data Enables the First Scaling Laws for LLMs in Recommendationllm-systems2602.07298AI at MetaFeb 7, 2026score 4~120 min
- FebScaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgmentsscaling-laws2602.23234AppleFeb 26, 2026~104 min
- JanUnlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Textagents2601.10355LongCatJan 15, 2026score 9~111 min
- JanEvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experienceagents2601.15876meituanJan 22, 2026score 9~120 min
- JanMotion Attribution for Video Generationdata2601.08828NVIDIAJan 13, 2026score 2~115 min
- JanFine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelersevaluation2601.03211Jan 6, 2026~113 min
- JanEvasionBench: A Large-Scale Benchmark for Detecting Managerial Evasion in Earnings Call Q&Aevaluation2601.09142Jan 14, 2026score 2~119 min
- JanRethinking Video Generation Model for the Embodied Worldmultimodal2601.15282ByteDance SeedJan 21, 2026score 2~104 min
- JanShaping capabilities with token-level data filteringsafety2601.21571AnthropicJan 29, 2026score 3~110 min
- JanUncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuningtraining-methods2601.13697alibaba-incJan 20, 2026score 9~129 min
- JanGolden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Texttraining-methods2601.22975NVIDIAJan 30, 2026score 9~105 min
- JanDecouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-trainingtraining-methods2602.00747Jan 31, 2026~121 min
- JanMolmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Groundingvision2601.10611Jan 15, 2026score 9~97 min
2025
63- DecGR-Dexter Technical Reportagents2512.24210ByteDance SeedDec 30, 2025score 2~103 min
- DecDataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AIdata2512.16676Peking UniversityDec 18, 2025score 9~98 min
- DecEgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editingdiffusion2512.06065SnapDec 5, 2025score 2~105 min
- DecOmni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalizationmultimodal2512.10955Snap ResearchDec 11, 2025score 5~100 min
- DecInsight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Languagemultimodal2512.11251Dec 12, 2025~108 min
- DecNemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervisiontraining-methods2512.15489NVIDIADec 17, 2025score 8~103 min
- DecCosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modelinguncategorized2512.23162NVIDIADec 29, 2025score 2~99 min
- DecDreamOmni3: Scribble-based Editing and Generationvision2512.22525ByteDanceDec 27, 2025score 5~119 min
- NovScaling Agent Learning via Experience Synthesisagents2511.03773Nov 5, 2025~120 min
- NovFara-7B: An Efficient Agentic Model for Computer Useagents2511.19663Microsoft ResearchNov 24, 2025score 6~124 min
- NovLong Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scaledata2511.05705NVIDIANov 7, 2025score 9~105 min
- NovNVIDIA Nemotron Parse 1.1multimodal2511.20478NVIDIANov 25, 2025score 9~127 min
- NovReusing Pre-Training Data at Test Time is a Compute Multiplierpretraining2511.04234AppleNov 6, 2025~116 min
- NovDRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generationtraining-methods2511.06307OpenAINov 9, 2025score 8~114 min
- NovInstruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Datasetvision2511.15186KAIST AINov 19, 2025score 2~116 min
- NovHunyuanVideo 1.5 Technical Reportvision2511.18870Tencent HunyuanNov 24, 2025score 5~112 min
- OctWebscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levelsdata2510.06499SalesforceOct 7, 2025score 10~115 min
- OctDocReward: A Document Reward Model for Structuring and Stylizingdata2510.11391Microsoft ResearchOct 13, 2025score 6~95 min
- OctAgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesisdata2510.24695TongyiLabOct 28, 2025score 9~113 min
- OctOmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLMmultimodal2510.15870NVIDIAOct 17, 2025score 10~108 min
- OctolmOCR 2: Unit Test Rewards for Document OCRmultimodal2510.19817Ai2Oct 22, 2025score 9~115 min
- OctDaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agentstraining-methods2510.19336OPPOOct 22, 2025score 8~104 min
- OctHigh-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splattinguncategorized2510.10637DAMO AcademyOct 12, 2025score 3~114 min
- OctTowards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculumvision2510.27571Alibaba DAMOOct 31, 2025score 8~135 min
- SepScaling Generalist Data-Analytic Agentsagents2509.25084QwenSep 29, 2025score 9~103 min
- SepEmbeddingGemma: Powerful and Lightweight Text Representationsllm-systems2509.20354DeepMindSep 24, 2025~107 min
- SepLLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Trainingmultimodal2509.23661Sep 28, 2025score 9~125 min
- SepApertus: Democratizing Open and Compliant LLMs for Global Language Environmentspretraining2509.14233Swiss AI InitiativeSep 17, 2025score 9~107 min
- SepThinking Augmented Pre-trainingpretraining2509.20186Sep 24, 2025~103 min
- SepDA$^{2}$: Depth Anything in Any Directionscaling-laws2509.26618Tencent HunyuanSep 30, 2025score 9~100 min
- SepWinning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuningtraining-methods2509.23873alibabaSep 28, 2025score 9~104 min
- SepSocratic-Zero : Bootstrapping Reasoning via Data-Free Agent Co-evolutiontraining-methods2509.24726alibaba-incSep 29, 2025score 9~105 min
- SepManipulation as in Simulation: Enabling Accurate Geometry Perception in Robotsuncategorized2509.02530ByteDance SeedSep 2, 2025score 2~106 min
- AugEvaluating, Synthesizing, and Enhancing for Customer Support Conversationdata2508.04423Qwen DianJinAug 6, 2025score 3~117 min
- AugOpen Data Synthesis For Deep Researchdata2509.00375Aug 30, 2025~139 min
- AugDeep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMspretraining2508.06601Aug 8, 2025score 6~110 min
- JulWebShaper: Agentically Data Synthesizing via Information-Seeking Formalizationdata2507.15061Alibaba DAMOJul 20, 2025score 9~109 min
- JulMeta CLIP 2: A Worldwide Scaling Recipemultimodal2507.22062Jul 29, 2025score 10~113 min
- JulThe Imitation Game: Turing Machine Imitator is Length Generalizable Reasonerreasoning2507.13332Intern Large ModelsJul 17, 2025score 9~106 min
- JunOpenThoughts: Data Recipes for Reasoning Modelsdata2506.04178DeepSeekJun 4, 2025score 9~121 min
- JunInfini-gram mini: Exact n-gram Search at the Internet Scale with FM-Indexdata2506.12229Jun 13, 2025score 9~123 min
- JunMiniCPM4: Ultra-Efficient LLMs on End Devicesllm-systems2506.07900OpenBMBJun 9, 2025score 9~122 min
- JunThe Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Textpretraining2506.05209Jun 5, 2025score 10~119 min
- JunAccurate and scalable exchange-correlation with deep learningpretraining2506.14665MicrosoftJun 17, 2025~145 min
- JunMERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Queryretrieval2506.03144ByteDanceJun 3, 2025score 8~123 min
- MayGranary: Speech Recognition and Translation Dataset in 25 European Languagesdata2505.13404NVIDIAMay 19, 2025~86 min
- MayREASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewardsreasoning2505.24760May 30, 2025score 9~109 min
- AprDataDecide: How to Predict Best Pretraining Data with Small Experimentsdata2504.11393Apr 15, 2025~133 min
- AprOpen-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resourcesmultimodal2504.00595Apr 1, 2025score 9~95 min
- AprNemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-trainingpretraining2504.13161Apr 17, 2025score 10~127 min
- AprOLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokenstraining-methods2504.07096Apr 9, 2025score 9~92 min
- AprON LINEAR REPRESENTATIONS AND PRETRAINING DATA FREQUENCY IN LANGUAGE MODELSuncategorized2504.12459Apr 16, 2025~95 min
- MarPokéChamp: an Expert-level Minimax Language Agentagents2503.04094Mar 6, 2025score 4~130 min
- MarCube: A Roblox View of 3D Intelligencearchitecture2503.15475RobloxMar 19, 2025~109 min
- FebReformulation for Pretraining Data Augmentationdata2502.04235ByteDance SeedFeb 6, 2025score 9~95 min
- FebScaling Pre-training to One Hundred Billion Data for Vision Language Modelsmultimodal2502.07617DeepMindFeb 11, 2025score 9~135 min
- FebolmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Modelsmultimodal2502.18443Feb 25, 2025~135 min
- FebSmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Modelpretraining2502.02737Hugging Face Smol ModelsFeb 4, 2025score 10~133 min
- Jan2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretrainingdata2501.00958Jan 1, 2025score 10~113 min
- JanTowards Best Practices for Open Datasets for LLM Trainingdata2501.08365Jan 14, 2025~104 min
- JanPeople who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textevaluation2501.15654Jan 26, 2025score 10~110 min
- JanOmni-RGPT: Unifying Image and Video Region-level Understanding via Token Marksmultimodal2501.08326NVIDIAJan 14, 2025score 9~104 min
- JanCosmos World Foundation Model Platform for Physical AIpretraining2501.03575NVIDIAJan 6, 2025score 5~137 min
2024
47- DecLAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotationsdata2412.08580Dec 11, 2024score 6~88 min
- Dec2 OLMo 2 Furiouspretraining2501.00656Allen Institute for AIDec 31, 2024~111 min
- DecHunyuanProver: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Provingreasoning2412.20735Dec 30, 2024score 9~119 min
- DecPhi-4 Technical Reporttraining-methods2412.08905Microsoft ResearchDec 12, 2024~112 min
- DecRobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Responsetraining-methods2412.14922Dec 19, 2024score 10~90 min
- NovOPENCODER: THE OPEN COOKBOOK FOR TOP-TIER CODE LARGE LANGUAGE MODELScode2411.04905Nov 7, 2024~110 min
- NovREAD MOREdata2411.12372Together AINov 19, 2024~127 min
- OctTLDR: Token-Level Detective Reward Model for Large Vision Language Modelsalignment2410.04734Meta LlamaOct 7, 2024score 8~119 min
- OctSelf-Boosting Large Language Models with Synthetic Preference Dataalignment2410.06961Oct 9, 2024score 9~118 min
- OctMOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languagesdata2410.01036NVIDIAOct 1, 2024score 2~112 min
- OctUndesirable Memorization in Large Language Models: A Surveysafety2410.02650Oct 3, 2024~133 min
- OctDocument Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extractionvision2410.21169Oct 28, 2024score 10~114 min
- SepSource2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sourcesdata2409.08239Sep 12, 2024score 10~112 min
- SepEuroLLM: Multilingual Language Models for Europellm-systems2409.16235Sep 24, 2024score 9~103 min
- SepMolmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Modelsmultimodal2409.17146Allen Institute for AISep 25, 2024score 9~121 min
- AugUnleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Modelsdata2408.02085Aug 4, 2024score 10~182 min
- AugCogVideoX: Text-to-Video Diffusion Models with An Expert Transformerdiffusion2408.06072Zhipu / GLMAug 12, 2024score 3~121 min
- AugSelf-Taught Evaluatorsevaluation2408.02666Aug 5, 2024score 9~103 min
- AugTo Code, or Not To Code? Exploring Impact of Code in Pre-trainingpretraining2408.10914Aug 20, 2024~120 min
- AugBaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Competitive Large Language Model Baselinepretraining2408.15079Aug 27, 2024score 9~122 min
- AugBuilding and better understanding vision-language models: insights and future directionsvision2408.12637Aug 22, 2024score 10~96 min
- JunYODAS: Youtube-Oriented Dataset for Audio and Speechdata2406.00899NVIDIAJun 2, 2024~96 min
- JunDataComp-LM: In Search of the Next Generation of Training Sets for Language Modelsdata2406.11794AppleJun 17, 2024score 10~93 min
- JunThe FineWeb Datasets: Decanting the Web for the Finest Text Data at Scaledata2406.17557FineDataJun 25, 2024score 10~123 min
- JunScaling Synthetic Data Creation with 1,000,000,000 Personasdata2406.20094Jun 28, 2024~88 min
- JunFrom Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipelineevaluation2406.11939Jun 17, 2024score 9~105 min
- JunInstruction Pre-Training: Language Models are Supervised Multitask Learnerspretraining2406.14491Jun 20, 2024score 10~97 min
- MayArctic-Embed: Scalable, Efficient, and Accurate Text Embedding Modelsdata2405.05374SnowflakeMay 8, 2024~111 min
- MayBiMix: A Bivariate Data Mixing Law for Language Model Pretrainingpretraining2405.14908May 23, 2024score 9~105 min
- MayDeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Datareasoning2405.14333DeepSeekMay 23, 2024score 9~101 min
- AprEagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrencearchitecture2404.05892Apr 8, 2024score 9~133 min
- AprExtending Llama-3's Context Ten-Fold Overnightcontext-optimization2404.19553Apr 30, 2024score 9~103 min
- AprPhi-3 Technical Report: A Highly Capable Language Model Locally on Your Phonepretraining2404.14219Microsoft ResearchApr 22, 2024score 9~114 min
- AprAdvancing LLM Reasoning Generalists with Preference Treesreasoning2404.02078OpenBMBApr 2, 2024score 9~115 min
- MarLLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compressioncontext-optimization2403.12968Microsoft ResearchMar 19, 2024score 8~113 min
- MarGecko: Versatile Text Embeddings Distilled from Large Language Modelsdata2403.20327DeepMindMar 29, 2024score 9~98 min
- MarDeepSeek-VL: Towards Real-World Vision-Language Understandingmultimodal2403.05525DeepSeekMar 8, 2024score 9~88 min
- MarYi: Open Foundation Models by 01.AIscaling-laws2403.0465201-AIMar 7, 2024score 9~73 min
- FebOLMo: Accelerating the Science of Language Modelsagents2402.00838Allen Institute for AIFeb 1, 2024score 10~113 min
- FebStarCoder 2 and The Stack v2: The Next Generationcode2402.19173RobloxFeb 29, 2024score 9~111 min
- FebData Engineering for Scaling Language Models to 128K Contextdata2402.10171Feb 15, 2024score 9~112 min
- FebOmniPred: Language Models as Universal Regressorspretraining2402.14547DeepMindFeb 22, 2024score 5~101 min
- FebDeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Modelsreasoning2402.03300DeepSeekFeb 5, 2024score 10~130 min
- FebDo Membership Inference Attacks Work on Large Language Models?safety2402.07841Feb 12, 2024~125 min
- JanAutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agentsagents2401.12963DeepMindJan 23, 2024score 3~112 min
- JanRephrasing the Web: A Recipe for Compute and Data-Efficient Language Modelingdata2401.16380AppleJan 29, 2024score 9~109 min
- JanDolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Researchdata2402.00159Allen Institute for AIJan 31, 2024score 10~137 min
2023
27- DecGenerative agent-based modeling with actions grounded in physical, social, or digital space using Concordiaagents2312.03664DeepMindDec 6, 2023score 6~96 min
- DecMagicoder: Source Code Is All You Needcode2312.02120Dec 4, 2023score 9~82 min
- DecLLM360: Towards Fully Transparent Open-Source LLMspretraining2312.06550Dec 11, 2023score 9~104 min
- NovRoboVQA: Multimodal Long-Horizon Reasoning for Roboticsuncategorized2311.00899DeepMindNov 1, 2023score 5~140 min
- OctOpenWebMath: An Open Dataset of High-Quality Mathematical Web Textdata2310.06786EleutherAIOct 10, 2023~112 min
- OctProperty-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generationdata2310.12371NVIDIAOct 18, 2023~106 min
- OctIN-CONTEXT PRETRAINING: LANGUAGE MODELING BEYOND DOCUMENT BOUNDARIESpretraining2310.10638Oct 16, 2023~111 min
- OctDetecting Pretraining Data from Large Language Modelssafety2310.16789Oct 25, 2023score 9~111 min
- SepWhen Less is More: Investigating Data Pruning for Pretraining LLMs at Scalepretraining2309.04564Sep 8, 2023score 9~110 min
- SepTextbooks Are All You Need II: phi-1.5 technical reportpretraining2309.05463Microsoft ResearchSep 11, 2023~92 min
- SepLMSYS-CHAT-1M: A LARGE-SCALE REAL-WORLD LLM CONVERSATION DATASETsafety2309.11998Sep 21, 2023~126 min
- AugSIMPLE SYNTHETIC DATA REDUCES SYCOPHANCY IN LARGE LANGUAGE MODELSalignment2308.03958Aug 7, 2023~107 min
- AugSelf-Alignment with Instruction Backtranslationalignment2308.06259Aug 11, 2023score 9~109 min
- JulScaling TransNormer to 175 Billion Parametersllm-systems2307.14995Jul 27, 2023score 9~114 min
- JunThe RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Onlydata2306.01116Jun 1, 2023score 9~120 min
- JunGAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Explorationdata2306.01481EleutherAIJun 2, 2023~101 min
- JunTextbooks Are All You Needdata2306.11644Microsoft ResearchJun 20, 2023~112 min
- JunRead Morepretraining2306.02254EleutherAIJun 4, 2023score 5~108 min
- MayUnified Embedding: Battle-Tested Feature Representations for Web-Scale ML Systemsarchitecture2305.12102May 20, 2023~102 min
- MayDoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretrainingdata2305.10429May 17, 2023score 9~98 min
- MayEnhancing Chat Language Models by Scaling High-quality Instructional Conversationsdata2305.14233OpenBMBMay 23, 2023score 9~99 min
- MayLet's Verify Step by Stepreasoning2305.20050May 31, 2023~124 min
- AprLLaVA: Visual Instruction Tuningmultimodal2304.08485Apr 17, 2023~91 min
- AprDINOv2: Learning Robust Visual Features without Supervisionpretraining2304.07193Meta AI / FAIRApr 14, 2023~104 min
- AprInstruction Tuning with GPT-4training-methods2304.03277Allen Institute for AIApr 6, 2023~113 min
- AprSegment Anythingvision2304.02643Meta AI / FAIRApr 5, 2023~122 min
- JanThe Flan Collection: Designing Data and Methods for Effective Instruction Tuningtraining-methods2301.13688Allen Institute for AIJan 31, 2023~111 min
2022
3- OctLAION-5B: An open large-scale dataset for training next generation image-text modelsdata2210.08402LAIONOct 16, 2022~105 min
- JanDatasheet for the Piledata2201.07311EleutherAIJan 13, 2022~124 min
- JanBLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generationmultimodal2201.12086SalesforceJan 28, 2022~104 min
2021
3- NovLAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairsdata2111.02114EleutherAINov 3, 2021~95 min
- NovScaling Law for Recommendation Models: Towards General-purpose User Representationsscaling-laws2111.11294Nov 15, 2021~105 min
- SepAn Empirical Exploration in Quality Filtering of Text Datadata2109.00698EleutherAISep 2, 2021~106 min
2020
4- DecMLS: A Large-Scale Multilingual Dataset for Speech Researchdata2012.03411NVIDIA340 citesDec 7, 2020~137 min
- DecThe Pile: An 800GB Dataset of Diverse Text for Language Modelingdata2101.00027Stability AIDec 31, 2020~117 min
- JulCoVoST 2 and Massively Multilingual Speech-to-Text Translationdata2007.10310NVIDIAJul 20, 2020~136 min
- JulMultiWOZ 2.2 : A Dialogue Dataset with Additional Annotation Corrections and State Tracking Baselinesdata2007.12720Jul 10, 2020~101 min
2019
12018
12017
22016
3- SepGoogle's Neural Machine Translation System: Bridging the Gap between Human and Machine Translationarchitecture1609.08144Sep 26, 2016~114 min
- JulEnriching Word Vectors with Subword Informationpretraining1607.04606Meta AI / FAIRJul 15, 2016~124 min
- MarMastering the game of Go with deep neural networks and tree searchdata1603.07185Mar 23, 2016~115 min