Transformer foundations
86 papers in this thread, across 18 domains.
Progress0 of 29
- 2025VoxtralMultimodal2507.13264MistralJul 17, 2025~111 minMultimodal
- 2025Tensor Product Attention Is All You NeedArchitecture2501.06425math-aiJan 11, 2025score 9~113 minArchitecture
- 2024GPT-4o System CardSafety2410.21276OpenAIOct 25, 2024score 9~137 minSafety
- 2024Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERTRetrieval2402.07440Together AIFeb 12, 2024~97 minRetrieval
- 2023Gemini vs GPT-4V: A Preliminary Comparison and Combination of Vision-Language Models Through Qualitative CasesMultimodal2312.15011Dec 22, 2023score 5~118 minMultimodal
- 2023A Challenger to GPT-4V? Early Explorations of Gemini in Visual ExpertiseVision2312.12436Dec 19, 2023score 9~92 minVision
- 2023Orca: Progressive Learning from Complex Explanation Traces of GPT-4Reasoning2306.02707Stability AIJun 5, 2023score 9~91 minReasoning
- 2023Morphosyntactic probing of multilingual BERT modelsno summary yetcs cl2306.06205AllenAI8 citesJun 9, 2023cs cl
- 2023MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsMultimodal2304.10592Apr 20, 2023~99 minMultimodal
- 2023Instruction Tuning with GPT-4Training Methods2304.03277Allen Institute for AIApr 6, 2023~113 minTraining Methods
- 2023Supporting Qualitative Analysis with Large Language Models: Combining Codebook with GPT-3 for Deductive Codingno summary yetcs cl2304.10548Microsoft Research206 citesApr 17, 2023cs cl
- 2023Sparks of Artificial General Intelligence: Early experiments with GPT-4Evaluation2303.12712Mar 22, 2023~117 minEvaluation
- 2023GPT-4 Technical ReportPretraining2303.08774Mar 15, 2023~118 minPretraining
- 2023Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settingsno summary yeteess as2303.05737Google Research1 citesMar 10, 2023eess as
- 2023A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPTno summary yetcs ai2302.09419Salesforce153 citesFeb 18, 2023cs ai
- 2022MarkBERT: Marking Word Boundaries Improves Chinese BERTno summary yetcs cl2203.06378Tencent0 citesMar 12, 2022cs cl
- 2022RescoreBERT: Discriminative Speech Recognition Rescoring with BERTno summary yeteess as2202.01094Amazon36 citesFeb 2, 2022eess as
- 2021DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPretraining2111.09543Microsoft ResearchNov 18, 2021~106 minPretraining
- 2021ReasonBERT: Pre-trained to Reason with Distant Supervisionno summary yetcs cl2109.04912Google Research0 citesSep 10, 2021cs cl
- 2021Want To Reduce Labeling Cost? GPT-3 Can Helpno summary yetcs cl2108.13487Microsoft Research27 citesAug 30, 2021cs cl
- 2021Noise Stability Regularization for Improving BERT Fine-tuningno summary yetcs cl2107.04835Baidu0 citesJul 10, 2021cs cl
- 2021CoBERL: Contrastive BERT for Reinforcement Learningno summary yetcs lg2107.05431Google Research8 citesJul 12, 2021cs lg
- 2021Better than BERT but Worse than Baselineno summary yetcs cl2105.05915Baidu1 citesMay 12, 2021cs cl
- 2021BBAEG: Towards BERT-based Biomedical Adversarial Example Generation for Text Classificationno summary yetcs cl2104.01782Microsoft Research0 citesApr 5, 2021cs cl
- 2021Natural Language Understanding with Privacy-Preserving BERTno summary yetcs cl2104.07504Google Research40 citesApr 15, 2021cs cl
- 2021Probing Across Time: What Does RoBERTa Know and When?no summary yetcs cl2104.07885AllenAI0 citesApr 16, 2021cs cl
- 2021Cross-Attention is All You Need: Adapting Pretrained Transformers for Machine Translationno summary yetcs cl2104.08771Microsoft Research0 citesApr 18, 2021cs cl
- 2021Disfluency Detection with Unlabeled Data and Small BERT Modelsno summary yetcs cl2104.10769Google Research1 citesApr 21, 2021cs cl
- 2021Thinking Aloud: Dynamic Context Generation Improves Zero-Shot Reasoning Performance of GPT-2no summary yetcs cl2103.13033AllenAI16 citesMar 24, 2021cs cl
- 2021NewsBERT: Distilling Pre-trained Language Model for Intelligent News Applicationno summary yetcs cl2102.04887Microsoft Research4 citesFeb 9, 2021cs cl
- 2021What Makes Good In-Context Examples for GPT-3?Prompting2101.06804Jan 17, 2021~101 minPrompting
- 2021First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERTno summary yetcs cl2101.11109AllenAI29 citesJan 26, 2021cs cl
- 2021EmpathBERT: A BERT-based Framework for Demographic-aware Empathy Predictionno summary yetcs lg2102.00272Adobe3 citesJan 30, 2021cs lg
- 2020BinaryBERT: Pushing the Limit of BERT QuantizationLow Precision2012.15701Dec 31, 2020~104 minLow Precision
- 2020Improving BERT with Syntax-aware Local Attentionno summary yetcs cl2012.15150Tencent7 citesDec 30, 2020cs cl
- 2020It's not Greek to mBERT: Inducing Word-Level Translations from Multilingual BERTno summary yetcs cl2010.08275AllenAI0 citesOct 16, 2020cs cl
- 2020On the Transformer Growth for Progressive BERT Trainingno summary yetcs cl2010.12562Google Research4 citesOct 23, 2020cs cl
- 2020Commonsense knowledge adversarial dataset that challenges ELECTRAno summary yetcs cl2010.13049Alibaba2 citesOct 25, 2020cs cl
- 2020To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Taggingno summary yetcs cl2010.14042Amazon2 citesOct 27, 2020cs cl
- 2020RecoBERT: A Catalog Language Model for Text-Based Recommendationsno summary yetcs ir2009.13292Microsoft Research0 citesSep 25, 2020cs ir
- 2020CokeBERT: Contextual Knowledge Selection and Embedding towards Enhanced Pre-Trained Language Modelsno summary yetcs cl2009.13964Tencent36 citesSep 29, 2020cs cl
- 2020Parsing with Multilingual BERT, a Small Corpus, and a Small Treebankno summary yetcs cl2009.14124AllenAI2 citesSep 29, 2020cs cl
- 2020DeVLBert: Learning Deconfounded Visio-Linguistic Representationsno summary yetcs cv2008.06884Alibaba63 citesAug 16, 2020cs cv
- 2020BERT Loses Patience: Fast and Robust Inference with Early Exitno summary yetcs cl2006.04152Microsoft Research45 citesJun 7, 2020cs cl
- 2020BERTology Meets Biology: Interpreting Attention in Protein Language Modelsno summary yetcs cl2006.15222Salesforce53 citesJun 26, 2020cs cl
- 2020schuBERT: Optimizing Elements of BERTno summary yetcs cl2005.06628Amazon2 citesMay 9, 2020cs cl
- 2020Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representationno summary yeteess as2005.08575Amazon39 citesMay 18, 2020eess as
- 2020FashionBERT: Text and Image Matching with Adaptive Loss for Cross-modal Retrievalno summary yetcs ir2005.09801Alibaba6 citesMay 20, 2020cs ir
- 2020BERTweet: A pre-trained language model for English Tweetsno summary yetcs cl2005.10200NVIDIA73 citesMay 20, 2020cs cl
- 2020ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTRetrieval2004.12832Apr 27, 2020~118 minRetrieval
- 2020MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devicesno summary yetcs cl2004.02984Google Research91 citesApr 6, 2020cs cl
- 2020TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented Dialogueno summary yetcs cl2004.06871Salesforce53 citesApr 15, 2020cs cl
- 2020VD-BERT: A Unified Vision and Dialog Transformer with BERTno summary yetcs cv2004.13278Salesforce30 citesApr 28, 2020cs cv
- 2020What Happens To BERT Embeddings During Fine-tuning?no summary yetcs cl2004.14448Google Research15 citesApr 29, 2020cs cl
- 2020PhoBERT: Pre-trained language models for Vietnameseno summary yetcs cl2003.00744NVIDIA35 citesMar 2, 2020cs cl
- 2020BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingTraining Methods2002.02925Feb 7, 2020~113 minTraining Methods
- 2020Self-Distillation Amplifies Regularization in Hilbert Spaceno summary yetcs lg2002.05715Google Research98 citesFeb 13, 2020cs lg
- 2019Q8BERT: Quantized 8Bit BERTInference Optimization1910.06188Oct 14, 2019~91 minInference Optimization
- 2019Ludwig v0.3 Introduces Hyperparameter Optimization, Transformers and TensorFlow 2 supportPretraining1910.01108UberOct 2, 2019~108 minPretraining
- 2019Thieves on Sesame Street! Model Extraction of BERT-based APIsno summary yetcs cl1910.12366Google Research32 citesOct 27, 2019cs cl
- 2019A Simple but Effective BERT Model for Dialog State Tracking on Resource-Limited Systemsno summary yetcs cl1910.12995Adobe4 citesOct 28, 2019cs cl
- 2019ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsArchitecture1909.11942984 citesSep 26, 2019~104 minArchitecture
- 2019Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTInference Optimization1909.05840Sep 12, 2019~112 minInference Optimization
- 2019TinyBERT: Distilling BERT for Natural Language UnderstandingTraining Methods1909.10351Sep 23, 2019~106 minTraining Methods
- 2019Extremely Small BERT Models from Mixed-Vocabulary Trainingno summary yetcs cl1909.11687Google Research15 citesSep 25, 2019cs cl
- 2019End-to-End Resume Parsing and Finding Candidates for a Job Description using BERTno summary yetcs ir1910.03089Adobe38 citesSep 30, 2019cs ir
- 2019Sentence-BERT: Sentence Embeddings using Siamese BERT-NetworksArchitecture1908.10084MozillaAug 27, 2019~110 minArchitecture
- 2019ViLBERT: Pretraining Task-Agnostic Visiolinguistic RepresentationsMultimodal1908.02265Aug 6, 2019~109 minMultimodal
- 2019VisualBERT: A Simple and Performant Baseline for Vision and LanguageMultimodal1908.03557Aug 9, 2019~119 minMultimodal
- 2019Patient Knowledge Distillation for BERT Model Compressionno summary yetcs cl1908.09355Microsoft Research85 citesAug 25, 2019cs cl
- 2019Small and Practical BERT Models for Sequence Labelingno summary yetcs cl1909.00100Google Research17 citesAug 31, 2019cs cl
- 2019Giving BERT a Calculator: Finding Operations and Arguments with Reading Comprehensionno summary yetcs cl1909.00109Google Research9 citesAug 31, 2019cs cl
- 2019SpanBERT: Improving Pre-training by Representing and Predicting SpansPretraining1907.10529179 citesJul 24, 2019~119 minPretraining
- 2019RoBERTa: A Robustly Optimized BERT Pretraining ApproachPretraining1907.11692Meta AI / FAIRJul 26, 2019~113 minPretraining
- 2019How multilingual is Multilingual BERT?no summary yetcs cl1906.01502Google Research141 citesJun 4, 2019cs cl
- 2019Visualizing and Measuring the Geometry of BERTno summary yetcs lg1906.02715Google Research166 citesJun 6, 2019cs lg
- 2019BERTphone: Phonetically-Aware Encoder Representations for Utterance-Level Speaker and Language Recognitionno summary yetcs cl1907.00457Amazon28 citesJun 30, 2019cs cl
- 2019BERT with History Answer Embedding for Conversational Question Answeringno summary yetcs ir1905.05412Alibaba149 citesMay 14, 2019cs ir
- 2019BERT Rediscovers the Classical NLP Pipelineno summary yetcs cl1905.05950Google Research50 citesMay 15, 2019cs cl
- 2019SciBERT: A Pretrained Language Model for Scientific Textno summary yetcs cl1903.10676AllenAI48 citesMar 26, 2019cs cl
- 2019Transformer-XL: Attentive Language Models Beyond a Fixed-Length ContextArchitecture1901.02860Jan 9, 2019~110 minArchitecture
- 2019Assessing BERT's Syntactic Abilitiesno summary yetcs cl1901.05287AllenAI295 citesJan 16, 2019cs cl
- 2019A BERT Baseline for the Natural Questionsno summary yetcs cl1901.08634Google Research97 citesJan 24, 2019cs cl
- 2018BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingPretraining1810.04805DeepMind45,647 citesOct 11, 2018~131 minPretraining
- 2017Attention Is All You NeedArchitecture1706.03762NVIDIAJun 12, 2017~118 minArchitecture
- 2014Unconstrained Online Linear Learning in Hilbert Spaces: Minimax Algorithms and Normal Approximationsno summary yetcs lg1403.0628Google Research36 citesMar 3, 2014cs lg