ArXiv: 2108.07258
π― Pitch
A single foundation model like BERT now underpins nearly all state-of-the-art NLP, meaning flaws like biases are inherited wholesale by downstream systems. This report warns that the same homogenization fueling AIβs rapid progress creates dangerous single points of failure we barely understand.
1. Executive Summary
This report introduces and systematically analyzes the emerging paradigm of foundation models β models trained on broad data at scale that can be adapted to a wide range of downstream tasks β examining their capabilities across language, vision, robotics, reasoning, and interaction, their applications in healthcare, law, and education, and their technical underpinnings from modeling through interpretability. The core conceptual contribution is the characterization of this paradigm through two interacting forces: emergence (behavior implicitly induced by scale rather than explicitly constructed, such as GPT-3βs in-context learning) and homogenization (the consolidation of methodology where a single foundation model like BERT underpins nearly all state-of-the-art NLP systems), which together provide powerful leverage but create single points of failure where defects in the foundation model are inherited by all adapted derivatives. The report establishes that while foundation models demonstrate transformative potential β from improving access to legal services through few-shot document retrieval to enabling multimodal biomedical discovery β their emergent properties remain poorly understood, and their homogenizing effect demands caution, as the same representational biases (e.g., Anglocentric perspectives encoded in BERT, pernicious stereotypes in GPT-3) propagate across myriad applications, making rigorous interdisciplinary evaluation, documentation, and auditing essential before widespread deployment.
2. Context and Motivation
The Core Problem: We Lack a Coherent Framework for Understanding a Paradigm Shift in AI
The fundamental problem this report addresses is not a single technical challenge but rather a conceptual and analytical vacuum: AI is undergoing a paradigm shift, driven by the rise of models like BERT, GPT-3, DALL-E, and CLIP, that the existing language and frameworks of the field fail to adequately capture. The authors observe that while the technical components behind these models β deep neural networks, transfer learning, self-supervised learning β have existed for decades, the sociological and practical significance of what these models represent has outpaced our ability to describe, analyze, and govern them.
The gap is both terminological and structural. Existing terms like "pretrained model" or "self-supervised model" describe how these models are built, but fail to convey their distinctive role in the AI ecosystem: they are intermediary assets β unfinished, broadly capable models that serve as the common basis from which myriad task-specific models are built via adaptation. This role is unprecedented in its degree of leverage (a single model can improve hundreds of downstream applications) and its degree of risk (flaws in that single model propagate to all those applications). Without a shared vocabulary and analytical framework, the report argues, researchers, practitioners, policymakers, and the public cannot productively discuss the opportunities and risks of this shift.
This gap is significant for reasons that span scientific understanding, engineering practice, and societal governance:
-
Scientific understanding: The emergent properties of these models β capabilities like in-context learning that were neither explicitly trained for nor anticipated β challenge our ability to reason about what these models can and cannot do. As the report states in Section 1.1, "since the power of foundation models comes from their emergent qualities rather than their explicit construction, existing foundation models are hard to understand." This undermines the scientific goal of building predictable, reliable AI systems.
-
Engineering practice: The homogenization effect means that algorithmic monoculture is becoming entrenched. In NLP, "almost all state-of-the-art NLP models are now adapted from one of a few foundation models, such as BERT, RoBERTa, BART, T5, etc." (Section 1.1). While this provides enormous leverage β improvements to the foundation model immediately benefit all of NLP β it also means that any biases, failure modes, or security vulnerabilities in those few models become systemic across the entire field.
-
Societal governance: Foundation models are already being integrated into real-world deployments affecting billions of users (e.g., Google Search's use of BERT, Section 1.2). Yet the report emphasizes that "we do not fully understand the nature or quality of the foundation that foundation models provide; we cannot characterize whether the foundation is trustworthy or not" (Section 1.1.1). This creates a governance crisis: technologies with far-reaching consequences are being deployed without adequate understanding of their capabilities, limitations, or failure modes.
Why This Problem Matters: The Interplay of Emergence and Homogenization
The report's central analytical move is to identify emergence and homogenization as the two defining forces of this paradigm, and to show that their interaction creates a uniquely challenging risk profile.
Emergence means that the behavior of a system is implicitly induced rather than explicitly constructed. The report traces this as the culmination of a 30-year trend in AI (Figure 1, Section 1.1): machine learning replaced hand-crafted rules with patterns inferred from data (how a task is performed emerges from examples); deep learning replaced hand-engineered features with learned representations (high-level features emerge from training); and now foundation models replace task-specific architectures with a single adaptable model where "even advanced functionalities such as in-context learning emerge" (Section 1.1). This is "both the source of scientific excitement and anxiety about unanticipated consequences" (Section 1.1) β emergence generates substantial uncertainty about what a model is actually capable of and where it might fail.
Homogenization means the consolidation of methodologies so that a wide range of applications are powered by the same underlying technology. The report identifies this as an accelerating trend: machine learning homogenized algorithms (logistic regression across many applications), deep learning homogenized architectures (CNNs for vision, LSTMs for language), and foundation models homogenize the model itself β there are now "a few foundation models" underpinning entire fields (Section 1.1). This consolidation is spreading across modalities: similar Transformer-based approaches are being applied to text, images, speech, tabular data, protein sequences, and reinforcement learning (Section 1.1), suggesting "a possible future where we have a unified set of tools for developing foundation models across a wide range of modalities."
The critical insight of the report is that emergence and homogenization compound each other's risks (Section 1.1): "Homogenization and emergence interact in a potentially unsettling way." Homogenization means that any flaws in a foundation model are "blindly inherited by all adapted models." But because those flaws arise from emergent properties rather than explicit design choices, they are hard to anticipate, detect, and correct. You cannot audit a specification that was never written; you can only discover behaviors through extensive testing, and the space of possible behaviors may be vast. The report captures this with the pithy formulation: "aggressive homogenization through these models is risky business. Derisking is the central challenge in the further development of foundation models" (Section 1.1).
Where Prior Approaches Fall Short
The report identifies multiple dimensions along which existing frameworks, terminology, and institutional structures are inadequate for addressing this new paradigm.
Inadequate terminology. The authors explicitly discuss their choice of the term "foundation models" in Section 1.1.1, rejecting several alternatives. "Pretrained model" and "self-supervised model" capture technical dimensions but "fail to capture the significance of the paradigm shift in an accessible manner for those beyond machine learning." "Language model" is "simply too narrow": the scope of these models extends well beyond language to vision, robotics, protein folding, and code generation. "General-purpose model" or "multi-purpose model" capture the multi-task nature but "fail to capture their unfinished character and the need for adaptation." The term "foundation" was deliberately chosen to specify both the role (an incomplete basis from which task-specific models are built) and the normative aspiration (foundations must be stable, safe, and secure β "poorly-constructed foundations are a recipe for disaster"). The existence of this naming problem is itself evidence of the gap: the AI community lacked language adequate to the phenomenon.
Insufficient disciplinary scope. The report argues that research on building foundation models "has occurred almost exclusively in industry β big tech companies such as Google, Facebook, Microsoft, or Huawei, or startups such as OpenAI or AI21 Labs" (Section 1.3). This concentration means that the development of these models is driven primarily by commercial incentives and technical perspectives, without adequate input from "humanists and social scientists." The report explicitly calls for infusing "social considerations and ethical design deeply into the technological development of foundation models and their surrounding ecosystem from the start," rather than relying on "post-hoc audits of ethical and social consequences" (Section 1.3). It notes that academic institutions are "unique in that they host the widest set of disciplines under one roof," and argues for their "crucial role" in developing foundation models in ways that promote social benefit and mitigate harm.
Loss of accessibility and reproducibility. The report identifies a concerning reversal of the open science trends that accelerated deep learning progress (Section 1.3). While the deep learning era saw increasing norms of public code and dataset release, facilitated by frameworks like TensorFlow and PyTorch, foundation models "start to roll back this positive trend." Some models are not released at all (GPT-3 was initially available only via API to limited partners), some datasets are not released, and even when trained models are available, "the actual training of foundation models is unavailable to the vast majority of AI researchers, due to the much higher computational cost and the complex engineering requirements." This creates a fundamental access asymmetry: meaningful research on emergent behaviors like in-context learning requires scale, so "scale is needed to even ask the right questions" (Section 1.3). The report warns that the "fundamental centralizing nature of foundation models means that the barrier to entry for developing them will continue to rise."
Inadequate evaluation frameworks. Standard machine learning evaluation β train on a dataset, test on a held-out split for a specific task β is "not designed explicitly for the setting of foundation models" (Section 4.4). Foundation models are intermediary assets, not task-specific systems, so traditional evaluation cannot directly characterize them. The report identifies this as requiring new paradigms: intrinsic evaluation that directly measures properties of the foundation model itself (linguistic capabilities, biases, reasoning competencies) and extrinsic evaluation that accounts for the resources and adaptation methods used when measuring downstream task performance.
Unresolved tensions in fairness and accountability. Standard frameworks for algorithmic fairness assume a specific system deployed for a specific purpose with identifiable stakeholders. Foundation models complicate every step of this: they are task-agnostic, can be adapted by different entities for unforeseen purposes, and their intrinsic biases are latent until manifested in specific applications (Section 5.1). The report argues that this requires distinguishing "intrinsic biases" (properties of the foundation model that portend harm) from "extrinsic harms" (harms arising in specific applications), and developing new mechanisms for tracing harms to their sources across the adaptation chain.
Absence of professional norms. The report notes that "even the professional norms β what Robert Merton calls the ethos of science β around foundation models are underdeveloped" (Section 1.3). Questions such as when models are safe to release, how to respond to methodological misconduct, and what documentation standards should apply "are underdeveloped," leaving a regulatory and normative vacuum as these models are deployed.
How This Paper Positions Itself
The report positions itself not as advancing a single technical contribution but as providing a comprehensive, interdisciplinary map of an emerging paradigm. Its scope is deliberately broad: 26 sections spanning capabilities, applications, technology, and society, written by over 100 authors from diverse disciplines including computer science, law, philosophy, economics, medicine, and education (Section 1.4). The authors explicitly describe this as an experiment in interdisciplinary collaboration, acknowledging that "from the very beginning, the community included not just AI researchers, but those eager to apply foundation models to their domain (e.g., healthcare and law), as well as those who were interested in societal concerns (e.g., ethics and economics)" (Section 1.4).
The report's intellectual stance is characterized by several key positions:
Ecosystem thinking. Rather than analyzing foundation models in isolation, the report consistently situates them within a broader pipeline: data creation β data curation β training β adaptation β deployment (Section 1.2, Figure 3). This ecosystem view enables it to distribute responsibility and analysis across stages: data quality issues are data curation problems, harmful outputs may be addressed at adaptation time through constraints, and deployment decisions are "separate from [model] construction." The mantra is "think ecosystem, act model" β while social impact depends on the whole pipeline, researchers whose purview is restricted to training still need "surrogate metrics for a representative set of potential downstream evaluations."
Disciplinary pluralism as essential, not optional. The report argues that addressing the challenges of foundation models requires deep collaboration across disciplines "commensurate with their fundamentally sociotechnical nature" (Section 1.1). This is not framed as a nice-to-have ethical overlay but as a core requirement for getting the technology right β understanding capabilities requires philosophy of language (Section 2.6), deploying in healthcare requires engaging with privacy law and clinical workflows (Section 3.1), and reasoning about fairness requires social science methods for source tracing and harm documentation (Section 5.1).
The centrality of evaluation, documentation, and auditing. Throughout the report, there is a recurring theme that the path to responsible development runs through better measurement and transparency: intrinsic evaluation to characterize capabilities (Section 4.4), documentation akin to "nutrition labels" or model cards to communicate limitations (Section 5.6), and independent auditing β ideally through staged release programs with neutral third-party oversight β to uncover biases and failure modes that closed-door testing misses (Section 5.6). The report consistently frames evaluation not as a technical afterthought but as the critical infrastructure for making emergence tractable and homogenization safe.
A call to shape incentives, not just technology. The report explicitly engages with the political economy of foundation model development (Section 1.3, Section 5.6), arguing that market incentives alone will not produce socially beneficial outcomes β they will underinvest in applications for marginalized populations, ignore negative externalities like environmental costs, and concentrate power. It calls for public investment in shared computing infrastructure (the "National Research Cloud" initiative), community-led open-source efforts (EleutherAI, HuggingFace's BigScience), and professional norms and legal standards that create accountability. The report's position is that "the suitability of a technology relying largely on emergent behavior for widespread deployment to people is unclear," and that "we need to be cautious, and that now is the time to establish the professional norms that will enable the responsible research and deployment of foundation models" (Section 1.3).
3. Technical Approach
3.1 Reader Orientation
This report is not a single technical system but rather a comprehensive analytical framework: it constructs a conceptual architecture β a set of definitions, taxonomies, evaluation protocols, and design principles β for understanding and responsibly developing foundation models. The problem it solves is the absence of shared vocabulary and systematic analysis tools for an AI paradigm characterized by emergence and homogenization; the "shape" of the solution is a multi-layered sociotechnical map that traces foundation models from their technical underpinnings (data, architectures, training) through their capabilities and applications to their societal consequences, with evaluation, documentation, and interdisciplinary collaboration as the integrating threads.
3.2 Big-Picture Architecture (Diagram in Words)
Imagine a nested structure of analysis, organized into four concentric layers, with evaluation and societal impact cutting across all of them:
Layer 1: The Foundation Model Core. At the center sits the foundation model itself β a deep neural network (typically a Transformer) trained on broad, multimodal data using self-supervised objectives. This is the "unfinished" intermediary. Its properties are determined by five interacting forces: model architecture (expressivity, scalability), training procedures (objectives, optimization), data (scale, composition, curation), systems (hardware, parallelism), and adaptation mechanisms (fine-tuning, prompting).
Layer 2: Capabilities. These are the behaviors and skills that foundation models acquire through training β language understanding and generation, visual perception, robotic control, logical reasoning, and interactive abilities. Capabilities are emergent: they arise from the interaction of scale, data, and architecture, not from explicit programming. They are what make foundation models useful across many tasks.
Layer 3: Applications. These are the real-world domains where foundation models are deployed, each with unique requirements, data modalities, and constraints: healthcare and biomedicine (multimodal patient data, privacy regulations, need for explainability), law (long documents, logical reasoning over shifting case law, high-stakes decisions), and education (understanding student cognition, adaptive instruction, pedagogical expertise).
Layer 4: Societal Impact. This outermost layer examines how foundation models affect people and institutions: inequity and fairness (intrinsic biases propagating through homogenized systems), misuse (cheap, personalized, high-quality generation enabling disinformation and harassment), environmental cost (carbon emissions from training and inference), legal frameworks (liability, intellectual property, privacy), economic effects (productivity, wage inequality, concentration of power), and ethics (homogenization creating single points of failure, surveillance, erosion of accountability).
Cross-cutting threads: Evaluation, Interpretability, Theory, and Security. These run through all layers β how do we measure capabilities (evaluation), understand mechanisms (interpretability), provide guarantees (theory), and protect against attacks (security)? The report argues that progress on these cross-cutting challenges is essential for managing the risks of emergence and homogenization.
3.3 Roadmap for the Deep Dive
I will walk through the technical approach in six stages, building from the abstract framework outward to concrete mechanisms:
-
First, the conceptual architecture: emergence, homogenization, and the ecosystem. This is the report's core analytical innovation β defining the two forces that characterize the paradigm and the pipeline view that distributes responsibility. Everything else builds on these definitions.
-
Second, the modeling principles that underpin foundation models. The report identifies five desired properties of model architectures β expressivity, scalability, multimodality, memory, and compositionality β that collectively explain why Transformers have become the dominant substrate and what future architectures must achieve.
-
Third, the training and adaptation framework. This covers the design space of self-supervised objectives (generative vs. discriminative, input representation choices), the adaptation mechanisms that convert a foundation model into a task-specific tool (fine-tuning, prompting, lightweight adaptation), and the resource trade-offs involved.
-
Fourth, the data and systems infrastructure. Foundation models depend critically on data at unprecedented scale and on compute systems that push hardware limits. The report's data hub proposal and systems co-design analysis provide the engineering foundations.
-
Fifth, the evaluation and interpretability frameworks. These are the mechanisms for making emergence tractable: new intrinsic evaluation paradigms for directly measuring foundation model properties, resource-aware extrinsic evaluation, and the "one modelβmany models" framework for interpretability.
-
Sixth, the integration of societal analysis into technical design. The report's distinctive contribution is treating societal considerations β fairness, robustness, security, safety, legality β not as post-hoc audits but as design constraints that must be addressed within the technical development process itself.
3.4 Detailed, Sentence-Based Technical Breakdown
This is fundamentally an analytical and framework-building paper whose core idea is that foundation models represent a paradigm shift characterized by emergence and homogenization, and that responsibly developing this paradigm requires a comprehensive, interdisciplinary framework spanning technical design, capability analysis, application-specific requirements, and societal impact assessment, all integrated through a shared ecosystem view and centered on evaluation, documentation, and auditing as the primary risk-management tools.
The Conceptual Architecture: Emergence, Homogenization, and the Ecosystem Pipeline
The report grounds its entire analysis in two interacting forces that distinguish foundation models from previous AI paradigms. Emergence is defined in operational terms: it "means that the behavior of a system is implicitly induced rather than explicitly constructed" (Section 1.1). The report traces this as a historical progression through three eras of AI, each characterized by what emerges and what is homogenized:
- Machine learning (1990s): How a task is performed emerges from data rather than being hand-crafted in rules; learning algorithms are homogenized (logistic regression applied across many domains).
- Deep learning (2010s): High-level features emerge from training on raw inputs rather than being hand-engineered; model architectures are homogenized (CNNs for vision, RNNs/LSTMs for sequences).
- Foundation models (2020s): Advanced functionalities like in-context learning emerge from scale; the model itself is homogenized β a single BERT or GPT-3 becomes the substrate for nearly all NLP tasks.
The crucial property of emergence in the foundation model era is that it is both unanticipated and unpredictable: "in-context learning... was neither specifically trained for nor anticipated to arise" (Section 1.1). This is not a marginal phenomenon β it is the primary source of capability and the primary source of risk. The same scale that enables GPT-3 to perform arithmetic without explicit training also enables it to generate toxic content, encode stereotypes, or produce plausible-sounding falsehoods, and because these behaviors are emergent, they cannot be eliminated by inspecting a specification β they can only be discovered through extensive empirical probing.
Homogenization is the consolidation force β the tendency for "a single model to be useful for such a wide range of tasks" that it becomes the default starting point for all applications in a domain (Section 1.1). The report documents this empirically: in NLP, "almost all state-of-the-art NLP models are now adapted from one of a few foundation models, such as BERT, RoBERTa, BART, T5, etc." The homogenization is spreading across modalities: "similar Transformer-based sequence modeling approaches are now applied to text, images, speech, tabular data, protein sequences, organic molecules, and reinforcement learning" (Section 1.1). The trend toward multimodal models β trained jointly on language and vision β further intensifies this consolidation.
The interaction mechanism between emergence and homogenization is what the report identifies as the central challenge. Emergence means that capabilities and defects are not specified; homogenization means that whatever capabilities and defects exist are propagated widely. The report states this explicitly: "Homogenization could potentially provide enormous gains for many domains where task-specific data is quite limited... on the other hand, any flaws in the model are blindly inherited by all adapted models" (Section 1.1). This creates a leverage asymmetry: improvements to the foundation model yield widespread benefits, but defects in the foundation model cause widespread harm. The report's normative conclusion is that "derisking is the central challenge."
The ecosystem pipeline is the framework the report uses to distribute responsibility and analysis across stages rather than treating the foundation model as a monolithic black box. The pipeline consists of five sequential stages (Section 1.2, Figure 3):
-
Data creation: "All data is created by people and most data is at least implicitly about people." Data can be human-generated content (emails, photos, articles) or measurements of people (genomic data) or their environment (satellite images). The key principle: "all data has an owner and is created with a purpose (where that purpose may or may not include training a foundation model)."
-
Data curation: Raw data is filtered, selected, and organized into training datasets. The report emphasizes that "there is no single natural distribution of data; even the most permissive Internet crawl requires some selection and post-filtering." This stage involves legal and ethical judgments about data quality and consent that are often underappreciated in research but critical in practice.
-
Training: The computationally intensive process of optimizing the foundation model's parameters on curated data. This is "the celebrated centerpiece in AI research, though it is only one of many stages."
-
Adaptation: Converting the foundation model into a task-specific system. This may involve fine-tuning on labeled data, prompting, or incorporating additional modules (classifiers, rules, validation against external sources). The report emphasizes that "the extra application-specific logic is crucial for mitigating harms" β a problematic foundation model might be tolerable if appropriate downstream constraints are applied.
-
Deployment: The direct social impact occurs when the adapted system is deployed to users. The report advocates for "gradual releases, where deployment happens to an increasing fraction of users" to mitigate potential harms.
This pipeline view enables the report to distribute ethical and technical responsibility across stages. For example, data quality issues are primarily a curation problem; harmful outputs can be partially addressed at adaptation time through output filters; and deployment decisions can be made independently of model construction decisions. The report's formulation is "think ecosystem, act model" β researchers and practitioners whose purview is restricted to one stage (typically training) need surrogate metrics that predict downstream impact, and they need documentation standards that communicate their model's properties to downstream adapters.
Modeling Principles: Five Properties That Give Rise to a Foundation Model
The report identifies five properties that a model architecture should possess to serve as an effective foundation model (Section 4.1, Figure 17). These are not implementation specifications but desiderata β design goals that guide architectural choices and explain the dominance of the Transformer.
Expressivity concerns "the theoretical and practical capacity of a network to model the data distribution it is trained over and represent it in a flexible manner" (Section 4.1.1). The report traces how increasing depth (more stacked non-linear layers) enhances expressivity by enabling "powerful hierarchical and distributed representations." Different inductive biases suit different modalities: CNNs capture spatial invariance for images, RNNs capture sequential dependencies for text and time-series, and attention mechanisms in Transformers capture long-range dependencies and pairwise interactions between elements. The report identifies the attention mechanism's multiplicative interaction as particularly flexible: "dynamically adapting the computation to the input at hand" rather than using "rigid fixed-weight computation." A concrete example illustrates this: given the sentence "She ate the ice-cream with the X," an attention-based model can adapt its computation depending on whether X is "spoon" (updating the representation of "ate") or "strawberries" (linking to "ice-cream"). This context-dependent computation is what enables language models to handle the compositional flexibility of natural language.
The report identifies a fundamental trade-off between task-specialization and expressivity: "models with stronger structural priors can leverage them to improve sample efficiency... while conversely, models that integrate weaker inductive biases learn more slowly, but can in turn scale to higher volumes of data and adapt to a diverse set of domains." This explains the historical shift from RNNs and CNNs (strong inductive biases) to Transformers (weaker inductive biases, higher expressivity): as data and compute have become more abundant, the benefits of letting the data determine the appropriate structure have outweighed the sample-efficiency benefits of baked-in architectural assumptions. The report frames this as "a more promising approach for future research."
Scalability means that the model must be "easy-to-train" and "easy-to-adapt" while being "practically efficient" (Section 4.1.2). On the training side, this requires resilience to optimization instabilities like vanishing or exploding gradients. On the adaptation side, it requires overcoming catastrophic forgetting and supporting few-shot learning. Crucially, scalability is about hardware compatibility: "much of the transformers' great success over the previously dominating recurrent approach was driven by their higher degree of parallelism." Recurrent networks process tokens sequentially, creating a computational bottleneck; Transformers process all tokens in parallel through self-attention, mapping efficiently to GPU hardware designed for parallel matrix operations. The report envisions future architectures that are "amenable to schemes such as distributed training" and that "leverage properties such as sparsity" β predictions that align with subsequent developments like mixture-of-experts models.
Multimodality is "a key component of intelligence... a crucial factor for the development of both thorough and broad comprehension of the world" (Section 4.1.3). The report observes that "language learning is more effective when occurring in a grounded environment" and that "language encourages the emergence of abstractions that link between low-level perceptual signals and statistics to semantic concepts." The architectural design choices for multimodality center on the degree and timing of cross-modal interaction: late-fusion models like CLIP maintain "fully separate encoders for each data source, and compare their spaces only at the ultimate computation stage, using a simple dot product," while early-fusion models like ViLBERT "jointly reason over multiple modalities" throughout the model. The optimal stage for merging modalities remains an open research question.
Memory refers to the mechanisms for "access, storage, retrieval and manipulation of particular items or memories" (Section 4.1.4). The report advocates for separating computation from memory β "disentanglement between memory and computation has been a recurring goal" β rather than encoding all knowledge implicitly in network weights. This separation has multiple advantages: it "mitigates the inflation in models' size and number of parameters needed to store the growing quantities of knowledge," it "improves models' trust and reliability by increasing their knowledge provenance" (you can inspect what fact was retrieved), and it enables "memory update, manipulation or adaptation" without retraining the entire model. The report contrasts this with fully implicit knowledge: GPT-3 stores facts in its 175 billion parameters, making it impossible to update a single fact without retraining. Retrieval-based models like REALM and RAG store knowledge in explicit text databases, enabling fact updates by simply replacing text passages. However, the report notes a trade-off: "over-reliance on retrieval reduces the opportunities to learn how to represent information in compact and abstract manners" β models that can always look up facts may not develop the abstraction capabilities that implicit representations force. The in-context learning abilities of GPT-3 are hypothesized to "possibly emerge as a by-product of enforcing the network to represent the input sequential data through its bounded memory architecture."
Compositionality means that "the meaning of the whole is derived from the meaning of its constituent parts, and the rules applied to combine them" (Section 4.1.5). This principle underlies "capabilities to plan, reason and learn readily and efficiently from a handful of examples" and "may hold the key to achieve out-of-distribution β or specifically β combinatorial generalization." Compositionality can manifest at multiple levels: the model architecture (modular sub-networks), the computation (specialized experts), the training data (decomposable into subsets), and the learned representations (structured, object-oriented). However, the report also notes that compositionality "can also hinder the expressivity of the representation, and impede its capacity to account for idiosyncrasies, exceptions, and contextual correlations" β the classic example being that "red wine is not the same as red onion." The goal is "a better balance between contextuality and compositionality."
Training and Adaptation Framework
The training and adaptation framework is the engine that converts raw data and model architecture into useful capabilities. The report structures this around a central distinction: training objectives determine what the model learns, and adaptation mechanisms determine how that learned knowledge is specialized for specific tasks.
Training objectives are "mathematical functions describing how to transform a model architecture and large amount of broad data into a foundation model" (Section 4.2). The report defines three key goals for training objectives:
-
Leveraging broad data: Self-supervised learning "unlocked the power of internet-scale datasets which would be intractable to annotate by hand." Different data modalities require "bespoke self-supervised algorithms that leverage the unique structure within each kind of data" β for text, masked language modeling (predict missing words) or autoregressive language modeling (predict next word); for images, contrastive learning (make representations of augmented versions of the same image similar, and different images dissimilar); for multimodal data, matching images to their captions.
-
Domain completeness: Solving the pretraining task "requires capabilities that are broadly useful for downstream tasks." Language modeling may require models to "acquire capabilities as wide-ranging as coreference, sentiment and translation as the model learns to predict the next word in a document," whereas "a supervised learning task like sentiment classification may lead to a more narrow set of capabilities." The report notes that "it is not obvious a priori what tasks will result in domain complete capabilities, or even how to evaluate the full breadth of a model's capabilities."
-
Scaling and compute efficiency: Training procedures must "reliably convert data, a model architecture, and compute into a broadly capable model." The efficiency of training objectives varies dramatically β the report cites "4x for ELECTRA vs BERT" and "12x for contrastive vs generative approaches to CLIP training" β meaning that for a fixed compute budget, the choice of training objective can produce models with substantially different capabilities.
The report identifies three fundamental design trade-offs in current self-supervised methods:
Input representation abstraction level: Models can operate at "the level of raw bytes" (treating text as sequences of Unicode characters, images as pixel values), but "this high dimensionality may cause the model to focus on predicting less semantic aspects of the input" (like audio compression artifacts rather than word meanings) and becomes computationally intractable for Transformer architectures whose costs grow quadratically with input length. Tokenization β splitting text into subword units, images into patches β reduces the input space but "may jettison possibly-useful information in the input" (like character-level patterns needed for rhymes or puns).
Generative vs. discriminative training: Generative approaches train models to learn joint or conditional distributions over inputs, enabling flexible interaction β autoregressive models can generate continuations, denoising models can fill in gaps. Discriminative approaches train models to produce useful representations without generation capabilities, but "may enable more efficient learning for classification- or regression-based tasks in high-dimensional continuous settings like images." The report identifies capturing "the best of both approaches" as an open research avenue.
Capturing multimodal relationships: Multimodal training can take different forms based on the intended use. CLIP-like models encode modalities separately, enabling cross-modal retrieval and classification. ViLBERT-like models process modalities jointly from an early stage, enabling tasks that require fine-grained cross-modal reasoning like visual question answering. The report notes that "multimodal foundation models remain a nascent research area; much is still unexplored."
The report looks forward to three developments: "out-of-the-box SSL" that works across domains without bespoke design per modality, "obtaining a richer training signal" through objectives orders of magnitude more efficient than current ones, and "goal-directed training of foundation models" where the ability to understand and carry out goals is part of the pretraining objective itself β for example, training sequence models on goal-directed trajectories so that "prompting" them with a desired outcome produces appropriate behavior.
Adaptation mechanisms convert the unfinished foundation model into a task-specific system (Section 4.3). The report identifies three factors that govern adaptation choices:
Compute budget (storage and memory): For "foundation models with billions or trillions of parameters, fine-tuning all model parameters may demand prohibitively large memory." The report surveys "low-storage adaptation" methods that freeze most parameters and learn only a small number of task-specific ones. These include: tuning only the bias vectors (BitFit), learning low-rank residuals to weight matrices (LoRA), learning "soft prompts" β continuous vectors that are prepended to inputs rather than discrete text prompts β or interleaving small trainable adaptor modules between frozen layers. The report notes that these lightweight methods "sometimes achieve comparable performance to full fine-tuning, despite updating 1000Γ fewer parameters," and that "the performance gap between full fine-tuning and lightweight adaptation vanishes as the model size increases."
Data availability: When labeled adaptation data is scarce, "combining prompting and fine-tuning has been shown to be a promising direction." The report cites evidence that "a well-tuned prompt can be worth around 100 training examples" β reframing the task through careful prompt design can substitute for labeled data.
Access to foundation model gradients: This is a new consideration specific to the foundation model paradigm. On one extreme, if only model outputs are accessible (black-box API access), adaptation is limited to "in-context learning" β conditioning the frozen model on a prompt containing task instructions or examples. On the other extreme, with full gradient access, standard fine-tuning is possible. As a middle ground, with gradient access only to inputs (not parameters), lightweight methods like prefix tuning (optimizing continuous prompts while keeping the model frozen) become available.
The report expands the scope of adaptation beyond the standard task specialization framing to include: (a) temporal adaptation β updating models as the world changes (new facts, language evolution) without full retraining, perhaps through retrieval-based architectures; (b) domain specialization β adapting to a specific domain (legal text, medical images) through an intermediate stage of continued pretraining on unlabeled domain data before task-specific fine-tuning; (c) local model editing β correcting a model's behavior for specific inputs without changing its behavior for unrelated inputs, useful for fixing particular errors or removing memorized personal information; and (d) applying constraints β adapting models to satisfy privacy constraints (machine unlearning), remove toxic outputs, or comply with regulations.
The long-term vision is continual adaptation: enabling a foundation model to be "trained continuously on a non-stationary stream of data from different tasks, domains, or time periods" without catastrophic forgetting β a grand challenge that "requires closing the performance gap between a foundation model trained continuously on a non-stationary stream... and the same foundation model trained from i.i.d. data from the aggregate mixture" (Section 4.3.3).
Data and Systems Infrastructure
The report treats data and systems not as peripheral concerns but as foundational infrastructure that determines what foundation models can practically achieve.
Data management is addressed through a proposed architecture called the data hub (Section 4.6). The starting observation is that "current practices in foundation model development are generally ad-hoc across the entire lifecycle from data curation and data documentation to model monitoring and patching." The data hub is a holistic data management solution that addresses four desiderata:
-
Scalability: Foundation model datasets are massive (WuDao 2.0 trained on 4.9 TB of multimodal data) and growing, requiring "standard data management solutions such as infrastructure to store and maintain large-scale datasets as they change over time and scalable interfaces to query, select, and filter datasets."
-
Data integration: The value of foundation models often comes from combining modalities and structured/unstructured data. The hub "should incorporate data integration as a first class citizen" β the report cites examples of models that "use unstructured text data with structured entity knowledge or image data," and growing needs to "integrate datasets across diverse modalities such as text, video, eye-tracking, and robotic simulations."
-
Privacy and governance controls: Training data "may risk the violation of the privacy of data subjects; their data may be disclosed, collected, or used without their consent or outside the context for which consent was originally given." The report notes that "the issue of consent and use is especially relevant for foundation models where downstream applications cannot always be anticipated." The data hub needs tooling to support "diverse documentation, e.g., dataset sheets or data statements" and mechanisms to "help identify licensing violations and mitigate the impact of any governance violation."
-
Data quality monitoring: Training data "can contain different types of biases... and consist of poisoned, false, or duplicated information." The data hub should provide tooling for "slice finding" to identify subpopulations where models underperform, "model validation on relevant subsets," and "data valuation" to understand which training examples contribute most to model behavior. The vision includes supporting "community around sharing useful metrics and analysis pipelines."
The report explicitly raises open questions about data versioning (how to update datasets while maintaining reproducibility), provenance (how to trace where data came from while respecting privacy), and community governance (who stores the data, who pays for compute, who is liable if licensing is violated).
Computer systems are identified as "one of the largest bottlenecks to developing foundation models" (Section 4.5). The scale of the challenge is quantified: "the compute and memory requirements of state-of-the-art language models have grown by three orders of magnitude in the last three years, and are projected to continue growing far faster than hardware capabilities" (Figure 19). The report analyzes three dimensions of the systems challenge:
-
Performance through co-design: Training the largest models requires "careful co-design across algorithms, models, software, and hardware." The key parallelization strategies are: data parallelism (replicating the model across GPUs, each processing different data), pipeline parallelism (splitting model layers across GPUs in a pipeline), and tensor model parallelism (splitting individual layer computations across GPUs). These strategies have limits: data parallelism is limited by batch size, pipeline parallelism by the number of layers, and tensor model parallelism by the number of GPUs in a single server. The report argues that "to facilitate the next major leap in model capacity... it will be increasingly critical to co-design training algorithms, models, software, and hardware, because many of the avenues to dramatically increase performance alter the semantics of the training computation." Examples include lower-precision arithmetic (fp16), weight sparsity (only computing on non-zero weights), and novel architectures that map more efficiently to hardware (sparse attention, mixture-of-experts).
-
Retrieval-based models as a systems innovation: These models "store knowledge outside the model parameters in the form of text passages" and "use scalable top-k search mechanisms to extract knowledge pertinent to each input, while keeping the DNN model itself small" (Section 4.5.1). This design "improves computational efficiency as well as maintainability of the model in production: for example, developers can update the knowledge of the model just by replacing a text passage, without needing to retrain a large DNN." The report notes that this approach demands "functionality not readily supported by popular ML frameworks and nearest-neighbor indexes."
-
Automated optimization: The space of possible optimizations (combining different parallelization strategies, precision formats, compression techniques) is combinatorially large, and "manual experimentation is extremely expensive and time-consuming at the scale of thousands of GPUs." The report calls for "new software tools, libraries, and compilers to automatically identify compositions of optimizations that target comprehensive metrics like time-to-accuracy."
-
Productionization: Deploying foundation models requires addressing inference latency, model compression (distillation, quantization, pruning), and lifecycle management including "automated dataset curation and model quality assurance." The report advocates for "behavioral testing and model assertions" that "provide analogs to unit tests, runtime monitoring... and continuous model improvement" for machine learning systems.
A novel proposal concerns execution and programming models for the multi-model ecosystem that foundation models create (Section 4.5.3). Because many adapted models share the same underlying foundation model, there are opportunities for sharing computation and storage: "two models prefix-tuned from the same pretrained model can share the same model 'stem,' reducing the storage footprint... while also making it possible for execution to be shared and batched across the prefix-tuned models." The report argues that current frameworks do not provide "annotations to specify that various adapted models are derived from the same pretrained model" or that "various components of two models share parameters" β information that could enable substantial optimization.
Evaluation and Interpretability Frameworks
The report treats evaluation and interpretability as the primary mechanisms for making emergence tractable and homogenization safe. Without them, foundation models are opaque and their risks are unknowable.
Evaluation is reconceptualized around the intermediary nature of foundation models (Section 4.4). The report distinguishes two classes:
Intrinsic evaluation directly measures properties of the foundation model itself, "divorced from a specific task due to the task-agnosticity of these models." This is motivated by two limitations of evaluation tied to the training objective (e.g., perplexity for language models): it "lacks generality" (models with incompatible objectives cannot be compared) and it "relies upon a proxy relationship" that "likely will break down when assessing more diverse capabilities." The report proposes two complementary approaches: (a) "imputing intrinsic evaluation from broad extrinsic evaluation" β adapting the foundation model to a wide range of tasks and inferring its properties from aggregate performance, as in meta-benchmarks like SuperGLUE β and (b) "direct evaluation of intrinsic properties" β measuring specific capabilities (e.g., syntactic knowledge, social biases) through targeted probes, potentially inspired by psycholinguistic and psychological measurement methods. The value of intrinsic evaluation is demonstrated through the case study of in-context learning: "traditional task-based extrinsic evaluation does not provide a clear means by which in-context learning could have been identified; directly interacting with the foundation model appears to be necessary."
Extrinsic evaluation measures the performance of adapted, task-specific models, but the report argues that existing task-specific evaluation "often amounts to unfair comparisons between task-specific models produced with different (and, potentially, unequal) resources." The proposed solution is to account for adaptation resources: "all data used to adapt the foundation model and the data used to choose the adaptation method," as well as "the level of access required to adapt the foundation model" (black-box API access vs. gradient access). This enables fair comparison between adaptation methods: methods that achieve similar accuracy but require less data or less access are demonstrably better.
The report proposes reforms to evaluation design (Section 4.4.4): because foundation models enable sample-efficient adaptation, benchmarks can be "much smaller (since far less data needs to be provided as 'training', i.e., adaptation, data) and are far more diverse," shifting emphasis from quantity to quality and diversity. Leaderboards should report "measurements across diverse fronts" beyond accuracy β "robustness, fairness, efficiency and environmental impact" β and should allow stakeholders to "interact and manipulate how the ranking is done to align with their values."
Interpretability is structured around what the report calls the "one modelβmany model" paradigm (Section 4.11, Figure 23). A foundation model can be viewed as "one model" that "utilizes some set of generalizable model mechanisms to perform well across tasks and domains" β in this case, interpretability means identifying and characterizing these shared mechanisms. Or it can be viewed as "a large collection of independent expert models, each tailored to a specific task" β in this case, "explanations of model behavior in one task are therefore not necessarily informative about behavior in other tasks." The report argues that "understanding where foundation models lie on this spectrum between one and many models will be central to understanding their behavior."
The report proposes three levels of understanding, inspired by Marr's levels of analysis:
Characterizing behavior (the "what"): Identifying the full range of capabilities a foundation model possesses. This is made challenging by the "myriad unforeseen behaviors and tasks that these models are capable of performing" and by the sensitivity of behaviors to prompt formulation β "slight variations in prompts can result in meaningful changes of model behavior." The report advocates for controlled evaluations designed by domain experts, similar to psycholinguistic tests that determine whether a language model can distinguish grammatical from ungrammatical sentences.
Explaining behavior (the "why"): Providing explanations in terms of causes in the data. Current approaches β feature attribution, training example influence β provide local or global explanations but "do not necessarily provide insights into the model's behaviors for other (even seemingly similar) inputs, let alone other tasks and domains." The report cautions against over-reliance on self-explanations generated by models themselves, since "language models, and now foundation models, are exceptional at producing fluent, seemingly plausible content without any grounding in truth."
Characterizing model mechanisms (the "how"): Understanding the internal representations and algorithms that produce behavior. This is the deepest level: "by understanding individual model mechanisms, we can build up a compositional understanding of complex behaviors." The concrete example traces how one might investigate whether a model's mechanism for addition is reused when solving word problems β if it is, that increases confidence in the model's generalization; if it is not, that suggests the model is using a different, potentially less reliable heuristic. The report emphasizes that "establishing evidence of such a mechanism in a foundation model and its use can support a moral or legal responsibility to ban the model from tasks like predictive policing, marketing, loan applications, and surveillance at large."
The report concludes the interpretability section by highlighting a critical tension: "work aimed at interpreting foundation models is a double-edged sword." Incremental advances in interpretability "can be exaggerated to 'ethics-wash' and continue use of models as though they have achieved interpretability, belying the reality that they remain far below traditional standards of algorithmic interpretability." The responsibility of interpretability researchers is to ask "whether one is working toward making foundation models interpretable to researchers and model owners or interpretable to everyone."
Integrating Societal Analysis into Technical Design
The report's most distinctive technical contribution is its framework for integrating societal considerations β fairness, robustness, security, safety, legality β into the technical design process rather than treating them as separate ethical concerns.
Fairness is decomposed into intrinsic and extrinsic dimensions (Section 5.1): intrinsic biases are "properties of the foundation model that indirectly but pervasively affect downstream applications" β misrepresentation through stereotypes, underrepresentation or erasure of marginalized groups, overrepresentation of dominant perspectives. Extrinsic harms are "harms that arise in the context of specific downstream applications" β performance disparities, toxic outputs, psychological harms. The critical analytical contribution is the source tracing framework: "in order to fully characterize and properly intervene on the harms of foundation models, we must be able to trace their source to the properties of the foundation model and the adaptation process, and further decompose to the roles of individual sources of biases." Sources are categorized into data (training data, adaptation data, user interaction data), modeling decisions (training objectives, architectures, adaptation methods), and modelers (lack of diversity in development teams, community values). The report argues that "attributing the sources for bias and harm is fundamental for questions of intervention and responsibility; attribution requires new technical research to be done reliably." The framework distinguishes proactive intervention (data-centric and model-centric changes to prevent harm) from reactive recourse (mechanisms for feedback, accountability, and legal responsibility when harm occurs).
Robustness is analyzed through the lens of distribution shifts (Section 4.8). The report makes a nuanced argument: foundation models are "a particularly promising approach to robustness" because "pretraining on a large and diverse unlabeled dataset" improves performance across a wide variety of distribution shifts, "in contrast to many robustness interventions which are constrained to narrow types of distribution shifts." The evidence cited includes CLIP achieving "6% higher accuracy on ImageNetV2 and 35% higher accuracy on ImageNet Sketch" compared to standard ResNet models with equivalent in-distribution accuracy. However, the report also identifies persistent challenges: foundation models "may exacerbate or mitigate the effects of spurious correlations, but this depends on the nature of the particular downstream task" β they can help by "quickly learning from counterexamples to the spurious correlations" in diverse training data, but can also "exacerbate the issue by introducing biases present in the foundation model training data." The report warns that "foundation models cannot be assumed to automatically extrapolate within a given modality" β zero-shot transfer in CLIP "suffers greatly in satellite image domains," and ImageNet pretraining "does not substantially improve the performance of large models on medical images."
Security and privacy are reframed around the single-point-of-failure problem (Section 4.7). The report posits that "the security role of foundation models in future machine learning systems will be akin to the role played by the operating system in traditional software systems." This is a double-edged characterization: a foundation model "may become a single point of failure and thus a prime target for attacks against applications derived from this model," but also, "a foundation model imbued with strong security and privacy properties could form the backbone for the design of a variety of secure and reliable ML applications." Concrete threats include: data poisoning (adversarial manipulation of training data, demonstrated in the CLIP setting where "modifying as little as two out of 3 million training examples" could change model behavior), function creep (using models beyond their intended purposes, like repurposing CLIP for facial recognition despite its model card explicitly placing surveillance as out-of-scope), multimodal inconsistencies (adversaries exploiting discrepancies across modalities, like wearing clothes with imprinted text to evade visual recognition systems), and memorization of training data enabling extraction attacks. The opportunities include: "cheaper private learning" β foundation models pretrained on public data can be adapted for sensitive tasks with "significantly less confidential data" β and a potential path to adversarial robustness through the combination of scale and unlabeled data.
AI safety is analyzed through the lens of emergent goal-directed behavior (Section 4.9). Traditional AI safety concerns focused on reinforcement learning agents with explicitly specified reward functions, where the central challenge is value alignment β preventing reward hacking and ensuring corrigibility. Foundation models introduce a new challenge: "goal-directed behavior may emerge despite not being explicitly optimized for." The report gives the example of large language models trained on corpora "where agents use language in goal-directed ways, such as in persuasive text" β "to predict the next token well, a model may acquire a general capability to reason and produce arguments, which could emerge with suitable contexts." This means that safety concerns previously reserved for explicitly goal-directed RL agents may become relevant for self-supervised foundation models. The report identifies three challenges in characterizing capabilities: "the generality of foundation models means that they can be applied to countless different kinds of applications in unexpected ways"; "model capabilities are emergent: they grow and change in unexpected ways as models scale"; and "even within a particular application and scale, a model's capabilities are not easy to characterize" β small prompt variations can dramatically alter performance. The report flags potential catastrophic risks including "correlated failures that span multiple critical functions or failsafes" if a single foundation model is integrated into multiple critical systems, and the risk of "optimizing misaligned yet easy-to-specify goals" at scale (Goodhart's Law).
Legal analysis (Section 5.4) maps the uncertain legal landscape: training data collection may implicate the Computer Fraud and Abuse Act, copyright law (the "transformative use" doctrine applied to model training), and privacy laws (Illinois Biometric Information Privacy Act, GDPR, CCPA with its "right to be forgotten"); model outputs raise questions of liability under tort law and civil rights law (disparate treatment claims); and legal protections for outputs involve First Amendment questions about "AI speech" and copyright questions about ownership of machine-generated content. The report does not resolve these questions but systematically identifies the areas of legal uncertainty that foundation model development must navigate.
The ethics of scale (Section 5.6) provides a framework for decision-making about foundation model development and deployment. The report discusses homogenization as an ethical risk: "if the same model is used across a variety of domains with minimal adaptation, the strengths, weaknesses, biases, and idiosyncrasies of the original model will be amplified." The concept of algorithmic monoculture is introduced β "employing many adaptations of the same foundation model for multiple automated decision-making tasks means that decision subjects may face a more homogeneous set of judgments" β with the consequence that individuals may be "consistently and arbitrarily rejected, mis-classified, or ill-treated" across multiple systems. The report discusses surveillance and power concentration: the data demands of foundation models incentivize "aggressive data collection, even when that pursuit is legally questionable or contrary to user expectations," and the computational costs mean that "the organizations most capable of producing competitive foundation models will be the most well-resourced: venture-funded start-ups, already-dominant tech giants, and state governments."
The report then proposes concrete norms and mechanisms: documentation standards (model cards, data statements, nutrition labels), reporting structures for downstream feedback to propagate to foundation model developers, release and auditing protocols (staged release with neutral third-party oversight to ensure diverse auditing), and institutional mechanisms for deciding "when not to build" β the report argues that "technologies reflect a set of choices made by humans; human agency shapes the technological frontier" and that developers should "make deliberate and judicious choices about what is worth the time, financial resources, expertise, and energy use to build."
Environmental impact (Section 5.3) is quantified through a cost-benefit framework:
where $V(M)$ is the net value of a model, $S(M)$ is the net social and environmental benefit, $C(M)$ is the social cost of carbon from energy use (the EPA upper bound estimate was E(M)O(M)$` captures second-order environmental effects (chip manufacturing impacts, compounding climate effects, strain on chip production). The framework is explicitly not about precise dollar valuation but about "the existence of and relative importance of each of these effects." The report emphasizes that "addressing such emissions is an imperative" given accelerating climate change, and recommends mitigation strategies: training in low-carbon-intensity energy grids, more efficient hardware and architectures (mixed-precision training, quantization, sparse models, distillation), and systematic reporting of carbon and energy impacts.
Economic analysis (Section 5.5) frames foundation models as "what economists refer to as a general-purpose technology" β like the steam engine or electricity β with three key characteristics: "pervasiveness, improvement over time, and ability to spawn complementary innovations." The economic effects are analyzed through three dimensions: productivity and innovation (foundation models can increase output per unit input and potentially enhance creativity itself), wage inequality (the technology can either substitute for or complement human labor, with different effects on employment and wages), and centralization (the high costs of training create barriers to entry that concentrate ownership and power). The report highlights that the economic outcomes "are not dictated solely by technology or economics, but by the choices and actions of technologists, policymakers, managers, workers, and other members of society."
4. Key Insights and Innovations
Innovation 1: A Diagnostic Framework Based on Emergence and Homogenization
The report's most fundamental intellectual contribution is not a model or algorithm but a diagnostic vocabulary β two concepts, emergence and homogenization, that together explain why foundation models represent a paradigm shift rather than an incremental advance, and why this shift creates a qualitatively new risk profile. This is a conceptual innovation that reframes what the AI community is actually building and what it should be worried about.
Prior to this report, the dominant narratives about large-scale models oscillated between breathless capability reporting ("GPT-3 can do arithmetic!") and broad ethical critique ("stochastic parrots"). What was missing was a framework that linked the technical mechanism (scale-induced, self-supervised learning on broad data) to the practical consequence (a single model underpinning hundreds of applications) in a way that clarified both opportunity and risk. The terms "pretrained model" and "self-supervised learning" described the how but not the so what.
The concept of emergence captures the fact that behaviors like in-context learning, arithmetic, and code generation were "neither specifically trained for nor anticipated to arise" (Section 1.1). This is not merely a claim about capability β it is a claim about epistemology: we cannot predict a foundation model's full capability surface from its training objective or architecture alone. The report positions this as the endpoint of a 30-year trend (Figure 1, Section 1.1): machine learning made how tasks are performed emerge from data; deep learning made features emerge from training; foundation models make entire functionalities emerge from scale. Each stage removes another layer of explicit human specification, making the system progressively harder to understand through inspection of its design.
The concept of homogenization captures the sociological consolidation β "almost all state-of-the-art NLP models are now adapted from one of a few foundation models" (Section 1.1). This is significant because it converts model quality from a local property (this specific classifier is biased) to a systemic property (all classifiers built on BERT inherit its Anglocentric similarity metric, Section 5.6.1). The report documents that this homogenization is spreading across modalities β the same Transformer architecture is "now applied to text, images, speech, tabular data, protein sequences, organic molecules, and reinforcement learning" (Section 1.1) β suggesting convergence toward a unified modeling substrate.
The critical intellectual move is the coupling of these two concepts. The report's thesis is that "homogenization and emergence interact in a potentially unsettling way" (Section 1.1). Emergence means we don't fully know what the model does; homogenization means whatever it does propagates widely. This interaction creates what the report calls the "central challenge": "derisking." This framing is innovative because it converts a diffuse set of concerns β bias, fairness, robustness, safety, interpretability β into a coherent diagnostic problem: given that emergence creates opacity and homogenization creates leverage, how do we systematically reduce the risk that unknown defects in a few models cause widespread harm? The answer, threaded throughout the report, is evaluation, documentation, and auditing β making emergence tractable through rigorous measurement and making homogenization safe through transparency and accountability.
This framing is fundamental rather than incremental: it provides a unified explanation for why foundation models are simultaneously the most promising and most concerning development in AI, and it organizes the report's sprawling scope (26 sections) around a coherent analytical spine.
Innovation 2: The Ecosystem Pipeline as a Framework for Distributed Responsibility
The report introduces the ecosystem view β a five-stage pipeline (data creation β data curation β training β adaptation β deployment) β as the analytical framework for understanding foundation model impact. This is a diagnostic innovation that addresses a structural problem in AI governance: foundation models are "unfinished intermediate objects that can be adapted to many downstream applications, sometimes by an entirely different entity for unforeseen purposes" (Section 1.2), making it impossible to assign responsibility to a single stage or actor.
Prior frameworks for algorithmic accountability typically assumed a specific system deployed for a specific purpose β you audit the hiring algorithm, you test the recidivism prediction tool. Foundation models break this model because they are adaptable: the same BERT model that powers a question-answering system for medical information also powers a toxic comment classifier. The harms that manifest in one application may originate in training data decisions made years earlier, or in adaptation choices made by a different organization, or in deployment contexts that the original model developers never anticipated.
The ecosystem view solves this by distributing analysis across stages and identifying the interfaces between stages as the critical points for intervention. The report's formulation β "think ecosystem, act model" β captures the dual requirement: researchers focused on training need surrogate metrics that predict downstream impact, and downstream deployers need documentation that communicates upstream properties. The proposal for "surrogate metrics for a representative set of potential downstream evaluations" (Section 1.2) is a concrete operationalization: if you cannot test every possible adaptation, you need a representative battery of tests whose results provide a reasonable bound on likely behavior.
The significance of this contribution lies in its rejection of both extremes β the techno-optimist view that model developers bear no responsibility for downstream misuse, and the precautionary view that any model with potential for harm should not be built. Instead, the pipeline creates a gradient of responsibility: data curators are responsible for documentation, model trainers for intrinsic evaluation and capability disclosure, adapters for task-specific safety measures and output constraints, and deployers for monitoring and gradual rollout. This is not a solution to the accountability problem, but it is the necessary conceptual infrastructure for any solution to be articulated β you cannot assign responsibility across stages until you have identified the stages.
The pipeline also makes explicit that "deployment of adapted foundation models is a decision separate from their construction" (Section 1.2, Figure 3). This distinction between research and deployment β a recurring theme β creates space for scientific investigation while acknowledging that actual social impact requires additional safeguards. The report argues that "research models are often not extensively tested and might have unknown failure modes; warning labels should be placed on research models that are not fit to deploy. On the other hand, deployed foundation models that actually affect people's lives should be subject to much more rigorous testing and auditing" (Section 1.2). This is a pragmatic, stage-gated approach that balances the value of open research with the imperative of safety.
Innovation 3: "One ModelβMany Models" as an Interpretability Principle
The report introduces the "one modelβmany models" paradigm (Section 4.11, Figure 23) as a framework for understanding foundation model behavior. This is a conceptual innovation that clarifies what interpretability means for foundation models and why existing interpretability methods, developed for task-specific models, are insufficient.
The core question is: does a foundation model use "some set of generalizable model mechanisms to perform well across tasks and domains" (one model) or is it better understood as "a large collection of independent expert models, each tailored to a specific task" (many models)? The answer has profound implications. If a foundation model is fundamentally one model with shared mechanisms, then understanding its behavior on a few representative tasks provides insight into its behavior across all tasks β interpretability is feasible. If it is fundamentally many models with task-specific mechanisms, then "explanations of model behavior in one task are therefore not necessarily informative about behavior in other tasks" β interpretability is, in the limit, impossible, because each adaptation creates a de facto new system requiring independent analysis.
Prior interpretability work implicitly assumed the "one model" framework: feature attribution methods developed for image classifiers were applied to language models, probing methods assumed that linguistic knowledge was encoded in shared representations, and circuit analysis sought universal computational motifs. The "one modelβmany models" framing makes this assumption explicit and challenges it, creating a research agenda: determine where specific models fall on this spectrum, develop methods that can distinguish shared mechanisms from task-specific specializations, and understand the conditions (data, architecture, training) that push models toward one regime or the other.
The report argues that "understanding where foundation models lie on this spectrum between one and many models will be central to understanding their behavior" (Section 4.11). This is a diagnostic contribution β it identifies the right question to ask rather than providing the answer. Its significance is that it prevents the field from making category errors: applying interpretability methods that assume shared mechanisms to models where mechanisms are task-specific, or concluding that models are uninterpretable because task-specific analyses fail to generalize.
The report complements this with a three-level framework for understanding (what, why, how) that maps naturally onto the one-many distinction. Characterizing what a model does (what capabilities it has) is agnostic to the one-many question; explaining why it produced a particular output (causes in the data) depends on the assumption that explanations generalize across instances; and characterizing how it works internally (mechanisms and representations) requires identifying whether those mechanisms are shared or task-specific. This structured approach to interpretability is valuable because it separates tractable questions (what can the model do?) from harder ones (are the mechanisms for addition the same as those for translation?), preventing the conflation that often muddies interpretability discourse.
Innovation 4: Source Tracing as the Core Technical Challenge for Fairness
The report introduces source tracing as the central technical framework for fairness in foundation models β the idea that "in order to fully characterize and properly intervene on the harms of foundation models, we must be able to trace their source to the properties of the foundation model and the adaptation process" (Section 5.1.3). This reframes fairness from a measurement problem (are outcomes equal across groups?) to an attribution problem (which stage of the pipeline caused the disparity?).
Prior work on algorithmic fairness focused predominantly on the measurement and mitigation of disparities in model outputs β auditing classifiers for demographic performance gaps, applying constraints during training to equalize error rates, post-processing predictions to satisfy statistical parity. Foundation models complicate this picture because the same intrinsic bias (e.g., BERT's Anglocentric similarity metric) manifests differently across adaptations, and the same extrinsic harm (e.g., misgendering in translation) may originate from training data, model architecture, adaptation data, or deployment context. Without tracing the harm to its source, interventions are guesswork: you might debias the adaptation data when the root cause is in the foundation model's training corpus, or you might modify the foundation model when the problem is in how it was fine-tuned.
The report taxonomizes sources into data, modeling decisions, and modelers, and argues that "attributing the sources for bias and harm is fundamental for questions of intervention and responsibility; attribution requires new technical research to be done reliably." This is a significant reframing because it connects technical analysis (influence functions, causal probing, representational auditing) to legal and ethical questions (who is responsible when an adapted model causes harm?). The report notes that "significant technical, policy, and legal work is needed in order to develop frameworks for communicating data, model, and derivative contents... to attribute responsibility for harms; and to create avenues for recourse" (Section 5.6.3).
The innovation here is the integration of technical and normative analysis into a single framework. Source tracing is simultaneously a technical challenge (how do you trace a specific output back to specific training examples or architectural choices?) and a governance challenge (what do you do when you've identified the source?). The report does not solve either β it identifies the gap and argues that solving it requires interdisciplinary collaboration "commensurate with [foundation models'] fundamentally sociotechnical nature" (Section 1.1). This is a conceptual advance rather than a technical one, but it is foundational: it defines the problem that subsequent technical work on fairness in the foundation model paradigm must address.
Innovation 5: Rigorous Uncertainty About Understanding β The Philosophical Stance
The report makes a distinctive philosophical contribution in Section 2.6 by refusing to resolve the question of whether foundation models can understand language, instead systematically mapping the space of possible answers and their dependencies on unresolved philosophical disputes. This is intellectually valuable not because it provides an answer but because it prevents premature closure β it shows that confident claims in either direction ("they're just stochastic parrots" vs. "they genuinely understand") rest on contested metaphysical commitments that the field has not acknowledged.
The framework distinguishes three philosophical positions on understanding: internalism (understanding requires the right internal representational structures), referentialism (understanding requires knowing how words connect to things in the world), and pragmatism (understanding is constituted by appropriate use of language in context). Each position has different implications for what would constitute evidence of understanding in a foundation model, and each is compatible with different conclusions about whether current or future models achieve it. The report's conclusion is deliberately modest: "skepticism about the capacity of future foundation models to understand natural language may be premature. It is by no means obvious that foundation models alone could ever achieve understanding, but neither do we know of definitive reasons to think they could not."
What makes this an innovation rather than a literature review is its operational value: by clarifying what would count as understanding under different philosophical frameworks, it provides guidance for evaluation design. If you hold a pragmatist view, behavioral testing is the right approach β but the history of the Turing Test shows that behavioral benchmarks are invariably met with skepticism and goalpost-moving. If you hold a referentialist view, grounding in multimodal data (images, audio, sensor readings) is necessary, and "multimodal training regimes may well be the most viable strategy" (Section 2.6.4). If you hold an internalist view, structural probing and causal intervention methods are the path forward. The report thus converts a philosophical debate into a research program: rather than arguing about whether GPT-3 "understands," the productive question is "what would constitute evidence of understanding under each framework, and how do we design experiments to test for it?"
This is significant because the "do they understand?" debate was becoming a barrier to progress β it generated heat without light, and the lack of clarity about what was being debated made it impossible to resolve. The report's framework provides the necessary philosophical hygiene to make the debate productive: specify your metaphysical commitments, identify the corresponding evidentiary standards, and design experiments accordingly. It does not settle the question, but it provides the tools for settling it, which is a more fundamental contribution.
5. Experimental Analysis
Evaluation Methodology
-
Dataset. The report does not present a single unified experimental evaluation with standardized train/test splits, metrics, and baselines in the manner of a typical empirical ML paper. Rather, the report is a comprehensive survey that synthesizes findings from hundreds of prior works, each with their own experimental protocols. When specific experiments are cited β such as CLIP's robustness on ImageNet distribution shifts, GPT-3's few-shot learning performance, or BERT's fine-tuning results across NLP benchmarks β they refer to evaluations conducted in the original papers. The report itself does not introduce new empirical datasets or experimental benchmarks. It does reference specific evaluation frameworks for illustration: the MATH benchmark for mathematical reasoning (Section 2.4), the SuperGLUE meta-benchmark for language understanding (Section 4.4), and domain-specific datasets like CaseHOLD for legal NLP (Section 3.2), among many others. The "data" for the report's empirical claims is the collection of published results across the surveyed literature.
-
Base model(s). The report surveys a wide range of foundation models rather than conducting experiments on a single family. Central examples include BERT (Devlin et al., 2019) with its bidirectional encoder architecture and masked language modeling objective; GPT-3 (Brown et al., 2020) with 175 billion parameters and autoregressive language modeling enabling in-context learning; CLIP (Radford et al., 2021) as a multimodal vision-language model trained on 400 million image-text pairs; DALL-E (Ramesh et al., 2021) for text-to-image generation; and domain-specific models like MathBERT (Shen et al., 2021) for educational applications, LegalBERT (Chalkidis et al., 2020) for legal document processing, and AlphaFold (Jumper et al., 2020) for protein structure prediction. These models are chosen as representative exemplars of the foundation model paradigm across modalities and application domains.
-
Metrics. The report discusses multiple evaluation dimensions rather than a single metric. Performance metrics are task-specific and drawn from the original literature: accuracy on classification and question-answering tasks, BLEU and ROUGE for generation quality, FID (FrΓ©chet Inception Distance) for image generation quality, perplexity for language modeling, and pass@k metrics for code generation. Beyond standard accuracy, the report advocates for evaluation across additional desiderata: robustness to distribution shifts (effective robustness, defined in Section 4.8.1 as the gap between in-distribution and out-of-distribution performance), fairness metrics (demographic performance disparities, representational bias measures like the Implicit Association Test-inspired word embedding tests), efficiency (FLOPs, inference time, memory usage), and environmental impact (carbon emissions in kgCO2eq).
-
Baselines. Since the report is a survey rather than a primary empirical study, baselines are not systematically defined across experiments. The report does contrast foundation model approaches against prior paradigms: fully supervised models trained from scratch on task-specific data (the dominant approach before BERT), domain-specific feature engineering pipelines, and non-neural machine learning methods. Specific comparisons mentioned include: CLIP vs. standard ResNet models on ImageNet robustness benchmarks (Section 4.8.1); adapted foundation models vs. human performance on standardized exams (GPT-3 on various benchmarks, BERT-based systems moving from 73.1% to 91.6% on 8th grade science questions, Section 2.1.2); and retrieval-based models vs. fully parametric models for knowledge-intensive tasks (Section 4.5.1).
-
Generation budget / compute accounting. The report discusses computational cost primarily in qualitative terms, emphasizing the extreme scale of foundation model training. Specific figures cited include GPT-3's training requiring "> 1000 petaFLOP/s-days" (Section 4.5), and growth trends showing that "the compute and memory requirements of state-of-the-art language models have grown by three orders of magnitude in the last three years" (Section 4.5, Figure 19). The report advocates for compute-aware evaluation: comparisons between foundation models should account for total training FLOPs and adaptation resources (Section 4.4.3), and inference efficiency should be measured for deployment scenarios. The FLOPs accounting framework in Section 4.5.1 uses the standard approximation
$X = 6ND_{\text{pretrain}}$for pretraining FLOPs and$Y = 2ND_{\text{inference}}$for inference FLOPs, where$N$is parameter count. -
Cross-validation / statistical protocol. The report does not describe cross-validation or statistical testing procedures for its own claims, as it does not conduct independent empirical experiments. When discussing evaluation methodology (Section 4.4), it addresses broader issues of validity: the gap between test-set performance and real-world deployment (ecological validity), the need to account for all data used in adaptation including prompt selection data (Perez et al., 2021, cited in Section 4.4.3), and the danger of overfitting to benchmarks through repeated leaderboard optimization (Goodhart's Law, Section 4.4.4). The report argues for expanding evaluation beyond single-number leaderboard rankings toward multi-dimensional reporting that captures robustness, fairness, efficiency, and environmental cost.
Main Quantitative Results
The report synthesizes quantitative results from the surveyed literature rather than presenting original experiments. The following organizes the key empirical findings by domain, as presented in the report.
Language Capabilities and Adaptation Performance
The headline empirical pattern is the substantial and consistent improvement that foundation models provide over prior task-specific architectures across NLP benchmarks. In Section 2.1.2, the report cites a representative example: "the best system for answering open-ended science questions in 2018, before foundation models, could get 73.1% on the NY Regents 8th grade science exam. A year later in 2019, an adapted foundation model scored 91.6%" (Clark et al., 2019). This 18.5 percentage point improvement illustrates the paradigm-level gains.
The report characterizes the adaptation efficiency of foundation models. In Section 4.3.1, it cites Le Scao and Rush (2021): "a well-tuned prompt can be worth around 100 training examples, and fine-tuning a carefully prompted foundation model is significantly more data-efficient than fine-tuning an unconditioned foundation model." This quantifies the value of prompt engineering in low-resource settings. The report further notes that lightweight adaptation methods "sometimes achieve comparable performance to full fine-tuning, despite updating 1000Γ fewer parameters" (Section 4.3.1, citing Zaken et al., 2021; Li and Liang, 2021; Hu et al., 2021), and that "the performance gap between full fine-tuning and lightweight adaptation vanishes as the model size increases" (citing Lester et al., 2021).
On the negative side, GPT-3 shows substantial sensitivity to input formulation. Section 4.11.1 notes that for sentiment classification, GPT-3 "will exhibit different response accuracies for each prompt" depending on minor rewordings, and Section 4.9.2 reports that "the ability of GPT-3 to perform addition improves dramatically once commas are added to the inputs." This sensitivity demonstrates that reported accuracy numbers can be misleading β they may represent an upper bound achievable only through extensive prompt engineering rather than consistent model capability.
Robustness to Distribution Shifts
Section 4.8.1 provides quantitative evidence for foundation models' improved robustness. The report cites Radford et al. (2021): "both CLIP and a standard ResNet50 obtain 76% accuracy on ImageNet, but CLIP achieves 6% higher accuracy on ImageNetV2 and 35% higher accuracy on ImageNet Sketch." This 35 percentage point improvement on sketch-based recognition represents a qualitatively different capability β CLIP has acquired some degree of abstraction beyond texture-based recognition.
The report contrasts this with the limitations of other robustness interventions: "many other robustness interventions, such as adversarial training, invariant risk minimization, and using larger models have had little impact on effective robustness on these ImageNet tasks, especially without explicit knowledge of the distribution shift" (Section 4.8.1, citing Taori et al., 2020; Santurkar et al., 2020; Radford et al., 2021; Miller et al., 2021). The key metric here is effective robustness β the gap between in-distribution and out-of-distribution performance β and the report argues that diverse pretraining data improves this metric more reliably than targeted robustness algorithms.
However, the report documents significant failures. Section 4.8.2 notes that "zero-shot transfer in CLIP suffers greatly in satellite image domains" (citing Radford et al., 2021), and that "ImageNet pretraining does not substantially improve the performance of large models on medical images" (citing Raghu et al., 2019; Ke et al., 2021). These domain-specific failures reveal that robustness improvements from diverse pretraining are not universal β they depend on the overlap between pretraining and target distributions.
Scaling and Compute Efficiency
Section 4.5 places specific numbers on the computational demands: GPT-3 required "> 1000 petaFLOP/s-days" for training. Figure 19 (Section 4.5) illustrates that "the compute and memory requirements of state-of-the-art language models have grown by three orders of magnitude in the last three years," a rate that "far exceeds the rate of increase in computational capacity of hardware (roughly 10Γ in four years)." This divergence implies that efficiency improvements from systems co-design are essential for continued scaling.
Section 4.2 notes that the efficiency of different training objectives varies substantially: "4x for ELECTRA vs BERT" and "12x for contrastive vs generative approaches to CLIP training." These numbers indicate that algorithm choice can match or exceed the gains from hardware scaling, making training objective design a critical research frontier.
For adaptation, Section 4.3.1 mentions that "lightweight adaptation techniques... sometimes achieve comparable performance to full fine-tuning, despite updating 1000Γ fewer parameters." This is a systems insight β for practitioners with limited storage, adapting a foundation model can be dramatically cheaper than storing full fine-tuned copies for each task.
Multilinguality and Low-Resource Languages
Section 2.1.3 provides nuanced quantitative patterns for multilingual models. The report notes that "multilingual models show better performance in languages that are similar to the highest-resource languages in their training data" and that "languages in multilingual models compete for model parameters." The GShard neural machine translation model is cited as showing "the largest gains over monolingual baselines for the lowest resource languages, with the gains increasing with model size" (Lepikhin et al., 2021). This suggests that scale can actually benefit low-resource languages β the opposite of what might be expected if high-resource languages dominate capacity. However, the report also documents limitations: for many languages, "the text data available is not enough to train a large-scale foundation model" β "there are over 65 million speakers of Fula, a West African language, but few if any resources available for NLP in Fula" (Section 2.1.3, citing Nguer et al., 2020).
Multimodal and Domain-Specific Results
The report cites domain-specific performance improvements from foundation model adaptation. In law (Section 3.2), LegalBERT (Chalkidis et al., 2020) "improves over BERT... on the downstream task of text classification and sequence tagging in legal documents," though the report notes that "continual foundation model training may perform worse than re-training from scratch in certain domains such as legal documents" β a negative result that complicates the narrative of universal transfer benefits.
In healthcare (Section 3.1), the report cites protein structure prediction and drug discovery applications: "successful applications range from predicting viral mutations that can escape a vaccine-induced immune response to predicting protein docking potential for better design of therapeutic antibodies" (citing Bepler and Berger, 2021; Hie et al., 2021; Tsaban et al., 2021; Wu et al., 2021; Rives et al., 2021). These are capability demonstrations rather than benchmark comparisons, reflecting the early stage of foundation model application in biomedicine.
Security Vulnerabilities
Section 4.7.1 provides quantitative evidence of security risks. Carlini and Terzis (2021) show that "targeted attacks against CLIP-style models require modifying as little as two out of 3 million training examples" β a poisoning rate of approximately 0.00007% β to change model behavior. Schuster et al. (2021) demonstrate that "a code auto-completion system trained with GPT-2 on Github data can be poisoned into suggesting insecure code snippets with the injection of only a few malicious files." These numbers indicate that the scale of training data does not provide immunity against data poisoning β the extreme specificity possible with targeted attacks means that a tiny fraction of poisoned data can compromise model behavior.
Effort and Cost of Content Creation
Section 5.2.1 provides economic context for misuse potential. The report cites a 2017 Russian influence operation with a budget of "75-$200 per article to American freelancers as part of a disinformation campaign." Foundation models will "lower these marginal costs" (Section 5.2.1), enabling scaled personalized content generation that was previously constrained by human writer costs.
Ablation Studies and Robustness Checks
The report itself does not present original ablation studies or controlled experiments. However, it surveys and synthesizes findings from the literature that serve the function of ablation analysis β isolating which factors contribute to foundation model performance and where claims break down.
Training objective efficiency varies dramatically: Section 4.2 reports that ELECTRA achieves approximately 4Γ the training efficiency of BERT, and contrastive approaches to CLIP training are approximately 12Γ more efficient than generative approaches. These are effectively ablation comparisons that isolate the impact of training objective choice while holding model architecture and data roughly constant. The magnitude of the difference β an order of magnitude or more β indicates that training objective design is not a marginal optimization but a first-order determinant of model capability per unit compute.
Domain specialization sometimes hurts rather than helps: Section 4.3.2 reports a finding from Cole et al. (2021): "fine-tuning a model pretrained only on the iNaturalist animal classification dataset provides better downstream performance than fine-tuning a model pretrained on iNaturalist along with 750K other images." This is a non-obvious negative result β more diverse pretraining data does not always improve downstream performance, and in some cases domain-specific pretraining on a narrower distribution yields better results. Similarly, LegalBERT (Chalkidis et al., 2020), which is "pretrained only on legal documents, improves over BERT, which is trained on a much more diverse training set, on the downstream task of text classification and sequence tagging in legal documents." The report notes that "continual foundation model training may perform worse than re-training from scratch in certain domains such as legal documents" (Section 4.3.2). These results serve as ablations on the diversity principle β they demonstrate that the relationship between pretraining data breadth and downstream performance is not monotonic.
Lightweight adaptation can match full fine-tuning despite 1000Γ fewer parameters: Section 4.3.1 surveys methods including BitFit (bias-only tuning), LoRA (low-rank weight residuals), prefix tuning, and adapter modules. The finding that these methods "sometimes achieve comparable performance to full fine-tuning" is a robustness check on the hypothesis that large-scale parameter updating is necessary for effective adaptation. The result that "the performance gap... vanishes as the model size increases" (Lester et al., 2021) suggests that larger models have more parameter redundancy, making lightweight adaptation increasingly viable.
Retrieval augmentation trades memorization capacity for generalization: Section 4.1.4 notes the trade-off in retrieval-based models: "over-reliance on retrieval reduces the opportunities to learn how to represent information in compact and abstract manners." The report hypothesizes that GPT-3's "in-context learning abilities... possibly emerge as a by-product of enforcing the network to represent the input sequential data through its bounded memory architecture" β without external retrieval, the model is forced to develop internal abstraction capabilities. This is a conceptual ablation on the role of external memory, suggesting that the choice between parametric and retrieval-based knowledge storage has consequences for the types of capabilities that emerge.
Adversarial robustness does not scale automatically with model size: Section 4.7.2 reports that "despite their unprecedented scale, current foundation models unfortunately see little gains in robustness to worst-case adversarial perturbations" (citing Fort, 2021; Wallace et al., 2019). This is a critical negative result β it means that scaling alone, without dedicated adversarial training or architectural innovations, does not solve the adversarial robustness problem. However, the report notes a more optimistic finding for non-adversarial robustness: "multimodal models such as CLIP are surprisingly robust to (non-adversarial) distributional shifts." The distinction between adversarial and natural distribution shifts reveals that robustness is not a unitary property β models can be simultaneously robust to natural variation and brittle to optimized perturbations.
Static model knowledge degrades over time: Section 4.3.2 cites work showing that "large language models become outdated" (Lazaridou et al., 2021; Hombaiah et al., 2021; Dhingra et al., 2021). Techniques like "re-weighting training data and dynamic evaluation... can partially alleviate, but not fully solve, this problem" (Section 4.3.2). This serves as a robustness check on the implicit assumption that a trained model's knowledge remains valid β the world changes, facts become outdated, and language evolves, and current methods for temporal adaptation are insufficient.
Bias mitigation methods are brittle: Section 5.1.4 surveys negative results in fairness interventions, noting that "methods that measure or combat intrinsic bias are brittle or ineffectual" (citing Gonen and Goldberg, 2019; Ethayarajh et al., 2019; Bommasani et al., 2020; Zhou et al., 2021; Antoniak and Mimno, 2021), and that there is "some evidence to suggest certain types of technical intervention may be simultaneously unsatisfiable, impossible, or may even exacerbate inequity" (citing Corbett-Davies and Goel, 2018; Kleinberg et al., 2017; Lechner et al., 2021; Xu et al., 2021). These negative results serve as robustness checks on the assumption that fairness is primarily a technical problem amenable to algorithmic solutions β the evidence suggests that current debiasing techniques often fail to address root causes and may provide false confidence.
Carbon costs vary dramatically with energy grid: Section 5.3 provides a sensitivity analysis: "QuΓ©bec has an extremely low carbon intensity due to its reliance on hydroelectricity, while Estonia's energy grid has an extremely high carbon intensity due to its reliance on shale oil" (citing Henderson et al., 2020). The top 5% of polluting power plants "contributed 73% of all electricity-based emissions" (citing Grant et al., 2021). This implies that the carbon footprint of training a foundation model can vary by orders of magnitude depending on where the computation is performed β a finding that serves as an ablation on the common practice of reporting a single emissions number without location context.
Critical Assessment
The report's central claims must be evaluated with careful attention to what kind of document this is. The report is a synthetic survey β it does not present new empirical results but rather organizes and interprets a vast body of existing work. This means standard experimental criticism (small test sets, missing baselines) applies to the individual studies cited rather than to the report itself. The appropriate critical question is: does the report's synthesis accurately represent the state of evidence, or does it overclaim what the evidence supports?
On the claim that emergence and homogenization characterize a paradigm shift: This is a conceptual claim, not an empirical one, and must be assessed on its analytical utility rather than experimental validation. The report provides substantial qualitative evidence for both phenomena. Emergence is demonstrated through specific examples β in-context learning in GPT-3, arithmetic ability arising without explicit training β that were documented in the cited literature. Homogenization is documented through the dominance of BERT, GPT-3, and similar models across NLP, and the spread of Transformer architectures across modalities. What the report does not provide is systematic quantitative evidence for the degree of homogenization β what fraction of deployed NLP systems actually use foundation models, how this fraction has changed over time, or whether the trend is accelerating. The evidence is suggestive and illustrative rather than exhaustive. This is a gap, but perhaps an unavoidable one: the report aims to characterize an ongoing shift, and comprehensive data on industry deployment practices is proprietary and inaccessible.
On the claim that foundation models improve robustness to distribution shifts: The evidence cited β CLIP's 35 percentage point improvement on ImageNet Sketch β is compelling but narrow. The report acknowledges this limitation: "foundation models cannot be assumed to automatically extrapolate within a given modality." The domain-specific failures (satellite imagery, medical imaging) serve as important qualifiers. A genuine empirical gap is the absence of systematic robustness comparisons across model families, data scales, and shift types β the existing evidence is scattered across papers with different protocols. The report does a service by synthesizing these findings, but a meta-analysis with harmonized metrics would be more informative than the current narrative synthesis.
On the claim that lightweight adaptation methods approach full fine-tuning performance: This is supported by specific citations but the report does not provide a comprehensive comparison. The finding that the gap "vanishes as the model size increases" (Lester et al., 2021) is based on experiments up to a particular scale (likely GPT-3 class models); whether it holds for future, larger models is unknown. Additionally, the report notes that most of this evidence comes from text domains β the transferability of lightweight adaptation effectiveness to vision or multimodal models is less established.
On bias and fairness claims: The report's taxonomy of intrinsic vs. extrinsic bias is conceptually valuable but empirically underdetermined. The report acknowledges fundamental measurement problems: "methods that measure or combat intrinsic bias are brittle or ineffectual" and "there is some evidence to suggest certain types of technical intervention may be simultaneously unsatisfiable, impossible, or may even exacerbate inequity." This is a more honest treatment than many fairness papers provide, but it also means that the report's framework for addressing bias β source tracing, proactive intervention, reactive recourse β is aspirational rather than validated. No experiment in the surveyed literature demonstrates that source tracing actually enables effective intervention at scale. The report identifies this as a research gap: "attributing the sources for bias and harm is fundamental for questions of intervention and responsibility; attribution requires new technical research to be done reliably." Until such research exists, the fairness framework is a conceptual scaffold awaiting empirical foundation.
On environmental impact: The report's cost-benefit framework $V(M) = S(M) - C(M) - E(M) - O(M)$ is analytically clear but practically uncomputable with current methods. The social cost of carbon has disputed values, the social benefit $S(M)$ of general-purpose models is impossible to quantify ex ante, and second-order effects $O(M)$ (chip manufacturing, compounding climate effects) lack established methodologies. The framework is valuable as a structure for thinking but not as an operational tool for decisions. The report is transparent about this, noting "uncertainty in which methodology to use when valuing each component" and that "the key takeaway of this cost-benefit analysis... is not the dollar valuation of each term... but rather the existence of and relative importance of each of these effects." This is intellectually honest but limits the framework's immediate practical utility.
On what the report does not cover: Several significant empirical questions are left unaddressed. The report provides almost no evidence on the failure modes of foundation models in production deployments β how often do they produce harmful outputs, under what conditions, and with what consequences? This reflects the state of the literature: systematic post-deployment auditing of foundation models was rare at the time of writing. The report's discussion of when models fail (Section 2.1.2 on language, Section 4.8 on robustness) relies primarily on benchmark evaluations and researcher probes, not real-world incident data. This is a significant evidentiary gap: the central concern about emergence and homogenization is that widespread deployment of poorly understood models could cause widespread harm, but the report cannot quantify this risk because the necessary monitoring infrastructure does not exist.
On the philosophical stance on understanding: Section 2.6 is a conceptual analysis, not an empirical one, and must be assessed accordingly. Its contribution is clarifying what would count as evidence of understanding under different philosophical frameworks. It does not attempt to determine whether current models achieve understanding, and it explicitly declines to resolve the underlying philosophical disputes. This is the appropriate stance for a survey report, but it means that the section provides no empirical guidance β it maps the space of possible positions without providing tools for discriminating between them. The claim that "multimodal training regimes may well be the most viable strategy" for achieving understanding is a plausible inference from the referentialist framework, but it is not empirically tested.
In summary, the report's central contribution is conceptual synthesis rather than experimental demonstration. Its empirical claims are largely drawn from the cited literature and are appropriately qualified. The most significant gaps are not in the experiments the report describes but in the experiments that do not yet exist: systematic post-deployment auditing of homogenized model ecosystems, validated source tracing methodologies for fairness, and computable environmental cost-benefit analyses. The report correctly identifies these as critical research directions, and its value lies as much in mapping what we do not know as in synthesizing what we do.
6. Limitations and Trade-offs
The Report Provides a Conceptual Framework, Not an Operational System
The assumption or constraint. The report is explicitly a survey, synthesis, and diagnostic framework, not a deployable system with measurable performance. It identifies "derisking" as the central challenge (Section 1.1), proposes intrinsic evaluation and source tracing as the necessary mechanisms (Sections 4.4, 5.1), and advocates for ecosystem-wide documentation and auditing (Section 5.6). However, the report does not implement or validate any of these mechanisms. It acknowledges this structural limitation implicitly throughout, framing itself as providing the "conceptual architecture" for understanding foundation models rather than a technological artifact whose performance can be benchmarked. The report states that "we currently lack a clear understanding of how they work, when they fail, and what they are even capable of due to their emergent properties" (Section 1.1), and this very lack of understanding is the problem the report diagnoses rather than solves.
The consequence. Without operational mechanisms for intrinsic evaluation, source tracing, or systematic auditing, the report's prescriptions remain aspirational. A practitioner who reads the report and concludes "we should do intrinsic evaluation before deployment" receives no validated protocol, no benchmark suite, and no evidence that intrinsic evaluation actually predicts downstream harm. The report argues compellingly that such mechanisms should exist, but cannot tell a deployer how to build them today. The gap between diagnostic framework and operational toolkit means that the report is most useful for structuring research agendas and policy discussions, but provides limited immediate guidance for engineers making concrete deployment decisions.
What evidence exists in the paper. The report is transparent about this throughout. Section 5.1.4 states that "technical mitigation of all forms at present is severely limited" and that "methods that measure or combat intrinsic bias are brittle or ineffectual" β the report is synthesizing negative results rather than providing solutions. Section 4.4.2 notes that "the mechanics of [intrinsic evaluation] are unclear" and offers only "general principles and considerations." Section 5.6.3 proposes documentation norms (model cards, nutrition labels) but acknowledges that "significant technical, policy, and legal work is needed" to make them effective. The limitations are not hidden β they are the report's central finding: we know enough to identify the problem but not enough to solve it.
Mitigation status. The report does not attempt to mitigate this limitation β it is inherent to the genre. A 200-page interdisciplinary survey cannot simultaneously be a validated engineering artifact. The report positions itself as the necessary first step (shared vocabulary, identified gaps, research agenda) that must precede operational solutions. Section 1.4 frames this explicitly: "we have attempted to clarify the nature of a paradigm that may only have just begun, rather than waiting for more to unfold or the dust to settle." The mitigation is the research community building on the report's framework to develop the missing operational tools.
Intrinsic Biases Are Hypothesized Rather Than Empirically Traced
The assumption or constraint. The report's fairness framework (Section 5.1) rests on the fundamental assumption that intrinsic biases in foundation models are the upstream causes of extrinsic harms in downstream applications, and that these harms can be "traced" to specific sources in the pipeline (training data, model architecture, adaptation method, modelers). The report states that "attributing the sources for bias and harm is fundamental for questions of intervention and responsibility; attribution requires new technical research to be done reliably" (Section 5.1.4). However, the report does not provide empirical evidence that such tracing has been successfully performed for any real-world harm. It points to correlations β BERT encodes an Anglocentric similarity metric, GPT-3 reproduces stereotypes β but does not demonstrate causal chains from specific training data or architectural choices to specific downstream harms in deployed systems.
The consequence. The source tracing framework may not be tractable in practice, even with improved tools. Foundation models are trained on terabytes of unlabeled internet data with minimal documentation; tracing a particular biased output to its origins may be computationally infeasible or fundamentally underdetermined (many possible sources could produce the same bias). If source tracing is impossible, then the report's fairness framework β which depends on attribution for "questions of intervention and responsibility" β collapses to a measurement framework: we can detect bias but cannot reliably determine its cause or assign responsibility. This would mean that the report's call for "proactive intervention" at specific pipeline stages cannot be targeted effectively β interventions would remain blunt (retrain with more diverse data, add output filters) without guidance on which intervention is appropriate for which harm.
What evidence exists in the paper. The report acknowledges this limitation explicitly, citing negative results throughout Section 5.1.4: "methods that measure or combat intrinsic bias are brittle or ineffectual" (with multiple citations), and there is "some evidence to suggest certain types of technical intervention may be simultaneously unsatisfiable, impossible, or may even exacerbate inequity." The report notes that "the relationship between the training data, along with associated data practices... and the intrinsic biases acquired by the foundation model remains unclear" (Section 5.1.3). These acknowledgments mean that the report's fairness framework is a hypothesis about how fairness could work in the foundation model paradigm, not a demonstration that it does work. No experiment in the surveyed literature validates the full source-tracing pipeline from training data to intrinsic bias to extrinsic harm to successful intervention.
Mitigation status. The report treats this limitation as a research agenda: "establishing scaling laws for bias, akin to those for accuracy metrics... may enable systematic study at smaller scales to inform data practices at larger scales" (Section 5.1.3). It also advocates for complementary approaches β documentation, auditing, participatory design β that do not depend on perfect source tracing. However, the report does not acknowledge the possibility that source tracing may be fundamentally impossible at foundation model scale, which would require a different approach to fairness (perhaps focusing exclusively on downstream monitoring and recourse rather than upstream intervention).
Difficulty Estimation Cost Is Unaddressed for Practical Deployment Decisions
The assumption or constraint. The report's evaluation framework (Section 4.4) proposes intrinsic evaluation β directly measuring foundation model properties like capabilities, biases, and failure modes β as a necessary complement to task-specific extrinsic evaluation. It also proposes the ecosystem pipeline view (Section 1.2) where each stage requires documentation and measurement. However, the computational and organizational cost of comprehensive intrinsic evaluation is never quantified. Running a battery of probes across a foundation model with hundreds of billions of parameters, testing for thousands of potential capabilities and biases, and repeating this for every model version β this cost is likely comparable to, or greater than, the cost of training the model itself. The report mentions the practical challenge only obliquely: the loss of accessibility means that "the actual training of foundation models is unavailable to the vast majority of AI researchers, due to the much higher computational cost and the complex engineering requirements" (Section 1.3). Evaluation at the same scale faces the same barrier.
The consequence. If comprehensive intrinsic evaluation costs as much as training, the economics of the ecosystem pipeline break down for all but the largest organizations. The report envisions a world where model providers document intrinsic properties, adapters report downstream behavior, and the feedback loop between them enables continuous improvement. But if only the organizations that can afford to train the model can also afford to evaluate it, then evaluation becomes an internal quality assurance process rather than a public accountability mechanism β and the report's vision of independent auditing by diverse stakeholders is undermined. The cost problem is particularly acute for the report's proposed "data hub" (Section 4.6): continuously monitoring data quality, detecting distribution shifts, and validating model behavior on fine-grained subpopulations requires infrastructure that may exceed the budgets of all but the largest industry labs.
What evidence exists in the paper. The report does not address evaluation cost directly. Section 4.6 discusses the data hub's requirements ("scalable interfaces to query, select, and filter datasets," "tooling to help identify licensing violations") but does not estimate the computational resources required. Section 4.5 details the extreme cost of training foundation models and the innovations needed to make training feasible, but does not provide analogous analysis for evaluation cost. The closest the report comes is acknowledging that "the ability to opt-out can be incorporated into the foundation model ecosystem at many stages" (Section 5.6.5) β but opt-out mechanisms require detection mechanisms, and detection mechanisms require computational resources. The absence of cost analysis for the evaluation infrastructure the report advocates is a significant gap.
Mitigation status. The report partially addresses this through its advocacy for community infrastructure: shared data hubs, public computing resources (the National Research Cloud initiative, Section 1.3), and open-source frameworks. However, these proposals address the distribution of cost (making it shared rather than borne by individual researchers) without addressing the magnitude of cost. If comprehensive intrinsic evaluation of a GPT-3-scale model costs millions of dollars in compute time, democratizing access to compute does not make the evaluation affordable β it just spreads the bill. The report does not explore whether lightweight, approximate evaluation methods could provide sufficient signals at lower cost, which would be necessary for the proposed ecosystem to function in practice.
The Analysis Binds to a Specific Moment in a Rapidly Evolving Field
The assumption or constraint. The report acknowledges at the outset that it is characterizing "a paradigm that may only have just begun" (Section 6). The models the report analyzes in depth β BERT, GPT-3, CLIP, DALL-E β represent a snapshot of the field circa 2021, and the report's claims about what foundation models can and cannot do are grounded in that snapshot. The report notes that certain capabilities (e.g., in-context learning) have "only been demonstrated in models of sufficient size" (Section 1.3), implying that the capability threshold could shift as models scale further. Similarly, the report's safety analysis (Section 4.9) explicitly notes that "current foundation models may be far from posing [catastrophic] risks; however, the breadth of their capabilities and potential applications is striking, and a clear shift from previous ML paradigms" β acknowledging that its risk assessment may not bound future risks.
The consequence. Several of the report's specific empirical claims may be outdated by the time the report is widely read, and more importantly, the report's framework for analyzing foundation models may need updating as the technology evolves. The report's analysis of capabilities (what foundation models can do) and limitations (where they fail) is tied to specific model architectures and scales. If new architectures enable fundamentally different capabilities β or if scaling continues to produce qualitative capability jumps β then the report's taxonomy of capabilities, its assessment of what is "hard" vs. "easy" for foundation models, and its risk analysis may require revision. The report's treatment of multimodal models as "nascent" (Section 4.1.3) and its expectation that "the scope of foundation models goes well beyond language" (Section 1.1.1) both anticipate this evolution but cannot predict its trajectory.
What evidence exists in the paper. The report provides considerable evidence of rapid capability evolution even within its snapshot. It notes that GPT-3's in-context learning was "neither specifically trained for nor anticipated to arise" (Section 1.1) β a capability that emerged between GPT-2 and GPT-3, roughly one year. It cites scaling laws showing "surprisingly regularity" (Kaplan et al., 2020) but notes that "due to the emergent nature of these foundation models, some functionalities like in-context learning have only been demonstrated in models of sufficient size, so scale is needed to even ask the right questions" (Section 1.3). This means the report cannot distinguish between "capabilities that foundation models lack in principle" and "capabilities that current-scale foundation models lack but larger ones may acquire." The report's Section 4.9.2 on AI safety discusses this problem explicitly: "model capabilities are emergent: they grow and change in unexpected ways as models scale. What the emergent properties of future foundation models will look like is unknown."
Mitigation status. The report partially addresses this gap through its forward-looking research agenda β identifying areas where current models fail and calling for investigation. The report also frames many of its claims as contingent: "at present," "current foundation models," "existing foundation models have the potential to." This careful hedging is appropriate but does not solve the fundamental problem: a diagnostic framework may be invalidated if the disease changes. The report's conceptual architecture (emergence, homogenization, ecosystem pipeline) is designed to be robust to specific model changes, but its detailed capability analysis and risk assessment are tied to the state of the art at the time of writing. The report does not propose a mechanism for updating its analysis as the field evolves, which means its utility as an ongoing reference depends on subsequent work extending and revising its claims.
The Report Provides No Resolution to the Tension Between Open Research and Safety
The assumption or constraint. The report identifies a fundamental tension between the benefits of open access (enabling diverse auditing, democratizing innovation, serving low-resource languages) and the risks of open access (enabling misuse by malicious actors, concentrating power in those who can develop models privately) β but does not resolve it. Section 5.6.4 asks "does the benefit of release outweigh the potential for harm from actors sophisticated enough to use a released model or API but not sophisticated enough to create their own?" and answers "we believe that the answer is yes," but acknowledges that this is a judgment call based on incomplete evidence. The report simultaneously argues that academia should have greater access to foundation models (Section 1.3, "loss in accessibility") and that release practices need careful risk assessment (Section 5.6.4, staged release), without providing a decision procedure for balancing these competing demands.
The consequence. Adopters of the report's framework face an unresolved dilemma. A researcher who follows the report's advice to pursue open science may inadvertently enable misuse; a company that follows the report's advice to implement staged release with neutral third-party oversight may slow innovation and concentrate power. The report provides the vocabulary for discussing this dilemma β emergence creates uncertainty, homogenization amplifies consequences β but does not provide a framework for making case-by-case decisions. The absence of a resolution means that different actors can cite the same report to justify opposite policies: "the report shows that open access enables auditing, therefore we should release" vs. "the report shows that misuse risks are real, therefore we should restrict access."
What evidence exists in the paper. The report provides evidence on both sides. On the benefits of release: multilingual models enable "cross-lingual transfer, which β when the models are open-sourced β may allow for adaptation to languages which otherwise would have too few texts available" (Section 5.6.4), and "the harms to be weighed against the benefits are those from less well-resourced actors who would not be able to create their own foundation model" β implying that the misuse risk from release is bounded because sophisticated actors can build their own models anyway. On the risks of release: the report documents that "foundation models will allow for the creation of content that is often indistinguishable from content created by humans" (Section 5.2.1), that they will "substantially decrease the costs of content creation" enabling scaled disinformation, and that they can "embarrass, intimidate, and extort victims" through deepfakes. The report does not attempt to quantify or compare the magnitudes of these competing effects.
Mitigation status. The report proposes institutional mechanisms that could partially address the dilemma β staged release boards with neutral third-party oversight, mandatory documentation, feedback mechanisms from downstream adapters to foundation model providers β but these are proposals for how to make the decision rather than what the decision should be. The report's Section 5.6.5 discusses "when not to build" as a moral question "rooted in context and values," and notes that "answering the question of when not to build is a matter of individual responsibility as well as a broader professional responsibility." This frames the dilemma as requiring collective deliberation and professional norm-setting, but does not provide the norms themselves. The report calls for "professional oaths," "regulatory bodies," and "official protocols for ethics review" (Section 5.6.5) β institutional infrastructure that did not exist at the time of writing and largely still does not exist. The unresolved tension between openness and safety remains the most consequential practical question for anyone developing or deploying foundation models, and the report's value is in clarifying why the question is so difficult rather than in answering it.
7. Implications and Future Directions
How This Work Changes the Landscape
This report is a taxonomic and diagnostic intervention, not an algorithmic one. It does not introduce a new model architecture, training objective, or benchmark β rather, it provides the shared vocabulary and conceptual scaffolding that the field lacked for discussing the paradigm shift already underway. Its impact should be measured by how it reshapes research agendas, institutional practices, and policy discourse, not by accuracy metrics on a downstream task.
The magnitude of the shift is disciplinary, not technical. The report's central contribution is the emergenceβhomogenization diagnostic: the insight that foundation models are defined by the interaction of two forces β capabilities that arise unpredictably from scale, and a consolidation of methodology that amplifies both the benefits and defects of a few models across thousands of applications (Section 1.1). This diagnostic reframes what was a diffuse set of observations (GPT-3 can do arithmetic, BERT dominates NLP, CLIP transfers across tasks) into a coherent framework with clear normative implications: because emergence creates opacity and homogenization creates leverage, the central challenge is derisking through evaluation, documentation, and auditing. This is a paradigm-level reframing of the field's relationship to its own artifacts.
Prior to this report, discourse about large-scale models was fragmented. NLP researchers studied BERT's linguistic capabilities (Section 2.1), computer vision researchers documented CLIP's robustness properties (Section 2.2), ethicists critiqued the environmental and bias costs of large models (Sections 5.1, 5.3), and economists analyzed automation potential (Section 5.5). Each community operated with its own vocabulary and concerns. The report's integration of these perspectives β the claim that all of these phenomena are manifestations of the same underlying shift β is what makes it transformative. A researcher studying bias in language models and a researcher studying few-shot adaptation in vision models are, in the report's framework, studying the same thing: the consequences of emergence and homogenization in different modalities. This unification creates intellectual coherence where there was fragmentation.
The report resolves (or at least reframes) several prior contradictions. The debate about whether large language models "understand" language (Section 2.6) had become a philosophical stalemate: one side pointed to fluent generation as evidence of understanding, the other pointed to brittle failures as evidence of "stochastic parrots." The report's framework does not settle this debate but makes it productive by mapping it onto philosophical positions (internalism, referentialism, pragmatism) and identifying what evidence would count under each. This moves the conversation from "do they understand?" (an unanswerable question without shared criteria) to "under what definition of understanding, with what evidence, could we determine whether they do?" β a tractable research program.
Similarly, the report reconciles conflicting findings about the effectiveness of robustness interventions. Section 4.8.1 documents that "many other robustness interventions, such as adversarial training, invariant risk minimization, and using larger models have had little impact on effective robustness," while foundation models trained on diverse data show substantial gains (CLIP's 35 percentage point improvement on ImageNet Sketch). The report's framework explains this discrepancy: targeted robustness algorithms make narrow assumptions about the nature of distribution shifts, while broad pretraining provides a general-purpose robustness prior β but only for shifts that overlap with the pretraining distribution. This explains both the successes (ImageNet Sketch shares visual concepts with internet photos) and the failures (satellite imagery and medical images are far from CLIP's training distribution).
The report also provides a framework for understanding negative transfer results that had been puzzling in isolation. Section 4.3.2 reports that "fine-tuning a model pretrained only on the iNaturalist animal classification dataset provides better downstream performance than fine-tuning a model pretrained on iNaturalist along with 750K other images" (Cole et al., 2021) and that LegalBERT "improves over BERT" on legal tasks despite being trained on less data (Chalkidis et al., 2020). These findings β that more diverse pretraining does not always help β are explained by the report's analysis of domain specialization: broad pretraining provides useful priors only when the target domain falls within the support of the pretraining distribution. When it does not, continued pretraining on domain-specific data is necessary. This is not a failure of the paradigm but a refinement of when transfer works.
The report redirects research attention toward evaluation and interpretability as first-class problems, not afterthoughts. Section 4.4.2 introduces the distinction between intrinsic and extrinsic evaluation and argues that the field has systematically underinvested in intrinsic evaluation β directly measuring what a foundation model knows and what biases it encodes, independent of any specific task. The report's framework implies that improving intrinsic evaluation methodology is more fundamental than improving any particular downstream task performance, because without reliable intrinsic evaluation, we cannot predict how a model will behave when adapted for new purposes. This reorients the NLP research community's relationship to benchmarks: SuperGLUE scores are signals of one kind of capability, but they are insufficient for characterizing a model's full behavior profile.
The "one modelβmany models" interpretability framework (Section 4.11, Figure 23) similarly redirects research: rather than developing ever-more-sophisticated feature attribution methods (which assume task-specific behavior), the priority should be determining whether foundation models use shared mechanisms across tasks or develop task-specific specializations. If models are fundamentally "one model" with generalizable mechanisms, interpretability is feasible and should focus on identifying those mechanisms. If models are "many models" with independent task-specific behaviors, interpretability is fundamentally more difficult and the field should adjust its expectations accordingly. This framework converts an unarticulated assumption (that models use shared mechanisms) into a testable hypothesis.
The report makes certain research directions more attractive and others less so. On the more attractive side: (a) intrinsic evaluation methodology, including psycholinguistic probes and psychological bias measures adapted for model auditing (Section 4.4.2); (b) source tracing for fairness, developing influence functions and causal attribution methods that can connect specific training data to specific downstream harms (Section 5.1.3); (c) retrieval-based architectures that separate memory from computation, enabling fact updates without retraining and providing provenance for outputs (Section 4.1.4); (d) low-storage adaptation methods (prompt tuning, adapters, BitFit) that reduce the cost of specializing foundation models (Section 4.3.1); and (e) systems co-design that treats efficiency as a first-class architectural constraint rather than an optimization after the fact (Section 4.5).
On the less attractive side: (a) developing bespoke architectures for individual tasks β the foundation model paradigm makes this increasingly unnecessary for tasks within the pretraining distribution's support; (b) treating fairness as purely a post-hoc measurement problem without engaging the pipeline view β the report's source tracing framework implies that fairness auditing without causal attribution is insufficient; and (c) assuming that scaling alone will solve robustness, safety, or bias problems β the report documents multiple cases where larger models amplify rather than mitigate these issues (e.g., Section 5.6.1 on algorithmic monoculture, Section 4.7.1 on data poisoning affecting even the largest models).
Follow-Up Research This Work Enables
1. Developing and validating intrinsic evaluation benchmarks that predict downstream harm. The report identifies intrinsic evaluation β directly measuring foundation model properties β as necessary but acknowledges that "the mechanics of such evaluation are unclear" (Section 4.4.2). A concrete follow-up would construct a suite of intrinsic probes measuring representational biases (e.g., Implicit Association Test-style word embedding tests, Section 5.1.3), toxic generation propensity, factual accuracy, and reasoning capabilities, and then test whether performance on these probes predicts downstream harm in specific applications (e.g., does higher intrinsic bias predict higher demographic performance disparities in adapted hate speech classifiers?). The key question is whether intrinsic evaluation provides actionable signal or merely correlates with phenomena we could measure directly in adapted models. A negative result β finding that intrinsic bias measures do not predict extrinsic harm β would fundamentally challenge the report's fairness framework and suggest that fairness auditing should focus exclusively on deployed systems rather than upstream models.
2. Causal source tracing from training data to downstream harm. The report's fairness framework depends on the hypothesis that harms can be traced to specific sources in the pipeline β particular training examples, architectural choices, or adaptation decisions (Section 5.1.3). A concrete experiment would take a known downstream harm (e.g., GPT-3 generating stereotypical associations between demographics and occupations, documented in Section 5.1.2) and attempt to trace it backward: which training examples are most influential? Would removing or reweighting those examples reduce the harm? Does the bias originate in pretraining data, adaptation data, or somewhere else? Influence function methods (Koh and Liang, 2017, cited in Section 5.1.3) provide a technical starting point, but scaling them to foundation models with hundreds of billions of parameters and terabytes of training data is an open challenge. A positive result β successfully attributing a specific harm to a specific data source β would validate the source tracing framework; a negative result β finding that harms are distributed across millions of training examples with no clear causal origin β would suggest that fairness interventions must be systemic (better data curation overall) rather than targeted (removing specific examples).
3. Testing the one modelβmany models hypothesis through mechanistic interpretability. The report's interpretability framework (Section 4.11) poses the question: do foundation models use shared mechanisms across tasks, or do they develop task-specific specializations? A concrete experiment would identify a specific capability (e.g., arithmetic in GPT-3, cited in Section 4.11.3) and determine whether the same model components are used for addition, multiplication, and word-problem solving. One could use causal mediation analysis (Vig et al., 2020, cited in Section 2.6.3) to ablate specific attention heads or MLP layers and measure the effect on different tasks. If the same components are causally necessary across tasks, the model is functioning as "one model" with shared mechanisms, and interpretability efforts can focus on characterizing those shared components. If different components are activated for different tasks, the model is "many models," and interpretability must be task-specific. This experiment would provide the first empirical evidence on where current models fall on the one-to-many spectrum, directly informing whether interpretability research should pursue universal mechanism discovery or task-specific analysis.
4. Measuring the environmental cost amortization curve for foundation models. The report proposes that foundation models' environmental costs should be amortized over downstream adaptations (Section 5.3.2, Figure 27), but provides only a hypothetical calculation. A concrete follow-up would instrument the full lifecycle of a specific foundation model β training energy, adaptation energy for $N$ downstream tasks, and inference energy β and compare against the energy required to train $N$ task-specific models from scratch. The question is: at what $N$ does the foundation model approach become more energy-efficient? The hypothetical Figure 27 suggests a crossover point around 80 tasks for BERT-base, but this depends on adaptation method efficiency, hardware, and energy grid. Real measurements across multiple model scales (BERT-base, BERT-large, GPT-3-scale models) and adaptation methods (full fine-tuning, lightweight adaptation, in-context learning) would provide the empirical basis for the report's environmental cost-benefit framework, converting it from a conceptual tool to an operational one.
5. Auditing the effectiveness of staged release with neutral third-party oversight. The report proposes staged release with a "neutral third party" deciding who receives early access, arguing that this would prevent "charges of favoritism, selective distribution, and manipulating public perception" (Section 5.6.4). A concrete experiment would compare staged release programs with and without third-party oversight: do third-party-overseen programs produce more diverse auditor pools, more comprehensive documentation of model failures, and faster identification of biases? The key metric would be the number and severity of issues discovered during staged release that were not discovered during internal testing. A positive result β showing that third-party oversight substantially increases the discovery of harmful behaviors β would provide evidence for the report's proposed institutional mechanism. A negative result β finding that internal testing is comparably effective β would question whether the administrative overhead of third-party oversight is justified, or whether the bottleneck is not who does the testing but what testing methodologies are used.
6. Characterizing the difficulty-dependence of foundation model benefits. Although this report predates the formal study of difficulty-conditioned test-time scaling (later developed by works building on this framework), a key implication of the report's capability analysis is that foundation model benefits are not uniform across tasks. Section 2.1.3 documents that multilingual models "show better performance in languages that are similar to the highest-resource languages," and Section 4.8.2 documents that "zero-shot transfer in CLIP suffers greatly in satellite image domains." A concrete experiment would systematically characterize how foundation model adaptation performance varies as a function of the distance between pretraining and target distributions, across multiple modalities and task types. The goal would be to produce a "transferability map" that predicts, for a given target task, whether a foundation model will help, hurt, or have no effect compared to training from scratch. This would operationalize the report's call for better understanding of "the types of distribution shifts for which foundation models are effective" (Section 4.8.2) and provide practical guidance for when to use foundation models versus when to build task-specific systems.
Practical Applications and Downstream Use Cases
1. Multilingual NLP for low-resource languages in government and legal services. The report documents that multilingual foundation models enable "cross-lingual transfer, which β when the models are open-sourced β may allow for adaptation to languages which otherwise would have too few texts available" (Section 5.6.4). A concrete deployment scenario is government service provision in linguistically diverse regions: the report notes that there are "over 65 million speakers of Fula, a West African language, but few if any resources available for NLP in Fula" (Section 2.1.3). A multilingual foundation model adapted with a few hundred translated legal documents could enable basic document understanding β extracting key information from court filings, translating public health announcements, or routing citizen queries to appropriate agencies. The benefit is specific and measurable: the report cites GShard neural machine translation showing "the largest gains over monolingual baselines for the lowest resource languages, with the gains increasing with model size" (Lepikhin et al., 2021, cited in Section 2.1.3). This means deploying larger multilingual models directly translates to better service for the most underserved language communities.
2. Efficient adaptation for healthcare applications with limited labeled data. The report argues that foundation models "could be adapted to perform specific tasks with significantly less confidential data" than training from scratch (Section 4.7.2), and that the cost of expert annotation in healthcare makes sample efficiency critical (Section 3.1.3). A concrete deployment scenario is clinical NLP in a hospital with a specific documentation format and vocabulary: rather than requiring thousands of physician-annotated notes to train a named entity recognition system from scratch, a foundation model pretrained on biomedical text (BioBERT, cited in Section 3.1.1) and adapted with a few hundred annotated examples could achieve comparable accuracy. The report's quantitative justification comes from Section 4.3.1: "a well-tuned prompt can be worth around 100 training examples, and fine-tuning a carefully prompted foundation model is significantly more data-efficient." In practical terms, a hospital that can afford to annotate 200 clinical notes rather than 2,000 saves approximately $18\times$ in expert annotation cost (assuming roughly $100-200 per hour for physician annotators). This is the difference between feasibility and infeasibility for many clinical NLP deployments.
3. Staged auditing infrastructure for deployed foundation models. The report proposes staged release with neutral third-party oversight as a mechanism for discovering biases and failure modes before widespread deployment (Section 5.6.4). A concrete implementation would be a foundation model provider (e.g., a company deploying a language model for customer support automation) contracting with an independent auditing organization that recruits diverse testers β spanning different demographic groups, linguistic backgrounds, and use contexts β to probe the model for harmful behaviors during a pre-deployment phase. The report's quantitative context is the finding that "targeted attacks against CLIP-style models require modifying as little as two out of 3 million training examples" (Section 4.7.1, citing Carlini and Terzis, 2021) β a poisoning rate of approximately 0.00007% β and that models can exhibit "different response accuracies for each prompt" depending on minor rewordings (Section 4.11.1, citing Zhao et al., 2021). These findings mean that internal testing by a homogeneous team is unlikely to surface the full range of failure modes; diverse external auditors probing the model systematically would discover issues that developers miss. The practical benefit is preventing the deployment of models that would produce discriminatory outputs for underrepresented groups β a harm that the report documents across multiple modalities and applications (Section 5.1.2).
4. Carbon-aware model selection for large-scale inference deployment. The report's environmental cost framework and emphasis on amortization (Section 5.3.2, Equation 7, Figure 27) provide a practical decision tool for organizations deploying models at scale. A concrete scenario: a company running a commercial translation service must decide between deploying a large foundation model (high one-time training cost, low per-task adaptation cost) versus training task-specific models (no upfront training cost, higher per-task cost). Using the report's amortization framework, the company would estimate the number of language pairs it expects to serve ($N$), measure the energy cost of training the foundation model once versus training $N$ bilingual models, and compute the crossover point where the foundation model approach becomes more carbon-efficient. The report provides the FLOPs accounting formula ($X = 6ND_{\text{pretrain}}$, Section 4.5.1) and notes that efficient adaptation methods could shift the crossover point dramatically β lightweight methods that train only $1000\times$ fewer parameters (Section 4.3.1) make the foundation model approach carbon-efficient at much smaller $N$. The practical outcome is a data-driven decision rather than relying on intuition about whether foundation models are "greener" β a question the report shows cannot be answered without specifying the deployment context.