ArXiv: 2311.06477
🎯 Pitch
Generative AI appears to invert a foundational principle of copyright law, turning the idea-expression dichotomy “upside down” and challenging core legal assumptions. This cross-disciplinary workshop reveals that today’s litigation merely scratches the surface of the legal upheaval to come, arguing that a shared technical-legal vocabulary is the essential first step for effective governance.
1. Executive Summary
This report synthesizes the key takeaways from the inaugural Workshop on Generative AI and Law (GenLaw), convened in July 2023 to bring together cross-disciplinary experts in machine learning and law—with an emphasis on U.S. law—to address the technical, doctrinal, and policy challenges raised by generative-AI systems. Rather than proposing new technical methods, the workshop and this report identify four essential needs for productive interdisciplinary progress: building a shared knowledge base across disciplines (exemplified by a glossary defining terms like "memorization" and "pre-training" that carry different technical and colloquial meanings in each field), clarifying the distinctive technical capabilities of generative-AI systems relative to prior AI (notably the transition to open-ended generative models trained at massive scale via multi-stage pipelines), constructing a logical taxonomy of legal issues spanning intent in torts, privacy, misinformation, and intellectual property (where, for instance, the report observes that Generative AI "seems to turn the idea-expression dichotomy upside down" by inverting the traditional relationship between prompt and generated expression), and articulating a concrete research agenda that includes difficult problems like machine unlearning as an inadequate analogue for notice-and-takedown and the absence of agreed-upon evaluation metrics for legal concepts like memorization. A central finding is that copyright concerns—while the focus of current litigation—represent only a small fraction of the legal issues Generative AI will raise, establishing that meaningful legal analysis requires understanding technical design choices throughout the generative-AI supply chain as well as the business models (B2C hosted services, B2B integration, open-model derivatives, and specialized supply-chain actors) through which these systems reach users.
2. Context and Motivation
The Core Problem: Two Disciplines Talking Past Each Other
The fundamental issue this workshop report addresses is not a single technical or legal problem, but rather a communication and coordination failure between two communities—machine learning (ML) and law—that must urgently collaborate to address the societal implications of generative AI. As generative-AI systems rapidly advance and deploy, they raise legal questions that cut across intellectual property, privacy, torts, free speech, and criminal law. Yet the experts best positioned to analyze these questions—legal scholars and ML researchers—often lack the shared vocabulary, conceptual frameworks, and mutual understanding of each other's technical methods needed to produce rigorous, actionable analysis.
The report frames this gap explicitly in terms of four interconnected needs (Section 1):
-
A shared knowledge base providing a common conceptual language for experts across disciplines—without which terms like "memorization," "privacy," and "pre-training" create persistent misunderstandings that derail collaboration before it begins.
-
Clarification of what makes generative AI distinct from prior AI and computing technologies—because legal analysis that treats generative AI as "just another software system" will miss essential features of how these models are trained, deployed, and used.
-
A logical taxonomy of legal issues these systems raise—moving beyond the current copyright-centric public conversation to map the full landscape of legal doctrines that generative AI intersects.
-
A concrete research agenda to promote collaboration and knowledge-sharing on emerging issues—identifying specific problems (like machine unlearning, evaluation metrics for legal concepts, and centralized-versus-decentralized development pathways) where technical and legal expertise must be jointly applied.
This gap matters because law and technology do not evolve on independent tracks. The legal frameworks that govern generative AI—what constitutes infringement, what privacy protections apply, who bears liability for harmful outputs—will be shaped by arguments that draw on technical claims about how these systems work. If those technical claims are misunderstood (by lawyers) or if legally significant design choices are made without awareness of their legal implications (by technologists), the resulting legal rules and technical systems will both be worse for it.
Why This Problem Is Important: The Generativity Thesis
The report grounds its urgency in Jonathan Zittrain's theory of generative technologies (Section 2), arguing that generative AI hits what the authors call the "generativity jackpot" across all five of Zittrain's dimensions. Understanding this framework is essential because it explains why generative AI will create legal challenges of a scope and complexity comparable to those raised by computers and the Internet—Zittrain's two canonical examples from 2008.
The five dimensions, as applied to generative AI:
Leverage refers to a technology's capacity to make difficult tasks easier. Generative AI provides leverage across creativity, programming, knowledge synthesis, and automation of repetitive tasks—fundamentally changing the cost structure of producing text, images, code, and other expressive content. This leverage is what makes the technology economically significant enough that legal questions about it cannot be ignored by governments or markets.
Adaptability means the technology can be applied to a wide range of uses. The report emphasizes that generative AI has been applied to "programming, painting, language translation, drug discovery, fiction, educational testing, graphic design, and much more." This breadth means that legal issues will not be confined to a single domain (e.g., copyright for artists) but will cascade across many areas of law simultaneously.
Ease of mastery describes whether users without specialized training can adopt and adapt the technology. The report notes that while pre-training still requires technical skills, "the ability to use chat-style, interactive, natural-language prompting to control generative-AI systems greatly reduces the difficulty of adoption." This democratization of access means that the population of potential plaintiffs, defendants, and users whose behavior the legal system must regulate is vastly larger than for prior AI technologies that required programming expertise to use.
Accessibility concerns barriers to use—cost, regulation, secrecy, linguistic limits. The report observes that while creating cutting-edge models "requires enormous inputs of data, compute, and human expertise—currently limiting model creation to a handful of institutions," inference services are "widely available to the public" and "inexpensive for users making small numbers of queries." This asymmetry—concentrated creation, distributed use—creates a distinctive legal landscape where a small number of actors make design decisions that affect millions of users and rightsholders.
Transferability is the ease with which changes in the technology can be conveyed to others. The report notes that "once pre-trained or fine-tuned, generative-AI models can be easily shared, prompts and prompting techniques are trivially easy to describe, and systems built around generative-AI models can be made broadly available at increasingly low effort and cost." This means that legal interventions aimed at one deployment (e.g., requiring a hosted service to implement safeguards) must contend with the fact that the underlying model can be replicated and redeployed elsewhere.
The report's central claim here is both descriptive and predictive: generative AI is the first technology since the Internet to score highly on all five generativity dimensions, and this fact is what makes its legal challenges "comparable in scope, scale, and complexity" to those raised by computers and the Internet. This is not merely an analogy—it is an argument that the legal system should expect decades of doctrinal development, regulatory iteration, and interdisciplinary collaboration, rather than hoping for quick resolutions.
Where Prior Approaches Fall Short
The report identifies several categories of failure in how the ML and legal communities have engaged with each other and with the problems generative AI raises. These are not presented as a systematic literature review (the report is a workshop synthesis, not a research paper), but they emerge clearly from the roundtable discussions.
Communication Failures: Terminological Confusion
The most immediate practical barrier identified at the workshop was conflicting definitions of shared terms (Section 3). The report distinguishes between two types of misunderstanding:
Known terminological differences, where both communities are aware they use a term differently. The report's key example is "privacy." Technologists' formal definitions—such as differential privacy, which provides mathematical guarantees about what can be inferred about individual records in a dataset—do not encompass the full range of interests protected by privacy law, which defines privacy contextually based on "social norms and reasonable expectations." As the report notes, legal scholars have articulated real-world harms from misuse of private information; computer scientists have demonstrated attack vectors for leaking it. Both communities "have generally understood that they mean something different by 'privacy' and have read each other's work with a working understanding of these differences in mind."
Unknown terminological differences, where members of one community do not realize a loosely defined term in their field is a term of art in the other, or assume a meaning subtly different from how it is actually used. The report gives two concrete examples that emerged during the roundtable discussions:
-
"Pre-training": Technologists use this term to refer to "an early, general-purpose phase of the model training process" that produces a base model later adapted via fine-tuning. Legal scholars "assumed that the term referred to a data preparation stage prior to and independent of training." This is not a minor semantic quibble—it matters for legal analysis of who does what, when, and with what legal consequences. If pre-training is mischaracterized as data preparation rather than training, the entire copyright analysis shifts, because training is where copyrighted works are used as examples to update model parameters; data preparation (like deduplication or formatting) is a different category of activity with different legal implications.
-
"Harms": The report notes that "many technologists were not aware of the importance of harms as a specific and consequential concept in law, rather than a general, non-specific notion of unfavorable outcomes." In law, what counts as a cognizable harm determines standing to sue, the scope of available remedies, and whether conduct is regulated at all. Physical injury is the most widely accepted type of legal harm; other recognized forms include property damage, economic losses, and certain privacy violations. But "fear of future injury" or having one's data exposed in a breach without concrete downstream consequences may not constitute legally cognizable harm. A technologist who treats "harms" as synonymous with "any bad outcome the system might cause" is working with a concept that does not map onto the legal system's narrower, more structured framework.
These communication failures "hampered our ability to collaborate on assessing emerging issues." The report emphasizes that participants "found our way to common understandings only over the course of our conversations, and often only after many false starts"—implying that the current state of interdisciplinary discourse is characterized by unrecognized misunderstandings that waste time and produce confused analysis before productive collaboration can begin.
The Need for Shared Conceptual Infrastructure
Beyond terminology, prior approaches fall short by lacking shared mental models for how generative-AI systems actually work (Section 4). The report highlights a recurring question from legal scholars to ML experts during the roundtable: "What's so special about Generative AI?" This question reflects a genuine analytic need—lawyers need to understand which technical features of generative AI are genuinely novel (and thus require novel legal analysis) versus which are extensions of existing technology (and thus can be analyzed under existing frameworks).
The report's answer (detailed in Section 4) identifies three categories of novelty that prior approaches have not adequately captured: the transition from task-specific discriminative models to flexible generative ones (4.1), the multi-stage training pipeline and supply chain (4.2–4.3), and the massive scale at which these models operate (4.4). Without this characterization, legal analysis risks either (a) treating generative AI as legally equivalent to prior software, missing novel issues, or (b) treating it as entirely unprecedented, failing to apply relevant existing doctrine.
Legal Analysis Fragmented by Copyright Focus
The report identifies a significant gap in the scope of current legal engagement with generative AI: copyright dominates the conversation, but the legal issues extend far beyond IP (Section 5, and reiterated as a major takeaway in Section 7). The report is explicit that its own scope (privacy and IP, at the first workshop) "should be considered non-exhaustive, and the omission of other topics is not a judgment that they are unimportant."
The legal taxonomy in Section 5 covers intent in torts and criminal law (5.1), privacy (5.2), misinformation and disinformation (5.3), and intellectual property (5.4)—and even within IP, the discussion extends beyond copyright to trade secrecy, patent, and authorship eligibility. The implication is that the current legal conversation—dominated by high-profile copyright lawsuits and fair-use debates—represents "just the scratch the surface" of the issues generative AI will raise.
Moreover, the report argues that the copyright focus may lead to category errors in legal analysis. For example, in Section 6.1, the report discusses how the licenses under which datasets are released (e.g., MIT license for LAION) do not guarantee that the underlying data examples are licensed for use in model training. A lawyer focused only on copyright might look at an open-source dataset license and conclude the data is free to use, while missing that the constituent copyrighted works within the dataset may have their own restrictions that the dataset-level license does not override. The report frames this as a problem requiring both legal and technical innovation to resolve—better licenses, better provenance tracking, and potentially new legal frameworks—rather than something existing copyright doctrine can handle.
Technical Interventions Without Legal Grounding
The report identifies machine unlearning as a case study in how technical subfields can develop without adequate legal grounding (Section 6.3). Machine unlearning is a research area that "attempts to define the desired goals for removing an example [from a trained model] and to design algorithms that satisfy these goals"—motivated in part by the "Right to be Forgotten" in GDPR. But the report argues that "there is no straightforward analogue for simply removing a piece of data from a database" for generative-AI models, because "once a model has been trained, the impact of each data example in the training data is dispersed throughout the model and cannot be easily traced."
The gap here is twofold. First, "impact" and "removal" are not well-defined in the ML context—what does it mean to remove the influence of a training example from a model whose parameters encode statistical relationships across billions of examples? Second, even if ML researchers develop definitions and algorithms, those may not correspond to what the law actually requires. The legal standard for notice-and-takedown under Section 512 of the US Copyright Act, or for erasure under GDPR's Right to be Forgotten, was developed for systems where data can be identified and removed from databases. Applying these frameworks to models where training data influence is distributed and non-localized requires rethinking what compliance means.
The report also flags that "both machine unlearning and attribution are very young fields, and their strategies are (for the most part) not yet computationally feasible to implement in practice for deployed generative-AI systems." This means that even if the definitions could be aligned with legal requirements, the technical capacity to comply does not yet exist—a gap with significant practical consequences for companies facing takedown requests or regulatory obligations.
How This Paper Positions Itself Relative to Existing Work
The report positions itself primarily as a synthesis and agenda-setting contribution rather than as original doctrinal analysis or technical research. This is stated explicitly in the introduction: "Our intended audience is scholars and practitioners who are already interested in engaging with issues at the intersection of Generative AI and law." The report assumes baseline familiarity ("lawyers who have familiarity with terms like 'large language model'") and aims to "synthesize reflections from the workshop to highlight key issues that need to be addressed for successful research progress."
Within this genre, the report positions itself by:
Acknowledging existing resources while identifying their limits. The report references other works that have "made significant attempts to catalog" generative-AI concerns (citing Fergusson et al., 2023, for a broader harms taxonomy), but notes that its own taxonomy focuses "on highlighting the ways in which specifically legal issues may arise"—a narrower, legally-oriented lens that complements broader policy catalogs. It also cites existing legal scholarship on generative AI and copyright (Section 5.4) and defers to that work for doctrinal details, instead offering "a few high-level observations about current and impending IP issues."
Emphasizing cross-disciplinary infrastructure over policy proposals. The report explicitly declines to "delve into the technical details of specific generative-AI systems, the legal details of complaints and lawsuits involving those systems, or policy proposals for regulators." This is a deliberate scoping choice: rather than advocating for particular legal reforms or technical standards, the report focuses on building the capacity for productive analysis—the glossaries, metaphors, taxonomies, and research questions that enable others to do the detailed work.
Identifying the generative-AI supply chain as a key analytic framework. Drawing on prior work by some of the same authors (Lee et al., 2023, "Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain"), the report uses the supply chain as an organizing concept throughout. The supply chain perspective emphasizes that generative-AI systems involve multiple stages, actors, and design decisions—data collection, curation, pre-training, fine-tuning, alignment, deployment, and use—each of which may have different legal implications. This framework is positioned as a corrective to analyses that treat "the model" as a black box or "the company" as a monolithic entity, when in reality legal responsibility may attach differently to different actors at different stages.
Establishing ongoing, evolving engagement rather than a one-time analysis. The report is explicit that its glossary is "offered as a starting point, not a finish line" and that "the field is in flux; its terminology will evolve as new technologies and controversies emerge." The GenLaw organization is described as "growing into a nonprofit home for research, education, and interdisciplinary discussion" that will maintain and update resources. This positions the report as the first installment in an ongoing project rather than a static reference—consistent with the underlying argument that generative AI's legal challenges are comparable in scope and duration to those raised by the Internet, and will require sustained interdisciplinary engagement over years or decades.
Building on the Internet-law analogy without over-claiming it. The report uses Zittrain's generativity framework and the Internet-law analogy extensively (Section 2), but it treats this as a heuristic for understanding why generative AI will be legally significant, not as a claim that Internet-law doctrines will directly apply. The analogy serves to calibrate expectations (this will be big, complex, and long-lasting) and to motivate the need for shared vocabulary and mutual education between technologists and lawyers—parallel to what was required when Internet law emerged as a field.
In summary, the report positions itself as filling a meta-level gap: not analyzing any single legal issue in depth, but rather creating the conditions under which rigorous, cross-disciplinary analysis of many legal issues can proceed. This is a distinctive contribution because, as the workshop revealed, the absence of shared vocabulary, clear characterization of the technology, and frameworks for organizing legal questions was itself a primary obstacle to progress. The report does not solve the legal problems of generative AI—it argues that we need better tools and infrastructure before we can effectively start solving them.
3. Technical Approach
3.1 Reader Orientation
This paper is a workshop synthesis and agenda-setting document, not a traditional research paper with a single technical system. The "system" being built is a framework for interdisciplinary collaboration—a set of conceptual tools, shared vocabulary, taxonomies, and research questions that enable legal scholars and machine-learning researchers to productively analyze generative-AI issues together. The core problem it solves is that these two communities currently lack the common language and mental models needed for rigorous joint analysis, and the "shape" of the solution is a layered infrastructure: first defining terms precisely, then characterizing what makes generative AI technically distinctive, then mapping the legal issues that technical features raise, and finally identifying research problems where both forms of expertise are required.
3.2 Big-Picture Architecture (Diagram in Words)
The framework has four major components, arranged in an ascending order of complexity—each builds on the foundation laid by the previous ones:
-
Shared Knowledge Base (Section 3): A glossary of technical and legal terms of art, a set of instructive metaphors with their limitations explicitly marked, and an evolving map of business-model patterns. This component addresses the communication failure that prevents the two disciplines from getting started. Its primary outputs are definitions and analogies that travel across disciplinary boundaries.
-
Technical Characterization of Generative AI (Section 4): An analysis of what makes generative-AI systems distinct from prior AI—organized around flexibility (4.1), training pipelines (4.2–4.3), and scale (4.4). This component gives legal scholars the "theory of the machine" they need to map legal questions onto technical reality. Its primary output is a set of claims about which properties of generative AI are genuinely novel (and thus require novel legal analysis) versus extensions of existing technology.
-
Legal Taxonomy (Section 5): A structured mapping of legal issues—intent, privacy, misinformation, intellectual property—onto the technical features characterized in component 2. This component identifies where generative AI interacts with legal doctrines, not how those interactions should be resolved. Its primary output is an organized landscape of questions rather than answers.
-
Research Agenda (Section 6): A set of concrete, cross-disciplinary research problems (centralization vs. decentralization, rules and standards, machine unlearning, evaluation metrics) that require joint technical-legal investigation. This component identifies intervention points where design choices have legal consequences and legal frameworks constrain technical possibilities. Its primary output is a prioritized list of open problems.
Information flows linearly from component 1 through component 4: you cannot identify what is legally significant (component 3) without understanding the technology (component 2), and you cannot do that without a shared vocabulary (component 1). The research agenda (component 4) then selects specific problems from the legal taxonomy where technical and legal analysis must proceed in lockstep.
3.3 Roadmap for the Deep Dive
-
First, the shared knowledge base (Section 3 of the report), because before analyzing anything, we must establish what words mean. This covers the glossary (3.1), the role and limits of metaphors (3.2), and the business-model landscape that determines which actors are involved in generative-AI production and use (3.3). Understanding the terminological traps—like "pre-training" meaning training, not data preparation—is a prerequisite for everything that follows.
-
Second, the characterization of generative AI's unique features (Section 4), because legal analysis of this technology requires understanding what distinguishes it from prior AI. We walk through the transition from task-specific discriminative models to flexible generative ones (4.1), the multi-stage pre-training/fine-tuning pipeline (4.2–4.3), and the consequences of massive scale (4.4). This section answers the question that legal scholars repeatedly asked during the workshop: "What's so special about Generative AI?"
-
Third, the legal taxonomy (Section 5), which maps the distinctive technical features from Section 4 onto specific areas of legal doctrine. Each doctrinal area—intent, privacy, misinformation, intellectual property—receives a separate treatment that identifies how generative AI creates novel issues within that area. The taxonomy is explicitly non-exhaustive and scoped to IP and privacy (the workshop's stated focus).
-
Fourth, the research agenda (Section 6), which selects concrete problems—centralization/decentralization, rules versus standards for reasonableness, machine unlearning as inadequate notice-and-takedown, and evaluation metrics for legal concepts—where technical design choices and legal frameworks are tightly coupled. These are problems that cannot be solved by either discipline alone because the technical definition of success depends on legal standards, and legal standards must be informed by technical feasibility.
Throughout, I explain design choices: why define difficulty this way rather than borrowing the dataset's labels, why emphasize metaphors and their limits, why characterize the pre-training/fine-tuning distinction as an artifact of practice rather than an essential boundary, and why machine unlearning is flagged as a case study in misaligned technical and legal conceptions.
3.4 Detailed, Sentence-Based Technical Breakdown
This is a workshop synthesis paper whose core idea is that productive interdisciplinary work on generative AI and law requires building infrastructure before building analysis: shared vocabulary, clear technical characterization, mapped legal terrain, and identified intervention points where both forms of expertise are necessary. The paper does not propose a technical system, a legal doctrine, or a policy prescription—it proposes a framework for thinking that enables others to produce those things.
The Shared Knowledge Base: Terminology, Metaphors, and Business Models
Why a Glossary Is the First Intervention
The report identifies that the most immediate barrier to interdisciplinary collaboration is terminological confusion—and it distinguishes between two qualitatively different types of confusion that require different remediation strategies (Section 3). Understanding this distinction matters because it explains why the glossary (Appendix A) is structured the way it is and why it is offered as a starting point rather than an authoritative reference.
Known differences in terminology occur when both communities are aware they use a term differently. The canonical example is "privacy." In computer science, privacy often refers to formal definitions that are computationally tractable—most prominently differential privacy, which provides a mathematical guarantee about what can be inferred about individual records in a dataset. Specifically, differential privacy works by adding calibrated noise to data or query results such that the presence or absence of any single individual's data in the training set does not substantially change the probability distribution over outputs. The report notes that "the U.S. Census has sparked debate over its use of differential privacy: a technique that provides strong theoretical guarantees of privacy preservation. Critics question whether or not the definition of privacy reflected in differential privacy accords with the census's broader goals of privacy preservation."
In law, by contrast, privacy is typically defined contextually, based on social norms and reasonable expectations. As the report explains, "It is typically necessary to first identify which norms are at play in a given context, after which it is then possible to determine if those norms have been violated (and what to do about it). Such definitions of privacy are fundamentally nuanced; they resist quantification." A legal expert at the workshop provided a useful intuition for this tension: "Computer scientists often want to be able to quantify policy, including policy for handling privacy concerns; in the law, the mere desire to quantify complex concepts like privacy can itself be the source of significant problems."
The report argues that these known differences are manageable because "the two communities have generally understood that they mean something different by 'privacy' and have read each others' work with a working understanding of these differences in mind." The remediation strategy for known differences is translation: providing mappings between the technical and legal definitions so that each community can understand what the other means.
Unknown differences in terminology are more dangerous because they cause collaborations to derail without participants realizing why. The report gives two concrete examples:
-
"Pre-training": In machine learning, pre-training refers to "an early, general-purpose phase of the model training process" where a model learns broad patterns from large-scale data before being adapted (via fine-tuning) to specific tasks. Crucially, pre-training is training—it involves updating model parameters through an optimization algorithm, just like fine-tuning. But legal scholars at the workshop "assumed that the term referred to a data preparation stage prior to and independent of training." This confusion has direct legal consequences: the copyright analysis for data preparation (e.g., deduplication, formatting, tokenization) differs from the analysis for training (where copyrighted works are used as examples to update model parameters). If a lawyer mischaracterizes pre-training as data preparation, they may incorrectly conclude that it does not involve copying or creating derivative works in the copyright sense, when in fact the model parameters encode statistical information derived from the training examples.
-
"Harms": The report notes that "many technologists were not aware of the importance of harms as a specific and consequential concept in law, rather than a general, non-specific notion of unfavorable outcomes." In U.S. law, what counts as a cognizable harm determines whether a plaintiff has standing to sue in federal court, the scope of available remedies (damages, injunctions), and whether conduct falls within a regulatory regime at all. The glossary entry for harm (Appendix A.3) distinguishes: physical injury is the most widely accepted type of legal harm; other recognized forms include damage to property, economic losses, loss of liberty, restrictions on speech, and some kinds of privacy violations. But other cases have held that "fear of future injury is not a present harm," and "having one's personal information included in a data breach may not be a harm by itself—but out-of-pocket costs and hassle to cancel credit cards are recognized harms." A technologist who uses "harms" to mean "any bad outcome the model might cause" is operating with a concept that is simultaneously too broad (including outcomes the legal system does not recognize) and too shallow (missing the structured analysis that determines whether a harm is legally actionable).
The remediation strategy for unknown differences is explicit definition: creating a glossary that flags terms of art and provides baseline definitions suitable for non-experts in either discipline.
The Glossary's Structure and Goals
The glossary (Appendix A, entries marked in green throughout the report text and hyperlinked) has two primary goals stated explicitly (Section 3.1):
-
Identifying terms of art with technical or multiple meanings—both "words that also have general, colloquial meanings, such as 'attention' or 'harm'" and cases where "a technical term has taken on a broader meaning in society at large" (e.g., "algorithm" in public discourse refers to social-media ranking systems, but technically an algorithm is "a formal, step-by-step specification of a process" used for both training and inference).
-
Providing succinct definitions of critical concepts for a non-expert in either law or ML (or both). These definitions "are not intended to cover the full complexity of a concept or term from the expert perspective." The report gives the example that "one could write volumes on privacy—and many have, arguably for thousands of years"—instead, the glossary's purpose is "simply to show technologists that there is more to privacy than removing personally identifiable information (PII)."
The glossary is organized into four sections with alphabetized entries within each:
-
A.1: Concepts in Machine Learning and Generative AI—covering algorithm, alignment, API, architecture, attention, base model, checkpoint, context window, data curation/pre-processing, datasets, decoding, diffusion-based modeling, embedding, examples, generalization, generation, fine-tuning, foundation model, hallucination, hyperparameter, in-context learning, inference, language model, large language model, loss, memorization, model, multimodal, neural network, objective, parameters, pre-processing, pre-training/fine-tuning, prompt, regurgitation, reinforcement learning, reward, scale, supply chain, tokenization, transformer, training, vector representation, web crawl, and weights.
-
A.2: Open versus Closed Software—explaining that "open versus closed is not a binary distinction" and providing entries for closed dataset, closed model, closed software, open dataset, open model, and open software. The entry for open model is particularly notable because it identifies a tension within the ML community: "The machine-learning community has described a model as open-source when a trained checkpoint has been released with a license allowing anyone to download and use it, and the software package needed to load the checkpoint and perform inference with it have also been open-sourced"—yet pre-existing open communities like the Open Source Initiative have objected to this usage, arguing it fails to capture important qualities of openness as understood in the software community for decades (freedom to inspect, use for any purpose, modify, and distribute modifications).
-
A.3: Legal Concepts in Intellectual Property and Software—covering claims, copyright, copyright infringement, damages, fair use, the field of intellectual property, harm, idea vs. expression, license, non-expressive/non-consumptive use, patent, prior art, terms of service, and transformative use.
-
A.4: Privacy—covering anonymization, the California Consumer Privacy Act, consent, differential privacy, GDPR, personally identifiable information, privacy policy, privacy violation, the right to be forgotten, and tort.
A crucial design choice: the glossary is offered as a living document to be updated over time. The report states that "the field is in flux; its terminology will evolve as new technologies and controversies emerge. We will host and update this glossary on the GenLaw website." This is not merely aspirational—it reflects the paper's broader thesis that generative AI's legal challenges are comparable in scope to those raised by the Internet and will require sustained engagement over years, not a one-time reference document.
Metaphors as Tools for (Imperfect) Understanding
The report argues that well-chosen metaphors are useful for cross-disciplinary communication, but their value comes both from what they capture and from the ways they break down (Section 3.2). The paper treats metaphors as "sources of inspiration" and "rational frameworks for thinking through similarities and differences"—explicitly connecting this to the role of analogy in legal rhetoric, where "analogy and metaphor are central to legal rhetoric [and] provide a rational framework for thinking through the relevant similarities and differences between cases."
The report discusses two examples in detail and provides additional ones in Appendix B:
Metaphorical anthropomorphism: Machine-learning practitioners commonly use terms that anthropomorphize models—"learn," "respond," "memorize," "hallucinate." The report argues that these "should be considered terms of art—perhaps inspired by human actions, but grounded in technical definitions that bear little resemblance to human mechanisms." The critical analytic move is distinguishing inspiration from mechanism: while neural networks contain "neurons" that "fire" analogously to biological neurons, the claim is not that models learn or memorize in exactly the same way humans do.
The report identifies a specific risk from anthropomorphic metaphors: they "can lead people to conclude that a machine-learning system is completing such actions using the same mechanisms and thought processes that a human would." This risk has direct legal consequences. For example, if a judge or jury attributes intent to a generative-AI system because it "learned" or "memorized" certain training data, they may apply legal standards (like the mens rea requirement in criminal law or the volitional act requirement in copyright) that presuppose human-like cognitive processes the system does not possess.
The report acknowledges that some in the GenLaw community have advocated for different terminology—ones that do not "elicit such strong comparisons to human behavior" (citing Cooper et al., 2022, and related work). However, it takes the pragmatic position that "until we have better terms, understanding when a term is indeed a term of art, and the ways that it is inspired by (but not equivalent to) colloquial understandings, will remain a critical part of any interdisciplinary endeavour." This is a design choice: the report does not attempt to replace anthropomorphic terminology but to mark it as potentially misleading so that interdisciplinary collaborators can adjust their interpretations accordingly.
Memorization: The report identifies memorization and regurgitation as terms that cause particular confusion because "the connection to the colloquial meanings of these words can cause confusion." The technical definitions of memorization in machine learning are precise and quantifiable—"specific ways to measure the amount of memorization (as it is technically defined) present in a model or its outputs" (citing Carlini et al., 2023; Ippolito et al., 2023; Kudugunta et al., 2023; Anil et al., 2023 for different operationalizations).
But the colloquial understanding of "memorization" imports assumptions that do not hold for models. The report identifies three specific disconnects:
-
Some uses go beyond the technical definition: Discussion around text-to-image models "memorizing" an artist's style is "not equivalent to memorization in the technical sense of the word." Style similarity is an active research area (citing Casper et al., 2023), but it is a different phenomenon from the verbatim or near-verbatim reproduction that technical memorization metrics measure. This matters legally because claims about "style memorization" may be presented as technical facts when they are actually conceptual claims about artistic influence that existing technical definitions do not capture.
-
Intentionality is absent: "People deliberately memorize; for example, an actor will actively commit a script to memory. In contrast, models do not deliberately memorize; the training examples that end up memorized by a model were not treated any differently during training than the ones that were not memorized." The report explains that models are trained using an objective function that rewards producing data similar to the training data, but the goal is generalization—regularization and alignment techniques are used to prevent exact reproduction. The fact that some examples end up memorized is a property of the training dynamics, not a deliberate action by the model or the model creator.
-
No distinction between memorizing and remembering: "Humans distinguish between memorizing (which has intentionality) and 'remembering' (when a detail is recalled without the intent to memorize it). Generative-AI models have no such distinction. For humans, it is how we personally feel about a thought, action, or vocalization that leads us to call it 'memorized.' But for models, which lack intent or feeling, memorization is merely a property assigned to their outputs and weights through technical definitions."
The operational implication: when legal discussions invoke "memorization," participants must clarify whether they mean (a) technical memorization as defined by specific metrics in the ML literature (measurable, quantifiable, usually referring to verbatim or near-verbatim reproduction of training examples), (b) a broader colloquial sense that includes style similarity or factual recall, or (c) a legal concept like copying or unauthorized reproduction that may or may not align with either technical or colloquial definitions. The report's framework does not resolve which definition should apply in which legal context, but it identifies that failing to distinguish among these definitions is a primary source of confusion in current debates.
Additional metaphors in Appendix B: The report provides brief treatments of "models are trained" (comparing to dog training but noting the absence of a curriculum), "models learn like children do" (noting that children and models use "very different mechanisms" and that techniques that help models learn better, like increasing model size, "have no parallels in child development"), "generations are collages" (drawing on Lee et al., 2023, to argue that "a generative-AI system does not take several works and splice them together" and that "there is no author selecting, coordinating, or arranging training examples to produce the resulting generation"—quoting the statutory definition of "compilation" from 17 U.S.C. § 101), "large language models are stochastic parrots" (citing Bender et al., 2021, and acknowledging that critics "say that these competencies imply models understand meaning in a human-like way" while proponents "might argue that Generative AI passing a difficult standardized exam... is more about parroting training data than human-like skill"), and "large language models are noisy search engines" (noting that "the training data is seen during training, but models are used separately from the training data"—a key architectural distinction from systems that retrieve from a database).
Business Models as a Practical Necessity
Section 3.3 argues that "real, current information about business models can be very useful for understanding who is involved in the production, maintenance, and use of different parts of generative-AI systems." The motivation is that legal analysis requires knowing which actors are responsible for which decisions in the supply chain—a model is not a monolithic entity, and understanding the relationships between companies, users, and intermediaries is essential for determining liability, jurisdiction, and applicable legal standards.
The report identifies four business-model patterns (not exhaustive, but representative of the diversity observed at the time of the workshop in mid-2023):
-
Business-to-consumer (B2C) hosted services: Companies release direct-to-consumer applications and APIs for producing generations—cited examples include OpenAI's ChatGPT, Anthropic's Claude, Google's Bard, Midjourney's and Ideogram's text-to-image applications. The report characterizes these as providing "a mix of entry points to their systems and models, including user interfaces and APIs, often offered via subscription-based services," with the crucial note that "typically, the systems and models developed by these companies are hosted in proprietary services; users can access these services to produce generations but, with some notable exceptions (e.g., fine-tuning APIs), cannot directly alter or interact with the models embedded within them." This has legal significance because the company maintains control over the model and system behavior, which affects arguments about who is the "speaker" or "publisher" of generated content and who bears responsibility for harmful outputs.
-
Business-to-business (B2B) integration with hosted services: Models are integrated into other companies' products either through direct partnerships (e.g., ChatGPT integration into Microsoft Bing search) or through API usage (e.g., Poe developed via partnership between Anthropic and Quora). The report notes that the nature of these business relationships can affect legal analysis, but also that "often, we will not be able to tell the nature of the business relationship (unless disclosed publicly) between corporate partners"—creating opacity that complicates legal attribution.
-
Products derived from open models and datasets: Some companies operate with open-source product offerings (e.g., some versions of Stable Diffusion offered by Stability AI) or mixed offerings (e.g., Meta releases Llama model weights openly but keeps training data details closed). The report emphasizes that "open" is not a binary state—a model may have open weights but closed training data, or an open license for the model but restricted APIs for accessing it. The glossary entry for open model (Appendix A.2) elaborates on these gradations.
-
Companies that operate at specific points in the supply chain: The report identifies emerging business niches at different stages of the pipeline, offering three concrete examples: datasets (Scale AI works on data example annotation for training datasets), training diagnostics (Weights & Biases handles data analysis and diagnostics for model training dynamics), and training and deployment (MosaicML, acquired by DataBricks, and Together AI develop solutions for bespoke model training and serving). The report notes that while "it is generally tremendously costly to train and deploy large generative-AI models, advancements in open-source technology and at smaller companies... have helped make training custom models more efficient and affordable"—implying that the landscape of actors is diversifying rather than consolidating, which creates more complex legal relationships.
The key analytic point is that "there are many ways that generative-AI models may be integrated into software systems, and... many different types of business models associated with the training and use of these models." This matters for legal analysis because questions of liability, jurisdiction, and applicable law turn on who did what at which stage—a vertically integrated company that collects data, trains a model, and deploys a consumer-facing product faces a different legal profile than a company that only provides a fine-tuning API to developers who then build their own applications, which in turn differs from an open-source model released without any deployment infrastructure.
Characterizing Generative AI's Distinctive Features
The Question Legal Scholars Asked: "What's So Special About Generative AI?"
The report reports that during the roundtable discussions, "the legal scholars and practitioners had a recurring question for the machine-learning experts in the room: What's so special about Generative AI? Clearly, the outputs created by Generative AI today are better than anything we have seen before, but what is the 'magic' that makes this the case?" (Section 4). This question is not idle curiosity—it has direct legal significance. If generative AI is essentially an improved version of existing technology, then most legal issues can be analyzed under existing frameworks with minor adaptations. If it is genuinely novel in legally relevant ways, then existing frameworks may need substantial extension or revision.
The report's answer identifies three aspects "for which it can be productive to consider recent developments in AI as meaningfully novel or different in comparison to past technology." These are not exhaustive of all differences, but represent the dimensions that the workshop participants converged on as most legally salient.
The Transition from Discriminative to Flexible Generative Models
The first distinctive feature is a dual shift in what models are trained to do (Section 4.1). Prior to the current era, machine-learning models were predominantly trained for discriminative tasks: given an input, produce a simple classification or regression output. The canonical examples are "labeling an image according to the class of object it depicts [32, 33] or classifying the sentiment of a sentence as positive or negative [57]." These models output a class label or a scalar value—their output space is small and structured.
Modern generative-AI models change this in two ways:
Shift one: from discriminative to generative outputs. Instead of producing a class label, generative models "output complex content, such as entire images or paragraphs of text." The report gives the example: "given the input of cat, outputting a novel image of a cat, sampling from the near-infinite space of reasonable cat images it could create" (citing Lee et al., 2023, Part I.A for a longer treatment). This matters legally because the output of a generative model is itself a creative work (or something resembling one), which raises copyright, defamation, and other content-regulating legal issues that a classification label ("cat" or "dog") does not. A sentiment classifier that outputs "negative" does not defame anyone; an LLM that generates a paragraph falsely accusing someone of a crime might.
Shift two: from task-specific to general-purpose models. Instead of training separate models for sentiment analysis, summarization, part-of-speech tagging, etc., "many state-of-the-art systems today handle a wide variety of tasks using a single model." The report acknowledges that this is not an absolute rule ("it is rumored that the models underlying ChatGPT are actually an ensemble of on the order of 10 expert models"), but the trend is toward generality. This matters legally because a model that can be prompted to do thousands of different tasks is harder to characterize for legal purposes than a purpose-built classifier—what is the "intended use" of a general-purpose LLM, and how does that affect, for example, the fair-use analysis for the training data or the product-liability analysis for harmful outputs?
The report emphasizes that the scaling up of generative AI (discussed in Section 4.4) "has facilitated generativity across a wide range of applications and modalities" beyond the commonly discussed text-to-text and text-to-image applications. The non-exhaustive list includes "image captioning, music generation, speech generation and transcription, tools for lowering the barrier to learning to program, and research questions in the physical sciences (including on protein folding, drug design, and materials science)"—citing Agostinelli et al., 2023 (MusicLM), Le et al., 2023 (Voicebox), Radford et al., 2022 (Whisper), Yilmaz & Yilmaz, 2023 (programming education), and Corso et al., 2023 (molecular docking). The breadth of applications means that legal issues will not be confined to a single domain but will propagate across many areas of law simultaneously.
The Multi-Stage Training Pipeline
The second distinctive feature is the now-standard division of training into pre-training and fine-tuning stages (Sections 4.2–4.3), and more broadly the existence of a complex supply chain with multiple actors making design decisions at different stages.
Pre-training: The report characterizes pre-training as typically involving an "enormous, often web-scraped dataset, which instills a 'base' of knowledge about the world within the model." During pre-training, models "learn underlying patterns from their input data" (quoting Callison-Burch's 2023 Congressional testimony). For LLMs, this means learning "syntax and semantics, facts (and fictions) about the world, and opinions, which can be used to produce summaries and perform limited reasoning tasks." For image-generation models, it means learning "to produce different shapes and objects, which can be composed together in coherent scenes." The report appends a footnote: "Though, notably, not a collage! See Appendix B for a discussion of why the metaphor of a collage for generative-AI outputs can be misleading."
The flexibility conferred by pre-training is what makes base models reusable: "one can further train (i.e., fine-tune) the base model on domain-specific data (e.g., legal texts and case documents) to specialize the model's behavior" or "fine-tune the base model to understand a dialog-like format (as ChatGPT has done [74])."
Fine-tuning: The report emphasizes that fine-tuning uses smaller datasets than pre-training and is correspondingly faster and less expensive—"pre-training is very expensive, which means it only happens once (or a small handful of times), but subsequent fine-tuning tends to be much faster... so it can more tractably occur many times." This cost asymmetry has structural consequences for the supply chain: it is economically feasible for many actors to fine-tune a single pre-trained base model, but not to pre-train their own. This creates a dependency relationship where fine-tuners are downstream of (and legally affected by decisions made by) pre-trainers.
A crucial clarification about the pre-training/fine-tuning distinction: The report is explicit that "this division is not well-defined. It is predominantly an artifact of choices made regarding training, rather than an essential aspect of the training process. Both pre-training and fine-tuning are just training (though perhaps configured differently)." The reason for distinguishing them is purely practical: researchers choose to divide training into stages, and the supply chain often distributes responsibility across actors (one actor pre-trains, another fine-tunes). The report flags this because "during the GenLaw roundtable discussions, it became apparent that legal experts shared some misconceptions about the roles of pre-training and fine-tuning"—specifically, the misconception that pre-training is "a data preparation stage prior to and independent of training." The report reiterates that "pre-training is training, and not a data preparation stage."
The supply chain: The report defers to Lee et al., 2023, Part I.C for a detailed treatment of the supply chain but summarizes that "there are numerous decisions and intervention points throughout the system, which extend to elements beyond choices in pre-training and fine-tuning." The key conceptual move is to treat the model not as a single artifact but as the product of a chain of design decisions made by potentially different actors.
The report gives a concrete example to illustrate supply-chain complexity: consider the choice of training data. Decisions include (1) which data examples to include, (2) where data will be stored, (3) how long data will be retained, and (4) where the resulting model will be deployed. These choices are interdependent: "a model training on a private user data has a very different privacy-risk profile if such a model were never to leave the user's personal device, compared to if it is to be shared across many users' devices." The point is that legal analysis of, say, privacy implications cannot treat "the model" as a uniform category—it must account for specific design decisions made at specific points in the supply chain.
The report further emphasizes that "not all design choices are about models and how they are trained." Models are embedded within systems that include components like "prompt input filters, generation output filters, rate limiting... access controls, terms of use, use-case policies for APIs, user interface and experience design... and so on." Each of these involves design decisions "that can have their own legal implications"—for example, a content filter that removes certain types of generated outputs might affect a platform's Section 230 immunity analysis or its obligations under content-moderation laws.
The Role of Massive Scale
The third distinctive feature is scale (Section 4.4)—not just quantitative improvement but qualitative changes in capabilities and in the economics of model development. The report states that "state-of-the-art models today are an order of magnitude larger and trained on significantly more data than the biggest models from five years ago."
The report identifies several consequences of scale that are legally relevant:
Emergent capabilities: The report notes that "many experts have studied methods for collecting and curating massive, web-scraped datasets, as well as the 'emergent behaviors' of models trained at such large scales [109, e.g.]." The term "emergent behaviors" refers to capabilities that appear only above certain scale thresholds—a model with 1 billion parameters may not be able to perform a task that a model with 100 billion parameters handles competently. This matters legally because it means that restricting model size (e.g., via regulation) could foreclose entire categories of capability, and that the same legal framework applied to small and large models may produce different practical outcomes.
Concentration of model creation: "One of the implications of scale is that machine-learning practitioners are training fewer state-of-the-art models today than were being trained in the past." When models were small, it was common to re-train many times with different hyperparameters to find the best configuration. Today, "the cost of training just one state-of-the-art model can be hundreds of thousands or even millions of dollars" (citing Bekman, 2022; BigScience, 2022; Lee et al., 2023). This "further incentivizes the push toward general-purpose models" (Section 4.1) because the high fixed cost of pre-training is only recouped if the base model can be adapted to many uses.
The techniques are not new, but the scale is: "Language models, for example, have existed since at least the 1980s [87]. The difference is that, in recent years, we have figured out how to scale these techniques tremendously (e.g., modern language models use context windows of thousands of input tokens, compared to the 5 to 10 input tokens used by language models in the early 2000s)." This is legally significant because it means the novelty is quantitative rather than architectural—arguments that generative AI is "completely different" from prior technology may overstate the case, while arguments that it is "just more of the same" may understate how quantitative change produces qualitative legal differences.
Efficient architectures and systems for scale: The report notes that scaling has been enabled not just by bigger computers but by "research into more efficient neural architectures and better machine-learning systems for handling model training and inference at scale" (citing Ratner et al., 2019 on the MLSys field). This means that scaling was a sociotechnical achievement—the result of deliberate research and engineering decisions—not an inevitable consequence of Moore's Law. This has potential legal significance for arguments about foreseeability, reasonable care, and the allocation of responsibility for model behaviors that emerge at scale.
The Legal Taxonomy
Section 5 of the report presents what it calls "progress toward a taxonomy of the legal issues that Generative AI raises." The framing is deliberately provisional: "the initial analysis presented here is very much an interim contribution as part of an ongoing project." The taxonomy is scoped to privacy and intellectual property (the workshop's stated focus) and focuses on "highlighting the ways in which specifically legal issues may arise" rather than cataloging all possible concerns.
A foundational observation precedes the taxonomy: Generative AI inherits all the issues of AI/ML technology more generally, because it can be used to perform the same tasks that purpose-built models perform. For example, rather than using a dedicated sentiment-analysis model, "one might simply prompt an LLM with labeled examples of text and ask it to classify text of interest." The resulting classifications "may or may not be as reliable as ones from a purpose-built model, but insofar as one is using a machine-learning model in both cases, any legal issues raised by the purpose-built model are also present with the LLM."
Similarly, "any crime or tort that involves communication could potentially be conducted using a generative-AI system"—fraud, blackmail, defamation, spam, deepfakes, false advertisements. "Almost any speech-related legal issue is likely to arise in some fashion in connection with Generative AI." This observation is not a specific legal claim but a scoping principle: existing legal frameworks for AI, for communication, and for content regulation provide the starting point, and the novel contribution of generative AI is what additional issues it raises beyond those that were already present.
The taxonomy then walks through four areas: intent (5.1), privacy (5.2), misinformation/disinformation (5.3), and intellectual property (5.4). I treat each in order.
Intent in Torts and Criminal Law
The report frames intent as an area where generative AI will "force us to rethink" how the law operates, because "Generative AI can cause harms that are similar to those brought about by human actors but without human intention." The legal system uses intention pervasively—in criminal law, mens rea is often an element of a crime; in tort law, some torts require intent (fraud, defamation) while others are strict liability (products liability). Where intent is required, "a defendant's lack of wrongful intent means they cannot be convicted or held liable."
The operational challenge is that "in contrast to most prior types of ML systems, Generative AI can cause harms that are similar to those brought about by human actors but without human intention." The example given: "an LLM might emit false and derogatory claims about a third party—claims that would constitute defamation if they had been made by a human." The traditional defamation analysis asks about the speaker's state of mind (did they know the statement was false? Were they reckless about its truth?). That analysis does not map cleanly onto a system that has no state of mind.
The report describes one discussion from the workshop about whether a respondeat superior model could address this: treating the generative-AI system as the "employee" and ascribing responsibility to the "employer" (the user or system operator). Respondeat superior is strict liability as to the employer—if the employee commits a tort within the scope of employment, the employer is liable regardless of the employer's own intent. This seems to "sidestep the need to deal with intent."
But the report notes a counterpoint raised during the discussion: "in the usual application of respondeat superior, there is still an embedded notion of intention. That is because the employee's intentions are still relevant in determining whether a tort has been committed at all; only if there has can liability then also be placed on the employer." In other words, respondeat superior does not eliminate the intent inquiry—it assumes an intentional tort by the employee and then adds a second layer of (strict) liability for the employer. If the "employee" is an AI system that lacks intent, it is unclear whether the first step of the analysis can be satisfied.
Another line of discussion considered whether generative AI might lead to "a greater focus on the human recipients of generative-AI outputs." The report cites Collins & Skover, 2018, who argue that AI systems create a world of "intentionless free speech" where communications should be assessed "purely based on their utility to the listener." This framework would establish a First Amendment right to use generative AI but would also raise "difficult questions about how to protect users from Generative AI in cases of false or harmful outputs."
The report does not resolve these questions—it flags them as areas where "there is unlikely to be a simple across-the-board answer as to how the 'intent' of a generative-AI system should be measured, in part because the legal system uses intention in so many ways and so many places."
Privacy
The privacy analysis (5.2) builds directly on the cross-disciplinary translation work from Section 3. The report re-emphasizes that "'privacy' is a notoriously difficult term to define" and that "different disciplines rely on different definitions that simplify the concept in different ways, which can make it very difficult to communicate about privacy across fields." The characterization of the computer-science approach (mathematical formalisms, often differential privacy) versus the legal approach (contextual, based on social norms and reasonable expectations, resisting quantification) is restated with an added observation from a GenLaw legal expert: "Computer scientists often want to be able to quantify policy, including policy for handling privacy concerns; in the law, the mere desire to quantify complex concepts like privacy can itself be the source of significant problems."
The report frames generative AI as making existing privacy challenges harder along several dimensions:
Training data scale and composition: Generative-AI models are "typically trained on large-scale web-scraped datasets" that "can contain all sorts of private information (e.g., PII)" (citing Brown et al., 2022). This information can be memorized by the model and then "leaked in generations" (citing Carlini et al., 2023; Somepalli et al., 2022).
Novel linking of information: Whereas "a traditional search engine only locates individual data points," a generative-AI model could "link together information in novel ways that reveal sensitive information about individuals." This is a qualitative difference: search engines can surface information that exists on the web, but generative models can synthesize across sources to produce new combinations that were not previously accessible in a single location.
Adversarial extraction: "Adversarially designed prompts can extract other sensitive information, such as internal instructions used within chatbots" (citing Edwards, 2023 on prompt injection attacks against Bing Chat). The report notes parenthetically that this "is potentially also a trade-secret issue (Section 5.4), depending on the nature of the information leaked"—illustrating how privacy, security, and IP concerns are not cleanly separable in practice.
The report does not attempt to resolve the definitional tension between computer-science and legal notions of privacy. Instead, it characterizes the landscape: "we understand how difficult these privacy challenges are only because of decades of research in law and computer science," and "Generative AI is poised to make these privacy challenges even harder." The implication for the research agenda (Section 6) is that privacy-preserving techniques developed for prior technologies (differential privacy, anonymization) may not be sufficient for generative AI, but the report does not specify what new techniques would be needed.
Misinformation and Disinformation
The misinformation analysis (5.3) distinguishes between misinformation (false or misleading material regardless of intent) and disinformation (deliberately false or misleading material, often with manipulative purpose). The report observes that generative AI "can be used to produce plausible-seeming but false content at scale," which has implications for any area of law that prohibits false speech.
The report identifies two distinct mechanisms by which generative-AI systems can produce false content, corresponding roughly to the misinformation/disinformation distinction:
From misinformation in training data: Generative-AI models are "very sensitive to their training data, which may itself include misinformation or disinformation." The specific example: "in August of this year, it was discovered that a book about mushroom foraging, which was produced with the assistance of an LLM, contained misinformation about which mushrooms are poisonous (likely due to inaccurate information learned from the data on which the LLM was trained)" (citing Cole, 2023, 404 Media). This is a passive mechanism—the model reproduces false information it encountered during training, without any adversarial intent by the model creator or user.
The report also cites two related phenomena observed in LLMs: sycophancy, where "a model answers subjective questions in a way that flatters their user's stated beliefs," and sandbagging, where "models are more likely to endorse common misconceptions when their user appears to be less educated" (citing Bowman, 2023; Perez et al., 2022). These are not exactly the same as reproducing training-data errors—they are emergent behaviors of models trained to be helpful and responsive to user cues—but they produce similar results (the model outputs false or misleading information).
From deliberate manipulation: Models "can be deliberately manipulated through adversarially selected fine-tuning data or through alignment" to "skew models to deliberately produce misleading content." This is an active mechanism where an actor with adversarial intent shapes the model's behavior. The report frames this as a disinformation capability distinct from (and additive to) the misinformation from training data.
A specific discussion point from the workshop illustrates how generative AI blurs the boundary between misinformation and disinformation, and between those and privacy: "generated disinformation about individuals (e.g., deepfakes) will potentially contribute to new types of defamation-related harms." But "other experts, particularly those with a privacy background in computing, questioned whether such harms could also be classified as intimate privacy violations (citing Citron & Solove, 2022)." The report notes that "in many respects, the harms caused by sufficiently convincing forgeries are very similar to those caused by truthful revelations" (citing Zipursky & Goldberg, 2023). This blurs the line between defamation (false statements that harm reputation) and privacy violations (true statements that disclose private information)—a legally significant distinction that generative AI complicates.
The lawyers at the workshop responded that "this would likely not constitute a cognizable privacy harm under the law, although it would still be actionable as defamation or false light." This exchange illustrates the report's broader point about the need for cross-disciplinary conversation: the privacy experts in computing thought of deepfake harms as privacy issues, while the legal experts pointed out that existing privacy doctrine would not classify them that way—but both agreed that existing legal categories might not cleanly map onto the harms generative AI enables.
Intellectual Property
The IP analysis (5.4) is deliberately high-level, noting that "given the recent spate of lawsuits about copyright and Generative AI (as well as the stated thematic focus for the first GenLaw Workshop), IP was a frequent topic" and that it "has also been one of the first generative-AI subjects explored in detail by scholars" (citing Callison-Burch, 2023; Lee et al., 2023; Sag, 2023; Samuelson, 2023; Vyas et al., 2023 among others). The report "leave[s] discussion of the doctrinal details to their work" and instead offers "a few high-level observations about current and impending IP issues—many of which also have applications beyond IP."
The observations are organized into seven categories:
Volition: The report identifies that "human volition plays an important and subtle role in defining IP infringement." Copyright infringement normally requires that "a human intentionally made a copy of a protected work, but not that the human was consciously aware that they were infringing." The concern raised at the workshop was that it may be "easy to deflect the role of human-made design choices (Section 4.2) by making such choices seem 'internal' to the system (when, in fact, such choices are typically not foregone conclusions or strict technical requirements" (citing Cooper et al., 2022 on accountability). Since "purely internal" copies tend to be fair use, characterizing a design choice as an automatic or inevitable result of the training process—rather than a choice made by developers—could serve as "a copyright liability shield." The report flags this as something "legal experts will need to contend with in their analysis of generative-AI systems."
Market externalities: Generative AI raises concerns about "mass labor displacement, significant market changes, and the concentration of market power"—issues that "extend beyond IP, but necessarily invoke related questions of ownership." The report notes that these are "matters for labor law, international trade law, and other areas of law. But they are also IP issues, because doctrines such as fair use invite courts to consider such societal effects in weighing the propriety of particular copying." This is a specific structural observation: fair use (17 U.S.C. § 107) considers "the effect of the use upon the potential market for or value of the copyrighted work" as one of four statutory factors, meaning that market-harm arguments are built into the copyright analysis itself, not just external policy concerns.
Trade secrecy: Fine-tuning on proprietary data is "poised to become a potentially useful pattern in the adoption of generative-AI technology." But "existing generative-AI models are known to memorize their training data [17, 18]," which "raises the possibility that an adversarial user could extract proprietary information in training data, thereby presenting issues related to trade secrecy" (citing Edwards, 2023 on prompt injection). This creates a tension: using generative AI on proprietary data may be valuable, but doing so may risk losing trade-secret protection if the model can be prompted to reveal the data—and once trade-secret protection is lost (through public disclosure), it cannot be regained.
Scraping: "The legality of scraping training data is inextricable from the IP treatment of Generative AI." The report notes that "Generative-AI companies both rely on scraped data as an input and take measures (both technical and legal) to prevent outputs from their systems from being used as inputs to other systems without permission." This creates an asymmetry: companies that benefit from scraping others' content may simultaneously restrict others from scraping their outputs—a tension that may be legally relevant.
Authorship: The report notes that "IP law may need to reconsider authorship eligibility in light of Generative AI." The issue is not new ("computer authorship is not a new topic of analysis in the law," citing Grimmelmann, 2016; Samuelson, 1985), but "Generative AI is likely to present new variations on old themes." Specifically, "purely computer-generated works are not currently covered by copyright. However, some argue that this situation is not sustainable (e.g., Lee, 2023, Ars Technica). Where AI-generated works have significant value, there will be strong economic pressures on courts to give users copyright in those works." The report does not take a position on what the law should be, but identifies the economic pressure toward recognizing some form of protection for AI-assisted or AI-generated works.
Patent: Generative-AI techniques have applications in the physical sciences—drug design, protein folding—which "raises implications for patent law." U.S. patent law requires a human inventor as a condition of patent eligibility, and "just as copyright's human authorship requirement has been challenged (but so far upheld [103, e.g.]), similar challenges may arise with respect to patents."
Idea-expression dichotomy: This is the report's most theoretically interesting IP observation. Copyright law distinguishes between ideas (not copyrightable) and expressions of ideas (copyrightable). The report argues that "Generative AI seems to further blur the already often-murky line between idea and expression." The specific twist is an inversion of the typical relationship: "one could attempt to analogize the prompt to an idea and the associated generation to its expression, but this presents several problems." First, "it suggests that the AI, rather than the human, is responsible for the creative expression (which is not currently protectable by copyright law)." Second, "there may be sufficient creativity for copyrightability of the prompt itself, even if it is ultimately (by the prior analogy) responsible for the idea in the resulting generation." Third, "there is a tenable argument that the human prompter and generative-AI system are acting in concert to produce the resulting generation, and that the way that an idea is expressed in a prompt makes it inextricably indivisible from the resulting expressive generation."
The report concludes, citing Lemley, 2023: Generative AI "seems to turn the idea-expression dichotomy 'upside down.'" In traditional copyright, a human has an idea and expresses it in a fixed form (the expression is copyrightable). With generative AI, a human expresses an idea in a prompt, and the AI generates an output—but is the prompt the idea (with the output as expression, making the AI the "author") or is the prompt itself the expression (with the output as a derivative or commissioned work)? The traditional framework does not cleanly resolve which human contributions count as "authorship" and which AI contributions count as "mere tools."
The Research Agenda: Centralization, Standards, Unlearning, and Evaluation
Section 6 identifies four research areas that require joint technical-legal investigation. The report frames these as "just a sample of emerging research areas" that "gives a flavor of how these two disciplines can concretely inform each other."
Centralization and Decentralization
The report frames centralization versus decentralization as a question that is "partly a technical question" (can we train powerful models with less centralized infrastructure?), "partly a business question" (how are investment and business models evolving?), and "most significantly... a fundamentally legal question" (how do licensing, competition, and antitrust law shape the landscape?).
The technical dimension: "current methods require centralized pre-training at scale based on datasets typically gathered from highly decentralized creators." The open question is whether either of these constraints will change—"improvements in training algorithms may reduce the investment required to pre-train a powerful base model, opening it up to greater decentralization," while "improvements in synthetic data may enable well-resourced actors to generate their own training data, partially centralizing the data-collection step."
The report gives a concrete example of the legal-technical entanglement: the "controversies over the use of closed-licensed data (within web-scraped datasets) as training data for generative-AI models." Datasets like LAION are often released with open licenses (e.g., MIT license), but "this does not guarantee that the associated and constituent data examples in those datasets can be licensed for use in this way [62]. Many examples within datasets have closed licenses." The footnote elaborates: "this is particularly complex for datasets used to train multimodal models, like text-to-image models; the examples to train text-to-image models are image-caption pairs, where for each pair the image and the text caption could be subject to their own copyrights (and even hypothetically could be subject to a copyright as a compilation)."
The legally significant observation is that dataset-level licenses (the terms under which a curated dataset like LAION-5B is distributed) are distinct from and do not override the licenses on individual data examples within the dataset. A dataset could be distributed under the MIT license (which permits use and copying) while containing examples that are individually under restrictive licenses or copyrighted without any license at all. The report frames this as requiring "significant technical innovation, both in techniques for collecting such datasets at scale while respecting licensing conditions and also in training models that make best use of the limited materials available in them"—and notes that "current attempts to train models on such openly licensed datasets have yielded mixed results in terms of generation quality" (citing Gokaslan et al., 2023).
The legal dimension: "competition and antitrust law are likely to play a major role going forward. Every important potential bottleneck in Generative AI—from copyright ownership to datasets to compute to models and beyond—will be the focus of close scrutiny." The report lists possible policy interventions ranging from antitrust enforcement to "government subsidies, open-access requirements, 'public option' generative-AI infrastructure, export restrictions, and structural separation"—and argues that "these questions cannot be discussed intelligently without contributions from both technical and legal scholars."
Rules, Standards, Reasonableness, and Best Practices
The report observes that "since the technological capabilities of today's generative-AI systems are so new, it is unclear what duties the creators and users of these systems should be." The legal system has a repertoire of approaches for specifying duties: bright-line rules (like HIPAA's specification of which data is protected), flexible standards (like reasonableness), and industry best practices that develop even in the absence of legal mandates.
The report argues that the legal system will need to articulate duties for generative-AI creators and users across all of these modalities, and that the expectations "have always been technology-specific and necessarily change over time as technology evolves." Two specific examples illustrate the dynamism:
- Under the Uniform Trade Secrets Act, information must be "the subject of efforts that are reasonable under the circumstances to maintain its secrecy." What counts as reasonable security has changed over time as cybersecurity practices have evolved.
- The FTC monitors the state of the art in cybersecurity and has been willing to argue that companies engage in "unfair and deceptive trade practices by failing to implement widely used cost-effective measures"—meaning that what is reasonable is defined partly by what is technically feasible and what peers are doing.
The report identifies an important asymmetry: "the definition of reasonableness is contextual; what is considered reasonable for a large company (e.g., in terms of system development practices) is typically different than what is considered reasonable for smaller actors." This matters for generative AI because the supply chain includes actors of very different sizes and resources—from large companies pre-training foundation models to individual developers fine-tuning open models—and a uniform standard of reasonableness would have drastically different effects on different actors.
The report frames the research need as bidirectional: "legal scholars urgently need to study—and technical scholars urgently need to explain—which generative-AI safety and security measures are recognized as efficient and effective." This is not a one-time exchange but an ongoing process, because "what is currently an effective countermeasure against extracting memorized examples from models may fail completely in the face of newly developed techniques." The legal system must be "attuned to the dynamism of generative-AI development"—but so must the technical community, because "new techniques of training and alignment may be developed that are so clearly effective that it is appropriate to expect future generative-AI creators to employ them."
A specific pathway identified: "once technology begins to stabilize, it becomes easier to define concrete standards (e.g., safety standards). Accordingly, by definition, compliance with such standards is sufficient for meeting the bar of reasonableness. Until there is some stability, when harms occur, there will necessarily be some flexibility; there will be some deference to system builder's self-assessments of whether their design choices reflected reasonable best efforts to construct safe systems. In turn, today's best efforts will guide future standards-setting and determination of best practices."
The report's call to action is that "both the legal and machine-learning research communities should face this reality head-on; they should take hold of the opportunity to actively engage in research and public policy regarding today's generative-AI systems, such that they can help shape the development of future standards." The pathway is: current practices → informal best practices → formalized standards → legal requirements, and the transition between each stage is shaped by technical feasibility, demonstrated effectiveness, and stakeholder advocacy. If the ML community does not participate in this process, standards may be set by actors without deep technical understanding, producing requirements that are either ineffective (too weak) or infeasible (too strong).
Machine Unlearning as Inadequate Notice-and-Takedown
Section 6.3 makes a specific, critical intervention: machine unlearning is not a drop-in replacement for legal notice-and-takedown mechanisms, and treating it as such misunderstands both the technical limitations of unlearning and the legal requirements of notice-and-takedown.
What notice-and-takedown requires: The report references Section 512 of the US Copyright Act, which provides a safe harbor for online service providers who remove allegedly infringing material upon receiving proper notice. The operational model is: (1) a rightsholder identifies specific infringing content hosted on a platform; (2) the rightsholder sends a notice; (3) the platform removes that specific content. This model assumes that content is discrete, identifiable, and removable—that there is a specific file, stored at a specific location, that can be deleted or access-restricted.
Why this does not translate to generative-AI models: Once a model has been trained, "the impact of each data example in the training data is dispersed throughout the model and cannot be easily traced." To remove an example's influence, "one must either track down all the places where the example has an impact and identify a way to negate its influence, or re-train the entire model." The report flags that "'impact' is not well defined, and neither is 'removal.'" The alternative framing—defining "what it means to 'take down' a training example from a generative-AI model"—is itself "an ill-defined problem."
The report characterizes the state of technical research: "the subfield of 'machine unlearning' [11, 16] attempts to define the desired goals for removing an example and to design algorithms that satisfy these goals." Another line of work "attempts to quantify data-example attribution and influence; it seeks to define 'attribution' and then attribute generations from a model to specific data examples in the training data." The report emphasizes that "both machine unlearning and attribution are very young fields, and their strategies are (for the most part) not yet computationally feasible to implement in practice for deployed generative-AI systems."
The legal-technical gap is therefore twofold: (1) the technical community has not converged on definitions of "removal" or "attribution" that are computationally tractable for large-scale generative models, and (2) even if such definitions existed, it is not clear that they would satisfy what the law requires—because the legal concept of notice-and-takedown was designed for a different technological context where removal is conceptually straightforward.
The report frames this as an active area of research with "intense (and growing) investment," but its current state does not provide a reliable mechanism for complying with takedown requests or Right-to-be-Forgotten obligations for training data embedded in model parameters. The implication is that alternative legal or technical mechanisms may be needed—perhaps filtering model outputs rather than modifying model parameters, or requiring retraining with data excluded rather than attempting post-hoc removal—but the report does not specify what those alternatives should be.
Evaluation Metrics
Section 6.4 addresses the "clear need for useful metric definitions for Generative AI," particularly for legal concepts that do not have obvious quantitative formulations. The report identifies several specific challenges:
The tension between legal and ML evaluation: "The force of legal rules depends on how they are implemented and interpreted. Many decisions are made on a case-by-case basis, taking into account specific facts and context." In contrast, "machine-learning practitioners evaluate systems at scale. It is common practice to define metrics that can be applied directly to every situation (or at least a large majority of them)." These metrics "necessarily use pre-specified sets of features that may leave out considerations that may be important to forming a decision that appropriately accounts for broader context." The report frames this observation as not new—"this is hardly a new observation; it has had significant influence in machine-learning subfields, such as algorithmic fairness"—but emphasizes that it "holds true for Generative AI" and that there are "specific complexities for Generative AI that have not been so readily apparent in prior work in machine learning (e.g., in other areas of machine learning, there are accepted (though imperfect) notions of 'ground truth' labels, which are absent in Generative AI)."
Multiple, purpose-dependent definitions: The report uses memorization as an example of how "researchers create different, precise, definitions of memorization for different purposes." The definition of memorization for an image-generation model "will differ greatly from a code-generation model or a text-generation model." Similarly, "since 'removing' the impact of a training example from a trained model is an ill-defined problem, researchers may develop different metrics for quantifying whether or not training data points are successfully removed." The legal significance is that which metric is chosen can determine the legal conclusion—a model might satisfy one definition of "successfully unlearned" while failing another, and the law may not specify which definition controls.
Evaluation is dynamic, not static: "Many metrics are defined in terms of technical capabilities. For example, the evaluation of the amount of memorized training data in a model depends on the ability to extract and discover the memorized training data [18, 53]. As the techniques for data extraction improve, the evaluation of the model will change." This means that a model that appears safe (low memorization) under today's extraction techniques may be revealed to be unsafe under tomorrow's better attacks—and the legal system must decide whether "reasonable care" requires anticipating future extraction techniques or only protecting against currently known ones.
Supply-chain dependencies: "The way that the supply chain is constructed for a particular generative-AI model may alter the way that the systems (in which that model is embedded) can and should be evaluated." The report gives a concrete example: "the analysis may change depending on whether a model is aligned or not. In turn, as another example, it is possible that some actors may not have the relevant information to perform necessary evaluations; some actors may not even know if a particular model is aligned or not." This points to a supply-chain transparency problem: if a downstream fine-tuner receives a base model from an upstream provider, they may not know what evaluations were performed, what safety measures were applied, or even whether the model underwent alignment training—yet they may bear legal responsibility for the model's behavior in deployment.
Compositional effects: The report speculates about a scenario where "a machine-unlearning method may be applied to a model to remove the effect of a specific individual's data. However, that model may later be fine-tuned on additional data that is very similar to the removed individual's data. This may cause the individual's data to effectively 'resurface.'" The report acknowledges this is "presently speculation, though not all-together baseless," citing recent research showing that "the effects of alignment methods may be negated through the course of fine-tuning" (Qi et al., 2023). The broader point is that evaluations performed at one stage of the supply chain may not remain valid after subsequent stages—a model that passes a privacy audit before fine-tuning may not pass the same audit after fine-tuning on new data, and vice versa. This creates a legal challenge for certification or compliance regimes that assume a model can be evaluated once and then treated as safe.
The report's meta-point about evaluation is that designing useful metrics for generative-AI systems' legal compliance requires understanding both the technical properties being measured and the legal standards being operationalized—and that currently, neither community has a clear picture of what the other requires. The research agenda item is to develop that understanding, not (yet) to propose specific metrics.
4. Key Insights and Innovations
Innovation 1: Generative AI Is a "Generativity Jackpot" — and That Changes How Law Must Engage With It
The paper's most conceptually foundational move is not a legal argument or a technical result, but a diagnosis of why generative AI will be so legally consequential. By applying Jonathan Zittrain's five-dimensional theory of generative technologies (leverage, adaptability, ease of mastery, accessibility, transferability) to contemporary generative-AI systems, the paper makes a claim that is simultaneously descriptive and predictive: generative AI is the first technology since the Internet to score highly on all five dimensions simultaneously, and this fact means the legal challenges it raises will be "comparable in scope, scale, and complexity to those raised by computers and the Internet" (Section 2).
This framing matters because it calibrates expectations. Prior public and legal discourse has oscillated between treating generative AI as merely the latest in a long line of software innovations (and thus manageable under existing frameworks with minor adjustments) and treating it as fundamentally unprecedented (and thus requiring wholesale legal reinvention). The paper charts a middle path grounded in a pre-existing theoretical framework: the individual capabilities are not entirely new, but their simultaneous presence at massive scale creates a qualitatively different technology from a regulatory perspective. The Internet-law analogy provides a template for what sustained interdisciplinary engagement looks like — decades of doctrinal development, not a one-time legislative fix.
What distinguishes this from a casual analogy is the structured application of all five dimensions to specific, cited capabilities of generative AI (Section 2). For instance, the paper does not merely assert that generative AI is "accessible" — it observes the specific asymmetry that model creation requires enormous compute and expertise (concentrating it in a handful of institutions) while model use via natural-language prompting is inexpensive and widely available (distributing access to millions). This asymmetry has direct legal consequences: a small number of actors make design decisions with consequences for vast populations of users and rightsholders, which shapes the appropriate regulatory posture (who should bear what duties) in ways that differ from technologies where creation and use are more symmetrically distributed.
This is a fundamental diagnostic contribution, not an incremental refinement. The paper is essentially saying: before we argue about the answers, we need to agree on the scope and nature of the problem, and Zittrain's framework gives us a principled way to do that. The rest of the paper — the glossary, the taxonomies, the research agenda — follows from the recognition that this is an Internet-scale challenge requiring Internet-scale institutional responses.
Innovation 2: The Pre-Training/Fine-Tuning Distinction Is a Social Construct — But One With Legal Teeth
The paper makes a subtle but important intervention in how the technical community communicates the model-training pipeline to legal audiences. Section 4.2 identifies a conceptual trap that emerged during the workshop: legal scholars assumed "pre-training" referred to "a data preparation stage prior to and independent of training," when in fact pre-training is training — it involves updating model parameters through optimization, just like fine-tuning.
The paper's corrective is to articulate that the pre-training/fine-tuning distinction is an artifact of practical choices, not an essential technical boundary: "Both pre-training and fine-tuning are just training (though perhaps configured differently). The reasons we differentiate between these two stages have to do with how large-scale model training is done in practice; the distinction is only meaningful because researchers frequently choose to divide stages of training along these lines" (Section 4.2).
This is more than a terminological cleanup. It has direct consequences for legal analysis. If pre-training is mischaracterized as data preparation rather than training, the copyright analysis shifts: training involves using copyrighted works as examples to update model parameters (potentially implicating reproduction and derivative-work rights), while data preparation (deduplication, formatting, tokenization) is a different category of activity. The report's insistence that "pre-training is training, and not a data preparation stage" (Section 4.2) is a corrective to mischaracterizations that could serve as a liability shield — if training is framed as a neutral, automatic, purely technical process rather than a set of design choices with legal consequences.
The paper connects this to a broader accountability concern raised in Section 5.4: some design choices may be "made to seem 'internal' to the system (when, in fact, such choices are typically not foregone conclusions or strict technical requirements)" — citing Cooper et al., 2022, on accountability in algorithmic systems. Because "purely internal" copies tend to be fair use, characterizing training as an automated, developer-invisible process rather than as a human-designed pipeline with specific, contestable choices could affect the fair-use analysis.
This is an incremental but practically significant innovation in cross-disciplinary communication. It does not propose a new legal theory or a new technical method — it identifies a specific, recurring misunderstanding that can derail legal analysis before it starts, and provides the precise conceptual fix (pre-training is training; the distinction is a social convention, not a natural kind) that legal scholars need to avoid the trap.
Innovation 3: The Supply Chain Is the Unit of Legal Analysis, Not "The Model"
The paper's most significant framing innovation for legal analysis is its insistence that generative-AI systems should be analyzed through the lens of their supply chain — the sequence of stages (data collection, curation, pre-training, fine-tuning, alignment, deployment, use) and actors (data annotators, model trainers, API providers, application developers, end users) that produce and deploy these systems. This is not the paper's invention — it credits Lee et al., 2023, for the detailed treatment — but the paper elevates the supply-chain perspective to a central organizing principle for the entire interdisciplinary agenda.
Prior legal analysis of AI systems has often treated "the model" or "the company" as a monolithic entity — asking, for instance, "does this AI system infringe copyright?" or "is this company liable for this output?" The supply-chain perspective reveals that these questions are dangerously under-specified. Different actors at different stages make different decisions with different legal consequences: the entity that scrapes training data may be different from the entity that pre-trains the base model, which may differ from the entity that fine-tunes it for a specific application, which may differ from the entity that deploys it behind a consumer-facing API. Each actor's legal exposure depends on which decisions they made at which stage.
The paper demonstrates the analytic payoff of this framing in several concrete ways:
- Licensing complexity (Section 6.1): A dataset like LAION may be distributed under an open MIT license, but the individual image-caption pairs within it may each have their own (potentially restrictive) copyrights. A supply-chain analysis distinguishes between the dataset-level license (governing distribution of the curated collection) and the example-level rights (governing use of individual images for training), preventing the category error of treating the former as permission for the latter.
- Verifier quality depends on the chain (Section 4.3): The paper notes that decisions made early in the supply chain affect downstream behavior and evaluation: a model fine-tuned on proprietary data has a different trade-secrecy risk profile than a model that never leaves a user's device; a model that passes a privacy audit before fine-tuning may not pass the same audit after fine-tuning on new data.
- Centralization incentives (Section 6.1): The high cost of pre-training creates a structural dependency where many actors can afford to fine-tune but few can afford to pre-train, concentrating certain legal and economic decisions (about training data composition, alignment targets, and safety measures) in a handful of entities while distributing deployment decisions across many. The supply-chain perspective makes this dependency visible and subject to legal scrutiny.
This is a fundamental conceptual contribution to the emerging field of generative-AI law. It is not a technical method or a legal doctrine — it is a way of organizing analysis that makes visible the distribution of responsibility, the interdependence of design choices, and the mismatches between existing legal categories and the multi-actor, multi-stage reality of how generative-AI systems are produced and deployed. The paper validates this framing by showing that it yields concrete insights (the licensing-asymmetry observation, the centralization-incentive analysis) that a monolithic "model" perspective would miss.
Innovation 4: Machine Unlearning as a Case Study in Misaligned Technical and Legal Concepts
Not every innovation in a paper needs to be a solution. The paper's treatment of machine unlearning (Section 6.3) is a negative result with diagnostic power: it identifies a case where a technical subfield has developed with goals that do not correspond to what the law actually requires, and uses this misalignment to illustrate the broader need for bi-directional communication between the ML and legal communities.
The paper makes a specific, falsifiable claim: machine unlearning is not a drop-in replacement for legal notice-and-takedown under Section 512 of the US Copyright Act or for the "Right to be Forgotten" under GDPR. The claim rests on two observations. First, the operational model of notice-and-takedown assumes that content is discrete, identifiable, and removable — a specific file at a specific location that can be deleted. In a trained generative-AI model, "the impact of each data example in the training data is dispersed throughout the model and cannot be easily traced." Second, the legal concept of "removal" presupposes that we can define what successful removal means, but in the ML context, "'impact' is not well defined, and neither is 'removal'" (Section 6.3). The paper notes that defining "what it means to 'take down' a training example from a generative-AI model" is itself "an ill-defined problem."
The significance of this observation extends beyond the specific case of unlearning. It is a methodological argument about how technical and legal research should interact: if the ML community develops algorithms for "unlearning" without engaging legal scholars about what the law requires, and the legal community assumes "unlearning" provides compliance without understanding its technical limitations, both communities will converge on approaches that are simultaneously technically insufficient and legally non-compliant. The paper uses unlearning as a canonical example of the coordination failure it diagnoses throughout: two communities working on the same problem with incompatible definitions of success.
The paper reinforces this by noting that both "machine unlearning and attribution are very young fields, and their strategies are (for the most part) not yet computationally feasible to implement in practice for deployed generative-AI systems" — so even if the definitions could be aligned, the technical capacity to comply does not yet exist. This is a practical contribution for policymakers and legal practitioners: it warns against assuming that technical solutions for compliance are available, and identifies machine unlearning specifically as a field where regulatory demands should be informed by current (and realistically projected) technical feasibility rather than aspirational assumptions.
Innovation 5: Generative AI Inverts the Idea-Expression Dichotomy
Within the paper's broader legal taxonomy (Section 5.4), a specific observation stands out as theoretically distinctive: generative AI may turn copyright's idea-expression dichotomy "upside down." The paper credits Lemley, 2023, for this framing, but uses it to anchor a concrete set of analytic puzzles that illustrate why generative AI strains existing copyright categories in ways that prior technologies did not.
The traditional relationship in copyright is straightforward: a human author has an idea, expresses it in a fixed tangible form (the expression), and the expression receives copyright protection — but the underlying idea does not. With generative AI, several things about this sequence break. If we analogize the prompt to the idea and the generation to the expression, then the AI — not the human prompter — appears to be the one responsible for the creative expression, and purely computer-generated works are not currently copyrightable. But this analogy is unstable: the prompt itself may exhibit sufficient creativity to be independently copyrightable as a literary work, and there is a "tenable argument that the human prompter and generative-AI system are acting in concert to produce the resulting generation, and that the way that an idea is expressed in a prompt makes it inextricably indivisible from the resulting expressive generation" (Section 5.4).
What makes this a genuine conceptual innovation rather than a restatement of known problems is its precision about what specifically breaks. The paper identifies not just that generative AI creates authorship puzzles (this has been discussed since at least Samuelson, 1985, and Grimmelmann, 2016), but that it inverts the structural logic that copyright assumes about the relationship between human creativity and fixed expression. In the traditional model, the human provides the creative spark and the medium (pen, brush, keyboard) executes it. In generative AI, the human specifies desiderata (the prompt) and the system generates executions that may or may not correspond in detail to what the human envisioned. The relationship between idea-provider and expression-generator is scrambled — and copyright law, which allocates rights based on this relationship, has no straightforward mechanism for unscrambling it.
This is a theoretical contribution to the copyright scholarship on AI authorship. It does not propose a doctrinal resolution — the paper explicitly defers that to legal scholars — but it provides a crisp formulation of the problem that makes clear why it resists easy solutions, and why the "AI as tool" versus "AI as author" binary is insufficient. The observation generalizes beyond copyright: the same inversion logic may apply to questions of intent in torts and criminal law (if the AI generates harmful content the human did not specify) and to speaker attribution for First Amendment purposes (if the "speech" is generated by a system rather than formulated by a person).
5. Experimental Analysis
Evaluation Methodology
-
Dataset. The paper reports on discussions from the GenLaw workshop rather than on computational experiments. There is no quantitative dataset in the traditional machine-learning sense—no training set, test set, or benchmark over which models are evaluated. The "data" are the perspectives, expertise, and collaborative insights of the approximately forty workshop participants, drawn from computer science and law across 25 different institutions, as captured in the roundtable discussions on July 30, 2023. The workshop's topical scope was explicitly limited to intellectual property and privacy, with most discussion focused on U.S. law, though participants acknowledged the non-exhaustive nature of this scoping (Section 1). The report further notes that "not all capabilities, consequences, risks, and harms of Generative AI are legal in nature, so this taxonomy is not a complete guide to generative-AI policy" (Section 5), and that "the omission of other topics is not a judgment that they are unimportant" (Section 5).
-
Base model(s). The paper does not train, evaluate, or compare computational models. The "base models" in this context are the conceptual frameworks brought by participants from their respective disciplines: machine-learning researchers contributed technical understanding of model architectures, training procedures, scale, and supply chains (Section 4); legal scholars contributed doctrinal frameworks from copyright, privacy, torts, criminal law, and competition law (Section 5). The paper explicitly identifies PaLM 2-S*, GPT-4, ChatGPT, Stable Diffusion, LLaMA, DALL-E, Midjourney, MusicLM, Whisper, and others as examples of generative-AI systems discussed during the workshop, but does not subject any of them to systematic evaluation. The choice of which systems to reference reflects their prominence in public discourse and litigation at the time of the workshop in mid-2023—not a controlled experimental selection.
-
Metrics. The paper does not report quantitative metrics. Its "evaluation" is qualitative synthesis: identifying recurring themes, points of consensus, areas of confusion, and open research questions from the roundtable discussions. The paper's own characterization of this process is that it "reflect[s] the takeaways from the roundtable discussions" (Section 1) and that participants converged on "the most urgently needed contributions to the research area" organized into five broad headings. There is no quantitative measure of agreement among participants, no formal consensus-building procedure (e.g., Delphi method), and no inter-rater reliability statistics for the synthesis. The paper's validity rests on the credibility of the participants as domain experts and the plausibility of the synthesized takeaways, not on statistical rigor. The authors acknowledge this limitation implicitly through the disclaimer that "all of the listed authors contributed to the workshop upon which this report is based, but they and their organizations do not necessarily endorse all of the specific claims in this report" (Abstract)—signaling that the synthesis reflects the organizers' interpretation of discussions, not unanimous participant agreement.
-
Baselines. The paper does not compare against quantitative baselines. Its implicit conceptual baselines are the state of cross-disciplinary discourse prior to the workshop, characterized by specific communication failures: legal scholars misunderstanding "pre-training" as data preparation rather than training (Section 3), technologists treating "harms" as a colloquial notion rather than a legally structured concept (Section 3.1), and the dominance of copyright-centric analysis to the exclusion of other legal areas (Sections 5, 7). These are not formal baselines in the sense of a controlled experiment—they are characterizations of a pre-existing condition that the workshop was designed to improve. The paper also implicitly compares against other attempts to catalog generative-AI risks and policy concerns, citing Fergusson et al., 2023, as an example of a broader policy catalog, and positions its contribution as complementary: focusing "on highlighting the ways in which specifically legal issues may arise" (Section 5) rather than providing a general risk taxonomy.
-
Generation budget / compute accounting. There is no computational budget to account for. However, the paper does identify a form of attention budget that shaped the workshop: the two-day format (one public day at ICML, one off-site day for roundtables with approximately forty participants) and the deliberate scoping to IP and privacy created boundaries on what could be discussed and synthesized. The paper is explicit about these scope limitations: "for this first workshop, most discussion was limited to considerations of U.S. law" (Section 1), and the taxonomy is "non-exhaustive" (Section 5). The paper also identifies a key missing "expenditure" in its own analysis: "the GenLaw workshop was explicitly scoped to privacy and intellectual property (IP) issues, so this analysis should be considered non-exhaustive, and the omission of other topics is not a judgment that they are unimportant" (Section 5). This is analogous to acknowledging that a computational experiment was run on a limited benchmark and may not generalize—but without the ability to quantify the limitation.
-
Cross-validation / statistical protocol. There is no statistical protocol. The paper's mechanism for ensuring that the synthesis reflects more than the organizers' preconceptions is multi-participant deliberation: approximately forty experts from different institutions and disciplines engaged in roundtable discussions, and the reported takeaways are "organized into five broad headings, reflecting the participants' consensus about the most urgently needed contributions" (Section 1). However, the paper does not describe a formal consensus-building methodology—no voting, no thematic coding with inter-coder reliability, no member-checking where participants reviewed and endorsed the synthesis before publication. The disclaimer that listed authors "do not necessarily endorse all of the specific claims in this report" (Abstract) indicates that the synthesis is the organizers' interpretive work rather than a collectively authored document. This is a legitimate format for a workshop report, but it means the findings are best understood as informed expert judgment rather than as empirically validated results in the traditional experimental sense.
Main Quantitative Results
The paper reports no quantitative experimental results—no accuracy percentages, no F1 scores, no statistical tests, no tables of benchmark performance, no figures with plotted data. This is not a research paper with a computational evaluation; it is a workshop synthesis. The "results" are conceptual findings: identified communication failures (the terminological confusions documented in Section 3), characterizations of generative AI's distinctive technical features (Section 4), a taxonomy mapping those features onto legal doctrines (Section 5), and a prioritized set of research problems (Section 6).
Within this qualitative framework, the paper does organize its findings into logically distinct groupings, which can be characterized as "results" in the sense of synthesized takeaways:
Finding group 1: Terminology and metaphor failures as barriers to collaboration (Section 3). The paper reports that during the roundtable discussions, participants repeatedly encountered misunderstandings caused by terms that carry different meanings in ML and law. The specific examples cited are:
-
"Pre-training": Legal scholars "assumed that the term referred to a data preparation stage prior to and independent of training," while technologists use it to mean "an early, general-purpose phase of the model training process" that is itself training (Section 3). The paper reports that this confusion was identified and resolved only through extended conversation: "we found our way to common understandings only over the course of our conversations, and often only after many false starts" (Section 3).
-
"Harms": "Many technologists were not aware of the importance of harms as a specific and consequential concept in law, rather than a general, non-specific notion of unfavorable outcomes" (Section 3.1). The paper reports that this distinction matters because legal standing, remedies, and regulatory scope depend on legally cognizable harm—not just any unfavorable outcome.
-
"Memorization": The paper reports that the connection between technical memorization definitions (precise, quantifiable, referring to verbatim or near-verbatim reproduction of training examples) and "the colloquial meanings of these words can cause confusion" (Section 3.2). Specifically, "some outside of the machine-learning community misinterpret 'Generative AI memorization' to include functionality that goes beyond what machine-learning practitioners are actually measuring with their precise definitions," such as "discussions around text-to-image generative models 'memorizing' an artist's style"—which "is not equivalent to memorization in the technical sense of the word" (Section 3.2, citing Casper et al., 2023 for style similarity measurement as distinct from memorization measurement).
The paper's response to these findings is the glossary (Appendix A), which is presented as "a starting point, not a finish line" (Section 3.1).
Finding group 2: What makes generative AI technically distinctive (Section 4). The paper reports that legal scholars at the workshop had "a recurring question for the machine-learning experts in the room: What's so special about Generative AI? Clearly, the outputs created by Generative AI today are better than anything we have seen before, but what is the 'magic' that makes this the case?" (Section 4). The synthesized answer identifies three features:
-
Flexibility: The shift from task-specific discriminative models to "single, general-purpose models to solve many different tasks, rather than employing a model customized to each task we would like to perform" (Section 4.1). The paper characterizes this as applying "across a wide range of applications and modalities" including text-to-text, text-to-image, image captioning, music generation, speech generation, transcription, programming education, and scientific applications like protein folding and drug design.
-
Multi-stage training pipelines and supply chains: The paper reports that clarifying the pre-training/fine-tuning distinction and the multi-actor supply chain was an important outcome of the discussions, because "it became apparent that legal experts shared some misconceptions about the roles of pre-training and fine-tuning" (Section 4.2). The paper's corrective position—"pre-training is training, and not a data preparation stage"—is reported as a takeaway from the cross-disciplinary exchange.
-
Scale: The paper reports that "state-of-the-art models today are an order of magnitude larger and trained on significantly more data than the biggest models from five years ago" (Section 4.4), and that the techniques are not fundamentally new—"language models, for example, have existed since at least the 1980s"—but the scale at which they are now deployed produces emergent behaviors and economic dynamics (concentrated model creation, distributed use) that are legally salient.
These findings are not quantified. The paper does not, for example, provide parameter counts for models across years, FLOPs comparisons, or benchmark performance trajectories. The "results" are conceptual characterizations intended to equip legal scholars with a working mental model of generative-AI technology—a model that the paper argues is more accurate than the mental models legal scholars brought to the workshop (as evidenced by their initial questions and misconceptions).
Finding group 3: Legal taxonomy highlighting non-obvious issues (Section 5). The paper reports that the roundtable discussions produced progress toward a taxonomy organized around intent, privacy, misinformation, and intellectual property. The key substantive observations that emerged include:
-
Intent: The paper reports that "there is unlikely to be a simple across-the-board answer as to how the 'intent' of a generative-AI system should be measured, in part because the legal system uses intention in so many ways and so many places" (Section 5.1). A specific discussion about applying respondeat superior (employer liability for employee torts) to the AI-as-employee scenario revealed that "in the usual application of respondeat superior, there is still an embedded notion of intention... the employee's intentions are still relevant in determining whether a tort has been committed at all" (Section 5.1), so the analogy does not sidestep the intent problem as neatly as it first appeared.
-
Privacy: The paper reports a meta-level observation about cross-disciplinary tension: "Computer scientists often want to be able to quantify policy, including policy for handling privacy concerns; in the law, the mere desire to quantify complex concepts like privacy can itself be the source of significant problems" (Section 5.2). Generative AI "is poised to make these privacy challenges even harder" because of web-scraped training data containing PII, the potential for models to "link together information in novel ways that reveal sensitive information about individuals" (going beyond what a search engine can do by surfacing existing data), and adversarial prompt-based extraction of sensitive information.
-
Misinformation/disinformation: The paper reports a discussion about whether deepfake harms should be classified as defamation, privacy violations, or both. Privacy experts in computing suggested deepfakes could be "intimate privacy violations" (citing Citron & Solove, 2022); legal experts responded "that this would likely not constitute a cognizable privacy harm under the law, although it would still be actionable as defamation or false light" (Section 5.3). The paper reports that this exchange "raised questions about whether Generative AI could create new types of harms that blur current conceptions of disinformation and privacy harms" (Section 5.3). This is not a settled finding—it is a reported disagreement that the paper identifies as an open question requiring further research.
-
Intellectual property: Within the IP discussion, the paper reports observations about volition (the risk that design choices are characterized as "internal to the system" to deflect liability, Section 5.4), market externalities (that fair-use analysis already considers market effects, so labor displacement and market concentration are internal to copyright analysis, not external policy concerns), trade secrecy (the tension between fine-tuning on proprietary data and the risk of extraction through prompting), and the idea-expression dichotomy (the inversion where "the AI, rather than the human, is responsible for the creative expression" if the prompt is treated as the idea and the generation as the expression). The paper explicitly defers to existing legal scholarship for doctrinal details (citing Callison-Burch, 2023; Lee et al., 2023; Sag, 2023; Samuelson, 2023; Vyas et al., 2023, among others).
Finding group 4: Research agenda items (Section 6). The paper reports that participants identified four research areas where technical and legal expertise must be jointly applied. These are not experimental results but agenda items—problems the paper argues are important and under-studied:
-
Centralization vs. decentralization (6.1): The paper reports that "it remains to be seen whether courts will rule that the use of such datasets [web-scraped, with closed-licensed examples] constitutes fair use" and that an alternative path is to "invest in producing open, permissively licensed datasets that avoid the alleged legal issues of using web-scraped data." Current attempts at this have "yielded mixed results in terms of generation quality" (citing Gokaslan et al., 2023). This is reported as an active area requiring both technical innovation (training models on openly licensed data) and legal innovation (developing licenses that function as intended across jurisdictions).
-
Rules, standards, reasonableness, and best practices (6.2): The paper reports on the dynamic relationship between technical feasibility and legal standards: "once technology begins to stabilize, it becomes easier to define concrete standards (e.g., safety standards) ... accordingly, by definition, compliance with such standards is sufficient for meeting the bar of reasonableness. Until there is some stability, when harms occur, there will necessarily be some flexibility" (Section 6.2). The paper frames this as a call for both communities to "take hold of the opportunity to actively engage in research and public policy regarding today's generative-AI systems, such that they can help shape the development of future standards."
-
Machine unlearning vs. notice-and-takedown (6.3): The paper reports that "for generative-AI models, there is no straightforward analogue for simply removing a piece of data from a database" because "once a model has been trained, the impact of each data example in the training data is dispersed throughout the model and cannot be easily traced." Both machine unlearning and attribution "are very young fields, and their strategies are (for the most part) not yet computationally feasible to implement in practice for deployed generative-AI systems." This is reported as a case where technical and legal conceptions of "removal" are misaligned, requiring joint redefinition of what compliance means.
-
Evaluation metrics (6.4): The paper reports that "there is a clear need for useful metric definitions for Generative AI," particularly for legal concepts that resist quantification. The specific complexity identified is that "in other areas of machine learning, there are accepted (though imperfect) notions of 'ground truth' labels, which are absent in Generative AI" (Section 6.4)—meaning that evaluating generative-AI systems for legal compliance requires defining both what "correct" behavior means (a legal question) and how to measure it at scale (a technical question).
Ablation Studies and Robustness Checks
The paper does not contain ablations in the computational-experiment sense—there is no model component removed, no hyperparameter varied, no alternative configuration compared. However, the paper does contain several structural features that serve an analogous function: they test whether the reported findings depend on specific assumptions or conditions, or they identify boundaries and limitations that a less careful analysis might elide.
Scope disclaimer as a robustness check: The paper repeatedly identifies what it is not covering, which functions as a form of boundary specification. The workshop "was explicitly scoped to privacy and intellectual property (IP) issues" (Section 5); "most discussion was limited to considerations of U.S. law" (Section 1); and "this analysis should be considered non-exhaustive, and the omission of other topics is not a judgment that they are unimportant" (Section 5). These disclaimers mean the paper's taxonomy is explicitly acknowledged as partial—not a claim to comprehensiveness. This is methodologically significant because it preempts the criticism that the paper fails to cover important legal areas (e.g., labor law, international law, environmental law) by stating upfront that it does not attempt to.
Business-model diversity as a robustness check: Section 3.3 identifies four different business-model patterns (B2C hosted services, B2B integration, open-model derivatives, supply-chain specialists) and emphasizes that this landscape "is likely to continue to evolve as new business players enter the field." The paper does not claim that its characterization of business models is exhaustive or stable. This is analogous to testing whether findings hold across different deployment contexts—the paper acknowledges that different business models create different legal relationships and that findings from one context may not generalize to others.
"Known vs. unknown" terminological differences as a sensitivity analysis: Section 3 distinguishes between terminological differences that communities are aware of (and can adjust for) versus those they are not (and thus cannot). The "known" differences—like privacy—are less dangerous because "the two communities have generally understood that they mean something different." The "unknown" differences—like pre-training and harms—are more dangerous because they cause "collaborations to derail without participants realizing why." This distinction is a form of sensitivity analysis: it tests whether the conclusion "terminological confusion is a problem" depends on which terms are confused, finding that the nature and severity of the problem depends on whether the confusion is recognized by the participants. The paper's recommendation (glossary entries that explicitly flag terms of art) follows from this analysis: known differences require translation; unknown differences require explicit definition.
Metaphor limitations as a robustness check: The paper's discussion of metaphors (Section 3.2 and Appendix B) includes explicit boundary conditions for each metaphor discussed. For example, the "models are trained like dogs" metaphor is acknowledged as imperfect because "unlike training a dog, model training does not typically have a curriculum; there is no progression of easier to harder skills to learn." The "LLMs are stochastic parrots" metaphor includes the counterpoint that "critics of the stochastic-parrot analogy say that it undervalues the competencies that state-of-the-art language models have." The "generations are collages" metaphor is explicitly analyzed for where it breaks down: "a generative-AI system does not take several works and splice them together... there is no author 'selecting, coordinating, or arranging' training examples." This structured analysis of metaphor failure is functionally equivalent to an ablation that identifies when a simplifying assumption (the metaphor) ceases to be useful for legal reasoning.
Cross-disciplinary disagreement as a natural robustness check: The paper reports at least two instances where workshop participants disagreed, and the disagreement itself is reported as a finding rather than suppressed as noise:
-
On whether deepfake harms constitute privacy violations: computing privacy experts argued yes (citing intimate privacy violation frameworks); legal experts responded that existing privacy doctrine would likely not classify them that way, though defamation or false light might apply (Section 5.3). The paper reports this as an unresolved question, not as a settled consensus.
-
On whether respondeat superior sidesteps the intent problem: one participant proposed this as a solution; another participant noted that the employee's intent is still relevant in the first step of the analysis (Section 5.1). The paper reports both positions and the tension between them, rather than resolving it.
These reported disagreements function as robustness checks on the claim that the workshop produced "consensus"—they show that the paper is not falsely representing uniform agreement where expert opinion actually diverges. The existence of reported disagreement strengthens credibility because it demonstrates that the synthesis did not paper over genuine conflicts in expert judgment.
Predictive speculation flagged as speculation: The paper includes a forward-looking claim about machine unlearning: "a machine-unlearning method may be applied to a model to remove the effect of a specific individual's data. However, that model may later be fine-tuned on additional data that is very similar to the removed individual's data. This may cause the individual's data to effectively 'resurface'" (Section 6.4). The paper explicitly flags this as "presently speculation, though not all-together baseless," citing Qi et al., 2023 on how alignment effects can be negated through fine-tuning. This is methodologically analogous to acknowledging that a hypothesized result has not been experimentally verified and that the supporting evidence is indirect—a form of intellectual honesty about the limits of current knowledge.
The glossary as a living document: The paper explicitly states that the glossary is "offered as a starting point, not a finish line" and that it "will host and update this glossary on the GenLaw website" because "the field is in flux; its terminology will evolve as new technologies and controversies emerge" (Section 3.1). This is akin to acknowledging that model evaluation on a static test set may not capture performance on future data distributions—the paper builds in a mechanism for updating its own reference frame as the technology and legal landscape change.
Critical Assessment
The fundamental challenge in assessing whether this paper's experiments support its claims is that the paper makes no experimental claims in the traditional sense. It reports a synthesis of expert discussions, not computational results. Assessing it by the standards of a quantitative ML paper would be a category error—equivalent to faulting a legal brief for lacking test-set accuracy. The appropriate assessment framework is whether the paper achieves what it sets out to do: provide a useful synthesis of interdisciplinary discussions that identifies barriers to collaboration, characterizes generative AI for a legal audience, maps legal issues, and sets a research agenda.
Within that framework, the paper achieves its stated goals with reasonable credibility but has identifiable limitations that readers should weigh when deciding how heavily to rely on its conclusions.
Strengths of the approach:
The paper's decision to include concrete, specific examples of terminological confusion, rather than making vague claims about "communication challenges," substantially increases its usefulness. The identification of "pre-training" as a term that legal scholars misunderstood as "data preparation" rather than "training" is not just an abstract observation—it is a specific, actionable finding that lawyers reading the paper can use to check their own understanding, and that technologists can use to adjust how they communicate with legal colleagues. This is the paper's strongest contribution: it names specific failure modes rather than gesturing at general difficulties.
Similarly, the paper's treatment of metaphors as simultaneously useful and dangerous represents a genuine analytic contribution. Rather than either rejecting metaphors as misleading (which would discard a valuable communication tool) or embracing them uncritically (which would propagate misunderstandings), the paper models a third approach: identify each metaphor, explain what it captures, and—crucially—articulate where it breaks down. The glossary in Appendix A and the metaphor discussion in Appendix B operationalize this approach. A lawyer who reads the "generations are collages" entry and its deconstruction will be better equipped to avoid the copyright errors that arise from taking that metaphor literally.
The disagreement-reporting—where the paper acknowledges that workshop participants disagreed about specific legal classifications (deepfakes as privacy vs. defamation) and about the viability of specific legal analogies (respondeat superior)—is methodologically honest and practically useful. It prevents the reader from incorrectly inferring expert consensus where it does not exist, and it identifies the precise fault lines where future research is needed.
Weaknesses and limitations:
Participant selection and representativeness are not addressed. The paper reports that "approximately forty participants conducted a series of roundtable discussions" on the second day (Section 1) and that participants came from "25 different institutions" (Section 7). But there is no information about how participants were selected—were they invited by the organizers? Through an open call? What criteria were used? The disciplinary balance, the range of perspectives represented (e.g., plaintiff-side vs. defendant-side legal views, industry vs. academic affiliations, different subfields within ML), and potential selection biases that might skew the synthesis toward particular conclusions are not discussed. A reader cannot assess whether the reported "consensus" reflects broad expert agreement or the views of a particular subset of the relevant expert community. This is analogous to a computational paper reporting results on a dataset without describing how the dataset was constructed—the conclusions may be valid for the sample studied, but the reader cannot judge their generalizability.
No formal consensus methodology. The paper reports that the five organizing headings "reflect the participants' consensus about the most urgently needed contributions" (Section 1), but provides no information about how consensus was determined. Were there votes? Did the organizers independently code the discussions and check inter-coder reliability? Did participants review and endorse the synthesized findings? The disclaimer that authors "do not necessarily endorse all of the specific claims in this report" (Abstract) indicates that the synthesis is the organizers' interpretive work, not a document that all participants co-authored or formally approved. This is a legitimate format for a workshop report—many conference and workshop reports are written this way—but it means the paper's epistemic status is informed expert opinion by the organizers, informed by discussion, not empirically validated expert consensus as determined by a systematic method.
The "generativity jackpot" claim is underexamined. Section 2 applies Zittrain's five dimensions of generativity to generative AI and concludes that it "hits the generativity jackpot"—scoring highly on all five dimensions. This is a plausible claim, and it structures the paper's argument about why generative AI will be legally transformative. But the paper does not consider potential counterarguments: Are there dimensions where generative AI scores less highly than computers or the Internet, and does this affect the analogy? Is "ease of mastery" really high when effective prompt engineering for complex tasks requires substantial skill? Is "accessibility" really high when cutting-edge models are controlled by a handful of companies and access is mediated by APIs with terms of service that restrict use? The paper acknowledges the creation-use asymmetry—"currently limiting model creation to a handful of institutions" (Section 2)—but does not explore whether this asymmetry undermines the generativity thesis by concentrating control over the technology's evolution. A more critical application of the framework would strengthen rather than weaken the paper, by clarifying the scope conditions under which the Internet-law analogy is most (and least) applicable.
The taxonomy's scope limitations are deeper than acknowledged. The paper states that the taxonomy is limited to IP and privacy and that "the omission of other topics is not a judgment that they are unimportant" (Section 5). This is fair. But within the topics covered, there are also significant gaps. The IP discussion (Section 5.4) touches on copyright, trade secrecy, patent, and authorship, but does not address trademark (generative AI producing outputs that infringe on protected marks, or generating confusingly similar branding), right of publicity (generation of likenesses without consent, intersecting with the deepfake discussion in 5.3 but receiving separate legal treatment), or moral rights (attribution and integrity rights that exist in some jurisdictions and may be implicated by AI-generated derivatives). The privacy discussion (Section 5.2) does not address sectoral privacy laws (HIPAA for health data, FERPA for educational data, COPPA for children's data) that impose specific obligations beyond general privacy principles. The paper's disclaimer that it is "non-exhaustive" is accurate—but the specific omissions may affect practical relevance for lawyers working in regulated sectors where these additional frameworks apply.
The research agenda items are high-level; specific research designs are not proposed. Section 6 identifies four research areas, but for each, the paper characterizes the problem and its importance without specifying what a research project in that area would look like. For example, Section 6.4 identifies the need for "useful metric definitions for Generative AI" for legal concepts, and observes that "evaluation is dynamic, not static" because extraction techniques improve over time. But the paper does not propose specific metrics, describe what a useful metric would need to satisfy, or identify concrete legal standards that could be operationalized. This is not necessarily a weakness—the paper explicitly frames these as agenda items for future work, not as completed analyses—but a reader hoping for actionable guidance on how to develop such metrics will find only problem framing, not solutions.
The U.S.-centric focus limits direct applicability elsewhere. The paper acknowledges that "for this first workshop, most discussion was limited to considerations of U.S. law" (Section 1). This is a significant limitation given that many generative-AI systems are deployed globally and are subject to multiple legal regimes simultaneously. The GDPR discussion (Appendix A.4, Section 6.3 on the Right to be Forgotten) acknowledges European law in passing, but the paper's legal analysis (Section 5) is overwhelmingly organized around U.S. doctrinal categories: fair use rather than fair dealing, U.S. constitutional standing requirements rather than other jurisdictions' approaches to justiciability, U.S. tort law rather than other countries' delictual frameworks. The paper's value for non-U.S. legal audiences depends on how well U.S. legal concepts translate to their jurisdictions—and the paper does not address this translation problem (ironically, given its central concern with cross-disciplinary translation). The planned global expansion (Section 7: "we are actively focusing on expanding our engagement globally") acknowledges this limitation implicitly.
Missing: cost-benefit analysis of the proposed interventions. The paper argues for building shared knowledge bases, developing metaphors, creating taxonomies, and pursuing cross-disciplinary research. But it does not assess the cost of these activities relative to alternatives. For instance, is the effort required to maintain an evolving glossary (with "frequently updated to keep pace with technological changes," Section 7) better spent than, say, funding specific legal scholarship or developing specific technical standards? The paper's implicit theory of change is that better communication and shared conceptual frameworks will lead to better legal and technical outcomes—a plausible theory, but one the paper does not defend against the alternative view that the urgent problems require immediate doctrinal or technical solutions, and that investing in communication infrastructure is a luxury that can wait. A more critical paper would acknowledge this tension and argue explicitly for why infrastructure-first is the right prioritization.
Overall assessment: The paper does not make experimental claims that can be "supported" or "not supported" in the traditional sense. What it does is identify real communication failures that occurred during the workshop, synthesize those failures into a structured diagnosis of what prevents productive interdisciplinary work on generative AI and law, and propose infrastructure (glossaries, taxonomies, metaphors-with-boundary-conditions, research agendas) designed to address those failures. The credibility of this diagnosis rests on the specificity of the examples and the plausibility of the synthesis, not on quantitative validation. The paper's value for a reader depends on whether they recognize the communication failures it describes—either from their own experience or from finding the examples convincing—and whether they find the proposed infrastructure useful for their own work. For a reader who does recognize these failures, the paper provides a vocabulary for naming them and a starting framework for addressing them. For a reader who does not, the paper may read as overgeneralized workshop reflections. The paper's greatest limitation is that it cannot, by its nature, demonstrate that its proposed interventions work—that the glossary improves legal-technical collaboration, that the taxonomy produces better legal analysis, that the research agenda leads to productive projects. These are empirical questions that the paper leaves to future research, including, presumably, future GenLaw workshops that can compare outcomes against the baseline the current paper establishes.
6. Limitations and Trade-offs
Limitation 1: The Workshop's Scope Was Limited to IP and Privacy Under U.S. Law
The assumption or constraint. The paper explicitly narrows its substantive focus: "for this first workshop, most discussion was limited to considerations of U.S. law" (Section 1), and "the GenLaw workshop was explicitly scoped to privacy and intellectual property (IP) issues, so this analysis should be considered non-exhaustive, and the omission of other topics is not a judgment that they are unimportant" (Section 5). The legal taxonomy in Section 5 covers intent, privacy, misinformation, and IP—but does not address torts beyond intent (e.g., negligence, products liability in detail), criminal law beyond mens rea, labor law, trademark, right of publicity, sectoral privacy regulations (HIPAA, FERPA, COPPA), international law, human rights frameworks, or competition/antitrust law beyond brief mention in Section 6.1.
The consequence. A practitioner reading this report as a comprehensive guide to generative AI's legal landscape will miss entire categories of legal risk. For instance, the paper does not analyze how products liability doctrine—which holds manufacturers strictly liable for defective products regardless of intent—might apply to generative-AI systems that produce harmful outputs, despite the fact that Section 5.1 discusses intent at length and notes that "some crimes and torts are 'strict liability' (e.g., a manufacturer is liable for physical harm caused by a defective product regardless of whether they intended that harm)." The paper identifies the existence of strict liability but does not develop the analysis, leaving practitioners without guidance on what may become a primary litigation avenue. Similarly, the U.S.-centric focus means that companies operating globally—which includes essentially all major generative-AI providers—receive no analysis of GDPR compliance beyond the glossary entry and a brief mention of the Right to be Forgotten in Section 6.3, or of the EU AI Act, China's generative-AI regulations, or other emerging international frameworks.
What evidence exists in the paper. The paper is transparent about its scope limitations through explicit disclaimers. Section 5 states that "this analysis should be considered non-exhaustive," and Section 7 acknowledges that "while our first event and materials have had a U.S.-based orientation, we are actively focusing on expanding our engagement globally." The paper does not provide any evidence that its findings do generalize to other legal domains or jurisdictions—it simply reports what was discussed and flags what was not.
Mitigation status. The paper partially addresses this by framing itself as a first workshop report ("interim contribution as part of an ongoing project," Section 5) and announcing plans for global expansion (Section 7). However, no timeline or mechanism for expanding scope is specified. The paper suggests this is a direction for future GenLaw events rather than a gap that can be filled by other means. For practitioners needing guidance on non-U.S. law or on legal areas beyond IP and privacy, the paper provides framing tools (the supply-chain analysis, the glossary, the technical characterization) but no doctrinal analysis.
Limitation 2: Difficulty Estimation for Cross-Disciplinary Communication Has No Cost Accounting
The assumption or constraint. The paper's central intervention is building infrastructure for cross-disciplinary communication—a glossary, metaphors, taxonomies, and research agendas. Yet the paper reports that during the workshop itself, participants achieved common understanding only through intensive, high-bandwidth interaction: "we found our way to common understandings only over the course of our conversations, and often only after many false starts" (Section 3). The workshop format—two days, approximately forty experts, substantial organizer effort—represents a very high "cost" of communication that is not acknowledged as a scalability constraint.
The consequence. The paper's proposed solution—maintaining an evolving glossary, identifying and deconstructing metaphors, mapping legal taxonomies onto technical features—assumes that these artifacts can substitute for the intensive, interactive sense-making that occurred at the workshop. But it provides no evidence that a static (even if periodically updated) glossary or taxonomy document produces the same reduction in misunderstanding as forty experts arguing through definitions in real time. If the primary barrier to interdisciplinary collaboration is not lack of definitions but lack of awareness that a definition is needed—the "unknown unknowns" of terminology that the paper identifies as the most dangerous category (Section 3)—then written artifacts may be insufficient because they require the reader to recognize that they are confused before they consult the reference. The workshop format resolved this through conversation: someone used a term incorrectly, someone else noticed, and they worked through the misunderstanding. A glossary cannot replicate this detection-and-correction loop.
This limitation parallels a common failure mode in technical systems: building a tool that solves the problem as diagnosed while missing that the diagnosis captured only the symptoms visible in a high-resource, expert setting. The glossary may be most useful to people who are already sophisticated enough to know they should consult it—the very people who need it least.
What evidence exists in the paper. The paper provides no evaluation of whether the glossary or other artifacts actually improve cross-disciplinary communication outside the workshop context. There is no user study, no before-and-after comparison of legal analysis quality by lawyers who did versus did not read the glossary, and no measure of whether the identified terminological confusions (pre-training, harms, memorization) persist in subsequent GenLaw events. The paper's only "evidence" that the glossary is needed is the report of confusions that occurred at the workshop—which establishes the problem but not the efficacy of the proposed solution. The paper acknowledges that "the resources that we develop (such as those in this report) will need to be frequently updated to keep pace with technological changes" (Section 7), but updating a resource is different from validating that it works.
Mitigation status. The paper does not address this limitation directly. It frames the creation of shared knowledge base resources as an unalloyed good without examining whether the form (static documents) matches the function (resolving unrecognized misunderstandings in real time). The paper's commitment to updating the glossary on the GenLaw website and growing GenLaw as an organization (Section 7) suggests iterative improvement, but there is no proposed mechanism for evaluating whether the resources are effective. This is a significant gap for anyone deciding whether to invest in building or adopting similar infrastructure in their own cross-disciplinary work.
Limitation 3: The Paper Does Not Engage with Competing Frameworks or Alternative Diagnoses
The assumption or constraint. The paper frames the core problem as a communication and coordination failure between ML and legal communities, and proposes infrastructure (glossaries, metaphors, taxonomies) as the solution. It does not consider alternative diagnoses of why interdisciplinary collaboration on generative AI and law has been difficult. For example: the problem might not be primarily communication but incentives—ML researchers and legal scholars have different professional reward structures, publication norms, and timelines that discourage the kind of sustained collaboration the paper advocates. Or the problem might be asymmetric stakes: technology companies have strong economic incentives to shape legal outcomes, creating a structural power imbalance that no amount of shared vocabulary can equalize. Or the problem might be fundamentally political rather than intellectual—disagreements about what the law should be, not confusion about what the technology does.
The consequence. If the primary barrier to productive interdisciplinary work is not terminological confusion but misaligned incentives, power asymmetries, or political disagreement, then the paper's proposed solutions (glossaries, taxonomies, research agendas) address symptoms rather than causes. A lawyer who perfectly understands that pre-training is training may still face insurmountable obstacles to obtaining discovery about training data composition if the model developer claims trade-secret protection—a problem of legal procedure and corporate power, not terminology. A technologist who perfectly understands the legal standard for fair use may still be unable to predict whether a court will find their particular use fair, because fair use is a multi-factor, case-specific determination that does not yield to algorithmic specification. The paper's communication-infrastructure approach implicitly assumes that better mutual understanding leads to better outcomes, but it does not defend this assumption against the possibility that the obstacles are structural rather than informational.
What evidence exists in the paper. The paper provides no analysis of alternative explanations for the difficulty of cross-disciplinary collaboration. The workshop discussions presumably touched on some of these issues—Section 5.4 discusses the risk that design choices are characterized as "internal" to deflect liability (a power/incentive issue, not a communication issue), and Section 6.1 discusses centralization incentives (a structural-economic issue). But the paper does not integrate these observations into its diagnostic framework or consider whether they challenge the primacy of the communication-failure diagnosis. The paper treats communication barriers as the primary problem and structural/incentive issues as additional problems within the taxonomy, rather than considering whether structural issues might be more fundamental and communication barriers a downstream consequence.
Mitigation status. Not addressed. The paper does not acknowledge that its diagnosis might be incomplete or that alternative frameworks might yield different priorities. This is a limitation for readers who are skeptical that better glossaries and metaphors will meaningfully shift outcomes, or who believe that the most urgent interventions are political, regulatory, or economic rather than intellectual. The paper would be stronger if it acknowledged these alternatives and argued for why, despite their plausibility, the communication-infrastructure approach is the right starting point—or at minimum, identified the conditions under which it is versus is not sufficient.
Limitation 4: No Guidance on How to Resolve Conflicts Between Legal and Technical Definitions
The assumption or constraint. The paper identifies that legal and technical communities often define the same term differently—privacy, memorization, harms—and that these definitional differences are a barrier to collaboration. The glossary (Appendix A) provides both legal and technical definitions for many terms. However, the paper provides no framework for resolving situations where the legal and technical definitions conflict or where a legal standard requires a determination that technical definitions cannot provide.
The consequence. Practitioners operating at the intersection of generative AI and law face a specific, recurring problem: they must satisfy legal requirements that are defined in terms that do not map cleanly onto technical metrics. The paper names this problem but does not help solve it. For example, Section 5.2 observes that "computer scientists often want to be able to quantify policy, including policy for handling privacy concerns; in the law, the mere desire to quantify complex concepts like privacy can itself be the source of significant problems." Section 6.4 notes that "the challenges of operationalizing or concretizing societal concepts into math has been discussed at length in prior works." But the paper stops at identifying the tension—it does not offer principles for navigating it. When a company must comply with GDPR's requirement to implement "appropriate technical and organizational measures" to protect personal data, should they deploy differential privacy (which provides mathematical guarantees but may not satisfy European regulators' understanding of "appropriate")? When a fair-use determination requires assessing whether a use is "transformative," can memorization metrics (Section 3.2) provide evidence one way or another, or are they measuring something orthogonal to what the legal standard requires? The paper identifies that these are open questions but provides no framework for answering them.
This is particularly consequential because the paper's research agenda (Section 6) explicitly calls for work on evaluation metrics for legal concepts (6.4) and for standards of reasonableness (6.2)—both of which require bridging the legal-technical definitional gap. The paper tells us that the gap exists and needs bridging, but not how to bridge it, even at the level of methodological principles.
What evidence exists in the paper. The paper thoroughly documents the existence of the gap. Section 3 gives examples of definitional conflicts. Section 5 provides multiple instances where legal and technical conceptions diverge: privacy (5.2), what constitutes misinformation (5.3), authorship (5.4), the idea-expression dichotomy (5.4). Section 6.3 analyzes machine unlearning as a case where "there is no straightforward analogue for simply removing a piece of data from a database" and "'impact' is not well defined, and neither is 'removal'"—showing that the technical community has not converged on definitions that map to legal requirements. But the paper provides no methodology for resolving such conflicts, no criteria for preferring one definition over another in a legal context, and no examples of successful resolution it can point to as models.
Mitigation status. The paper frames this as an open research problem—"designing useful metrics will be an important, related area of research for Generative AI and law" (Section 6.2)—rather than as something the workshop resolved. This is an honest characterization but leaves the practitioner without actionable guidance. The paper's contribution is diagnostic (identifying the gap) rather than prescriptive (telling you how to cross it). For a reader hoping to operationalize legal standards in a technical system, or to explain technical limitations to a legal audience in a way that informs legal reasoning, the paper provides vocabulary for describing why the task is hard but no strategy for doing it.
Limitation 5: The Supply-Chain Framework Identifies Complexity Without Providing Simplifying Heuristics
The assumption or constraint. The paper's supply-chain perspective (Sections 4.2–4.3, drawing on Lee et al., 2023) emphasizes that generative-AI systems involve multiple stages, actors, and decision points, and that "since many actors can be involved in the generative-AI supply chain, and decisions made in one part of the supply chain can impact other parts of the supply chain, it can be useful to identify each intervention and decision point and think about them in concert" (Section 4.3). The paper uses this framework to argue against monolithic analysis that treats "the model" or "the company" as a uniform entity, and to show that specific legal conclusions (about licensing, about privacy, about evaluation) depend on which actor made which decision at which stage.
The consequence. The supply-chain framework, taken seriously, implies that every legal analysis of a generative-AI system must trace the entire chain of actors and decisions—a task of potentially unbounded complexity. The paper provides no heuristics for determining which supply-chain links are legally salient for which questions, or for bounding the analysis when full supply-chain information is unavailable. In practice, downstream deployers may not know what training data was used, what curation decisions were made, or whether the base model underwent alignment (Section 6.4: "it is possible that some actors may not have the relevant information to perform necessary evaluations; some actors may not even know if a particular model is aligned or not"). The paper identifies this transparency problem but does not provide guidance for how a legal analysis should proceed when supply-chain information is missing—should it assume the worst case? Treat the downstream deployer as bearing the risk of upstream opacity? Defer to whatever the upstream provider discloses?
This is a practical limitation because many of the actors most likely to need legal guidance—startups fine-tuning open models, companies integrating generative-AI APIs into their products, individual developers building applications on top of hosted services—sit at the end of supply chains they did not create and cannot fully inspect. If the supply-chain framework implies that their legal exposure depends on decisions made upstream that they cannot see, the framework creates analytical paralysis rather than actionable guidance.
What evidence exists in the paper. The paper demonstrates the value of supply-chain analysis through specific examples—the licensing complexity of datasets like LAION (Section 6.1), the compositional evaluation problem where fine-tuning can undo privacy protections (Section 6.4). But it does not demonstrate that the framework is tractable for practitioners with limited information. The paper acknowledges the information-asymmetry problem in passing (Section 6.4: "the analysis may change depending on whether a model is aligned or not... some actors may not even know if a particular model is aligned or not") but treats it as an additional complexity rather than as a potential failure mode for the entire framework.
Mitigation status. Not addressed. The paper does not propose transparency requirements, information-sharing protocols, or safe-harbor provisions that might make supply-chain analysis feasible for actors with incomplete information. The research agenda (Section 6) identifies centralization versus decentralization as a topic (6.1) and evaluation metrics as a topic (6.4), both of which touch on supply-chain transparency, but neither is framed as solving the tractability problem. The paper mentions that "effective standards will differ by model modality and other system capabilities" (Section 6.2) but does not discuss whether those standards should require supply-chain disclosure, or what downstream actors should do in the absence of such disclosure. For a practitioner deciding whether to deploy a generative-AI system, the paper's message is essentially: your legal exposure depends on the entire supply chain, good luck figuring it out.
Limitation 6: The Paper's Theory of Change Is Unspecified and Untested
The assumption or constraint. The paper proposes building shared knowledge bases, clarifying technical capabilities, creating legal taxonomies, and articulating research agendas. It argues that these are "urgently needed" (Section 1) and that they address the communication failures that "hampered our ability to collaborate on assessing emerging issues" (Section 3). The paper implicitly assumes that better mutual understanding between ML and legal communities will lead to better legal analysis, better technical design choices, and ultimately better societal outcomes—but it never articulates this causal chain or subjects it to scrutiny.
The consequence. Without a specified theory of change, practitioners cannot evaluate whether the paper's proposed interventions are likely to produce the desired outcomes, or whether alternative investments of effort would yield higher returns. Consider several plausible but competing theories:
- Theory A (the paper's implicit theory): Better shared vocabulary → better mutual understanding → better legal scholarship and technical design → better law and better systems → better societal outcomes.
- Theory B (incentives-first): Legal outcomes are primarily determined by economic and political power. Better communication is irrelevant if the actors with power have no incentive to act on improved understanding. Investment should go to changing incentives (regulation, liability, collective action) rather than to glossaries.
- Theory C (litigation-first): The law of generative AI will be determined primarily through adversarial litigation, where each side deploys experts who translate technical concepts into legal arguments. Glossaries and taxonomies may help judges and juries understand these arguments, but the primary bottleneck is the slow pace of litigation and the inconsistency of judicial decisions, not the lack of shared vocabulary among academics.
- Theory D (standards-first): What matters most is the development of concrete technical standards and best practices that can be referenced in regulation and contracts. The paper's conceptual infrastructure is a precursor to standards-setting but is insufficient without transitioning to formal standards bodies and compliance frameworks.
The paper does not distinguish among these theories, argue for one over others, or identify the conditions under which its proposed interventions would be necessary versus merely helpful. A funder deciding whether to support more GenLaw workshops or instead invest in standards development, litigation support, or regulatory advocacy receives no guidance from the paper about relative priorities.
What evidence exists in the paper. The paper provides anecdotal evidence that communication failures occurred at the workshop and that participants found resolving them valuable. It does not provide evidence that resolving such failures in a workshop setting translates to improved outcomes in legal scholarship, judicial decisions, technical design, or regulatory policy. No comparison is made to other approaches (standards-setting, litigation, regulation) and no metric is proposed for evaluating whether GenLaw's activities are succeeding. The paper's commitment to ongoing engagement ("we are growing GenLaw into a nonprofit home for research, education, and interdisciplinary discussion," Section 7) implicitly assumes that sustained interdisciplinary conversation produces better outcomes, but this assumption is not defended.
Mitigation status. Not addressed. The paper treats the value of its proposed interventions as self-evident—if communication is currently poor, improving it must be good. This is not an unreasonable starting point, but it is a limitation for readers who need to prioritize among competing demands on their attention, funding, or institutional energy. The paper would be stronger if it acknowledged that its theory of change is an assumption, identified the conditions under which better communication is most likely to translate into better outcomes, and proposed metrics for evaluating whether this translation is occurring. The fact that the paper's own glossary is "offered as a starting point, not a finish line" (Section 3.1) acknowledges that the infrastructure is provisional, but not that the entire approach of building conceptual infrastructure as a primary intervention is itself open to question.
7. Implications and Future Directions
How This Work Changes the Landscape
This paper is not a technical contribution that shifts the state of the art on a benchmark—it does not propose a new model architecture, training algorithm, or evaluation metric for generative AI. Rather, it is a meta-level intervention in how the ML and legal communities relate to each other. Its primary effect, if successful, is to change the conditions under which subsequent technical and legal contributions are produced. This is a different kind of significance than a typical research paper, and evaluating its landscape-shifting potential requires assessing whether it successfully identifies and begins to remediate the bottlenecks that prevent progress in the emerging field of generative AI and law.
The paper's central diagnostic contribution: naming the coordination failure. The paper establishes that the most immediate barrier to rigorous legal analysis of generative AI is not the absence of good legal arguments or good technical explanations, but the absence of shared vocabulary and mental models that would allow those arguments and explanations to connect. The specific, named examples of this failure—"pre-training" mischaracterized as data preparation rather than training, "harms" treated as synonymous with any bad outcome rather than as a structured legal concept determining standing and remedies, "memorization" conflated across technical metrics, colloquial understanding, and legal standards—are the paper's most durable contribution. They give researchers in both communities a checklist of specific misunderstandings to guard against, rather than a vague admonition to "communicate better."
This diagnostic reframes the problem from one that each discipline might naturally attribute to the other ("lawyers don't understand the technology" / "technologists don't understand the law") to a symmetrical, structural account: both communities use terms of art that carry different meanings in different contexts, and neither community systematically recognizes when it is exporting a technical definition into a context where a different definition applies. The paper's identification of "unknown unknowns" in terminology—terms where one community doesn't realize a loosely defined word is a term of art in the other—is particularly valuable because it explains why well-intentioned collaboration can derail without participants understanding the source of the confusion.
The paper reconceptualizes "understanding generative AI" as a legal prerequisite. A recurring question from legal scholars at the workshop was, essentially, "what's so special about generative AI?" (Section 4). The paper's answer—organized around flexibility (the shift from task-specific discriminative models to open-ended generative ones), multi-stage training pipelines and supply chains, and massive scale—provides legal scholars with a theory of the machine that goes beyond "it's a big neural network trained on web data." This theory is not novel to ML researchers, but its structured articulation for a legal audience is a contribution because it identifies which technical features are legally salient and why. For instance, the observation that the pre-training/fine-tuning distinction is "predominantly an artifact of choices made regarding training, rather than an essential aspect of the training process" (Section 4.2) is a specific, actionable corrective that prevents legal analysis from treating pre-training as legally inert data preparation. The observation that generative AI inverts the idea-expression dichotomy (Section 5.4) is not just a description—it is a diagnosis of why existing copyright categories break, which can guide legal scholars toward more productive reformulations of the authorship question.
The paper resolves a nascent contradiction about whether existing law is adequate. Prior to this workshop, one could find two competing positions in the discourse: (1) generative AI is a new technology requiring fundamentally new legal frameworks, and (2) generative AI raises issues that existing law can handle with appropriate adaptation, just as prior technologies have been accommodated. The paper does not resolve this tension in favor of either position—but it clarifies the terms of the debate. By mapping generative AI's distinctive technical features (Section 4) onto specific legal doctrines (Section 5), the paper shows that the answer is doctrine-dependent: existing copyright's idea-expression dichotomy may need fundamental rethinking (it "seems to turn the idea-expression dichotomy 'upside down,'" Section 5.4), while existing privacy law's contextual, norm-based approach may already have the flexibility to address some generative-AI privacy issues without radical doctrinal change (Section 5.2). The paper's contribution is not a verdict on the "new law vs. old law" question but a framework for asking it precisely: which specific features of generative AI interact with which specific legal doctrines in which specific ways? This disaggregation prevents the debate from collapsing into unproductive generalities.
The paper redirects attention from copyright-centrism to a broader legal landscape. The report's observation that copyright lawsuits represent "just the scratch the surface of potential issues" (Section 7)—and its structured walk-through of intent, privacy, misinformation/disinformation, and multiple dimensions of IP beyond copyright—is a corrective to the public conversation, which at the time of the workshop (and still, at the time of this writing) is dominated by high-profile copyright litigation. This is not an original doctrinal contribution—legal scholars have been writing about AI and privacy, AI and torts, and AI and free speech for years—but it is a refocusing contribution for the cross-disciplinary community the workshop convened. By providing a common taxonomy (Section 5) and a research agenda (Section 6), the paper makes it harder for the interdisciplinary conversation to be captured entirely by the copyright question, and easier for researchers interested in privacy, competition law, or standards-setting to find their footing in the emerging field.
The paper establishes the supply chain as the unit of analysis. The supply-chain perspective (drawing on Lee et al., 2023) is not original to this paper, but the paper elevates it to a central organizing principle for the entire interdisciplinary research agenda. This has two consequences. First, it makes it impossible to responsibly analyze generative AI's legal implications by treating "the model" or "the company" as a black box—analysis must specify which actor at which stage made which decision with which legal consequences. Second, it reveals that many legal questions are under-specified: "does training on copyrighted data constitute infringement?" is not a well-formed question until you specify whose training, at which stage, using which data, for which downstream deployment. The supply-chain framework provides the structure needed to formulate these questions precisely, even when it does not provide the answers.
Which directions become more attractive, and which less so.
The paper makes less attractive the pursuit of technical "solutions" that purport to resolve legal problems without engaging legal scholars about what the law actually requires. Machine unlearning (Section 6.3) serves as the paper's canonical example: a technical subfield that developed around "the Right to be Forgotten" without clarifying what "removal" means in either a technical or legal sense. After this paper, a machine-learning researcher proposing an unlearning algorithm as a GDPR compliance mechanism should expect to be asked: "under what legal definition of removal does this satisfy the obligation, and how do you know?" The paper does not kill unlearning research—it identifies it as an area of "intense (and growing) investment" (Section 6.3)—but it redirects it toward engagement with legal definitions rather than purely technical optimization.
The paper makes more attractive several specific research directions that had not previously been framed as cross-disciplinary priorities. These are discussed in detail in the next subsection, but they include: developing evaluation metrics that operationalize legal concepts (rather than purely technical ones), studying the relationship between supply-chain transparency and downstream legal liability, investigating whether the pre-training/fine-tuning distinction should matter for legal analysis (or whether courts should treat all training stages uniformly), and building "difficulty estimators" for legal compliance—tools that would help a practitioner assess, given a proposed use of generative AI, which legal domains are likely to be implicated and what the key analytic questions are within each.
Follow-Up Research This Work Enables
Empirically testing whether the glossary reduces cross-disciplinary misunderstanding. The paper identifies a problem (terminological confusion) and proposes a solution (a glossary of terms of art with both ML and legal definitions). But it provides no evidence that the solution works—no before-and-after comparison of legal scholars' understanding of ML concepts or ML researchers' understanding of legal concepts. A strong follow-up study would recruit a cohort of legal scholars and ML researchers, administer a baseline assessment of their understanding of the paper's identified "trap" terms (pre-training, harms, memorization, privacy, algorithm, hallucination), randomly assign half to read the glossary (Appendix A) and half to a control condition, and re-assess understanding. The study should also measure overconfidence—the paper's distinction between "known unknowns" and "unknown unknowns" in terminology implies that the more damaging confusions are those where participants do not realize they are confused. A design that asks participants not only to define terms but to rate their confidence in those definitions could detect whether the glossary reduces overconfidence, which is arguably more important than increasing knowledge. The GenLaw organizers are positioned to run this study with participants at their next workshop, or with law and CS students at their respective institutions.
Measuring whether the pre-training/fine-tuning distinction has legal consequences in practice. The paper's key corrective—that "pre-training is training, and not a data preparation stage" (Section 4.2)—is presented as an important clarification for legal analysis. But it remains an open empirical question whether this distinction actually affects legal outcomes. A follow-up study could analyze the pleadings, motions, and judicial opinions in the current wave of generative-AI copyright lawsuits (e.g., Authors Guild v. OpenAI, Getty Images v. Stability AI, New York Times v. Microsoft/OpenAI) and code whether and how the parties characterize pre-training—as training, as data preparation, as an automated process, as a set of human design choices, or in other terms. The study would then track whether these characterizations appear to influence judicial reasoning about fair use, direct infringement, or other doctrinal questions. If courts systematically characterize pre-training as something other than training—and if this characterization correlates with findings of non-infringement—that would validate the paper's concern about the legal significance of this terminological distinction and suggest a need for corrective expert testimony or amicus briefing.
Developing a "legal difficulty estimator" for generative-AI supply chains. The paper's supply-chain framework (Sections 4.2–4.3) demonstrates that the legal exposure of an actor in the generative-AI ecosystem depends on which decisions they make at which stage, and that different actors face different legal profiles. This suggests the possibility of a practical tool: given a description of an actor's position in the supply chain, what data they use, what models they deploy, and how they deploy them, which legal domains are relevant, and what are the key analytic questions within each? A strong follow-up would be a decision tree or questionnaire—not an automated legal advisor (which would raise unauthorized-practice-of-law concerns), but a structured map that helps a startup founder, a product manager, or an academic researcher identify which legal experts they need to consult and what information those experts will need. The paper already provides the components for this tool (the business-model patterns in Section 3.3, the supply-chain stages in Section 4.3, and the legal taxonomy in Section 5). A follow-up would operationalize them into a structured protocol, then test it with actual practitioners to see whether it surfaces legal issues they had not previously considered.
Stress-testing the generativity-jackpot claim with a systematic comparison to prior technologies. The paper argues that generative AI "hits the generativity jackpot" by scoring highly on all five of Zittrain's dimensions (Section 2), and uses this to motivate the claim that its legal challenges will be "comparable in scope, scale, and complexity to those raised by computers and the Internet." This is a plausible but untested claim. A follow-up study could systematically code the legal histories of prior generative technologies—personal computers, the World Wide Web, smartphones, social media—along Zittrain's five dimensions, and compare the trajectory of legal development for each technology (how many years did it take for major doctrinal questions to stabilize? which areas of law were implicated first, and which later? did the technology's generativity profile predict the complexity of its legal challenges?). This study would not "test" the paper's claim in a falsificationist sense—the claim is more of a calibrated prediction than a hypothesis—but it would refine it by identifying where the Internet-law analogy holds and where generative AI may differ in legally relevant ways. This matters for practitioners and funders deciding whether to model their engagement strategies on the Internet-law playbook or to develop new approaches.
Creating a taxonomy of evaluation metrics that maps legal standards to technical measurements, with explicit gaps. The paper identifies the absence of agreed-upon evaluation metrics for legal concepts as a major research need (Section 6.4) and observes that "in other areas of machine learning, there are accepted (though imperfect) notions of 'ground truth' labels, which are absent in Generative AI." A concrete follow-up would inventory existing evaluation metrics used in the generative-AI literature (memorization metrics, extraction metrics, hallucination metrics, fairness metrics, toxicity metrics), and for each, analyze (1) what the metric actually measures in technical terms, (2) what legal standard it has been proposed (or could plausibly be proposed) to operationalize, and (3) where the gap is—does the metric over-claim relative to what it measures? Does it capture only part of the legal concept? Does it measure something orthogonal? The output would be a structured reference that tells legal scholars "if you see someone claim their model satisfies legal standard X because of metric Y, here are the three questions you should ask" and tells ML researchers "if you want your metric to be legally relevant, here is what the legal standard actually requires." The paper's glossary of machine-learning terms (Appendix A.1) and legal terms (Appendices A.3–A.4) provides the starting vocabulary; a follow-up would connect the vocabularies at the level of specific metric-standard pairs.
Studying whether downstream users of generative AI can realistically assess their legal exposure given current supply-chain opacity. The paper notes that "some actors may not have the relevant information to perform necessary evaluations; some actors may not even know if a particular model is aligned or not" (Section 6.4), but it treats this as an observation rather than an empirical question. A follow-up study could attempt to answer: for a representative sample of generative-AI models available through common deployment channels (commercial APIs like OpenAI's, open-weight models from Hugging Face, fine-tuning services), what information about the supply chain is actually disclosed? What training data was used? What curation decisions were made? What safety evaluations were performed? The study would then assess: can a downstream deployer, using only publicly available information, determine whether using this model for a given application raises risks under specific legal frameworks (e.g., copyright, privacy, defamation, trade secrecy)? The null hypothesis is that for most models and most legal frameworks, publicly available information is insufficient to assess legal exposure—and if this is the case, the paper's supply-chain framework implies that many downstream deployers are operating in a state of unavoidable legal uncertainty, which would motivate transparency requirements or safe-harbor provisions that the current legal landscape does not provide.
Practical Applications and Downstream Use Cases
Curriculum development for cross-disciplinary education in generative AI and law. The paper provides the raw materials—a glossary (Appendix A), a characterization of the technology (Section 4), a legal taxonomy (Section 5), and a set of metaphors with explicitly marked limitations (Section 3.2 and Appendix B)—that can be directly assembled into a syllabus for a law-school seminar on generative AI, a computer-science course on responsible AI, or an executive education program for policymakers. The paper's explicit disclaimer that the glossary is "offered as a starting point, not a finish line" (Section 3.1) makes it suitable as an evolving course text. The identification of specific terminological traps (pre-training, harms, memorization) could form the basis for the first week's reading, with subsequent weeks organized around the legal taxonomy (one week on IP, one on privacy, one on torts/intent, one on misinformation/speech). The benefit for educators is that the paper was written with exactly this audience in mind: "scholars and practitioners who are already interested in engaging with issues at the intersection of Generative AI and law" (Section 1). It assumes baseline knowledge but not deep expertise in either discipline, making it suitable for a mixed-classroom setting where law students need technical grounding and CS students need legal grounding. A course built around this paper would not produce experts in either field, but it would—per the paper's own goals—produce graduates who can recognize when they need to consult an expert from the other discipline and who share enough vocabulary to make that consultation productive.
Due-diligence framework for companies integrating generative-AI APIs or fine-tuning open models. The paper's supply-chain perspective (Sections 4.2–4.3) and business-model analysis (Section 3.3) can be adapted into a due-diligence questionnaire for legal counsel at companies that are integrating generative-AI functionality into their products. Rather than asking the generic question "are we exposed to copyright liability?", the framework guides counsel to ask: at which stage(s) of the supply chain are we operating? Are we an API customer (B2B integration, Section 3.3, pattern 2), in which case our exposure depends significantly on the API provider's terms of service, indemnification provisions, and data-handling practices? Are we fine-tuning an open model (pattern 3), in which case we need to assess the provenance of the training data, the licenses under which the base model checkpoints were released, and whether fine-tuning on proprietary data creates trade-secrecy risk through memorization (Section 5.4)? Are we hosting a model ourselves (pattern 4), in which case notice-and-takedown obligations and machine unlearning limitations (Section 6.3) become directly relevant? The paper does not provide legal answers—it is not a substitute for legal advice—but it structures the questions in a way that makes them specific, actionable, and connected to the technical reality of how generative-AI systems are built and deployed. A legal team using this framework would be less likely to miss issues (like trade-secrecy risk from fine-tuning) that a generic "are we exposed to IP liability?" question would not surface.
Prioritization framework for research funders and policy organizations. The paper's research agenda (Section 6) identifies four broad areas, but the structured walk-through of each area provides enough specificity to guide funding decisions. A foundation or government agency deciding how to allocate a generative-AI-and-law research budget could use the paper's taxonomy to ensure their portfolio covers each of the major areas: centralization/decentralization (Section 6.1), standards and reasonableness (Section 6.2), machine unlearning and alternative compliance mechanisms (Section 6.3), and evaluation metrics for legal concepts (Section 6.4). The paper's identification of specific gaps—for example, that "both machine unlearning and attribution are very young fields, and their strategies are (for the most part) not yet computationally feasible to implement in practice for deployed generative-AI systems" (Section 6.3)—provides concrete justification for why funding basic research in these areas is necessary, rather than assuming that off-the-shelf solutions exist. The paper's observation that "legal scholars urgently need to study—and technical scholars urgently need to explain—which generative-AI safety and security measures are recognized as efficient and effective" (Section 6.2) could motivate a joint RFP from a law-and-technology funder and a computer-science funder for collaborative projects that pair a legal scholar and an ML researcher to study a specific safety measure.
A starting reference for judicial education. The paper is written at a level accessible to a legally trained reader who has "familiarity with terms like 'large language model'" (Section 1), which describes the audience for judicial education programs on AI. The glossary (Appendix A) and the technical characterization (Section 4) could be excerpted as a primer for judges who are encountering generative-AI issues in their courtrooms—whether in copyright cases, criminal cases involving AI-generated evidence, or regulatory challenges—and need a compact, accurate overview of how these systems work and what legal issues they raise. The paper's specific identification of metaphors that can mislead (anthropomorphism, memorization, collages, stochastic parrots, noisy search engines, Appendix B) is particularly valuable for judges, who are likely to encounter these metaphors in expert testimony and briefing and need to recognize when a metaphor is doing argumentative work rather than accurately describing a technical reality. The paper's discussion of machine unlearning as an inadequate analogy for notice-and-takedown (Section 6.3) could help a judge evaluate whether a defendant's claim to have "removed" copyrighted material from a model is technically meaningful or merely rhetorical. The paper does not provide the answer—it provides the framework for asking better questions of the experts who will testify about these issues.
When to Prefer This Paper's Approach
The paper does not propose a method that competes against named alternatives in the sense of a technical paper that positions its algorithm against baseline approaches. It is a diagnostic and agenda-setting document, not a tool with an alternative you could choose instead. However, the paper does articulate a distinctive approach to cross-disciplinary engagement that differs from other ways the ML and legal communities could interact, and some of the paper's own observations imply boundary conditions for when that approach is most and least appropriate. I draw these out below, grounding each in specific claims from the paper.
Prefer the "build shared infrastructure first" approach (this paper's implicit stance) when:
-
The participants in a cross-disciplinary collaboration do not yet realize they define key terms differently. The paper's distinction between "known unknowns" and "unknown unknowns" in terminology (Section 3) is its strongest argument for investing in shared vocabulary before attempting substantive analysis—because if the parties do not even know they are confused, the analysis will be built on unrecognized errors. A collaboration between an ML researcher and a legal scholar who have never worked together before is in precisely this state: each has terms of art the other may not recognize as technical, and neither has a checklist of which terms to clarify. In this situation, the paper's recommendation to start with glossary-building and metaphor-deconstruction is not optional—it prevents the collaboration from being derailed by definitional mismatches that surface only late in the process when they are costly to correct.
-
The goal is to build institutional capacity for sustained engagement, not to answer a one-off question. The paper's commitment to an evolving glossary, ongoing workshops, and iterative updates (Section 7: "the resources that we develop... will need to be frequently updated to keep pace with technological changes") assumes that the generative-AI-and-law interaction will be a decades-long engagement, comparable to the development of Internet law. If an organization is setting up a standing cross-disciplinary lab, center, or working group, the paper's infrastructure-heavy approach is appropriate: invest once in building shared vocabulary, taxonomies, and research agendas, and amortize that investment over many subsequent projects. The GenLaw organization itself is an example—it is "growing... into a nonprofit home for research, education, and interdisciplinary discussion" (Section 7) precisely because the problems it addresses are not one-offs.
-
The legal issues are genuinely novel, such that existing legal frameworks require adaptation rather than straightforward application. The paper identifies specific areas where existing law breaks—the idea-expression dichotomy turned "upside down" (Section 5.4), intent requirements in torts and crimes when the "speaker" is an AI system (Section 5.1), machine unlearning as inadequate for notice-and-takedown (Section 6.3). For these novel issues, the paper's approach—first clarify what the technology does, then identify where legal categories fail, then articulate a research agenda—is more productive than jumping to doctrinal arguments or technical fixes. In contrast, for legal issues where existing frameworks clearly apply (e.g., a generative-AI company that makes false claims in advertising faces the same FTC enforcement as any other company), the infrastructure-heavy approach may be overkill.
Prefer direct doctrinal analysis or technical development over the "build shared infrastructure first" approach when:
-
The legal question is urgent and answerable under existing frameworks, and the participants already share sufficient vocabulary. If a company receives a copyright takedown notice regarding a model's training data and needs to respond within the statutory timeframe, it cannot wait for the glossary to be updated and the research agenda to mature—it needs legal advice now, under existing law, even if that law is imperfectly adapted to generative AI. The paper's own observation that "both machine unlearning and attribution are very young fields, and their strategies are (for the most part) not yet computationally feasible to implement in practice" (Section 6.3) implies that for near-term compliance, alternatives to unlearning—contractual indemnification, output filtering, data provenance documentation, insurance—must be pursued despite being less satisfying from a conceptual standpoint.
-
The primary barrier is political or economic power, not intellectual confusion. The paper's supply-chain analysis (Sections 4.2–4.3) reveals that a handful of actors control the pre-training stage while many actors participate in fine-tuning and deployment. If the legal outcomes are being determined primarily by the economic and lobbying power of those concentrated actors, then no amount of shared vocabulary or improved mutual understanding will change the trajectory—the intervention needs to be political (regulation, antitrust enforcement, collective action) rather than intellectual. The paper itself hints at this in its discussion of centralization versus decentralization (Section 6.1): "competition and antitrust law are likely to play a major role going forward," and "every important potential bottleneck in Generative AI—from copyright ownership to datasets to compute to models and beyond—will be the focus of close scrutiny." A glossary will not substitute for antitrust enforcement; the paper's approach is necessary but not sufficient when structural power dynamics dominate.
-
The collaboration involves parties with adversarial interests where shared vocabulary may be weaponized. The paper's approach assumes good-faith collaboration: participants who want to understand each other and reach shared conclusions. But much of the generative-AI-and-law interaction occurs in adversarial contexts—litigation, regulatory proceedings, contract negotiations—where each side has incentives to define terms strategically. In a copyright lawsuit, the plaintiff and defendant may each prefer different definitions of "memorization" or "transformative use" because those definitions favor their legal positions. The paper's glossary may be useful for a judge trying to understand the technical claims, but it will not resolve the underlying adversarial conflict. In these contexts, the paper's approach is helpful for framing the disagreement (making explicit where definitions diverge) but cannot substitute for the adversarial process of argument, evidence, and judicial determination.
The key tradeoff the paper implies but does not make explicit: Building shared infrastructure (glossaries, taxonomies, research agendas) has a high fixed cost and a long time horizon to payoff. It is an investment in the capacity for future work rather than a contribution to any specific legal or technical question. This investment is most justified when (1) the problems are complex and novel enough that ad hoc approaches will produce errors, (2) the participants expect to work together repeatedly over time, and (3) the participants are operating in good faith with aligned interests in mutual understanding. It is least justified when (1) the legal questions are straightforward under existing frameworks, (2) the interaction is a one-off, or (3) the interaction is adversarial and definitional precision, while useful for clarity, will not resolve the underlying conflict. The paper's contribution is to argue, persuasively, that much of the generative-AI-and-law landscape falls into the first category—and to provide the initial infrastructure for the sustained engagement that landscape requires.