ArXiv: 2311.02462
🎯 Pitch
The paper reveals that today's most advanced AI, like ChatGPT, belongs in a newly defined category called 'Emerging AGI'—systems equal to or slightly better than an unskilled human at a range of tasks—demolishing the all-or-nothing myth and giving us a meaningful yardstick for progress. This framework maps systems on a matrix of performance and generality, showing we are already on the AGI spectrum and providing a vital tool to decouple raw capability from the autonomy decisions that introduce true risk.
1. Executive Summary
This position paper proposes a Levels of AGI ontology — a two-dimensional classification framework that jointly rates AI systems by performance (depth, from "Emerging" to "Superhuman") and generality (breadth, either "Narrow" or "General") — to replace ambiguous, single-endpoint definitions of artificial general intelligence with a graduated, operationalizable taxonomy. Analyzing nine historical definitions of AGI to distill six design principles (focus on capabilities over processes, joint emphasis on performance and generality, inclusion of metacognitive tasks such as learning new skills or knowing when to ask for help, evaluation of potential rather than deployment, ecological validity in benchmarking, and a leveled path rather than a binary threshold), the paper situates contemporary systems within this matrix — classifying frontier LLMs (ChatGPT, Bard, Llama 2, Gemini) as Level 1 "Emerging AGI" while placing systems like AlphaFold at Level 5 "Superhuman Narrow AI." The framework further decouples capability from deployment by introducing six Levels of Autonomy (from "AI as a Tool" through "AI as an Agent") that are unlocked but not dictated by advancing AGI levels, establishing that risk assessments must consider the interplay of performance, generality, and autonomy rather than any single dimension in isolation.
2. Context and Motivation
The Core Problem: We Have No Shared Language for AGI
The fundamental problem this paper addresses is not about how to build AGI, but about how to talk about AGI. The term "Artificial General Intelligence" is simultaneously one of the most consequential concepts in AI — shaping research goals, corporate charters, regulatory discussions, and risk assessments — and one of the most poorly defined. As the authors observe in the Introduction, AGI has evolved from "a subject of philosophical debate, to one which also has near-term practical relevance," driven by the rapid advancement of large language models. Yet despite its centrality to the field's ambitions and anxieties, there is no widely shared, operationalizable definition of what AGI means.
This is not an academic quibble. The ambiguity creates concrete, cascading problems across multiple domains:
For research and development: The absence of agreed-upon levels means different teams and organizations work toward different — and often implicit — targets. A company claiming to pursue AGI might be aiming for "human-level performance on cognitive tasks" (the Legg/Goertzel formulation), while another means "outperforming humans at most economically valuable work" (the OpenAI Charter). These are qualitatively different goals with different roadmaps, yet the same term obscures the distinction. The paper notes in Section 1 that AGI maps onto "goals for, predictions about, and risks of AI." Without a shared taxonomy, progress is unmeasurable — you cannot track advancement toward a destination, assess whether you are on course, or compare competing approaches if you cannot agree on what the destination looks like.
For risk assessment and safety: The paper argues that treating AGI as a binary threshold — a single point at which a switch flips and an AI system becomes "generally intelligent" — is actively dangerous for governance. If AGI is a single endpoint, then all risks (economic disruption, misuse, misalignment, existential threats) appear to emerge simultaneously and catastrophically. This framing makes it difficult to identify and prioritize incremental risks that appear at intermediate capability levels, and it complicates the design of graduated safety measures. The paper is explicit about this in Section 6.1: "A leveled approach to defining AGI enables a more nuanced discussion of how different combinations of performance and generality relate to different types of AI risk."
For policy and regulation: Policymakers need clear, measurable thresholds to write meaningful regulation, and they need a framework for understanding when new rules should take effect. The paper's parallel to autonomous vehicle standards — the SAE Levels of Driving Automation (SAE International, 2021) — is instructive. Before those levels were standardized, discussions about self-driving cars and the policies governing them were muddied by inconsistent terminology and vague capability claims. The SAE framework provided a common reference point that enabled industry, regulators, and the public to have coherent conversations about graduated deployment and graduated risk. The authors argue that AI governance needs the same scaffolding.
For communication across disciplines: The AGI conversation now involves not just AI researchers but economists (labor displacement), political scientists (geopolitical implications), ethicists (value alignment), and military strategists (Kissinger et al., 2022). Each discipline brings its own assumptions about what "general intelligence" entails. The paper's goal is to provide a lingua franca that bridges these communities — a shared ontology that is precise enough for technical benchmarking yet accessible enough for policy discussions.
The Conflict: Proliferating Definitions Without a Reconciliation Framework
The paper's analysis of nine historical definitions (Section 2) reveals a landscape of partially overlapping but mutually inconsistent formulations. This is not presented as a literature review in the traditional sense — the nine case studies serve as evidence of the problem rather than as prior work the paper aims to extend. Each definition captures something important about AGI, but each also fails along one or more dimensions that the paper's six principles are designed to address.
The tensions across these definitions are instructive:
Capabilities versus processes. The Turing Test (Case Study 1) operationalizes intelligence through a behavioral proxy — can a machine convince a human interlocutor that it is human? But as the paper notes, this "often highlights the ease of fooling people rather than the 'intelligence' of the machine," citing ELIZA (Weizenbaum, 1966) and the Eugene Goostman chatbot (Wikipedia, 2023a) as examples of systems that passed or nearly passed the test without genuine linguistic understanding. Searle's "Strong AI" (Case Study 2) goes in the opposite direction entirely, requiring not just behavioral indistinguishability but actual mental states — consciousness, understanding, intentionality. Gubrud's original 1997 definition (Case Study 3) similarly focuses on process ("rival or surpass the human brain in complexity") alongside capabilities. The paper's first principle — "Focus on Capabilities, not Processes" — cuts through this tension by arguing that what matters for practical purposes is what a system can do, not how it does it, and that process-based criteria like consciousness or brain-like complexity are definitional dead ends given our lack of scientific consensus on how to measure them.
Performance without generality, generality without performance. Some definitions emphasize breadth over depth. Agüera y Arcas and Norvig (Case Study 9) argue that mid-2023 LLMs already are AGIs because "generality is the key property of AGI" and these models demonstrate sufficient generality through their command of language, multimodal processing, and few-shot learning. The paper respectfully disagrees, arguing that generality must be paired with performance: "if an LLM can write code or perform math, but is not reliably correct, then its generality is not yet sufficiently performant." At the opposite pole, superhuman performance on a single task — AlphaFold (Jumper et al., 2021) — demonstrates astonishing capability depth but no breadth. The paper's two-dimensional matrix (Table 1) is designed precisely to capture both axes simultaneously, allowing us to say that an LLM is "Emerging in performance but General in breadth" while AlphaFold is "Superhuman in performance but Narrow in breadth."
What tasks? Which people? The Legg/Goertzel formulation (Case Study 4) — AGI as a system that can "do the cognitive tasks that people can typically do" — leaves two crucial dimensions unspecified: the set of tasks and the reference population. Does "people" mean all humans, or only skilled adults? Does it include learning new tasks, or only performing ones specified in advance? The Shanahan definition (Case Study 5) partially addresses this by explicitly requiring metacognitive ability ("can learn to perform as broad a range of tasks as a human"), but still leaves the task set and reference population unspecified. OpenAI's Charter (Case Study 6) specifies the population implicitly through the labor market ("economically valuable work"), but as the paper observes, this "does not capture all of the criteria that may be part of 'general intelligence,'" such as artistic creativity or emotional intelligence that lack clear economic value.
Deployment versus potential. OpenAI's definition requires not just capability but deployment — a system must actually perform economically valuable work to qualify. This conflates technical achievement with market adoption, and the paper identifies this as a critical weakness: "We may develop systems that are technically capable of performing economically important tasks but don't realize that economic value for varied reasons (legal, ethical, social, etc.)." The paper's fourth principle — "Focus on Potential, not Deployment" — is designed to decouple capability assessment from the non-technical factors that govern real-world deployment.
Single endpoints versus graduated paths. With the exception of Suleyman's "Artificial Capable Intelligence" proposal (Case Study 8), which the paper treats as an intermediate concept, nearly all existing definitions treat AGI as a binary threshold — you either have it or you don't. The paper argues this is analytically impoverished. Drawing an explicit parallel to the SAE autonomous driving levels, the authors contend that "there is value in defining 'Levels of AGI'" (Principle 6) that can support "clear discussions of policy and progress."
Where Existing Approaches Fall Short
The paper does not frame its contribution as a response to specific technical shortcomings of prior classification schemes — there is no prior classification scheme to critique in the way that a technical paper might identify flaws in an existing algorithm. Instead, the gap the paper identifies is the absence of any systematic framework at all. The nine case studies are not presented as competing systems to be evaluated; they are presented as evidence that the community has been operating without a shared conceptual vocabulary.
This is a fundamentally different kind of gap than what motivates most ML papers. The problem is not that existing AGI definitions are wrong but that they are incomplete and uncoordinated. Each captures a piece of the puzzle:
- The Turing Test captures the intuition that AGI should manifest behaviorally, but fails on reliability, comprehensiveness, and the confound of human gullibility.
- "Strong AI" captures the intuition that AGI might involve genuine understanding, but relies on unmeasurable properties.
- The Legg/Goertzel definition captures breadth of cognitive task performance but underspecifies the reference population and task set.
- OpenAI's Charter captures the economic stakes but conflates capability with deployment and excludes non-economic intelligence dimensions.
- Agüera y Arcas and Norvig capture the reality that generality is emerging now, but overclaim by dismissing performance as a criterion.
What is missing — and what the paper positions itself to provide — is a unifying framework that can accommodate the insights of each definition while addressing their individual weaknesses. The paper's six principles are derived from this analysis: they are the design requirements that any adequate AGI ontology must satisfy, extracted by identifying the failure modes of prior attempts.
How This Paper Positions Itself Relative to Existing Work
The paper is explicit about its genre: "We propose" and "we argue" (Section 1) and "position paper" (title and framing). It does not present experimental results, train models, or introduce new algorithms. Instead, it offers a conceptual framework — an ontology — designed to structure how the community thinks about, discusses, and measures progress toward AGI.
This positioning carries several implications for how the paper should be read:
It is normative, not descriptive. The paper is not merely observing how people currently use the term AGI; it is arguing for how they should use it. The six principles are explicitly prescriptive: "We argue that any definition of AGI should meet the following six criteria" (Section 3). The framework is offered as a proposal to the community, not as a discovery about the world.
It is structural, not empirical. The value of the paper lies in the structure it provides — the two-dimensional performance × generality matrix, the six autonomy levels, the decoupling of capability from interaction paradigm — rather than in empirical findings. This is appropriate for a position paper: the contribution is a way of organizing thought, not a measurement or a proof.
It is incomplete by design. The paper explicitly acknowledges the most significant incompleteness: it does not propose a benchmark. Section 5 is titled "Testing for AGI" and discusses the properties an AGI benchmark should possess (diverse cognitive and metacognitive tasks, ecological validity, living benchmark with evolving task sets, careful handling of dual-use capabilities), but the paper stops short of specifying tasks, thresholds, or evaluation protocols. The authors justify this: "Because of the immense complexity of this process, as well as the importance of including a wide range of perspectives (including cross-organizational and multi-disciplinary viewpoints), we do not propose a benchmark in this paper." This is both a strength (the framework can accommodate future benchmarks developed through community processes) and a limitation (the framework remains abstract until operationalized).
The paper also acknowledges the "open research question" of "what portion of benchmarking tasks at a given level demonstrate generality," leaving unspecified the precise proportion of tasks a system must pass to qualify for a given generality rating. This is not an oversight — it is a recognition that "it is impossible to enumerate the full set of tasks achievable by a sufficiently general intelligence" (Section 5) and that any threshold will necessarily involve judgment. The paper's pragmatic resolution — "Systems that pass the majority of the envisioned AGI benchmark at a particular performance level, including new tasks added by the testers, can be assumed to have the associated level of generality for practical purposes" — acknowledges that the framework is approximate but usable.
It connects to, but does not derive from, related taxonomies. The paper draws structural inspiration from the SAE Levels of Driving Automation but develops an independent scheme suited to the different nature of cognitive capabilities. The Autonomy Levels (Table 2) are compared to prior taxonomies of computer automation (Sheridan et al., 1978; Sheridan & Parasuraman, 2005; Parasuraman et al., 2000) but are distinguished by their "human-AI interaction style" framing rather than a computer-centric "how much control does the designer relinquish" perspective. The paper also connects its AGI levels to Anthropic's Responsible Scaling Policy (Anthropic, 2023b) and its AI Safety Levels (ASLs), suggesting that "including items matched to ASL capabilities in any AGI benchmark would connect points in our AGI taxonomy to specific risks and mitigations." This integration with concurrent safety frameworks reinforces the paper's positioning: the Levels of AGI is not a standalone proposal but is designed to be used in conjunction with other tools for governance and risk management.
It resolves the earlier conflict without declaring a single winner. The paper's most elegant positioning move is in Section 3's discussion of Principle 6: the leveled approach "supports the coexistence of many prominent formulations." The Agüera y Arcas and Norvig definition maps to Level 1 ("Emerging AGI"), OpenAI's to a higher level (roughly "Exceptional AGI"), and the Legg/Shanahan/Suleyman cluster maps to Level 2 ("Competent AGI"). The framework does not tell these authors they were wrong — it tells them they were operating at different points on a spectrum they had not yet articulated. This is a diplomatic as well as an analytical contribution: it transforms a confusing debate into a coherent map where each position finds its place.
In summary, the paper fills not an empirical gap but a conceptual vacuum. The problem is not that we have the wrong definition of AGI; it is that we have been using a term of enormous consequence without a shared framework for what it means, how to measure progress toward it, how to assess the risks that different levels of progress introduce, or how to design interaction paradigms appropriate to different capability thresholds. The paper's contribution is to propose that framework, derive it from an analysis of where prior attempts fall short, and provide sufficient structure that the community can begin the work of operationalizing it — building benchmarks, mapping risks, and designing governance — within a coherent, shared ontology.
3. Technical Approach
3.1 Reader Orientation
This paper develops a classification ontology — a structured system of named categories and rules for assigning AI systems to those categories — not a machine learning model or an algorithm. The problem it solves is the absence of a shared language for describing where an AI system sits on the path toward artificial general intelligence: rather than a single, binary "AGI or not" threshold that collapses all nuance, the ontology provides a two-dimensional grid with graduated performance levels and a generality axis, plus a separate scale for deployment autonomy, so that systems with very different capability profiles (like ChatGPT versus AlphaFold) can be meaningfully compared, risks can be assessed at intermediate capability levels, and progress can be tracked without waiting for a single magical threshold.
3.2 Big-Picture Architecture (Diagram in Words)
The framework has four major components:
-
The Performance × Generality Matrix (Table 1) — the core classification scheme. The rows are six performance levels (
Level 0: No AIthroughLevel 5: Superhuman), each defined by a human-percentile threshold. The columns are two generality categories (NarrowandGeneral). Every AI system gets placed in one cell of this 6×2 grid, producing labels like "Competent Narrow AI" or "Emerging AGI." -
The Six Design Principles (Section 3) — the normative criteria that were used to derive the matrix and that any adequate AGI definition must satisfy. They function as the "requirements document" for the ontology, explaining why the matrix has the structure it does.
-
The Benchmark Specification Blueprint (Section 5) — not an actual benchmark, but a set of properties that an AGI benchmark must possess (diverse cognitive and metacognitive tasks, ecological validity, living and evolving task set, appropriate handling of dual-use capabilities). This component defines how systems would be tested to determine their position in the matrix, even though the specific tests remain unspecified.
-
The Autonomy Levels (Table 2) — a separate six-level scale (
Level 0: No AIthroughLevel 5: AI as an Agent) that classifies how the system is deployed and interacts with humans, not its raw capability. This component is linked to but decoupled from the capability matrix: higher AGI levels "unlock" higher autonomy levels, but deployment at the maximum unlocked autonomy is optional and often undesirable.
Information flows as follows: an AI system's capabilities are evaluated against a (future) AGI benchmark covering diverse cognitive and metacognitive tasks → benchmark results determine the system's performance level (based on percentile relative to skilled adults) and generality category (based on the breadth of tasks meeting that performance threshold) → the performance × generality classification feeds into risk assessment, which also considers the deployed autonomy level (which may be lower than what the system is capable of) → the combined capability + autonomy assessment informs safety decisions, regulatory thresholds, and deployment choices.
3.3 Roadmap for the Deep Dive
- First, the six design principles (Section 3 of the paper), because they are the foundation: the matrix and autonomy levels are consequences of these principles, and understanding them first makes the design of the classification scheme logically transparent rather than arbitrary.
- Second, the Performance × Generality Matrix (Table 1), walking through each cell and the percentile thresholds that define it, since this is the central contribution.
- Third, how contemporary systems map onto the matrix, using the examples the paper itself provides (ChatGPT, AlphaFold, Deep Blue, etc.) to make the classification concrete.
- Fourth, Table 2's Autonomy Levels, explaining how they relate to but are not determined by the capability matrix and why this decoupling is essential for risk assessment.
- Fifth, the benchmark blueprint (Section 5), since it defines how the matrix would be operationalized — the properties a benchmark must satisfy, the open questions the paper leaves unresolved.
3.4 Detailed, Sentence-Based Technical Breakdown
This is a conceptual position paper whose core idea is that AGI should be defined through a two-dimensional, leveled taxonomy (performance depth × generality breadth) with a separate autonomy scale, rather than through any single binary threshold, and that this taxonomy should be derived from explicit design principles that correct for the failure modes of previous definitions.
The Six Design Principles as Derivation of the Framework
The paper does not present its Levels of AGI as an arbitrary invention. Section 3 analyzes nine prominent definitions of AGI (the case studies from Section 2), identifies what each captures well and what each misses, and distills six principles that an adequate ontology must satisfy. These principles function as the design constraints from which the matrix follows. Understanding them is essential because the matrix's features — its choice of two dimensions, its percentile-based performance thresholds, its inclusion of metacognitive tasks, its separation of capability from deployment — are all direct consequences of these constraints.
Principle 1: Focus on Capabilities, not Processes. This principle rules out definitions that require specific internal mechanisms. A system qualifies for a given AGI level based on what it can do — the tasks it can perform and at what level of proficiency — not on how it achieves that performance. This eliminates requirements for consciousness (Searle's "Strong AI"), human-like neural architecture (Gubrud's brain-complexity criterion), or genuine understanding (the philosophical interpretation of the Turing Test). The paper is explicit: "it is not a necessary precursor for AGI that systems possess qualities such as consciousness (subjective awareness) or sentience (the ability to have feelings), since these qualities have a process focus." This principle directly shapes the matrix: all classifications are based on task performance percentiles, which are observable and measurable, rather than on internal states, which are not.
A subtle implication: this principle allows process-based research (like mechanistic interpretability) to eventually inform capability assessment if it yields measurable correlates, but it does not require it. The paper notes in a footnote that "As research into mechanistic interpretability advances, it may enable process-oriented metrics. These may be relevant to future definitions of AGI." The principle is thus forward-compatible with improved measurement techniques while remaining grounded in what is currently measurable.
Principle 2: Focus on Generality and Performance. The analysis of existing definitions revealed a split: some emphasized breadth (Agüera y Arcas and Norvig argued that LLMs are AGIs because of their generality alone), while others emphasized depth without breadth (superhuman narrow systems like AlphaFold). The paper argues that both dimensions are necessary and that a definition must capture their interaction. This principle is why the matrix has two axes rather than one: performance (depth, how good the system is at a task relative to humans) and generality (breadth, the range of tasks across which that performance is demonstrated). Neither dimension alone is sufficient for classification.
The principle also implies that the relationship between the axes matters: a system can be "Narrow" with any performance level, but to be "General," it must demonstrate a given performance threshold across most cognitive tasks, not just some. The paper explicitly states that a Competent AGI "must have performance at least at the 50th percentile for skilled adult humans on most cognitive tasks, but may have Expert, Exceptional, or even Superhuman performance on a subset of tasks." This allows for uneven capability profiles — a general system is not required to be uniformly excellent at everything, only to meet the threshold broadly.
Principle 3: Focus on Cognitive and Metacognitive, but not Physical, Tasks. This principle addresses whether robotic embodiment is a prerequisite for AGI. The paper argues it is not: "the ability to perform physical tasks increases a system's generality, but should not be considered a necessary prerequisite to achieving AGI." This matters because physical capabilities lag behind cognitive ones (as of the paper's writing), and requiring embodiment would artificially delay the recognition of cognitive AGI.
The more significant part of this principle is the inclusion of metacognitive tasks as a requirement. Metacognition refers to capabilities about one's own cognitive processes, and the paper specifies three subtypes that an AGI benchmark must test: (1) the ability to learn new skills (since generality requires adapting to tasks not specified at training time), (2) the ability to know when to ask for help (since alignment and safe deployment require a system to recognize the limits of its own competence), and (3) social metacognitive abilities such as theory of mind (the ability to model what other agents know, believe, or intend). These are not optional add-ons; the paper treats them as "key prerequisites for systems to achieve generality."
This principle shapes the benchmark blueprint in Section 5, which explicitly calls for tasks measuring learning ability (citing Chollet, 2019), model calibration (knowing when to ask for help), and social cognition. The emphasis on metacognition reflects a design judgment that "generality" is not merely breadth of static task performance but includes the dynamic ability to acquire new competencies.
Principle 4: Focus on Potential, not Deployment. This principle draws a sharp line between capability (what a system could do if deployed) and deployment (what a system actually does in practice). The paper argues that classification should be based on capability alone, because deployment depends on non-technical factors — legal restrictions, ethical concerns, safety considerations, market dynamics — that are orthogonal to the system's underlying intelligence. A system that could theoretically perform economically valuable work at a superhuman level but is never deployed for safety reasons should still be classified as Superhuman AGI if benchmarked appropriately.
This principle directly motivates the separation of the AGI capability matrix from the Autonomy Levels. The matrix classifies what the system can do; the Autonomy Levels classify how humans choose to deploy it. The paper emphasizes this decoupling repeatedly: "AGI is not necessarily synonymous with autonomy" and "we may develop an AGI, but choose not to deploy it autonomously, or choose to deploy it with differentiated autonomy levels in distinct circumstances." This is not merely a taxonomic nicety — it has direct implications for risk assessment, because it means that safety measures can constrain deployment autonomy without altering the underlying capability classification.
Principle 5: Focus on Ecological Validity. A classification system is only as good as the benchmarks used to operationalize it. This principle requires that the tasks used to assess AGI levels be "real-world (i.e., ecologically valid) tasks that people value." The paper explicitly warns against relying on "traditional AI metrics that are easy to automate or quantify but may not capture the skills that people would value in an AGI," citing Raji et al. (2021) on the limitations of conventional benchmarks.
The term "ecological validity" is used here in the psychological sense: a task is ecologically valid if it corresponds to the kinds of challenges that arise in natural, real-world contexts, rather than being an artificial laboratory construct. The paper argues that complex, open-ended, interactive tasks — while difficult to benchmark — have greater ecological validity than easily scored multiple-choice or classification tasks. This principle drives the benchmark blueprint's call for "open-ended and/or interactive tasks that might require qualitative evaluation."
The paper also specifies that ecological validity should account for context-appropriate baselines. In benchmarking a self-driving car, the relevant comparison might not be "a human driving without any AI assistance" but "a human driving with the AI-assisted safety tools that are actually available" — because the real-world counterfactual includes those tools. This is a subtle but important methodological point: ecological validity is about fidelity to the actual deployment context, including the realistic baseline.
Principle 6: Focus on the Path to AGI, not a Single Endpoint. This is the principle that most directly produces the leveled structure. The paper draws an explicit analogy to the SAE Levels of Driving Automation, which "allowed for clear discussions of policy and progress relating to autonomous vehicles." A single AGI threshold would collapse all intermediate capability states — from "barely functional generalist" to "superhuman across all tasks" — into a single undifferentiated category, making it impossible to have graduated policy responses or to measure incremental progress.
This principle also resolves the tension among competing definitions. Rather than declaring one definition correct and the others wrong, the levels provide a framework where each prior definition finds a natural home at a different level. The paper explicitly maps: Agüera y Arcas and Norvig's LLMs-as-AGIs claim corresponds to Level 1 ("Emerging AGI"); the OpenAI Charter's labor-replacement criterion corresponds roughly to Level 4 ("Exceptional AGI"); the Legg/Shanahan/Suleyman cluster maps to Level 2 ("Competent AGI"). This is a diplomatic resolution — the framework absorbs rather than refutes prior work, reinterpreting disagreement as operating at different points on a shared spectrum.
The Performance × Generality Matrix (Table 1)
Table 1 is the paper's central contribution. It is a two-dimensional classification grid with six rows (performance levels, numbered 0 through 5) and two columns (generality categories: Narrow and General), yielding 12 possible classification cells. Each cell has a canonical label (e.g., "Emerging Narrow AI," "Competent AGI"), and the paper populates several cells with example systems to illustrate the schema.
The Performance Dimension (Rows). Performance is defined as the depth of a system's capabilities — how well it performs on a given task relative to a human reference population. The six levels are:
-
Level 0: No AI. The system has no meaningful AI capability. Examples: a calculator (Narrow Non-AI) or Amazon Mechanical Turk (General Non-AI — here "general" refers to the human workers, not the computing system itself, which merely coordinates). This level establishes the baseline: it captures tools and platforms that may be computationally assisted but do not exhibit autonomous intelligent behavior.
-
Level 1: Emerging. Performance is "equal to or somewhat better than an unskilled human." The reference population here is explicitly unskilled humans — people who have not received specialized training in the task. This is a lower bar than the higher performance levels, reflecting the "emerging" nature of the capability. Example systems: GOFAI (Good Old-Fashioned AI, i.e., symbolic rule-based systems like SHRDLU) for Emerging Narrow AI; ChatGPT, Bard, Llama 2, and Gemini for Emerging AGI.
-
Level 2: Competent. Performance is "at least 50th percentile of skilled adults." The reference population shifts from unskilled to skilled adults — people who possess the relevant skill. For English writing evaluation, the comparison group would be literate, fluent English-speaking adults, not the general population. Example systems: toxicity detectors (Jigsaw), smart speakers (Siri, Alexa, Google Assistant), and visual question-answering systems (PaLI) for Competent Narrow AI. The "Competent AGI" cell is explicitly "not yet achieved" — this is the level the paper identifies as the best catch-all for most historical AGI definitions.
-
Level 3: Expert. Performance is "at least 90th percentile of skilled adults." This represents the top decile of human performance within the skilled population. Example systems: grammar checkers like Grammarly and generative image models like Imagen and DALL-E 2 for Expert Narrow AI. Expert AGI is "not yet achieved."
-
Level 4: Exceptional. Performance is "at least 99th percentile of skilled adults." This represents the top 1% of human performers. Example systems: Deep Blue and AlphaGo for Exceptional Narrow AI. Exceptional AGI is "not yet achieved." (The paper notes that this level was originally called "Virtuoso AGI" in earlier versions but was renamed.)
-
Level 5: Superhuman. Performance "outperforms 100% of humans." This is defined as exceeding the performance of every human on the task, including the world's foremost experts. Example systems: AlphaFold (protein structure prediction), AlphaZero (chess/shogi/Go through self-play), and StockFish (chess engine) for Superhuman Narrow AI. The General AI cell at this level is labeled ASI (Artificial Superintelligence) and is "not yet achieved."
Important nuance about the percentile definitions. The paper specifies that "for all performance levels above 'Emerging,' percentiles are in reference to a sample of adults who possess the relevant skill." This means the reference population is task-dependent and excludes people who lack the prerequisite skill. This is a crucial methodological choice: it prevents a system from appearing competent at, say, calculus simply because it outperforms a population that has never studied calculus. The relevant comparison is to the population that can perform the task.
The Generality Dimension (Columns). The paper uses only two categories for generality, not a continuum:
-
Narrow: The system's capabilities are restricted to "a clearly scoped task or set of tasks." This is the column for systems that excel within a bounded domain. The examples (AlphaFold, Deep Blue, StockFish) all demonstrate that Narrow systems can reach the highest performance levels — Superhuman Narrow AI is already achieved for specific tasks.
-
General: The system's capabilities span "a wide range of non-physical tasks, including metacognitive tasks like learning new skills." The inclusion of metacognition in the definition of generality is not an afterthought: it is part of the definition of the General column. A system cannot achieve "General" status merely by being trained on a wide static set of tasks; it must demonstrate the ability to acquire new competencies, which is itself a metacognitive capacity.
The paper acknowledges that the boundary between Narrow and General is not sharp in all cases. Systems classified as General may have uneven performance — a Level 1 Emerging AGI might perform at Level 2 Competent or even Level 3 Expert on some tasks while remaining at Level 1 Emerging on most tasks. The classification is based on the minimum performance threshold met across most tasks, not the maximum achieved on any subset.
Why two levels of generality rather than a continuum? The paper does not explicitly defend this binary choice, but it follows from the practical difficulty of meaningfully distinguishing intermediate generality levels. "Clearly scoped" versus "wide range including metacognition" is a distinction that captures a qualitative shift — from static expertise to adaptive capability — without requiring a spurious precision about degrees of breadth. This is consistent with the paper's pragmatic orientation: the framework should support useful distinctions without over-claiming measurement precision.
Example classification exercise. The paper walks through DALL-E 2 to illustrate how the taxonomy handles real systems with uneven capability profiles. DALL-E 2 is classified as Level 3 Expert Narrow AI: its image generation quality exceeds what most people can draw (hence Expert-level performance), but its failure modes — incorrect numbers of fingers on hands, illegible text rendering — prevent it from reaching Exceptional. Crucially, the paper notes that the benchmarked performance (which might be Expert) may differ from deployed performance (which might be only Competent), because "prompting interfaces are too complex for most end-users to elicit optimal performance." This example serves three functions: it demonstrates classification of a real system, it shows how failure modes constrain performance level, and it previews the capability-vs-deployment distinction that the Autonomy Levels will formalize.
The Six Design Principles Revisited: How They Shape the Matrix
With the matrix now described, we can trace how each principle shaped specific features:
-
Principle 1 (Capabilities, not Processes): The matrix classifies by task performance percentiles, not by architecture, training method, or internal representations. AlphaFold and StockFish are both Level 5 Narrow AI despite being built on entirely different technical foundations (deep learning vs. classical search).
-
Principle 2 (Generality and Performance): The matrix has two axes rather than one. A system's classification requires specifying both its performance depth and its generality breadth. Neither axis alone is sufficient.
-
Principle 3 (Cognitive and Metacognitive, not Physical): The "General" column explicitly includes metacognitive tasks in its definition. Physical tasks (robotic manipulation, navigation) are optional — they can increase generality but are not required. This is why a disembodied language model can in principle achieve "General" classification.
-
Principle 4 (Potential, not Deployment): The matrix classifies based on benchmarked capability, not deployed performance. The DALL-E 2 example explicitly separates its theoretical Expert classification from its practical Competent deployment, without altering the matrix classification.
-
Principle 5 (Ecological Validity): This principle governs the benchmark (future work), not the matrix structure itself, but the matrix anticipates it by requiring that classification be based on ecologically valid task assessments.
-
Principle 6 (Path, not Endpoint): The six-level performance scale is the direct expression of this principle. Instead of a binary AGI/not-AGI threshold, the framework provides graduated levels that support incremental progress tracking and graduated policy responses.
Mapping Contemporary Systems onto the Matrix
The paper uses specific named systems as exemplars for populated cells. These assignments are approximate (the paper acknowledges this in Table 1's caption: "The assignment of example systems to cells is approximate. Unambiguous classification of AI systems will require a standardized benchmark of tasks") but serve to ground the abstract taxonomy in familiar reference points:
Narrow AI column (descending performance):
- Level 5 (Superhuman Narrow AI): AlphaFold, AlphaZero, StockFish. All three perform their specific tasks (protein folding, board games, chess) at a level exceeding all human experts. These are the clearest examples of achieved superhuman narrow capability.
- Level 4 (Exceptional Narrow AI): Deep Blue (chess, defeated Kasparov in 1997 but did not exceed all possible human performance) and AlphaGo (Go, before the AlphaZero generalization). The distinction from Level 5 is subtle: these systems outperform most humans and were historically significant, but may not strictly exceed 100% of humans.
- Level 3 (Expert Narrow AI): Grammarly (writing correction), Imagen and DALL-E 2 (image generation). These are at the 90th percentile of skilled human performance within their domains.
- Level 2 (Competent Narrow AI): Toxicity detectors (Jigsaw), smart speakers (Siri, Alexa, Google Assistant), visual QA systems (PaLI), Watson, and "SOTA LLMs for a subset of tasks" (short essay writing, simple coding). This is a heterogeneous category reflecting systems that achieve 50th-percentile skilled-adult performance on narrow task sets.
- Level 1 (Emerging Narrow AI): GOFAI systems and simple rule-based systems (SHRDLU). These represent early AI that could handle constrained tasks with some flexibility but fell far short of skilled human performance.
- Level 0 (Narrow Non-AI): Calculators and compilers. Deterministic tools with no learned or adaptive behavior.
General AI column (descending performance):
- Levels 2–5 (Competent through ASI): All explicitly marked "not yet achieved." This is the paper's central empirical claim about the state of the field: no system has yet achieved Competent-level performance on a general range of cognitive tasks.
- Level 1 (Emerging AGI): ChatGPT, Bard, Llama 2, Gemini. The paper justifies this classification by noting that these systems "exhibit 'Competent' performance levels for some tasks (e.g., short essay writing, simple coding), but are still at 'Emerging' performance levels for most tasks (e.g., mathematical abilities, tasks involving factuality)." The classification as Level 1 rather than Level 2 reflects the minimum performance across the task distribution — the system is only as "general" as its worst-performing task category allows.
- Level 0 (General Non-AI): Human-in-the-loop computing platforms like Amazon Mechanical Turk. Here the "general intelligence" is supplied by the human workers, not by the computing infrastructure.
Key insight about the frontier LLM classification. The paper's placement of ChatGPT et al. at Level 1 General AI — not Level 2, and not Narrow — is arguably its most consequential judgment call. It simultaneously rejects two opposing positions: the Agüera y Arcas and Norvig claim that these systems are already AGIs in the full sense (which would imply Competent-level performance, which they do not reliably demonstrate across most cognitive tasks), and the skeptic's claim that they are merely narrow tools (which would deny the breadth of their capabilities and their metacognitive features like few-shot learning). "Emerging AGI" captures the reality: the generality is real but the performance is not yet reliable for high-stakes deployment without human oversight.
The Benchmark Blueprint (Section 5)
Section 5 does not propose a concrete benchmark but instead defines the properties that an adequate AGI benchmark must possess. This is a blueprint — a specification of requirements — rather than an implementation. The paper explicitly defers benchmark creation to a future, multi-stakeholder process.
What the benchmark must measure. The paper specifies that an AGI benchmark must include "a broad suite of cognitive and metacognitive tasks," measuring:
- Linguistic intelligence
- Mathematical and logical reasoning (citing Webb et al., 2023)
- Spatial reasoning
- Interpersonal and intra-personal social intelligences
- The ability to learn new skills (citing Chollet, 2019)
- Creativity
The paper also suggests that the benchmark might draw from psychometric categories used in theories of intelligence from psychology, neuroscience, cognitive science, and education, but with an important caveat: existing psychometric tests must first be evaluated for suitability in this context, because "many may lack ecological and construct validity" when applied to computing systems (citing Serapio-García et al., 2023).
The centrality of metacognitive tasks. The paper is more specific about metacognitive requirements than about any other task category. Three types of metacognitive tasks are identified as essential:
-
Learning new skills. The ability to acquire competence on tasks not seen during training is "essential to generality, since it is infeasible for a system to be optimized for all possible use cases a priori." This requires sub-skills including strategy selection for learning (the ability to choose appropriate learning approaches for different types of tasks, citing Pressley et al., 1987).
-
Knowing when to ask for help. This is framed as both a safety requirement (supporting alignment and appropriate human-AI interaction) and a metacognitive one (requiring awareness of the limits of the model's own abilities, citing Demetriou & Kazi, 2006). It relates to model calibration — the system's ability to "proactively anticipate and retroactively evaluate how well it would do/did on certain tasks" (citing Liang et al., 2023).
-
Theory of mind. The ability to model end-users' mental states — what they know, believe, intend — is "a necessary component of alignment for AGI systems." The paper notes that theory of mind tasks are "sometimes considered metacognitive, though are sometimes classified separately as social cognition" (citing Gardner, 2011, on multiple intelligences).
Living benchmark requirement. The paper argues that it is "impossible to enumerate the full set of tasks achievable by a sufficiently general intelligence," and therefore any AGI benchmark must be a "living benchmark" — one that includes "a framework for generating and agreeing upon new tasks." This is not a minor implementation detail; it is a fundamental architectural requirement. A static benchmark would necessarily be incomplete, and systems might achieve high scores without genuine generality if they overfit to the fixed task set. The living benchmark concept means that new tasks are continually added, and a system's generality classification depends on its ability to handle unanticipated tasks as well as anticipated ones.
The open question: what proportion of tasks constitutes "generality"? The paper identifies but does not resolve a crucial threshold question: "what portion of benchmarking tasks at a given level demonstrate generality?" In other words, if a benchmark has 1,000 tasks across diverse cognitive domains, does a system need to meet the performance threshold on 80% of them? 90%? 99%? The paper's pragmatic resolution is:
"Systems that pass the majority of the envisioned AGI benchmark at a particular performance level, including new tasks added by the testers, can be assumed to have the associated level of generality for practical purposes (i.e., though in theory there could still be a test the AGI would fail, at some point unprobed failures are so specialized or atypical as to be practically irrelevant)."
The paper simultaneously acknowledges that "this will be a very high percentage" and that "it will probably not be 100%, since it seems clear that broad but imperfect generality is impactful (individual humans also lack consistent performance across all possible tasks, but are generally intelligent)." The unresolved nature of this threshold is explicitly flagged as "an open research question."
Dual-use capabilities in benchmarking. Whether an AGI benchmark should include tests for potentially dangerous capabilities (deception, persuasion, advanced biochemistry) is flagged as "controversial." The paper leans toward inclusion, arguing that "most such skills tend to be dual use (having valid applications to socially positive scenarios as well as nefarious ones)," and that dangerous capability benchmarking can be de-risked through Principle 4 (testing potential rather than deployment) by ensuring "benchmarks for any dangerous or dual-use tasks are appropriately sandboxed and not defined in terms of deployment." However, the paper also acknowledges the counterargument: including such tests in a public benchmark "may allow malicious actors to optimize for these abilities," and characterizes "understanding how to mitigate risks associated with benchmarking dual-use abilities" as an important open area for safety, ethics, and governance research.
Connection to Anthropic's Responsible Scaling Policy. The paper notes that Anthropic's RSP (Version 1.0, released concurrently) uses a levels-based approach inspired by biosafety levels to define AI Safety Levels (ASLs) with associated dangerous capabilities and containment measures. The paper suggests that "including items matched to ASL capabilities in any AGI benchmark would connect points in our AGI taxonomy to specific risks and mitigations," establishing interoperability between the capability classification framework and operational safety protocols.
The Autonomy Levels (Table 2)
Table 2 introduces a second, independent classification scheme: six Levels of Autonomy that characterize the human-AI interaction paradigm rather than the AI's raw capability. The paper is explicit that these are "correlated with" but "not determined by" the AGI levels — higher AGI levels unlock higher possible autonomy levels, but deployment can (and often should) occur at lower autonomy than the system is technically capable of.
The six autonomy levels, described from the human-AI interaction perspective:
-
Level 0: No AI. The human performs all tasks without AI assistance. Examples include sketching with pencil on paper or typing in a text editor without AI features. The paper notes that this paradigm remains important "for many contexts, including for education, enjoyment, assessment, or safety reasons," even when higher autonomy is technologically available.
-
Level 1: AI as a Tool. The human "fully controls task and uses AI to automate mundane sub-tasks." Examples include using a search engine for information-seeking, using a grammar checker to revise writing, or using a machine translation app to read a sign. The AI performs well-scoped sub-tasks at the human's direction; the human remains in full control of the overall task.
-
Level 2: AI as a Consultant. The AI "takes on a substantive role, but only when invoked by a human." The distinction from Level 1 is that the AI's contribution is substantive rather than mundane — summarizing documents, generating code, providing entertainment recommendations. The human is still the initiator and decision-maker, but the AI's role has expanded from automating sub-tasks to providing substantial intellectual contributions.
-
Level 3: AI as a Collaborator. This level involves "co-equal human-AI collaboration; interactive coordination of goals & tasks." The shift from Consultant to Collaborator is qualitative: rather than the human invoking the AI for specific sub-tasks, both agents work together interactively, coordinating and negotiating. Examples include training as a chess player through interactive analysis with a chess AI, or social entertainment through interactions with AI-generated personalities.
-
Level 4: AI as an Expert. The AI "drives interaction; human provides guidance & feedback or performs sub-tasks." This inverts the Level 1 relationship: the AI is now the primary driver of the task, and the human's role has shifted to oversight, guidance, and handling sub-tasks the AI delegates. An example is using an AI system to advance scientific discovery, where the AI proposes hypotheses, designs experiments, and interprets results, with human scientists providing direction and verification.
-
Level 5: AI as an Agent. The AI is "fully autonomous." The human is not in the loop during task execution; the AI acts as an independent agent. The example given is "autonomous AI-powered personal assistants," which the paper marks as "not yet unlocked."
The unlocking relationship between AGI Levels and Autonomy Levels. For each Autonomy Level, the paper specifies which AGI Levels make that paradigm "Possible" versus "Likely." For example, Level 2 Autonomy (AI as Consultant) is "Possible" with Competent Narrow AI but "Likely" with Expert Narrow AI or Emerging AGI. The asymmetry — higher AGI levels are needed for the paradigm to be likely than for it to be merely possible — reflects the uneven capability profiles of General systems. An Emerging AGI might be good enough at document summarization (a Level 2 capability for that specific task) to serve as a Consultant in that domain, even though its overall performance is Emerging. This is the "unlocking" logic: generality allows competence in specific sub-domains to reach levels that support particular interaction paradigms, even when the system's broad performance remains lower.
Risk assessment through the combined lens. The paper's key claim about risk is that it emerges from the interaction of capability level and autonomy level, not from either alone. The examples in Table 2 illustrate this:
- Level 1 Autonomy (Tool) with Competent Narrow AI introduces risks of "de-skilling (e.g., over-reliance)" and "disruption of established industries."
- Level 3 Autonomy (Collaborator) adds risks of "anthropomorphization (e.g., parasocial relationships)" and "rapid societal change."
- Level 5 Autonomy (Agent) introduces risks of "misalignment" and "concentration of power" — the classic AGI x-risk scenarios.
Crucially, the highest autonomy levels are associated with risks that emerge only when combined with high AGI capability. Level 5 Autonomy at Level 1 AGI capability (a fully autonomous Emerging AGI) would be incompetent and dangerous in a mundane way; Level 5 Autonomy at Level 5 AGI capability (a fully autonomous ASI) introduces qualitatively different existential risks. This interaction is what makes the combined framework more informative than either scale alone.
Comparison to prior automation taxonomies. The paper distinguishes its Autonomy Levels from earlier frameworks (Sheridan et al., 1978; Sheridan & Parasuraman, 2005; Parasuraman et al., 2000) by noting that those taxonomies "take a computer-centric perspective (framing automation in terms of how much control the designer relinquishes to computers)," whereas the paper characterizes autonomy "through the lens of the nature of human-AI interaction style." This is not merely a difference in labeling — it reflects a fundamentally different framing of the relationship. In the human-centric framing, autonomy is about how the human and AI interact and who drives which parts of the task, rather than about how much of the task the computer "takes over." The paper cites Shneiderman (2020) in support of the view that "automation is not a zero-sum game, and that high levels of automation can co-exist with high levels of human control."
The autonomy design choice is orthogonal to capability. The paper emphasizes that the choice of autonomy level for a given deployment is a design decision, not an inevitability. Higher AGI capability unlocks the possibility of higher autonomy, but humans — designers, deployers, regulators — choose whether to exercise that possibility. The self-driving car example makes this concrete: even when Level 5 Self-Driving technology is available, there remain valid reasons to use Level 0 (No Automation) for driver education, enjoyment by driving enthusiasts, driver's licensing exams, and safety during sensor-compromising conditions (technology failures, extreme weather). The same logic applies to cognitive AGI: we may develop an ASI and choose to deploy it only as a Consultant (Level 2 Autonomy) or Collaborator (Level 3 Autonomy) rather than as a fully autonomous Agent, depending on contextual risk-benefit tradeoffs.
Interface as a research priority. The paper concludes the autonomy discussion by arguing that "interfaces that support human-AI alignment through better task specification, the bridging of process gulfs, and evaluation of outputs are a vital area of research" (citing Terry et al., 2023). This connects the autonomy framework to a concrete research agenda: as capabilities advance, interaction design must advance in parallel to ensure that humans can effectively specify goals, understand AI processes, and evaluate AI outputs at each autonomy level. The paper frames the division of labor as: "The role of model building research can be seen as helping systems' capabilities progress along the path to AGI in their performance and generality... Conversely, the role of human-AI interaction research can be viewed as ensuring new AI systems are usable by and useful to people such that AI systems successfully extend people's capabilities" — the classic "intelligence augmentation" framing (citing Brynjolfsson, 2022; Englebart, 1962).
Summary of Design Choices and Their Justifications
- Two-dimensional classification (performance × generality) rather than a single scale: required by Principle 2, because existing definitions split on whether depth or breadth is primary, and both are necessary.
- Percentile-based performance levels (0th/50th/90th/99th/100th) rather than absolute performance metrics: enables comparison across tasks with different difficulty scales; the percentiles are defined relative to skilled adults for Levels 2+ (per Principle 5's ecological validity requirement).
- Binary generality axis (Narrow vs. General) rather than a continuous breadth score: reflects a qualitative distinction between bounded-task and open-domain systems; avoids spurious precision about degrees of breadth that cannot yet be measured reliably.
- Inclusion of metacognitive tasks in the definition of "General": required by Principle 3, and reflects the judgment that genuine generality requires the ability to acquire new competencies, not just static breadth.
- Capability classification based on potential, not deployment: required by Principle 4, and enables the framework to classify systems that are never deployed (for safety or other reasons) at their true capability level.
- Separate Autonomy Levels decoupled from AGI Levels: allows the framework to distinguish between what a system can do (classified by the matrix) and how humans choose to interact with it (classified by autonomy level); essential for nuanced risk assessment because risks depend on the interaction of capability and autonomy, not capability alone.
- Deferral of benchmark creation to a multi-stakeholder process: the paper treats the matrix as the framework and the benchmark as future work; this is a deliberate scoping decision motivated by the complexity of benchmark design and the importance of including cross-organizational and multidisciplinary perspectives.
- Living benchmark requirement: because enumerating all possible tasks for a general intelligence is impossible, the benchmark must evolve; a system's generality must be demonstrated against novel tasks, not just a fixed set.
- No fixed threshold for what proportion of tasks constitutes generality: the paper acknowledges this as an open research question, with a pragmatic default of "the majority including new tasks," reflecting the reality that perfect task coverage is neither achievable (given the infinite task space) nor required (given that humans also have uneven performance profiles).
4. Key Insights and Innovations
Innovation 1: Replacing the Binary AGI Threshold with a Two-Dimensional, Multi-Level Taxonomy
The paper's most fundamental conceptual contribution is the rejection of AGI as a single binary endpoint — a switch that flips from "not AGI" to "AGI" — in favor of a graduated, two-dimensional classification that independently tracks performance depth and generality breadth. This is not merely adding intermediate categories to an existing scale; it is a reframing of what kind of thing AGI is.
Prior to this work, the dominant mode of discussing AGI was implicitly binary. Definitions varied wildly in their thresholds — from the Turing Test's behavioral indistinguishability to OpenAI's economic labor replacement — but they shared the structural assumption that AGI was a destination, a point at which a system crosses a line and earns the label. The nine case studies in Section 2 are not presented as variations on a spectrum but as competing claims about where that single line should be drawn. The consequence of this binary framing was a set of practical pathologies: researchers talked past each other because they were aiming at different thresholds under the same name; risk discussions lurched between "AGI isn't here yet, so we don't need to worry" and "AGI might arrive suddenly, so we can't prepare incrementally"; and progress was unmeasurable because there was no way to say "we're 30% of the way to AGI" — you either had it or you didn't.
The paper's matrixed taxonomy (Table 1) dissolves these pathologies by making two structural moves simultaneously. First, it separates the two dimensions that prior definitions had conflated — performance (how good is the system at a task?) and generality (across how many tasks is it that good?) — recognizing that these can advance asynchronously. AlphaFold (Level 5 Superhuman Narrow AI) demonstrates that performance can reach the ceiling while generality remains at the floor; frontier LLMs demonstrate that generality can broaden substantially while performance remains unreliable. A single-dimensional scale cannot capture this asymmetry. Second, it introduces graduated performance levels with explicit human-percentile anchors (unskilled human, 50th percentile of skilled adults, 90th percentile, 99th percentile, 100% of humans), providing a measurement framework that is simultaneously precise enough to distinguish systems and flexible enough to accommodate different tasks with different difficulty scales.
The significance of this reframing extends beyond taxonomy. It changes the conversation from prediction (when will AGI arrive?) to assessment (what level are we at now, and what are the properties of the next level?), which is a more tractable and more useful question. It enables graduated policy responses — different regulatory thresholds for Emerging vs. Competent vs. Expert AGI, rather than a single regulatory cliff at whatever point the community agrees constitutes "true AGI." And it absorbs rather than refutes prior definitions: Agüera y Arcas and Norvig's claim that LLMs are already AGIs maps to Level 1 (Emerging AGI), which the framework treats as a real and important category rather than a category error. The paper explicitly makes this diplomatic move, noting that the leveled approach "supports the coexistence of many prominent formulations" (Section 3, Principle 6 discussion).
This is a fundamental conceptual shift, not an incremental refinement. The paper is not proposing a better definition of AGI within the existing binary paradigm; it is arguing that the binary paradigm itself is the problem. The parallel to the SAE Levels of Driving Automation is instructive precisely because that framework similarly transformed a binary question ("is this car self-driving or not?") into a graduated one, enabling the entire ecosystem of graduated regulation, graduated testing, and graduated deployment that the autonomous vehicle industry now relies on.
Innovation 2: Decoupling Capability from Autonomy as a Framework for Risk Assessment
The paper's second distinctive conceptual move is its insistence that what an AI system can do (its capability level in the Performance × Generality matrix) and how humans choose to deploy it (its Autonomy Level in Table 2) are independent design dimensions, and that risk assessment requires analyzing their interaction rather than conflating them.
The dominant discourse around AGI risk had largely treated capability and autonomy as synonymous or at least tightly coupled. The assumption — often implicit — was that more capable systems would naturally or inevitably be deployed with more autonomy, and that AGI risk was therefore a function of capability alone. This assumption shows up in definitions that conflate the two: OpenAI's Charter defines AGI in terms of "highly autonomous systems" (emphasis on autonomy alongside capability), and many risk scenarios (deception, misalignment, recursive self-improvement) implicitly assume that a highly capable system is also operating autonomously. The paper identifies this conflation as a category error with practical consequences: it makes risk assessment coarser than it needs to be, because it prevents distinguishing between "a highly capable system deployed with tight human oversight" and "a moderately capable system deployed with full autonomy" — two scenarios with very different risk profiles that a capability-only framework would treat identically.
The structural move the paper makes is to create a second, independent classification scale (the six Autonomy Levels in Table 2) and then analyze the unlocking relationship between the two scales: higher AGI levels make higher autonomy levels possible but do not mandate them. An Emerging AGI could be deployed at Autonomy Level 5 (as an Agent), but this would likely be unsafe and undesirable; a Competent AGI enables Autonomy Level 3 (Collaborator) to work well, because the system's metacognitive abilities — knowing when to ask for help, modeling human mental states — are sufficient for co-equal interaction. The paper's examples in Table 2 specify "Possible" vs. "Likely" unlocking AGI levels for each autonomy paradigm, encoding the insight that capability creates affordances but does not dictate design choices.
The significance of this decoupling for risk assessment is substantial. It means that safety measures can target autonomy level without necessarily constraining capability development — you can build a more capable system while maintaining it at a lower autonomy level, using capability improvements to increase reliability rather than independence. It means that risk taxonomies can be two-dimensional: Anthropic's Responsible Scaling Policy (ASL levels, cited approvingly in Section 6.1) defines risk based on dangerous capabilities, but the paper adds the dimension of deployed autonomy, creating a richer risk space. And it means that the paper can discuss extreme risks (x-risk, misalignment) as emerging specifically from the combination of high capability (Expert AGI and above) and high autonomy (Agent-level deployment), not from high capability alone — a more precise and actionable risk model than the undifferentiated "AGI is dangerous" framing.
This is a fundamental conceptual reframing, not an incremental addition. Prior work (Sheridan et al., 1978; Parasuraman et al., 2000) had developed taxonomies of automation, but these were computer-centric — they described how much control was delegated, not what kind of interaction was occurring. The paper's human-AI interaction framing (conscious of "the nature of human-AI interaction style" rather than "how much control the designer relinquishes") and its explicit argument that high automation and high human control can co-exist (citing Shneiderman, 2020) represent a qualitatively different perspective on the relationship between capability and deployment.
Innovation 3: Metacognition as a Prerequisite for Generality, Not an Optional Add-On
The paper makes a specific, non-obvious claim about what constitutes "generality" in an AI system: it is not merely breadth of static task coverage, but requires metacognitive capabilities — the ability to learn new tasks, to know when to ask for help, and to model other agents' mental states (theory of mind). This claim is embedded in the definition of the General column in Table 1 ("a wide range of non-physical tasks, including metacognitive tasks like learning new skills") and elaborated in Sections 3 (Principle 3) and 5 (benchmark requirements).
Prior definitions of AGI had treated metacognition inconsistently. Some (Legg, 2008; Goertzel, 2014) focused on performance across cognitive tasks without explicitly requiring learning ability, leaving open the possibility that a system with a sufficiently large static training set could qualify. Others (Shanahan, 2015) included learning as a requirement but did not specify what metacognitive sub-skills were necessary or how they should be measured. Marcus's (2022a) five tests of flexibility included tasks that implied metacognition (cooking in an arbitrary kitchen requires adapting to novel environments) without making the metacognitive requirement explicit. The paper's contribution is to elevate metacognition from an implicit desideratum to an explicit, operationalizable requirement with three specified sub-types.
The significance of this move is that it draws a bright line between "broad but static" systems and "genuinely general" ones. A system that achieves Expert-level performance on 1,000 predetermined tasks but cannot learn a 1,001st task without retraining would, under this framework, not qualify as General — regardless of its breadth. The requirement for learning ability is not a bonus criterion but a gate: you cannot reach the General column at any performance level without demonstrating it. Similarly, the requirement for "knowing when to ask for help" connects generality to safety — a system that is broadly capable but has no awareness of its own limitations is dangerous in ways that a metacognitively calibrated system is not. Theory of mind connects generality to alignment — a system that cannot model what its human users know, intend, or believe cannot interact safely at higher autonomy levels.
The paper's argument that "it is infeasible for a system to be optimized for all possible use cases a priori" provides the logical foundation: since the space of possible tasks is effectively infinite, no finite training set can cover it. Generality therefore requires the ability to acquire new competencies from limited experience — which is a metacognitive capacity, not merely a broader training distribution. This is a fundamental reframing of what "general" means, not an incremental addition of a new evaluation criterion. It changes the research agenda from "build bigger models trained on more tasks" to "build models that can learn new tasks from few examples and know the limits of their own competence" — a qualitatively different target.
The evidence for this claim is not empirical (the paper presents no experiments) but structural: the metacognition requirement is a design choice justified by the analysis of prior definitions' failure modes and the logic of what an adequate generality concept must entail. The paper strengthens this justification by connecting the metacognitive requirements to existing research traditions — Chollet (2019) on measuring learning ability, Demetriou & Kazi (2006) on self-awareness and metacognition, Liang et al. (2023) on model calibration — demonstrating that the concepts are not invented de novo but drawn from established literatures and adapted for the AGI context.
Innovation 4: A Unified Framework That Absorbs Rather Than Refutes Competing Definitions
The paper makes a distinctive rhetorical and analytical move that distinguishes it from typical position papers: instead of arguing that existing AGI definitions are wrong and proposing a correct one to replace them, it constructs a framework within which all major prior definitions find a coherent place at different levels, recasting disagreement as operating at different points on a shared spectrum.
This is both a substantive contribution and a strategic one. Substantively, it provides evidence that the framework is genuinely more expressive than any single prior definition — if the framework could not accommodate a major existing definition, that would be a weakness. The paper explicitly maps: Agüera y Arcas and Norvig's claim maps to Level 1 Emerging AGI; the Legg/Shanahan/Suleyman cluster maps to Level 2 Competent AGI; and OpenAI's labor-replacement criterion maps roughly to Level 4 Exceptional AGI (Sections 3 and 4). The framework does not tell these authors they were wrong about what constitutes AGI — it tells them they were defining different points on a spectrum that the community had not yet articulated. This resolution is more intellectually satisfying than simply picking one definition as correct, because it explains why reasonable people reached different conclusions: they were observing systems at different capability levels, or emphasizing different dimensions (performance vs. generality), or setting thresholds at different human-percentile comparisons.
Strategically, this move makes the framework harder to reject. A paper that proposes a novel definition of AGI invites counterarguments that the definition is too strict (excluding systems that many consider AGI) or too loose (including systems that are clearly not AGI). By accommodating existing definitions as points within the taxonomy rather than competitors to it, the paper sidesteps this debate. The question becomes not "is this the right definition?" but "does this framework provide a useful structure for organizing our existing (conflicting) definitions and identifying the relationships between them?" The answer to the latter question is more easily demonstrated: the framework shows that Emergent AGI (Level 1) and Competent AGI (Level 2) are different things, which explains why people who have been calling LLMs "AGI" and people who have been insisting "that's not AGI" are both partly right — they're using the same word for different levels.
This is a fundamental rhetorical and conceptual contribution rather than an incremental refinement of definitions. It changes the nature of the AGI definition debate from a competition between rival candidates to a collaborative project of building a shared coordinate system. The paper's six principles (Section 3) provide the design requirements for this coordinate system, derived by analyzing the strengths and weaknesses of the nine case studies. The matrix is the implementation. The mapping of prior definitions to matrix cells is the validation.
5. Experimental Analysis
Evaluation Methodology
Dataset. This paper does not use a traditional dataset in the empirical sense — it is a position paper that proposes a conceptual framework rather than training or evaluating models. The "data" it analyzes are nine historical definitions of AGI (Section 2) and a set of contemporary AI systems used as exemplars for the proposed classification cells in Table 1 (ChatGPT, Bard, Llama 2, Gemini, AlphaFold, AlphaZero, StockFish, Deep Blue, AlphaGo, DALL-E 2, Imagen, Grammarly, Siri, Alexa, Google Assistant, PaLI, Watson, Jigsaw toxicity detectors, SHRDLU, and calculator/compiler software). No training set, test set, or validation split exists. The paper does reference the MATH benchmark (Hendrycks et al., 2021) and other evaluation frameworks in Section 5 only as examples of the type of tasks that a future AGI benchmark might draw from, not as datasets used in the current work.
Base model(s). The paper does not build, train, or evaluate any model. It references contemporary systems — primarily large language models (ChatGPT from OpenAI, Bard and Gemini from Google, Llama 2 from Meta) and narrow superhuman systems (AlphaFold, AlphaZero, StockFish) — as exemplars for its classification cells. The choice of these systems is driven by their prominence and public familiarity: they serve as reference points that make the abstract classification scheme concrete. The authors are all from Google DeepMind, and the selection of exemplars includes several Google-developed systems (Bard, Gemini, PaLI, Imagen), but also includes major systems from competitors (ChatGPT, Llama 2, DALL-E 2, Grammarly, StockFish, Deep Blue), demonstrating an effort at cross-organizational coverage.
Metrics. The paper proposes but does not implement a measurement framework. The key metrics are:
-
Performance: Measured as the percentile of a system's task performance relative to a human reference population. For Level 1 (Emerging), the reference is "unskilled humans." For Levels 2–5, the reference is "a sample of adults who possess the relevant skill" — e.g., English writing ability is measured against literate, fluent English-speaking adults, not the general population. The specific percentiles are: Level 1 (no fixed percentile, "equal to or somewhat better than an unskilled human"), Level 2 (at least 50th percentile), Level 3 (at least 90th percentile), Level 4 (at least 99th percentile), Level 5 (outperforms 100% of humans).
-
Generality: Measured as the breadth of tasks for which a system reaches a target performance threshold. The paper proposes only two categories: "Narrow" (clearly scoped task or set of tasks) and "General" (wide range of non-physical tasks, including metacognitive tasks like learning new skills). The proportion of tasks that must be met to qualify as "General" is explicitly left as an open research question.
-
Autonomy: Measured on a separate six-level scale (Table 2) based on the human-AI interaction paradigm, from Level 0 (No AI, human does everything) through Level 5 (AI as an Agent, fully autonomous). This is not a metric in the quantitative sense but a qualitative classification of deployment mode.
No quantitative metric values are computed or reported for any system — all classifications are approximate judgments by the authors. The paper's Table 1 caption explicitly acknowledges this: "The assignment of example systems to cells is approximate. Unambiguous classification of AI systems will require a standardized benchmark of tasks."
Baselines. There are no baselines in the traditional ML sense, since the paper does not compare systems or methods. However, the nine historical AGI definitions analyzed in Section 2 function as conceptual baselines — they are the existing frameworks against which the proposed ontology is compared. These are:
- The Turing Test (Turing, 1950)
- Strong AI / systems possessing consciousness (Searle, 1980)
- Analogies to the human brain (Gubrud, 1997)
- Human-level performance on cognitive tasks (Legg, 2008; Goertzel, 2014)
- Ability to learn tasks (Shanahan, 2015)
- Economically valuable work (OpenAI Charter, 2018)
- Flexible and general — the "Coffee Test" and related challenges (Marcus, 2022a,b)
- Artificial Capable Intelligence (Suleyman, 2023)
- SOTA LLMs as generalists (Agüera y Arcas & Norvig, 2023)
Each is analyzed for strengths and limitations (Section 2), and the six design principles (Section 3) are derived from this analysis. The proposed ontology is then shown to accommodate all nine definitions at different points in its matrix, making the prior definitions reference points within the framework rather than competitors to be outperformed.
Generation budget / compute accounting. Not applicable. The paper does not involve model training, inference, or any computation whose cost needs to be measured or accounted for. This is a conceptual position paper that proposes a classification ontology; its "compute budget" is the analytical effort of deriving the framework from prior definitions and design principles.
Cross-validation / statistical protocol. Not applicable. The paper makes no empirical claims that require statistical validation. The classification of contemporary systems (e.g., placing ChatGPT at Level 1 Emerging AGI) is presented as the authors' judgment based on publicly observable capabilities, not as the result of a measurement protocol that could be cross-validated. The paper explicitly notes that "unambiguous classification of AI systems will require a standardized benchmark of tasks" (Table 1 caption), acknowledging that the current classifications are provisional and approximate.
Main Quantitative Results
This paper reports no quantitative results in the traditional sense — no accuracy numbers, no FLOPs comparisons, no ablation tables, no learning curves. It is a position paper proposing a conceptual framework. The "results" are the framework itself and its application to classifying existing systems. This section therefore describes the framework's structure as the paper's primary output and the classification assignments as its "findings."
The Performance × Generality Matrix (Table 1)
The 6×2 classification grid is the paper's central output. It defines 12 possible classification cells (six performance levels × two generality categories), each with a canonical label. The matrix structure embeds several specific claims:
Claim about achieved performance levels: The paper asserts that the highest achieved performance for Narrow AI systems is Level 5 (Superhuman), with AlphaFold, AlphaZero, and StockFish as exemplars. For General AI systems, the highest achieved level is Level 1 (Emerging), with ChatGPT, Bard, Llama 2, and Gemini as exemplars. Levels 2–5 in the General column are explicitly marked "not yet achieved." This constitutes a specific empirical claim about the state of the field as of September 2023: no system has demonstrated Competent-level (50th percentile of skilled adults) performance across a general range of cognitive tasks.
Claim about frontier LLM classification: The paper assigns frontier language models (ChatGPT, Bard, Llama 2, Gemini) to the "Emerging AGI" cell (Level 1 General AI). This is a specific substantive judgment that simultaneously rejects two alternative classifications:
-
It rejects the claim that these systems are already full AGIs at a higher performance level. The paper explicitly states: "Overall, current frontier language models would therefore be considered a Level 1 General AI ('Emerging AGI') until the performance level increases for a broader set of tasks (at which point the Level 2 General AI, 'Competent AGI,' criteria would be met)."
-
It rejects the claim that these systems are merely narrow tools. By placing them in the "General" column rather than the "Narrow" column, the paper acknowledges that their breadth of capability — spanning language, coding, mathematics, multimodal processing, multiple languages, and few-shot learning — qualifies as generality, even though performance remains unreliable across most tasks.
The justification for this classification is based on uneven performance: frontier LLMs "exhibit 'Competent' performance levels for some tasks (e.g., short essay writing, simple coding), but are still at 'Emerging' performance levels for most tasks (e.g., mathematical abilities, tasks involving factuality)." The classification as "Emerging" rather than "Competent" reflects the minimum performance threshold — a system is only as "general" as its weakest broad task categories allow.
Claim about the relationship between performance and generality in the classification logic: For General AI systems, the paper specifies that the classification level is based on the minimum performance across most tasks, not the maximum on any subset: "a Competent AGI must have performance at least at the 50th percentile for skilled adult humans on most cognitive tasks, but may have Expert, Exceptional, or even Superhuman performance on a subset of tasks." This asymmetric rule — the level is determined by the floor, not the ceiling — is a specific design choice with practical implications: a system with Superhuman performance on mathematics but Emerging performance on most other cognitive tasks would be classified as Emerging AGI, not Superhuman AGI.
Claim about the Narrow column examples (Table 1): The paper populates every Narrow AI cell from Level 0 through Level 5 with named exemplars, implicitly claiming that Narrow AI systems exist at all six performance levels. The specific assignments:
- Level 0 Narrow Non-AI: calculator software, compiler
- Level 1 Emerging Narrow AI: GOFAI (Good Old-Fashioned AI), simple rule-based systems like SHRDLU
- Level 2 Competent Narrow AI: toxicity detectors (Jigsaw), smart speakers (Siri, Alexa, Google Assistant), visual QA systems (PaLI), Watson, "SOTA LLMs for a subset of tasks" (short essay writing, simple coding)
- Level 3 Expert Narrow AI: spelling and grammar checkers (Grammarly), generative image models (Imagen, DALL-E 2)
- Level 4 Exceptional Narrow AI: Deep Blue, AlphaGo
- Level 5 Superhuman Narrow AI: AlphaFold, AlphaZero, StockFish
The distinction between Levels 4 and 5 for narrow systems is subtle. Deep Blue and AlphaGo are placed at Level 4 (Exceptional, 99th percentile) rather than Level 5 (Superhuman, 100%) — the paper's implicit claim is that while these systems defeated world champions, they may not strictly outperform 100% of all possible humans in their domains, whereas AlphaFold and AlphaZero do. This distinction is not rigorously defended but reflects a judgment about whether the performance ceiling has been definitively reached.
The Autonomy Levels Classification (Table 2)
Table 2 defines six autonomy levels with associated "unlocking" AGI levels, example systems, and example risks. This is a conceptual output rather than an empirical one, but it embeds specific claims about the relationship between capability and deployment:
Claim about the "unlocking" relationship: Each Autonomy Level specifies which AGI Levels make that paradigm "Possible" versus "Likely." For instance, Autonomy Level 2 (AI as Consultant) is "Possible" with Competent Narrow AI but "Likely" with Expert Narrow AI or Emerging AGI. The asymmetry — a lower AGI level is required for "Possible" than for "Likely" — reflects the judgment that General systems can achieve competence in specific sub-domains (unlocking higher effective autonomy for those tasks) even when their broad performance is lower.
Claim about the highest currently deployed autonomy: The example systems in Table 2 range from Autonomy Level 0 (sketching on paper, typing in a text editor) through Autonomy Level 4 (using AI for scientific discovery, e.g., protein folding with AlphaFold). Autonomy Level 5 (AI as an Agent, fully autonomous personal assistants) is marked "not yet unlocked," representing the paper's judgment that no deployed system currently operates at full autonomy with general capability.
Claim about risk progression: Each autonomy level is associated with specific example risks that are claimed to emerge at that level, independent of the specific AGI capability level. Level 1 (Tool) introduces "de-skilling" and "disruption of established industries." Level 2 (Consultant) adds "over-trust," "radicalization," and "targeted manipulation." Level 3 (Collaborator) adds "anthropomorphization (e.g., parasocial relationships)" and "rapid societal change." Level 4 (Expert) adds "societal-scale ennui," "mass labor displacement," and "decline of human exceptionalism." Level 5 (Agent) adds "misalignment" and "concentration of power." This progression constitutes a specific claim about the risk landscape: risks are not uniform across deployment modes but are qualitatively distinct at different autonomy levels, and the most severe risks (misalignment, power concentration) are specific to the highest autonomy level.
Mapping of Historical Definitions to the Framework (Sections 3 and 4)
The paper makes specific claims about where prior AGI definitions map onto its framework:
- Agüera y Arcas and Norvig's definition (Case Study 9) maps to Level 1 Emerging AGI.
- The Legg (2008), Shanahan (2015), and Suleyman (2023) formulations map to Level 2 Competent AGI.
- OpenAI's Charter definition (2018) maps roughly to Level 4 Exceptional AGI.
These mappings are not presented with data — they are interpretive judgments based on analyzing the performance thresholds and generality requirements implicit in each definition. The mapping of the Legg/Goertzel definition to Level 2 is particularly significant because that definition is the one that "popularized the term AGI among computer scientists" (Section 2, Case Study 4), suggesting that the historical default understanding of AGI corresponds to the "Competent" rather than "Emerging" level.
Ablation Studies and Robustness Checks
The paper is a conceptual position paper and conducts no experiments that could be ablated in the traditional sense. However, it does perform several forms of conceptual validation that function analogously to robustness checks in empirical work:
Alternative definition coverage (the nine case studies): The framework's ability to accommodate all nine major historical definitions of AGI at different points in its matrix serves as a form of coverage test. If the framework could not accommodate a prominent definition — for instance, if no level corresponded to the Turing Test's behavioral indistinguishability criterion — this would be evidence of insufficiency. The paper demonstrates coverage by mapping each definition to a specific level or combination of levels in the matrix. This is presented in Section 3 and the discussion of Principle 6, not as a formal ablation table, but as conceptual validation of the framework's expressiveness.
Principle adherence check (Section 3): The six design principles are derived from the analysis of prior definitions' failure modes. The paper implicitly demonstrates that its proposed framework satisfies all six principles:
- Principle 1 (Capabilities, not Processes): The matrix classifies by task performance percentiles, not by internal mechanisms.
- Principle 2 (Generality and Performance): The matrix has two axes explicitly tracking both dimensions independently.
- Principle 3 (Cognitive and Metacognitive, not Physical): The General column definition includes metacognitive tasks; physical tasks are explicitly optional.
- Principle 4 (Potential, not Deployment): The matrix classifies based on benchmarked capability; the Autonomy Levels handle deployment separately.
- Principle 5 (Ecological Validity): This governs the (future) benchmark design; the matrix is structured to accept ecologically valid task assessments.
- Principle 6 (Path, not Endpoint): The six performance levels provide the graduated structure.
This self-consistency check is analogous to verifying that a proposed system satisfies its own design requirements, but it is not an independent validation — it demonstrates internal coherence, not external correctness.
Deployment sensitivity test (the DALL-E 2 example): The paper uses DALL-E 2 to illustrate the distinction between benchmarked capability and deployed performance: "While theoretically an 'Expert' level system, in practice the system may only be 'Competent,' because prompting interfaces are too complex for most end-users to elicit optimal performance." This example tests whether the framework can distinguish between a system's theoretical classification (what it could do under optimal conditions) and its practical deployment level (what it achieves with typical users). The framework passes this test by placing DALL-E 2 at Level 3 Expert Narrow AI in the capability matrix while acknowledging that its deployed autonomy might reflect lower effective performance. This demonstrates that the framework can handle the capability-deployment gap that Principle 4 was designed to address.
Cross-domain exemplar coverage: The paper populates the matrix with exemplars spanning diverse domains — games (Deep Blue, AlphaGo, AlphaZero, StockFish), science (AlphaFold), natural language (ChatGPT, Bard, Llama 2, Gemini), vision (Imagen, DALL-E 2, PaLI), speech (Siri, Alexa, Google Assistant), and traditional software (calculator, compiler, Grammarly). This breadth tests whether the framework's categories are domain-agnostic — capable of classifying systems from fundamentally different technical traditions within the same ontology. The absence of contradictions (no system that clearly belongs in one cell but resists classification there) serves as informal validation of domain-independence.
Temporal stability check (explicit and implicit): The paper's classifications are timestamped: "as of this writing in September 2023." This acknowledges that classifications are provisional and that systems will move between cells as capabilities advance. The framework's leveled structure is designed to accommodate such movement — a system that is Emerging AGI today might become Competent AGI tomorrow without requiring a change to the classification scheme itself. This temporal flexibility is a robustness property: the framework is designed to remain stable as the technology evolves, with systems moving through its cells rather than the cells needing to be redefined.
Missing conceptual validations: Several robustness checks that would strengthen the framework are absent. The paper does not test whether independent raters would classify the same systems into the same cells (inter-rater reliability). It does not explore edge cases — systems that blur the boundary between Narrow and General, or between adjacent performance levels. It does not test whether the exemplar assignments would change if different task sets or reference populations were used. And it does not examine whether the framework can accommodate hypothetical future systems (e.g., a system with Superhuman performance on 40% of tasks and Emerging performance on 60% — where does it go in the matrix?). These are limitations of the current work rather than flaws in its internal logic, but they represent genuine gaps in the framework's validation.
Critical Assessment
This is a position paper that proposes a conceptual framework, not an empirical study that tests hypotheses through experiments. Evaluating whether the "experiments support the claims" requires a shift in analytical frame: the question is not whether measurements are statistically significant but whether the framework's structure is logically coherent, whether its classifications of existing systems are defensible, and whether its claims about the framework's utility are supported by the analysis presented.
Does the Two-Dimensional Matrix Genuinely Improve on Single-Dimensional Definitions?
The paper's central claim is that performance and generality are distinct dimensions that must be tracked independently. The matrix in Table 1 demonstrates this formally by providing cells that single-dimensional definitions cannot capture. AlphaFold (Level 5 Narrow AI) and ChatGPT (Level 1 General AI) are at opposite extremes on both axes: one has maximal performance with minimal generality, the other has emerging generality with uneven performance. No single number or binary threshold captures this asymmetry.
Supporting evidence: The nine case studies in Section 2 demonstrate that prior definitions systematically conflated these dimensions. Agüera y Arcas and Norvig emphasized generality at the expense of performance (classifying LLMs as AGIs based on breadth alone). The Superhuman Narrow AI examples (AlphaFold, StockFish) emphasize performance without generality. The OpenAI Charter emphasizes a specific performance threshold (economically valuable labor replacement) without specifying the breadth of tasks required. The framework absorbs all of these by giving each dimension its own axis.
Genuine weakness: The two-dimensional structure is not empirically validated — the paper does not demonstrate that performance and generality are in fact independent dimensions that advance asynchronously, as opposed to being correlated aspects of a single underlying capability. The asynchronous advancement claim is supported only by the existence of AlphaFold (high performance, low generality) and ChatGPT (moderate generality, low performance), but this is a sample size of two data points — more of an existence proof than a systematic demonstration. It is entirely possible that as systems advance toward higher generality, performance and generality become increasingly correlated, making the two-axis framework less informative at higher levels. The paper does not address this possibility.
Missing evidence: The framework would be strengthened by a systematic analysis of where existing AI systems actually fall in the matrix, not just a few hand-picked exemplars. How many systems occupy each cell? Are there cells that are empirically empty despite being theoretically possible? Are there systems that resist clean classification because they don't fit neatly into either the Narrow or General column? Such an analysis would test whether the matrix carves reality at its joints or imposes a structure that doesn't match the actual distribution of AI capabilities.
Is the Classification of Frontier LLMs as "Emerging AGI" Defensible?
This is the paper's most consequential single judgment. Placing ChatGPT, Bard, Llama 2, and Gemini at Level 1 General AI (Emerging AGI) rather than at Level 2 (Competent AGI) or in the Narrow column makes a specific claim about their capabilities: they are sufficiently broad to qualify as "General" (because they span language, coding, mathematics, multimodal processing, multiple languages, and demonstrate few-shot learning) but insufficiently reliable to qualify as "Competent" (because their performance falls below the 50th percentile of skilled adults on most cognitive tasks, particularly mathematics and factuality).
Supporting evidence: The paper cites specific failure modes: "mathematical abilities, tasks involving factuality" are areas where frontier LLMs remain at Emerging rather than Competent performance. This is consistent with published evaluations of these systems. However, the paper provides no quantitative evidence — no benchmark scores, no percentile calculations, no comparison to human performance distributions — to support the classification. The judgment is presented as the authors' assessment based on publicly observable behavior, not as a measurement.
Genuine weakness: The distinction between Narrow and General may be too coarse to capture important differences among frontier LLMs. ChatGPT, Bard, Llama 2, and Gemini differ substantially in their capabilities — in benchmark scores, in reasoning abilities, in multimodal processing, in factual accuracy. Placing all four in the same cell (Emerging AGI) obscures these differences. A framework with only two generality categories cannot distinguish between a system that performs well on 20% of cognitive tasks and one that performs well on 80% — both might be "General" if the task coverage is broad enough, even though their practical utility is vastly different. The paper acknowledges that the proportion of tasks constituting "generality" is an open research question but does not address how this ambiguity affects the classification of systems near the boundary.
Missing evidence: The classification depends on an implicit claim about "most tasks." If a benchmark of 1,000 diverse cognitive tasks existed, the paper could report the percentage on which frontier LLMs achieve Competent-level performance, and the percentage on which they achieve only Emerging-level performance. This would make the "Emerging AGI" classification falsifiable. Without such a benchmark, the classification is a judgment call that different observers could reasonably dispute.
Does the Decoupling of Capability from Autonomy Genuinely Enable Better Risk Assessment?
The paper claims that considering AGI Level and Autonomy Level jointly provides "more nuanced insights into risks associated with AI systems" (Section 6.3) than considering capability alone. Table 2 maps specific risks to specific autonomy levels, and the paper argues that the interaction of high capability with high autonomy introduces qualitatively different risks than either factor alone.
Supporting evidence: The risk progression in Table 2 is logically coherent. The risks escalate from mundane concerns at low autonomy (de-skilling from over-reliance on AI tools) through intermediate concerns (anthropomorphization from social AI collaborators, mass labor displacement from AI experts) to existential concerns at high autonomy (misalignment of fully autonomous agents, concentration of power). This progression makes intuitive sense and provides a structured way to think about graduated safety measures.
Genuine weakness: The mapping of risks to autonomy levels is based on the authors' judgment, not on empirical evidence about how risks actually manifest at different deployment modes. There is no demonstration that the identified risks are specific to their assigned autonomy levels — could anthropomorphization occur at Autonomy Level 2 (Consultant) rather than only at Level 3 (Collaborator)? Could mass labor displacement occur at Level 3 rather than only at Level 4? The framework asserts these mappings but does not validate them. This matters because if the risk-to-level mappings are incorrect, the framework's guidance for safety measures could be misleading — focusing attention on the wrong risks at the wrong autonomy levels.
Missing evidence: The framework would be strengthened by case studies or historical analysis demonstrating that specific AI deployments at specific autonomy levels have produced the predicted risks. For instance, have grammar checkers (Autonomy Level 1, Tool) actually produced de-skilling? Have recommender systems (Autonomy Level 2, Consultant) actually produced radicalization? Such evidence would move the risk mappings from plausible to empirically grounded.
Can the Framework Be Operationalized Given That No Benchmark Exists?
The paper explicitly defers benchmark creation to future work. Section 5 specifies properties a benchmark must have but does not propose one. The framework therefore remains at the level of conceptual proposal rather than operational tool.
The significance of this gap: Without a benchmark, the framework's classifications are unfalsifiable. The claim that frontier LLMs are "Emerging AGI" cannot be tested because there is no agreed-upon task set against which to measure their performance, no reference population against which to compute percentiles, and no threshold for what proportion of tasks must be met to qualify as "General." The framework provides a language for making such claims but not a method for verifying them. This is not a fatal flaw — position papers are allowed to propose frameworks without fully implementing them — but it means the framework's utility is currently limited to structuring discourse rather than enabling measurement.
The specific challenges the paper acknowledges but does not resolve:
- Enumerating the full set of tasks that constitute "generality" is "impossible" — the task space is infinite. The living benchmark concept partially addresses this but does not specify how to bootstrap the initial task set or how to adjudicate disputes about task inclusion.
- The proportion of tasks that must be passed to demonstrate generality is unspecified. The paper's pragmatic default ("the majority") is vague and could produce very different classifications depending on the task set.
- Metacognitive tasks — which the paper elevates to a requirement for generality — are particularly difficult to benchmark. How do you measure "knowing when to ask for help" or "theory of mind" in a way that is ecologically valid, scalable, and not gameable?
- The reference populations for percentile calculations are task-dependent and must be defined for each task in the benchmark, which is a massive methodological undertaking.
Missing evidence: The paper could have strengthened its case by providing even a small-scale proof-of-concept benchmark — say, 20–30 tasks spanning linguistic, mathematical, spatial, and social domains, with defined reference populations and percentile thresholds, applied to 2–3 contemporary systems to demonstrate that classification is feasible. Such a demonstration would not need to be comprehensive; it would only need to show that the framework can be operationalized in principle. Its absence means the framework remains a proposal rather than a demonstrated capability.
Cross-Organizational Validation
The paper's authors are all from Google DeepMind. The selection of exemplar systems for Table 1 includes Google-developed systems (Bard, Gemini, PaLI, Imagen) alongside competitor systems (ChatGPT from OpenAI, Llama 2 from Meta, DALL-E 2 from OpenAI, Grammarly, StockFish). While the paper appears to make a good-faith effort at balanced coverage, the classifications of Google's own systems (particularly the "Emerging AGI" classification of Bard and Gemini) could be perceived as self-serving — positioning Google's products at the frontier of AGI progress while stopping short of claiming they are fully "Competent." Independent validation of these classifications by researchers without organizational conflicts of interest would strengthen the framework's credibility. The paper acknowledges this indirectly by calling for "cross-organizational and multi-disciplinary viewpoints" in benchmark development (Section 5), but does not subject its own classifications to such scrutiny.
Summary Assessment
The paper's claims are logically coherent and well-structured, but they are not empirically validated. The framework provides a language for discussing AGI progress that is significantly more nuanced than the binary "AGI or not" framing it replaces, and this is a genuine contribution. The classification of frontier LLMs as "Emerging AGI" is a defensible judgment that captures both their impressive breadth and their unreliable performance. The decoupling of capability from autonomy is a useful conceptual move that enables more granular risk analysis.
However, the framework remains at the level of a proposal. Its classifications are based on author judgment rather than measurement. Its operationalization depends on a benchmark that does not exist and whose creation the paper identifies as extremely challenging. Its risk mappings are plausible but unvalidated. Its coverage testing (nine definitions, a dozen exemplar systems) is illustrative rather than systematic. These limitations are characteristic of position papers — the contribution is the framework itself, not its empirical validation — but they mean that the paper's claims should be understood as a structured proposal for how the community should think about AGI, not as a demonstration that this way of thinking has been shown to work in practice.
The paper's most significant gap is the absence of even a small-scale demonstration that the framework can be operationalized. Until a benchmark exists (even a partial one) and systems are classified against it (even approximately), the framework remains a conceptual tool rather than a practical one. The paper's value lies in providing the intellectual scaffolding for that benchmark development, not in providing the benchmark itself.
6. Limitations and Trade-offs
6.1 The Framework Is Not Operationalized — No Benchmark Exists to Classify Systems
The assumption or constraint. The entire Levels of AGI framework depends on the existence of a standardized benchmark that can measure AI system performance against human reference populations across a diverse set of cognitive and metacognitive tasks. The paper explicitly acknowledges that no such benchmark exists, deliberately defers its creation to future work, and characterizes benchmark design as involving "immense complexity" requiring "cross-organizational and multi-disciplinary viewpoints" (Section 5). The framework therefore provides a language for classification without providing a method for measurement.
This is not a minor omission — it is a fundamental gap between the framework's structure and its usability. The paper's key claims — that frontier LLMs are "Emerging AGI," that AlphaFold is "Superhuman Narrow AI," that Competent AGI is "not yet achieved" — are presented as author judgments rather than measurement outcomes. The caption to Table 1 is explicit: "The assignment of example systems to cells is approximate. Unambiguous classification of AI systems will require a standardized benchmark of tasks."
The benchmark blueprint in Section 5 identifies several unresolved methodological challenges that make operationalization particularly difficult:
-
Enumerating the task set that constitutes "generality" is acknowledged as "impossible" because the space of potential tasks is infinite. The "living benchmark" concept — a framework for continually generating and adding new tasks — is proposed as a partial solution, but the paper provides no mechanism for bootstrapping the initial task set or for adjudicating disputes about which tasks count as necessary for generality.
-
The proportion of tasks a system must pass to qualify as "General" is explicitly flagged as "an open research question" (Section 5). The paper's pragmatic default — "Systems that pass the majority of the envisioned AGI benchmark at a particular performance level, including new tasks added by the testers, can be assumed to have the associated level of generality" — is vague enough that different benchmark implementers could reach different classifications for the same system. Is "majority" 51%? 80%? 95%? The framework's precision at distinguishing capability levels is undermined by this unresolved threshold.
-
Metacognitive tasks, which the framework elevates to definitional requirements for generality (you cannot reach the "General" column without demonstrating learning ability, appropriate help-seeking, and theory of mind), are the most difficult to benchmark. How does one measure "knowing when to ask for help" in an ecologically valid, scalable, and non-gameable way? The paper identifies these as essential but provides no guidance on how to operationalize them.
-
Reference populations for percentile calculations must be specified for every task in the benchmark. For "Competent" performance (50th percentile of skilled adults), who counts as a "skilled adult" for, say, mathematical theorem proving or creative writing? These reference populations must be defined, sampled, and tested — a massive empirical undertaking that the paper does not address.
The consequence. Without an operational benchmark, the framework is unfalsifiable. The claim that ChatGPT is "Emerging AGI" cannot be tested because there is no agreed-upon task set, no reference population data, and no specified threshold for what constitutes "most tasks." Different researchers applying the same framework to the same system could reach different classifications — one might judge that ChatGPT's mathematical abilities are sufficient for Competent performance relative to the average literate adult, while another might judge them Insufficient. The framework provides no mechanism for resolving such disputes because it provides no measurement protocol.
This, in turn, limits the framework's practical utility for its stated goals. Policymakers cannot write regulations triggered by a system reaching "Competent AGI" if there is no agreed-upon way to determine when that threshold has been crossed. Risk assessments cannot use the framework to guide graduated safety measures if the classification of a given system is contestable. Researchers cannot track progress along the path to AGI if the benchmarks for measuring that progress do not exist. The framework succeeds at providing a more nuanced conceptual vocabulary for discussing AGI, but it does not (yet) enable the empirical tracking, regulatory triggering, or risk assessment that the paper argues are essential motivations for having such a vocabulary in the first place.
What evidence exists in the paper. The paper does not attempt to demonstrate operationalization — there is no proof-of-concept benchmark, even a small one, and no classification of any system based on quantitative measurement. All classifications in Table 1 are approximate, author-judgment-based placements. Section 5 is entirely prospective (what a benchmark should look like) rather than descriptive (what a benchmark does look like). The paper acknowledges the gap: "Because of the immense complexity of this process, as well as the importance of including a wide range of perspectives... we do not propose a benchmark in this paper" (Section 5).
Mitigation status. The paper does not mitigate this limitation — it explicitly defers benchmark creation to future work conducted through a multi-stakeholder process. The benchmark blueprint (Section 5) provides design requirements but no implementation. The acknowledgment is candid, but the limitation remains: the framework is a proposal for how to structure measurement, not a measurement instrument itself. A reader evaluating whether to adopt this framework for their organization's AGI tracking or risk assessment would need to supply (or wait for the community to supply) the benchmark that makes the framework operational. Until that benchmark exists, the framework is a conceptual contribution rather than a practical tool.
6.2 The Boundary Between Narrow and General Is Underspecified for Systems Near the Border
The assumption or constraint. The framework divides the generality axis into exactly two categories: "Narrow" (clearly scoped task or set of tasks) and "General" (wide range of non-physical tasks, including metacognitive tasks like learning new skills). There is no intermediate category, no continuous generality score, and no explicit decision rule for classifying systems that fall between "clearly narrow" and "clearly general." The paper provides only two criteria for the General column: breadth across cognitive domains and the presence of metacognitive capabilities. The proportion of tasks that must be met is left as an open research question, and the metacognitive requirement — while conceptually important — is itself difficult to measure without an operational benchmark.
This binary classification creates a cliff: a system is either Narrow or General, with no graduated way to express "broader than most narrow systems but not yet general." Given that the paper argues forcefully against binary thresholds for AGI as a whole (Principle 6: "Focus on the Path to AGI, not a Single Endpoint"), it is notable that the generality axis itself remains binary rather than graduated.
The consequence. The binary generality axis may obscure important distinctions among systems that are not clearly Narrow but also not unambiguously General. Consider a hypothetical system that achieves Competent performance on 60% of a diverse cognitive task benchmark and demonstrates some learning ability but struggles with theory of mind tasks. Is this Narrow or General? The paper's criteria ("wide range of non-physical tasks, including metacognitive tasks") could support either classification depending on how "wide range" is interpreted and whether partial metacognitive competence counts. Different assessors could reasonably place the same system in different columns, which undermines the framework's goal of enabling "clear communication among researchers, practitioners, and policymakers about systems' capabilities" (Impact Statement).
This ambiguity is particularly consequential because the "Narrow" vs. "General" distinction is what separates systems like Grammarly (Expert Narrow AI) from systems like ChatGPT (Emerging AGI). The practical difference between these systems — in terms of deployment flexibility, risk profile, and economic impact — is enormous. If the classification boundary between them is ambiguous, the framework's ability to guide deployment decisions and risk assessments is weakened.
The problem is compounded for systems that are genuinely in transition. The paper's own analysis suggests that frontier LLMs sit at an intermediate point: they demonstrate sufficient breadth that the paper classifies them as "General," but their metacognitive capabilities are uneven (they can learn from few-shot examples but their calibration — knowing when they are likely to be wrong — is poor). If a future system has the breadth of current LLMs but better calibration, does that make it "more General" or simply "more Competent"? The binary axis cannot express this distinction.
What evidence exists in the paper. The paper acknowledges the unresolved boundary implicitly through its treatment of the proportion-of-tasks question: "Determining what portion of benchmarking tasks at a given level demonstrate generality remains an open research question" (Section 5). It also acknowledges that individual systems have uneven capability profiles — a General system may perform at Level 2 on some tasks and Level 1 on others — but does not address how unevenness interacts with the binary Narrow/General classification. The paper's exemplar placements in Table 1 avoid ambiguous cases: the systems classified as Narrow are clearly narrow (AlphaFold does one task; Grammarly does one class of tasks), and the systems classified as General are frontier LLMs whose breadth is their defining feature. The framework is not tested against edge cases.
Mitigation status. Not addressed. The paper does not propose a graduated generality scale, does not specify a decision rule for the Narrow/General boundary, and does not discuss how to classify systems near the boundary. The binary structure appears to be a deliberate simplification — the paper argues that two categories capture the essential distinction — but the simplification's costs (loss of nuance at intermediate generality levels, classification ambiguity) are neither acknowledged nor justified. Future work could explore whether a continuous generality score (e.g., percentage of a benchmark's task categories where the system meets the target performance level) would provide more useful distinctions than the binary split.
6.3 The Classification of Frontier LLMs Rests on Author Judgment, Not Measurement, Creating a Credibility Problem for the Framework's Most Consequential Claim
The assumption or constraint. The paper's placement of ChatGPT, Bard, Llama 2, and Gemini at Level 1 General AI ("Emerging AGI") is its most publicly visible and consequential claim. It simultaneously affirms that these systems have crossed a significant threshold (they are genuinely "General," not merely "Narrow" tools) while denying that they have reached the level most historically associated with AGI ("Competent" — the level the paper identifies as corresponding to the Legg, Shanahan, and Suleyman definitions). This classification attempts to thread a needle: it acknowledges the impressive breadth of frontier LLMs without overclaiming their reliability.
However, this classification is based entirely on the authors' qualitative assessment of publicly observable system behavior. No benchmark scores are reported. No percentile calculations relative to human reference populations are provided. No systematic evaluation across cognitive domains is conducted. The paper references specific failure modes — "mathematical abilities, tasks involving factuality" (Section 4) — as evidence that performance remains at Emerging rather than Competent levels for most tasks, but provides no quantitative support. The judgment is presented as an observation about publicly available systems, not as a measurement.
The consequence. The credibility of the framework's most prominent classification depends entirely on whether readers accept the authors' judgment. A skeptic could reasonably argue that frontier LLMs do achieve Competent performance on most cognitive tasks that typical humans encounter — that their mathematical errors, while real, reflect performance that is still above the 50th percentile of the general adult population (which struggles with mathematics), and that their factual errors are comparable to the error rates of humans answering from memory. Under this interpretation, frontier LLMs would be classified as Competent AGI (Level 2) rather than Emerging AGI (Level 1) — a substantial difference, since the paper identifies Competent AGI as the level that "best corresponds to many prior conceptions of AGI" and that "may precipitate rapid societal change once achieved" (Table 1 caption).
Conversely, a different skeptic could argue that frontier LLMs should be classified as Narrow rather than General, on the grounds that their metacognitive capabilities — particularly theory of mind and calibrated help-seeking — are insufficient to meet the General column's definitional requirements. The paper's own analysis notes that model calibration (knowing when the system is likely to be wrong) remains poor in current systems, which could be interpreted as failing the metacognitive bar for generality.
The point is not that either of these alternative classifications is clearly correct — it is that the framework, as currently instantiated without a benchmark, provides no empirical basis for resolving the dispute. Different assessors applying the same framework to the same systems can reach different conclusions, which undermines the framework's claim to provide a "common language to compare models" (Abstract). A common language requires shared measurement procedures, not just shared vocabulary.
This is particularly problematic because the classification is not a neutral observation — the authors are all from Google DeepMind, and the systems classified as "Emerging AGI" include Google's own products (Bard, Gemini). While the paper makes a good-faith effort at cross-organizational coverage (also classifying ChatGPT and Llama 2 at the same level), the appearance of organizational self-interest — Google's products are described as having reached a significant milestone ("Emerging AGI") while stopping short of a classification ("Competent AGI") that might trigger heightened regulatory scrutiny — could undermine trust in the framework if not addressed through independent validation.
What evidence exists in the paper. None — the classification is presented as the authors' assessment without supporting data. Section 4 provides the classification and a qualitative justification (uneven performance: Competent on some tasks like short essay writing, Emerging on most tasks like mathematics and factuality), but no benchmark results, no human comparison data, and no systematic domain coverage analysis. The paper does not report having conducted evaluations of any of the classified systems.
Mitigation status. The paper acknowledges this limitation indirectly through the Table 1 caption ("The assignment of example systems to cells is approximate. Unambiguous classification of AI systems will require a standardized benchmark of tasks") and through the Section 5 discussion of the need for a benchmark. However, it does not acknowledge that the classification of its own employer's products might be perceived as self-serving, nor does it propose interim validation mechanisms (e.g., independent expert panels, partial benchmarks) that could strengthen the credibility of the classifications pending full benchmark development. The paper's call for "cross-organizational and multi-disciplinary viewpoints" in benchmark development (Section 5) applies to future work; it does not subject the paper's own classifications to such scrutiny.
6.4 The Risk-to-Level Mappings in Table 2 Are Asserted, Not Validated, Limiting the Framework's Utility for Practical Risk Assessment
The assumption or constraint. Table 2 maps specific example risks to specific autonomy levels, implying a causal or probabilistic relationship: particular risks emerge at particular autonomy thresholds. For example, "de-skilling (e.g., over-reliance)" is listed as a risk for Autonomy Level 1 (AI as a Tool); "over-trust, radicalization, targeted manipulation" for Level 2 (AI as a Consultant); "anthropomorphization (e.g., parasocial relationships)" for Level 3 (AI as a Collaborator); "societal-scale ennui, mass labor displacement, decline of human exceptionalism" for Level 4 (AI as an Expert); and "misalignment, concentration of power" for Level 5 (AI as an Agent).
Similarly, Section 6.1 asserts relationships between AGI capability levels and risk categories: "the 'Expert AGI' level is likely to involve structural risks related to economic disruption and job displacement," while "the 'Exceptional AGI' and 'ASI' levels are where many concerns relating to x-risk are most likely to emerge." The paper also suggests that "reaching 'Expert AGI' likely alleviates some risks introduced by 'Emerging AGI' and 'Competent AGI,' such as the risk of incorrect task execution" — a specific claim about risk reduction at higher capability levels.
These mappings are central to the paper's argument that the framework enables "more nuanced risk assessments" (Section 6.3), but they are presented as the authors' judgments rather than as empirically grounded or theoretically derived claims. The paper provides no evidence that de-skilling is primarily a Level 1 phenomenon rather than also occurring at higher autonomy levels, no evidence that anthropomorphization emerges specifically at Level 3, and no evidence that incorrect task execution risk decreases at higher capability levels (indeed, one could argue that more capable systems making errors in more consequential domains could increase rather than decrease that risk).
The consequence. If the risk-to-level mappings are incorrect — if, for example, mass labor displacement actually begins at Autonomy Level 3 (Collaborator) rather than Level 4 (Expert), or if anthropomorphization is a significant risk at Level 2 (Consultant) as well as Level 3 — then the framework could misdirect safety efforts. Organizations using the framework to guide their risk mitigation investments might focus on preventing misalignment at Autonomy Level 5 while underestimating the labor displacement risks that are already emerging at lower autonomy levels with current systems. Policymakers using the framework to determine regulatory thresholds might calibrate regulations to the wrong capability or autonomy levels.
The issue is not that the mappings are obviously wrong — many are intuitively plausible — but that they are untested. The paper's contribution is to propose a structure for thinking about the relationship between capability, autonomy, and risk; but without validation, the structure's specific predictions (which risks appear at which levels) cannot be relied upon for decision-making. This is a significant gap given that enabling better risk assessment is one of the paper's primary stated motivations (Section 1: "It is critical for the AI research community to explicitly reflect on what we mean by 'AGI,' and aspire to quantify attributes like the performance, generality, and autonomy of AI systems. Shared operationalizable definitions for these concepts will support... risk assessments and mitigation strategies").
The risk-alleviation claim is particularly consequential and particularly undersupported. The assertion that reaching "Expert AGI" might alleviate "the risk of incorrect task execution" associated with Emerging and Competent AGI implies a net safety benefit from capability advancement — that more capable systems are safer in at least some respects, not just more dangerous. If this claim is incorrect, and more capable systems are uniformly more dangerous (errors become more consequential even if less frequent), then the framework's implied safety narrative is misleading.
What evidence exists in the paper. None. The risk mappings in Table 2 and Section 6.1 are presented without citations to empirical studies, historical analyses, or theoretical models that would support the specific level assignments. The paper references general risk frameworks (Zwetsloot & Dafoe, 2019 on misuse, alignment, and structural risks; Shevlane et al., 2023 on extreme risks) and safety policies (Anthropic's Responsible Scaling Policy, 2023b) but does not draw on these to validate the specific mappings. The risk discussion is entirely prospective and judgment-based.
Mitigation status. The paper does not acknowledge this as a limitation, nor does it propose validation mechanisms. The risk mappings are presented as part of the framework's contribution rather than as hypotheses requiring testing. The paper suggests that "a more complete analysis of risk profiles associated with each level is a critical step toward developing a taxonomy of AGI that can guide safety/ethics research and policymaking" (Section 6.1), which implicitly acknowledges that the current analysis is incomplete. However, it does not flag the specific risk-to-level mappings as provisional or identify how they might be validated. A reader seeking to use the framework for risk assessment would need to supply their own evidence that the identified risks actually concentrate at the claimed levels.
6.5 The Framework Provides No Mechanism for Handling Systems with Highly Uneven Capability Profiles, Despite Acknowledging That Unevenness Is the Norm
The assumption or constraint. The classification logic for General AI systems specifies that the performance level is determined by the minimum performance across most cognitive tasks — "a Competent AGI must have performance at least at the 50th percentile for skilled adult humans on most cognitive tasks, but may have Expert, Exceptional, or even Superhuman performance on a subset of tasks" (Section 4). This "floor determines classification" rule is a specific design choice: a system's rating is only as high as its weakest broad task categories allow.
This rule works cleanly for the paper's primary exemplar — frontier LLMs, which are strong at some tasks (essay writing, simple coding) and weak at others (mathematics, factuality), producing a consistent classification at Level 1 (Emerging) based on the floor. However, the rule becomes problematic for systems with more extreme unevenness. Consider a hypothetical system that is Superhuman at theorem proving, Expert at code generation, Competent at natural language understanding, but Emerging at spatial reasoning and social cognition. Under the floor rule, this system would be classified as Emerging AGI (Level 1 General AI) — the same classification as a system that is Emerging at everything. Yet the first system is clearly more capable and more consequential than the second; its Superhuman theorem-proving ability might have enormous scientific and economic impact even if it cannot reason spatially. The floor rule collapses this distinction.
The paper acknowledges that uneven capability profiles are expected — "General systems that broadly perform at a level N may be able to perform a narrow subset of tasks at higher levels" (Table 1 note) — and even suggests that model cards "should detail this mixture of performance levels" (Section 4). But the framework's output is a single cell in the 6×2 matrix, which cannot express this internal heterogeneity. A model card that says "Emerging AGI (with Competent essay writing, Expert coding, Superhuman math)" communicates far more useful information than the matrix classification alone, but the information is in the unstructured description, not in the taxonomy.
The consequence. The framework's classification may be too coarse to distinguish systems that differ dramatically in their practical impact and risk profile. Two systems both classified as "Emerging AGI" could be radically different — one might be competent at 40% of cognitive tasks and emerging at 60%, while another is emerging at 95% and competent at only 5%. The first system might be deployable as a Consultant (Autonomy Level 2) for a substantial range of economically valuable tasks; the second might not be reliably usable for anything beyond Tool-level (Autonomy Level 1) applications. Yet they receive the same classification.
This coarseness limits the framework's utility for deployment decisions and risk assessment. A company deciding whether to deploy an AI system in a high-stakes domain needs to know not just that the system is "Emerging AGI" but which tasks it performs well on and which it does not. A policymaker designing regulations needs to know not just the overall AGI level but whether the system has dangerous capabilities in specific domains (e.g., Superhuman Narrow AI performance in chemical engineering paired with only Emerging performance in ethical reasoning — a combination the paper itself flags as potentially "a dangerous combination" in Section 4). The single-cell classification cannot capture these safety-relevant capability profiles.
The paper's suggestion that model cards should detail the mixture of performance levels is sensible but effectively admits that the matrix classification alone is insufficient. It implies that the full picture requires information beyond what the taxonomy captures — which raises the question of what the taxonomy adds beyond the model card information it supplements.
What evidence exists in the paper. The paper's own examples demonstrate the problem. DALL-E 2 is placed at Level 3 Expert Narrow AI, but the paper notes that it "has failure modes (e.g., drawing hands with incorrect numbers of digits, rendering nonsensical or illegible text)" that prevent Exceptional classification. This is an example of unevenness within a narrow domain — high quality on most image generation sub-tasks, poor quality on specific sub-tasks — and the framework captures this by lowering the overall classification from Exceptional to Expert. But for General systems, the unevenness is across domains rather than within a domain, and the floor rule may produce classifications that are misleadingly low relative to the system's capabilities in its areas of strength.
Frontier LLMs are described as "exhibit[ing] 'Competent' performance levels for some tasks (e.g., short essay writing, simple coding), but are still at 'Emerging' performance levels for most tasks" (Section 4). The classification as Level 1 captures the floor but loses the information that these systems are Competent or better at some economically significant tasks — information that a deployer would need to know.
Mitigation status. Partially acknowledged but not resolved. The paper recognizes that "the likely uneven performance of systems progressing along the path to AGI" is important and suggests model cards as a supplementary documentation mechanism (Section 4). It also notes, in the context of autonomy unlocking, that "such unevenness of capability for General AIs may unlock higher autonomy levels for particular tasks that are aligned with their specific strengths" (Section 6.2) — acknowledging that the effective autonomy level can be task-dependent even when the overall AGI classification is uniform. However, the framework itself does not incorporate this task-dependence. The matrix produces a single classification per system; the nuance lives outside the taxonomy in unstructured documentation. A more expressive framework might, for example, produce a capability profile (a vector of performance levels across task categories) rather than a single scalar classification, or might allow task-conditional AGI levels ("Emerging AGI overall, Competent AGI for text-generation tasks"). The paper does not explore these possibilities.
6.6 The Framework's Dependency on Human Reference Populations Creates Measurement Challenges That Are Acknowledged but Not Addressed
The assumption or constraint. The performance levels for Competent and above are defined by percentiles relative to "a sample of adults who possess the relevant skill" (Section 4). This means that to classify a system at Level 2 (Competent) or higher, one must: (1) define, for each task in the AGI benchmark, what constitutes "possessing the relevant skill"; (2) sample from that reference population; (3) measure the reference population's performance distribution on that task; and (4) compare the AI system's performance to that distribution to compute its percentile. This must be done for every task in the benchmark, and the reference populations are task-specific — the skilled adult population for English essay writing is different from the skilled adult population for mathematical theorem proving.
This creates a substantial empirical burden that the paper acknowledges only in passing. The reference population problem is fundamental to the framework's measurement model — without it, the percentile thresholds are undefined — but the paper provides no guidance on how to operationalize it. Who qualifies as "possessing the relevant skill" for, say, creative writing? Anyone who is literate? Only published authors? Only authors who have won literary prizes? The choice of reference population dramatically affects the percentile calculation. If the reference population for mathematics is "all adults who completed high school," the 50th percentile might be relatively low (basic algebra competence). If the reference population is "adults with undergraduate mathematics degrees," the 50th percentile is substantially higher. The same AI system could be classified as Competent under the first reference population and Emerging under the second.
The paper makes one specific choice about reference populations: for Level 1 (Emerging), the comparison is to "an unskilled human" rather than a skilled adult, establishing a lower bar. But for Levels 2–5, the reference population definition is underspecified. The paper's examples suggest intuitive reference classes — English writing compared to literate, fluent English-speaking adults — but this specifies only the inclusion criterion (literacy and fluency), not how to sample from that population, what sample size is needed, or how to handle the fact that different literate adults have very different writing abilities.
The consequence. The underspecification of reference populations means that the framework's performance level classifications are not uniquely determined by the benchmark tasks alone. Different implementers making different (reasonable) choices about reference populations could classify the same system at different performance levels — one at Competent, another at Expert — even if they agree on the benchmark tasks and the system's absolute performance. This introduces an additional degree of freedom that undermines the framework's claim to provide a "common language" for comparing models. If two organizations using the framework cannot agree on what "Competent" means for a given task because they have defined the reference population differently, the framework has not solved the coordination problem it was designed to address.
The problem is most acute for tasks that have no natural reference population. For well-defined professional skills (e.g., radiology, legal document review), the reference population might be board-certified professionals. But for broader cognitive tasks — "creativity," "spatial reasoning," "interpersonal intelligence" — who is the relevant reference population? The general adult population? People who use these skills professionally? The paper identifies these as AGI-relevant task categories (Section 5) but does not address how to define reference populations for them.
The cost and logistics of reference population measurement are also substantial. To establish performance distributions for even 100 tasks across their respective reference populations would require large-scale human subject testing — orders of magnitude more expensive and time-consuming than traditional ML benchmark development, which typically compares systems to static test sets rather than to human performance distributions. The paper does not acknowledge this cost, let alone propose how to manage it.
What evidence exists in the paper. Section 4 defines the percentile thresholds and specifies that "for all performance levels above 'Emerging,' percentiles are in reference to a sample of adults who possess the relevant skill." The DALL-E 2 example in Section 4 mentions that the "Expert" classification is based on the observation that "DALL-E 2 produces images of higher quality than most people are able to draw" — an implicit claim about the reference population (people who can draw) and the system's performance relative to it. But no methodology for defining or sampling reference populations is described. Section 5, on benchmarking, does not discuss the reference population problem at all.
Mitigation status. Not addressed. The paper acknowledges that benchmark development will be complex and require multi-stakeholder input, but does not flag reference population definition as a specific challenge within that broader complexity. This is a significant omission because the reference population problem is not merely an implementation detail — it is a conceptual choice that determines what the framework's performance levels actually mean. "Competent" relative to all literate adults is a very different claim from "Competent" relative to professional writers, and the framework provides no principles for making this choice across the diverse task categories an AGI benchmark would include. Future operationalization work would need to develop a methodology for reference population specification — including principles for when to use general population, skilled practitioner, or expert reference classes — that the current paper does not provide.
7. Implications and Future Directions
How This Work Changes the Landscape
This paper does not propose a new algorithm, train a model, or report empirical results. Its contribution is conceptual infrastructure — a shared coordinate system for discussions that had previously operated without one. The magnitude of change is therefore not "a new method outperforms baselines by X%" but rather "a conversation that was unstructured and often mutually incomprehensible now has a common vocabulary." This is a reframing contribution, not an empirical one. It changes how the community thinks rather than what it builds, and its impact will be measured by adoption — whether researchers, policymakers, and organizations start using "Emerging AGI" and "Competent Narrow AI" as meaningful categories in their own work — rather than by benchmark scores.
The most significant shift is the replacement of the binary AGI threshold with a graduated, two-dimensional matrix. Prior to this work, the dominant mental model for AGI was a switch: a system is either AGI or it is not, and the field's energy went into arguing about where to draw that single line. The paper demonstrates through its nine case studies that this binary framing created predictable pathologies — researchers talking past each other because they were using the same term for different capability thresholds, risk discussions oscillating between complacency ("AGI isn't here yet") and fatalism ("AGI might arrive suddenly and catastrophically"), and progress tracking that was essentially impossible because there was no way to say "we're 30% of the way there." The matrix dissolves these pathologies by providing a structured way to distinguish performance depth from generality breadth, to track asynchronous advancement on each axis, and to identify intermediate capability states that matter for deployment and risk assessment.
This reframing changes the nature of the AGI conversation from prediction (when will it arrive?) to assessment (where are we now, and what properties characterize the next level?). The prediction question has proven largely intractable — timelines vary by decades among serious researchers. The assessment question is more tractable, because it focuses on measurable current capabilities rather than extrapolation. The paper's framework provides the scaffolding for systematic assessment, even though the actual measurement instruments (benchmarks) remain to be built.
The framework resolves a specific, long-standing contradiction in the AGI literature. The paper's analysis reveals that prominent definitions of AGI were not so much disagreeing as operating at different points on spectra that the community had not yet articulated. Agüera y Arcas and Norvig (2023) argued that LLMs already are AGIs, emphasizing generality. Skeptics insisted that AGI requires reliable performance across tasks, emphasizing depth. The paper shows that both are partially right and partially wrong: LLMs have achieved generality at the Emerging performance level (Level 1 General AI), which is real progress, but they have not achieved the Competent performance level (Level 2) that most historical definitions of AGI implicitly assumed. The disagreement was not about facts but about definitions — and specifically about which dimension (generality vs. performance) to prioritize. By making both dimensions explicit axes in a single framework, the paper transforms a confusing debate into a coherent map where each position finds its place.
This reconciliation has practical consequences. It means that organizations can acknowledge the genuine breadth of frontier LLMs (they are "General" in a way that previous narrow systems were not) without overclaiming their reliability (they are "Emerging," not "Competent"). It means that safety discussions can calibrate to specific levels rather than to an undifferentiated "AGI" concept — the risks appropriate to Emerging AGI (misinformation, over-reliance) are different from those appropriate to Competent AGI (labor displacement at scale). And it means that the field can move past the sterile debate about whether current systems "are AGI" toward the more productive question of what capabilities they demonstrably have and what capabilities the next level requires.
The decoupling of capability from autonomy is a second, independent reframing that changes how risk is discussed. The paper's insistence that what a system can do (its AGI Level) and how humans choose to deploy it (its Autonomy Level) are separate design dimensions — correlated but not dictated — challenges the implicit assumption, widespread in AGI risk discourse, that more capable systems will naturally or inevitably be deployed with more autonomy. This decoupling has direct implications for safety governance: it means that capability advancement and deployment autonomy can be regulated separately, that safety measures can constrain autonomy without necessarily constraining capability research, and that the most severe risks (misalignment, power concentration) are specific to the combination of high capability and high autonomy, not to high capability alone. This is a more precise and actionable risk model than the undifferentiated "AGI risk" concept it replaces.
Research directions that become more attractive after this work:
-
Benchmark development for AGI levels becomes a first-order priority rather than an afterthought. The paper makes clear that the framework's utility depends on operationalization, and Section 5's benchmark blueprint provides design requirements that can guide that work. This shifts benchmark development from an ad-hoc activity (each lab building its own evaluation suite) to a community infrastructure project with clear design targets.
-
Metacognition research — particularly learning ability, calibrated help-seeking, and theory of mind — is elevated from an optional capability to a definitional requirement for generality. The paper argues that you cannot reach the "General" column without metacognitive capabilities, which redirects attention from simply scaling model size and training data toward building systems that can acquire new competencies and recognize their own limitations.
-
Human-AI interaction research is repositioned as a co-equal partner to model capabilities research. The paper's autonomy levels and its framing of interaction design as the mechanism that "ensures new AI systems are usable by and useful to people" (Section 6.2) implies that progress toward AGI requires advances in interface design, task specification, and evaluation support, not just in model performance.
Research directions that become less central:
-
Philosophical debates about consciousness, understanding, and "true intelligence" are explicitly sidelined by Principle 1 (Focus on Capabilities, not Processes). The paper argues that these are interesting philosophical questions but orthogonal to the practical task of measuring and governing AI progress. A researcher who insists that AGI requires consciousness is not contradicted — they are told that their concern belongs to a different conversation, and that the capability-based framework addresses the practical questions (deployment, risk, regulation) regardless.
-
Single-threshold AGI definitions are rendered obsolete by the framework's structure. If the community adopts the Levels of AGI, there is no longer a single "AGI threshold" to argue about — there are multiple levels, each with its own criteria, and the interesting questions are about which systems are at which level and what properties each level entails. The binary AGI debate is dissolved rather than resolved.
Follow-Up Research This Work Enables
A multi-stakeholder AGI benchmark developed through community process, with explicit task coverage, reference population definitions, and percentile thresholds. The paper's most significant gap is the absence of a benchmark to operationalize its framework. A strong follow-up would convene researchers from multiple organizations (addressing the paper's valid concern about cross-organizational perspectives), define an initial set of 50–100 cognitive and metacognitive tasks spanning linguistic, mathematical, spatial, social, and creative domains, specify reference populations for each task (e.g., "literate English-speaking adults" for essay writing, "adults with undergraduate mathematics training" for theorem proving), collect human performance distributions for those populations (likely through large-scale crowdsourcing with demographic controls), and then classify 5–10 prominent AI systems (frontier LLMs, narrow superhuman systems, mid-range models) against those benchmarks. The deliverable would be not just a classification table but a methodology paper describing how reference populations were defined, how percentile thresholds were computed, how the proportion-of-tasks question was resolved for the "General" classification, and what edge cases and ambiguities emerged. This would transform the framework from a proposal into a tool, and the edge cases it revealed (systems near the Narrow/General boundary, heterogeneous capability profiles) would directly inform refinement of the taxonomy.
A systematic analysis of asynchronous advancement between performance and generality using historical AI system data. The paper asserts that performance and generality are independent dimensions that can advance asynchronously, but supports this with only two data points (AlphaFold = high performance, low generality; ChatGPT = moderate generality, low performance). A rigorous follow-up would collect capability data on 20–30 historically significant AI systems (ELIZA, Deep Blue, Watson, AlphaGo, GPT-2, GPT-3, GPT-4, PaLM, Claude, Gemini, etc.), plot them on the 6×2 matrix at their respective release dates, and test whether the trajectory through the matrix follows a predictable pattern — do systems typically advance in generality first and then performance, or vice versa? Do narrow systems that later generalize (e.g., language models that started as text generators and acquired multimodal capabilities) follow a specific path through the matrix cells? This would validate or challenge the paper's claim that the two dimensions are genuinely independent, and would inform predictions about which direction future systems are likely to advance.
An empirical validation of the risk-to-autonomy-level mappings in Table 2 through systematic case studies of deployed AI systems. The paper asserts specific risks at specific autonomy levels (de-skilling at Level 1, over-trust at Level 2, anthropomorphization at Level 3, etc.) but provides no evidence. A validation study would select 3–5 deployed AI systems at each autonomy level, conduct systematic literature reviews or original user studies to measure the prevalence of the claimed risks, and test whether the risks are in fact specific to their claimed levels or appear across multiple levels. For example: do users of grammar checkers (Level 1, Tool) actually exhibit measurable de-skilling in writing ability over time? Do users of AI coding assistants (Level 2, Consultant) exhibit over-trust, defined as accepting incorrect AI suggestions without verification? Do users of social chatbots (Level 3, Collaborator) form parasocial relationships to a degree that measurably impacts their real-world social behavior? Results that confirmed the level-specificity of the claimed risks would strengthen the framework's utility for risk assessment; results that found risks bleeding across levels would indicate that the autonomy scale needs refinement (perhaps more levels, or risk categories that are not 1:1 with autonomy levels). Negative findings — e.g., no evidence of de-skilling from AI tools — would be equally valuable, suggesting that some widely-discussed risks may be overstated.
A "difficulty estimation" study that tests whether independent raters can reliably classify AI systems into the Levels of AGI matrix without a formal benchmark. The paper's classifications are based on author judgment. An inter-rater reliability study would recruit 10–20 AI researchers or knowledgeable practitioners who are not involved with the paper, provide them with the framework's definitions and a set of 10 AI systems (including both frontier LLMs and narrow systems), and ask them to independently classify each system into the 6×2 matrix and the Autonomy Levels. The outcome measures would be (a) Cohen's kappa or similar agreement statistic for the matrix classifications, (b) identification of which cells or which systems produce the most disagreement, and (c) qualitative analysis of the reasons raters give for their classifications (revealing which aspects of the framework are ambiguous in practice). Low inter-rater reliability would indicate that the framework cannot serve its stated goal of providing a "common language" without additional operationalization — it would mean that the vocabulary is shared but the referents are not. High reliability would provide preliminary evidence that the framework is usable as a qualitative assessment tool even pending full benchmark development. This study could be conducted now, without building a comprehensive benchmark, and would provide immediate feedback on the framework's practical clarity.
A "living benchmark" mechanism design specifying how new tasks are proposed, validated, and added to the AGI benchmark over time. The paper identifies that any AGI benchmark must be a living benchmark because the task space is infinite and a static test set would be gameable. A follow-up paper could propose and evaluate a specific governance mechanism for this: who can propose new tasks? What validation process ensures tasks are ecologically valid, not duplicative, and appropriately difficult? How are reference populations defined for novel tasks? How frequently are new tasks added, and how is the "proportion of tasks passed" threshold updated as the task set grows? The paper could test its proposed mechanism through simulation — generating a large synthetic task space, simulating AI systems with different capability profiles (some genuinely general, some overfit to existing tasks), and measuring whether the living benchmark mechanism correctly distinguishes general from narrow systems as new tasks are added. This addresses the "impossible to enumerate all tasks" problem not by solving it (which is impossible) but by demonstrating that a well-designed process for continuous task addition can make the benchmark practically sufficient even if formally incomplete.
A negative result study: attempt to classify systems that deliberately stress-test the framework's boundary conditions. The paper positions its framework as universally applicable, but tests it only on clear cases (clearly narrow systems like AlphaFold; clearly broad systems like ChatGPT). A stress-test study would construct or identify edge cases: a system with Superhuman performance on 40% of cognitive tasks and Emerging performance on 60% (where does it go in the matrix? The floor rule gives Emerging AGI, but the system is clearly more capable than a uniformly-Emerging system); a system that is clearly General in breadth but whose metacognitive capabilities are poor (does it fail the General column's definitional requirement?); a system that learns new tasks well but has no theory of mind (partial metacognition); a system deployed at Autonomy Level 3 that exhibits risks the framework claims are specific to Level 4. The study would not just classify these edge cases but would document where the framework produces ambiguous, contradictory, or clearly inadequate classifications, and would propose refinements — perhaps a graduated generality score, perhaps task-conditional capability profiles, perhaps a risk taxonomy that is not 1:1 with autonomy levels. Negative results (the framework fails to cleanly classify certain system types) would be as valuable as positive ones, because they would identify exactly where the taxonomy needs elaboration.
Practical Applications and Downstream Use Cases
Model card standardization across AI development organizations. The paper explicitly recommends that "documentation for frontier AI models, such as model cards, should detail this mixture of performance levels" (Section 4). A concrete deployment scenario: a major AI developer (Google DeepMind, OpenAI, Anthropic, Meta) adopts the Levels of AGI as the organizing framework for its model cards, reporting for each new system release (a) the overall AGI Level (e.g., "Level 1 Emerging AGI"), (b) the specific performance levels achieved on named task categories where the system exceeds the overall level (e.g., "Competent on short essay writing, Expert on Python code generation"), and (c) the metacognitive capabilities demonstrated and their limitations (e.g., "Demonstrates few-shot learning on novel tasks within language domain; calibration remains poor on mathematical reasoning, with overconfidence on incorrect answers"). This would transform model cards from narrative documents with heterogeneous content into standardized capability disclosures that enable direct comparison across organizations and over time. The benefit is not a number from the paper but a process change: model evaluation becomes structured around a shared taxonomy, making it harder for organizations to cherry-pick impressive benchmarks while obscuring weaknesses, and enabling policymakers to write regulations that reference specific AGI levels as triggers for additional oversight.
Regulatory threshold definition for AI governance frameworks. The paper's graduated levels provide a ready-made structure for defining regulatory triggers. A concrete scenario: a government AI safety body (e.g., the UK AI Safety Institute, the EU AI Office, or a future US equivalent) adopts the Levels of AGI to define which systems require which levels of pre-deployment testing, ongoing monitoring, and deployment restrictions. For example: "Systems classified as Level 1 Emerging AGI or below may be deployed with standard documentation requirements. Systems classified as Level 2 Competent AGI require third-party red-teaming and a published risk assessment before deployment. Systems classified as Level 3 Expert AGI require regulatory approval prior to deployment, with mandatory containment measures for any identified dual-use capabilities. Systems classified as Level 4 Exceptional AGI or above require international coordination and may be subject to deployment moratoria pending development of adequate safeguards." The Autonomy Levels provide a second dimension: even if a system is capable of Expert AGI, deployment at Autonomy Level 4 (Expert) or 5 (Agent) might require additional approvals beyond those required for deployment at Level 2 (Consultant). This regulatory structure is more nuanced than a single "AGI threshold" trigger, and it maps naturally onto the paper's framework. The benefit is that regulation can be proportionate — low-risk deployment modes for moderate-capability systems face lighter requirements than high-autonomy deployment of high-capability systems — and that the regulatory ladder can be climbed incrementally as capabilities advance, rather than activated all at once at some contested "AGI moment."
Investment and resource allocation decisions in AI research organizations. Organizations making multi-year research investments need to place bets on which capabilities will advance and when. The Levels of AGI framework provides a structured way to map those bets. A concrete scenario: a research lab's leadership uses the framework to assess that frontier LLMs have reached Level 1 (Emerging AGI) and that the primary bottleneck to reaching Level 2 (Competent AGI) is metacognitive capability (calibrated help-seeking, reliable learning of novel task types) rather than raw scale or additional static training data. This assessment redirects investment from scaling experiments (bigger models, more data) toward metacognition research (training procedures that produce calibrated uncertainty estimates, architectures that support efficient few-shot learning). Alternatively, a lab might assess that the path from Level 1 to Level 2 requires improved performance on specific task categories (mathematical reasoning, factual reliability) and invest accordingly in targeted data curation and fine-tuning for those categories. The framework does not tell organizations what to invest in, but it provides a shared language for articulating the investment thesis — "we are funding this research because it addresses the gap between Emerging and Competent AGI in the metacognition dimension" — that enables more structured portfolio planning and clearer communication with external stakeholders.
Public communication and expectation management around AI capabilities. A persistent challenge for AI organizations is communicating about system capabilities without overhyping (leading to public distrust when systems fail) or understating (leading to missed opportunities for beneficial deployment). The Levels of AGI framework provides a vocabulary for calibrated communication. A concrete scenario: when releasing a new model, an organization states "This system represents an advance from Emerging to Competent AGI, achieving at least 50th percentile of skilled adult performance on 65% of our internal evaluation tasks, up from approximately 30% in our previous release. It remains at Emerging performance on mathematical reasoning and factuality, and its metacognitive calibration is inconsistent — users should verify outputs in these domains. We are deploying it at Autonomy Level 2 (Consultant) for general use and at Level 1 (Tool) for high-stakes applications." This communication is simultaneously positive (genuine progress is acknowledged), honest (limitations are specified), and actionable (users know which domains to trust and which to verify). The benefit is that the public and policymakers receive a nuanced picture rather than binary "this is AGI" / "this is not AGI" claims, and that the framework's structure makes it harder to hide weaknesses — the matrix format naturally prompts disclosure of both performance level and generality breadth, and the expectation of uneven capability profiles makes it legitimate to acknowledge limitations without undermining the overall achievement.