ArXiv: 2510.12766

🎯 Pitch

Frequency of use—not abstract mental grammar—is the primary force shaping language, and LLMs' ability to synthesize coherent text from probabilistic analysis validates this empiricist thesis at scale. The critics' demand for 'deep structure' or 'grounding' is a theory-laden standard that generative grammar itself never passed, while LLMs succeed precisely because they master the totality of what is said and written. This reframes the stochastic parrot not as a deficiency but as a triumph of Mańczakian ornithology.


1. Executive Summary

This paper proposes a radical theoretical reorientation in how linguists evaluate language models, arguing for the empiricist principles of Witold Mańczak—who defines language as "the totality of all that is said and written" and identifies frequency of use as its primary organizing principle—over the abstract, dualistic frameworks inherited from de Saussure and Chomsky. Applying this Mańczakian lens, the paper reframes prior critiques of LLMs: the demand for "deep structure" is invalidated by the principle that synthesis must validate analysis (a test generative grammar never passed), the grounding objection is dismissed because meaning is predominantly relational and text-internal, and the poverty-of-stimulus argument is countered by decades of usage-based acquisition research showing children learn grammar through pattern recognition. The paper's central finding is that LLMs' capacity to synthesize coherent language from probabilistic analysis of text constitutes a large-scale validation of Mańczak's thesis—frequency is not peripheral but primary—and that critics' theory-laden standards interfere with useful analysis, establishing that LLMs are imperfect tools not because they fail to model language but precisely because they only model language as Mańczak defined it.

2. Context and Motivation

The Core Problem: Linguistics Lacks Validated Criteria for Evaluating Language Models

The fundamental gap this paper addresses is not a technical deficiency in LLMs but a methodological crisis in linguistics itself. The field, the paper argues, has no agreed-upon, empirically grounded criteria for determining what counts as valid linguistic knowledge or competence. This absence of validation standards becomes acutely visible—and practically consequential—when linguists attempt to evaluate whether LLMs "truly" model language.

The paper traces this crisis to what Mańczak identified as linguistics' deepest unexamined problem: the criteria of truth. As quoted in Appendix A:

"The fundamental problem of linguistics is that of the criteria of truth. Unfortunately, this problem is taboo. Given that linguistics has existed for two thousand years and that the Linguistic Bibliography recorded 21,000 works for the year 2001, it follows that linguists have published, in total, several hundreds of thousands of works, and yet none of these has been devoted to the criteria of truth."

This is not merely an academic concern. When prominent linguists publish critiques of LLMs—claiming they lack "deep structure" (Chomsky et al., 2023), cannot distinguish "between correctness and likelihood" (Fox and Katzir, 2024), or have no "access to meaning" despite experiencing only form (Bender et al., 2021)—they are applying criteria derived from specific, contestable theoretical frameworks. These frameworks, the paper contends, have never been independently validated through synthesis (reconstructing coherent language from their proposed components). The result is that evaluations of LLMs function less as objective assessments and more as defenses of a paradigm challenged by new empirical evidence.

Why This Problem Matters: Theoretical Paralysis Meets Practical Urgency

The significance of this gap operates on two levels.

Theoretical significance: A field unequipped to recognize its own validation. The paper argues that orthodox linguistics, dominated by Saussurean and Chomskyan frameworks, has spent decades developing increasingly elaborate theoretical constructs—deep structures, innate language organs, universal grammars, empty slots in phonological systems—without ever subjecting them to the test that Mańczak insisted upon: synthesis must validate analysis. This principle, drawn from the natural sciences, demands that any decomposition of language into theoretical components must be proven necessary by showing that those components are sufficient to reconstruct (synthesize) the original phenomenon.

The paper cites a devastating concrete example in Section 2.2. Chomsky, in Aspects of the Theory of Syntax, required approximately 10 pages to analyze the sentence "sincerity may frighten the boy." Yet to actually generate (synthesize) this sentence, the paper notes, Mańczak showed that only five simple positional rules are needed. The elaborate machinery of deep structure, transformations, and abstract underlying forms was not validated by synthesis—it was an analytical artifact that took on a life of its own. As Mańczak observed, the principle that synthesis validates analysis "was completely unknown to Chomsky."

This failure of synthesis extends to the entire generative enterprise. More than half a century after the "generative" turn, its adherents had "not yet written a single generative or transformational grammar of a concrete language" (Mańczak, 1996b). The roar of generative theory had produced not even a whisper of practical synthesis. When linguists needed to assess grammaticality, they searched corpora or polled native speakers—not consulted a generative rule system. They were, the paper argues, already operating in Mańczak's empirical world while their theories remained anchored in an unvalidated abstraction.

Practical urgency: LLMs have arrived as the synthesis that generative grammar never achieved. LLMs represent something unprecedented in the history of linguistics: a system that actually does synthesize coherent, contextually appropriate language at scale, and it does so through purely probabilistic analysis of text—precisely the kind of frequency-based approach Mańczak advocated. This is not a laboratory demonstration but a deployed technology producing millions of human-quality utterances daily.

The paper frames this as a moment of reckoning. When Bender et al. (2021) characterize LLMs as "stochastic parrots," they are applying criteria from a paradigm that never produced a working synthesis. The paper's response is pointed: "While Bender et al. advocate constraining the LLMs' tendency to act like 'stochastic parrots,' we call for a new science of ornithology that is equipped to understand what has actually taken flight" (Section 1). The parrot metaphor presupposes that mere statistical mimicry cannot constitute genuine linguistic competence—but this presupposition, the paper argues, is exactly what empirical evidence from LLMs challenges.

Where Prior Approaches Fall Short

The paper identifies specific deficiencies in the dominant frameworks that have been used to critique LLMs.

The Saussurean paradigm: Language as an abstract system. De Saussure's foundational move was to distinguish langue (the abstract system of signs) from parole (concrete utterances) and to declare the former the proper object of linguistic study. The paper notes a telling detail: "in the three hundred pages of [de Saussure's] Course in General Linguistics the term frequency of use does not appear even once" (Section 2.1, quoting Mańczak, 1969a). This absence is not incidental—it reflects a theoretical commitment to studying language as a static, idealized structure independent of its actual distribution in use.

The consequence, the paper argues, is a category error that persists in modern critiques of LLMs: the map is mistaken for the territory. The linguist's abstract analysis—the system of signs, the structural relationships, the rule systems—is treated as the functional prerequisite of language itself, rather than as a descriptive product of analysis. When critics then demand that LLMs demonstrate knowledge of these abstractions (e.g., "explain the rules of English syntax" as Chomsky et al. challenged), they are requiring that the model internalize the linguist's analytical framework, not that it demonstrate the capacity to produce and comprehend language. This conflates a particular descriptive apparatus with the phenomenon it describes.

The Chomskyan paradigm: Competence without performance validation. Chomsky's distinction between competence (idealized knowledge of language) and performance (actual use, with its errors and limitations) created a framework where the linguist's theoretical constructs could never be falsified by empirical data about actual language use. Linguistic theory aimed to characterize competence; performance data was, by definition, a degraded reflection of this idealized knowledge.

The paper argues this framework is fundamentally unfalsifiable in a way that violates Mańczak's criterion for scientific linguistics. When a theory posits an innate "language organ" or "deep structure" to explain language acquisition, but provides no synthesis that demonstrates these constructs are necessary or sufficient, it has failed the test that every other natural science requires: "synthesis validates analysis." The poverty-of-stimulus argument—that children cannot learn language from input alone because the input is too impoverished—is the key example. This argument was the primary justification for positing an innate language faculty, but as the paper notes (Section 2.3), it has been "challenged by evidence from cross-linguistic research and developmental psychology, causing many experts to abandon it." The usage-based alternative, which sees grammar as emergent from general cognitive tools applied to statistical patterns in input, aligns with Mańczak's position and has accumulated substantial empirical support from studies showing frequency effects at every level of language processing (Saffran et al., 1996; Romberg and Saffran, 2010; Ellis, 2002; Fló et al., 2025).

The grounding objection: Form without meaning. Bender and Koller (2020) and Bender et al. (2021) articulate a widely influential critique: LLMs manipulate linguistic form without access to meaning, because meaning is grounded in reference to the real world. An LLM trained purely on text, the argument goes, cannot truly understand what words mean—it only models their distributional patterns.

The paper addresses this through Mańczak's compositional semantics. Mańczak recognized that most words are defined relationally, through their connections to other words, with only a small set of axiomatic primitives serving as the foundation (Section 2.4). In this view, meaning for the vast majority of terms is their position in a relational network. An LLM that correctly uses "justice" does so because it has mastered the multidimensional web connecting that word to "fairness," "law," "equality," "crime," and other terms—not because it has grounded the concept in sensory experience.

The paper bolsters this with an analogy: "We do not demand that a calculator understands what '1+2' truly means to accept its utility. We do not dismiss the results of theorem-proving software because it cannot understand the philosophical basis of Zermelo-Fraenkel set theory" (Section 2.4). Advanced mathematics involves manipulation of formal systems according to rules; the "meaning" of an axiom lies in its role within that formal system. Demanding that LLMs meet a higher standard—requiring sensory grounding for linguistic meaning—is, the paper argues, an arbitrary requirement that does not reflect how linguistic meaning actually operates.

Furthermore, the paper notes that the grounding objection itself relies on a questionable assumption: that meaning must be grounded in physical reference. Piantadosi and Hill (2022) demonstrate that concepts without any possible referent—"perpetual motion machine," "king of San Francisco"—can exist meaningfully through purely relational networks. Mandelkern and Linzen (2024) argue that LLMs can use words to refer to real things because training texts already link words to the physical world; the question is whether the model counts as part of our speech community, not whether it has direct sensory access.

How This Paper Positions Itself

The paper's positioning is distinctive in several ways.

It is not a defense of LLMs on technical grounds but a reconstruction of the evaluative framework. Rather than arguing that LLMs satisfy existing linguistic criteria for competence—which would implicitly accept those criteria as valid—the paper argues that the criteria themselves are the problem. The Mańczakian framework is proposed not as an incremental refinement but as a "radical shift in perspective" (Abstract) that rejects the abstract, dualistic assumptions underlying prior critiques.

It draws on a figure largely unknown in computational linguistics but methodologically aligned with modern empiricism. Witold Mańczak (1924–2016) was a Polish historical and general linguist whose work anticipated many findings of usage-based linguistics by decades, reaching them through independent empirical analysis of textual corpora rather than through cognitive theory. The paper argues this makes his framework particularly suited to evaluating text-trained models: his focus on the structure of the input text itself—rather than on human cognitive processing—"aligns directly with how text-trained models actually function" (Section 4).

It identifies frequency as the bridge between Mańczak's linguistics and LLM architecture. The paper explicitly connects Mańczak's central thesis (frequency is the primary organizing force in language) to the mechanics of LLM training. Pretraining minimizes expected next-token surprisal, pushing the model's conditional predictions to match empirical next-token frequencies. As the training set grows, "estimation of the language's frequency structure improves and sharpens" (Section 2.2). This is not metaphorical—it is a direct mathematical correspondence between what Mańczak argued was linguistically primary and what LLMs optimize. The paper frames LLMs' success as "a large-scale validation of Mańczak's central thesis" (Section 2.2), citing the scaling laws literature (Kaplan et al., 2020; Hoffmann et al., 2022) as empirical evidence for the predictive power of frequency patterns.

It reframes "stochastic parrot" from criticism to description of what language actually is. The paper's most provocative move is to accept the empirical observation behind the "stochastic parrot" label—that LLMs model statistical patterns in text—but invert its normative force: "The 'stochastic parrots' do not merely mimic language but in fact reveal what language has been all along" (Section 3). In the Mańczakian framework, language is the totality of texts governed by frequency; an LLM that models this distribution is not a flawed approximation of language but a direct model of language itself. The limitation is not that LLMs fail to capture something essential about language—it is that they only capture what language is, without the grounding, intentionality, or world-models that critics demand but that Mańczak's definition excludes.

It acknowledges limitations explicitly while maintaining that the framework offers a constructive path forward. The paper's limitations section (Section 4) addresses that LLMs are trained on unrepresentative corpora ("very distant from 'the totality of all that is said and written'") and that the Mańczakian framework favors "principled, frequency-weighted corpus construction that reflects how language is actually used" as a corrective. It also acknowledges that meaning involves more than distributional relationships, but maintains that "the vast majority of the time 'meaning' can be inferred (and in the case of LLMs is inferred) solely from the relational structure of the text" (Section 4). This is presented not as a complete theory of semantics but as an empirical finding about how much linguistic competence distributional learning can achieve.

The paper thus positions itself at the intersection of theoretical linguistics, NLP, and philosophy of science: it uses the empirical success of LLMs to argue for a specific linguistic framework (Mańczakian) and uses that framework to reinterpret what LLMs are actually modeling. The ultimate claim is that the field needs, as the paper puts it in Section 1, "a new science of ornithology" that can study what has taken flight—a linguistics that treats text and frequency as primary data rather than as degraded reflections of an idealized abstract system.

3. Technical Approach

3.1 Reader Orientation

This paper is not building a computational system or model — it is a theoretical and methodological argument that proposes replacing the dominant framework linguists use to evaluate LLMs (abstract, rule-based, competence-focused) with an empiricist framework derived from the work of Witold Mańczak (text-focused, frequency-governed, synthesis-validated). The problem it solves is a meta-scientific one: current linguistic critiques of LLMs apply criteria from contestable theories that have never been independently validated, creating a situation where the field lacks rigorous, agreed-upon standards for assessing whether a system "models language"; the shape of the solution is to adopt Mańczak's definition of language ("the totality of all that is said and written") and his criterion of truth (statistical verification against observable textual distributions) as the evaluative framework, thereby reframing LLMs from "failed abstract systems" to "validated empirical models of language's primary organizing force — frequency."

3.2 Big-Picture Architecture (Diagram in Words)

The Mańczakian framework proposed in this paper has four interdependent components that together constitute an alternative paradigm for linguistic evaluation:

  1. The Language Definition Component — a naturalistic, operational definition that identifies language with the observable textual record rather than with an abstract system or mental faculty. Its input is the question "what is the object of linguistic study?" and its output is a boundary condition: any phenomenon not manifest in the totality of what is said and written falls outside linguistics proper.

  2. The Frequency Principle — the empirical claim that frequency of use is language's primary organizing force, governing everything from phonetic change to grammatical structure to the distinction between rules and exceptions. Its input is any linguistic phenomenon, and its output is a testable statistical hypothesis about how that phenomenon relates to occurrence counts in corpora.

  3. The Synthesis-Validates-Analysis Criterion — a methodological requirement that any decomposition of language into theoretical components (rules, structures, levels of representation) must be proven necessary by demonstrating that those components are sufficient to reconstruct the original linguistic phenomena. Its input is a proposed theoretical analysis, and its output is a pass/fail verdict: can the hypothesized components actually generate (synthesize) the language being analyzed?

  4. The Relational Semantics Component — a theory of meaning in which most words derive their semantic content from their position in a network of relations to other words, with only a small set of axiomatic primitives requiring external grounding. Its input is a claim about meaning, and its output is either (a) a decomposition into relational connections among terms, or (b) an identification of the term as axiomatic.

Information flows as follows: the Language Definition establishes what counts as linguistic data → the Frequency Principle generates testable predictions from that data → the Synthesis Criterion validates or rejects proposed analyses → the Relational Semantics component handles questions of meaning within the same text-internal framework.

3.3 Roadmap for the Deep Dive

  • First, the foundational definition: what Mańczak means by "language is the totality of all that is said and written," why this definition is radical in the context of Saussurean and Chomskyan linguistics, and what methodological consequences follow from it. This is the bedrock — every other component depends on it.
  • Second, the Frequency Principle and its corollaries: how frequency governs the grammar-lexicon continuum, the rule-exception distinction, and language change. This section explains the specific mechanisms by which frequency operates as an organizing force and connects them to LLM training objectives.
  • Third, the Synthesis-Validates-Analysis criterion: what it means operationally, how it exposes the failure of generative grammar, and why LLMs represent its vindication. This establishes the evaluative standard against which both linguistic theories and language models are measured.
  • Fourth, the Relational Semantics component and the grounding debate: how meaning operates in a text-internal framework, what work the axiomatic primitives do, and why the grounding objection to LLMs misidentifies what linguistic meaning requires.
  • Fifth, the framework's account of language acquisition and the analogy principle: how frequency-driven pattern recognition explains human language learning and why the same statistical mechanisms underlie LLM generalization through embedding-space relationships.

This order moves from ontology (what language is) to dynamics (how it changes) to validation (how we know we're right) to semantics (how meaning works) to acquisition (how it's learned), building a complete alternative paradigm layer by layer.

3.4 Detailed, Sentence-Based Technical Breakdown

This is a theoretical argument paper whose core idea is that adopting Mańczak's text-first, frequency-centric, synthesis-validated framework resolves the methodological crisis in evaluating LLMs by replacing unvalidated theoretical criteria with empirically grounded ones, and that LLMs' success constitutes a large-scale validation of Mańczak's thesis.


The Foundational Definition: Language as the Totality of Texts

The Mańczakian framework begins with an ontological claim that is simple in statement but radical in its implications for both linguistics and LLM evaluation. Mańczak defines language as:

"the totality of all that is said and written"

This definition is presented not as a theoretical insight but as a return to scientific fundamentals — a refusal to confuse the product of analysis with the object of study. The paper quotes Mańczak (1969a) directly on this point:

"The fundamental error of modern linguistics is the false equivalence between the product of a particular analysis and the object of the study itself."

What this definition rejects. To understand the definition's force, one must see what it deliberately excludes. De Saussure's foundational distinction between langue (the abstract system of signs shared by a speech community) and parole (individual utterances) privileged the former as the proper object of linguistics. Chomsky's competence-performance distinction similarly made the linguist's idealized reconstruction of a speaker's internalized grammar the target of study, treating actual utterances as degraded data contaminated by memory limits, attention shifts, and errors. Both frameworks locate language's essence in something invisible — a system, a competence, a structure — that is inferred from but not identical to observable speech and writing.

Mańczak's definition eliminates this dualism. Language is not a hidden system behind the observable data; language is the observable data in its entirety. There is no langue behind parole, no competence behind performance — there are only texts in their full distribution. Any claim about language that cannot be verified against this distribution is, in the Mańczakian framework, not a linguistic claim at all but speculation.

Methodological consequences. This definition carries immediate methodological requirements that the paper develops throughout Section 2. If language is the totality of texts, then:

  1. Linguistic investigation must be inductive and quantitative. The linguist's task is to discover patterns in the textual record through statistical analysis, not to impose abstract structures derived from theoretical commitments. The paper quotes Mańczak (1980, 1981, 1982, 1988a, 1996c) repeatedly on the need for "criteria of truth" — explicit, verifiable standards for distinguishing valid claims from invalid ones — and identifies statistics as the primary such criterion, with experiment as a secondary option for special cases.

  2. Frequency is necessarily primary. If language is the distribution of textual elements, then the most fundamental fact about any linguistic phenomenon is how often it occurs and in what contexts. Frequency is not a secondary, peripheral variable — it is the direct manifestation of language's structure. Any theoretical construct that ignores frequency is, by definition, ignoring the primary property of its ostensible object of study.

  3. The map must not be confused with the territory. A linguist's analysis — a grammar, a system of rules, a set of structural relations — is a description of patterns in the textual record. It is not the record itself, and it is certainly not a prerequisite for the record's existence. When critics demand that an LLM "explain the rules of English syntax," they are demanding that the model internalize a particular descriptive apparatus, not that it demonstrate the capacity that the apparatus describes.

The implicit critique of existing evaluation standards. The paper uses this definition to expose what it sees as a category error in standard linguistic critiques of LLMs. When Chomsky et al. (2023) claim LLMs fail to demonstrate linguistic competence because they don't "explain the rules," they are applying a criterion derived from a specific theoretical analysis (that language is governed by explicit rules that competent speakers can articulate) rather than from the observable phenomenon (that competent speakers produce and comprehend utterances). The Mańczakian response is that the ability to articulate rules is not a fact about language but a fact about a particular linguist's descriptive framework. The relevant test of an LLM's linguistic capacity is whether its output distributions match the distributions in the training corpus — i.e., whether it models the totality of texts — not whether it can reproduce a specific theoretical analysis of those texts.

This reframing is the paper's central methodological move. It changes the question from "does the LLM satisfy our theory of language?" to "does our theory of language accurately characterize what the LLM is doing when it successfully models textual distributions?" The answer the paper supplies is that Mańczak's theory does, while Saussure's and Chomsky's do not.


The Frequency Principle: Language's Primary Organizing Force

Once language is defined as the totality of texts, frequency emerges as the natural explanatory variable for virtually every linguistic phenomenon. The paper develops this principle across three domains: the grammar-lexicon relationship, the rule-exception distinction, and language change.

The grammar-lexicon continuum. The paper attributes to Mańczak a direct claim about the relationship between grammar and lexicon that dissolves what structuralist and generative traditions treat as a categorical distinction. The paper quotes Mańczak (1996h):

"grammar is 'the quintessence, condensation, abbreviation, or generalization of lexicography'"

In operational terms, this means that what linguists call "grammar" is simply the set of patterns that apply to high-frequency word classes — patterns that recur across so many lexical items that they can be abstracted and stated as general rules. What linguists call "lexicon" is the repository of information for all individual words, with lower-frequency items containing more idiosyncratic, item-specific information that resists generalization.

The paper makes this claim concrete. There is no categorical boundary where "lexicon" ends and "grammar" begins. Instead, there is a frequency continuum: at one extreme, patterns that apply to thousands of high-frequency words (e.g., English past tense -ed) are treated as grammar; at the other extreme, patterns specific to a single rare word (e.g., the idiosyncratic plural of a borrowed term) are treated as lexicon. Between these extremes lie patterns of intermediate generality — sub-regularities, semi-productive rules, constructional idioms — that traditional frameworks struggle to classify but that a frequency-based framework handles naturally as points on a continuum.

This has direct implications for LLM evaluation. An LLM that learns the distribution of word forms across contexts is not learning "surface statistics" separate from "grammatical rules" — it is learning the frequency patterns that constitute the grammar-lexicon continuum. The paper's claim is that Mańczak's framework predicts exactly what LLMs empirically achieve: mastery of both highly general patterns (grammar) and item-specific patterns (lexicon) through a single statistical learning mechanism applied to a sufficiently large and representative textual distribution.

Rules and exceptions as quantitative, not qualitative. Building on the continuum view, the paper extends the frequency principle to the distinction between "rules" and "exceptions." It states:

"high-frequency patterns are rules, while low-frequency patterns are exceptions"

This formulation is deceptively simple but carries substantial theoretical weight. In traditional generative grammar, rules and exceptions are categorically different: rules are generated by the computational system, while exceptions are listed in the lexicon and block rule application (e.g., the past tense of go is not *goed because the listed form went blocks the regular rule). In Mańczak's framework, this distinction collapses into a single quantitative dimension. A "rule" is simply a pattern with sufficiently high frequency that it dominates the distribution; an "exception" is a pattern with low frequency that coexists with the dominant pattern.

The paper connects this to Jaynes' (2003) probability theory by arguing that the classical logic of "true/false" statements is a special case of a more general probabilistic logic:

"a Mańczakian 'rule' corresponds to a high-plausibility inference, while an 'exception' represents a low-plausibility inference"

In operational terms, this means that a speaker (or an LLM) encountering a novel past-tense form does not consult a categorical rule system and then check for exceptions. It instead draws on the frequency distribution of past-tense formations in its experience, producing the most plausible form given the statistical evidence. The regularization of irregulars (e.g., dovedived) is predicted as the gradual increase in the plausibility of the regular pattern relative to the irregular one — a frequency effect, not a rule-change.

Language change as frequency-driven evolution. Perhaps the most vivid application of the frequency principle is to language change. The paper quotes Mańczak (1969a, "Critique du structuralisme"):

"the rules of grammar, once abstracted from texts, do not prevent the subsequent evolution of language. This evolution consists of errors which, if their frequency increases sufficiently, become new norms. Conversely, old norms, if their frequency diminishes, become errors."

This is a purely statistical theory of language change with no need for teleological or structural explanations. Language change has a simple mechanism:

  1. Speakers produce variants (including "errors" relative to current norms).
  2. Some variants increase in frequency (for reasons Mańczak attributes to various frequency-related pressures, including irregular phonetic development due to frequency).
  3. When a variant's frequency crosses some threshold, it becomes perceived as the norm.
  4. The old form, now declining in frequency, becomes perceived as an error or archaism.

The paper illustrates this with the evolution of Latin into Romance languages via Example [1], which provides three concrete cases:

  • (A) Analogical regularization of numerals: Latin had irregular ordinal numerals (ūnus "one" → prīmus "first"; duo "two" → secundus "second"). In Romance, high-frequency ordinals were gradually replaced by regularized forms (French deuxdeuxième rather than preserving the Latin irregular secundus). The paper presents this as a frequency effect: the analogical pattern (suffix -ième applied to the cardinal) increased in frequency until it became the norm.

  • (B) Divergent evolution depending on usage frequency: The paper shows that high-frequency Latin forms underwent irregular phonetic reduction (e.g., Latin hŏmo "man" → French on, a highly reduced, high-frequency indefinite pronoun), while lower-frequency cognates preserved more of their phonetic substance. This is Mańczak's theory of "irregular phonetic development due to frequency" — high-frequency items are phonetically eroded faster than low-frequency items because their predictability from context reduces the articulatory precision required.

  • (C) Grammaticalization from lexical verb to tense marker: Latin habēre "to have" was a lexical verb of possession. In Romance, it grammaticalized into a future/perfect auxiliary (French j'ai chanté "I have sung"), with phonetic reduction accompanying this functional shift (habēre → French avoir, but with forms like the future chanter-ai incorporating a reduced form of ai "I have"). The paper presents this as frequency-driven: a high-frequency lexical item gradually takes on grammatical function as its combination with other verbs becomes so frequent that it is reanalyzed as part of the tense system.

All three cases illustrate the same principle: frequency is not merely correlated with language change — it is the driver of change. The grammar of a language at any given moment is simply a snapshot of which patterns currently hold high frequency, and grammar evolution is the continuous redistribution of frequencies across variants.

The connection to LLM training. The paper draws an explicit mathematical connection between Mańczak's frequency principle and the LLM training objective. LLMs are trained to minimize expected next-token surprisal (cross-entropy loss):

L=1Tt=1TlogPθ(wtw<t)\mathcal{L} = -\frac{1}{T} \sum_{t=1}^{T} \log P_\theta(w_t \mid w_{<t})

where TT is the number of tokens in the training corpus, wtw_t is the token at position tt, w<tw_{<t} is the preceding context, and PθP_\theta is the model's predicted probability distribution over the vocabulary given that context.

Minimizing this objective pushes Pθ(wtw<t)P_\theta(w_t \mid w_{<t}) toward the empirical conditional frequency of wtw_t in context w<tw_{<t}. In other words, the model is learning to reproduce the frequency structure of the training corpus at every position, conditioned on prior context. This is not an analogy or a metaphor — it is a direct mathematical consequence of the training objective.

The paper supplements this with two empirical observations from the scaling laws literature:

  • Kaplan et al. (2020) and Hoffmann et al. (2022) show that LLM performance improves smoothly with pretraining data quantity, which is exactly what a frequency-based theory predicts: larger corpora provide more accurate estimates of the underlying frequency distribution, especially in the long tail of rare patterns.
  • Zhou et al. (2021) and Gong et al. (2018) show that LLMs naturally retain token frequency information in their embeddings, with the geometric relationships between embedding vectors reflecting co-occurrence statistics.

The paper's framing is that these are not incidental findings but direct confirmations of Mańczak's thesis: "LLMs' success is not a mystery or a 'stochastic parrot' trick. It is a large-scale validation of Mańczak's central thesis: language is text, and frequency is not a secondary, peripheral aspect, but its primary organizing force" (Section 2.2, emphasis added).


The Synthesis-Validates-Analysis Criterion

The frequency principle provides the explanatory framework, but the paper needs a validation standard — a way to determine whether a proposed linguistic theory or analysis is correct. Mańczak's criterion is stated early in Section 2.2:

"the principle that synthesis [reconstructing a coherent whole] is required to validate analysis [breaking things down into components] … should be applied to linguistics. Linguists who analyze text should look for only those components that are essential for its synthesis and must validate their analyses by means of synthesis."

This is presented not as a novel linguistic principle but as the application of standard scientific methodology to language. In the natural sciences, the test of whether you have correctly identified the components of a phenomenon is whether you can reconstruct the phenomenon from those components. If your analysis posits components X, Y, and Z as the building blocks of phenomenon P, you must demonstrate that X, Y, and Z are sufficient to produce P. If they are not, then either your components are wrong, incomplete, or include unnecessary elements.

What the criterion demands operationally. For linguistics, synthesis means generating or predicting the observable textual record from the proposed theoretical constructs. If a grammar posits a set of rules, those rules must be sufficient to generate the sentences that speakers actually produce and to distinguish grammatical from ungrammatical sequences. If a semantic theory posits a set of meaning primitives, those primitives must be sufficient to characterize the meaning distinctions that speakers actually observe.

Crucially, synthesis is the only validation that counts. The paper emphasizes that elaborate theoretical machinery, intuitive plausibility, and the agreement of authorities are all irrelevant if the proposed components cannot actually reconstruct the language. This is why the paper devotes so much attention to the failure of generative grammar to produce synthesis — it is not merely a rhetorical point but the application of the central methodological criterion.

The generative grammar failure as a worked example. The paper uses Chomskyan generative grammar as its primary case study of analysis without synthesis validation. The argument proceeds through several specific examples:

First, the paper cites Mańczak's observation about the scale mismatch between analysis and synthesis for even simple sentences. The sentence "sincerity may frighten the boy" required, in Chomsky's Aspects of the Theory of Syntax, approximately 10 pages of theoretical apparatus to analyze — including deep structure, surface structure, transformations, phrase structure rules, subcategorization frames, and selectional restrictions. Yet to actually generate this sentence from scratch, only five simple positional rules are needed (as illustrated in Example [2]):

  1. A sentence consists of a Noun Phrase followed by a Verb Phrase.
  2. A Noun Phrase can be filled by a determiner plus a noun, or by an abstract noun alone.
  3. A Verb Phrase can be filled by a modal auxiliary plus a transitive verb plus a Noun Phrase.
  4. The modal auxiliary "may" is selected.
  5. The specific lexical items "sincerity," "frighten," and "boy" occupy the appropriate slots.

The paper's point is not that Chomsky's analysis was wrong in its details, but that the enormous gap between the analytical complexity and the synthesis simplicity indicates that most of the analytical apparatus is unnecessary for the phenomenon it purports to explain. In Mańczak's terms, the components posited by the analysis (deep structure, transformations, etc.) are not validated by synthesis — they are artifacts of a particular analytical approach that has taken on a life of its own.

Second, the paper notes the broader failure of the generative enterprise to produce a working synthesis for any actual language:

"More than half a century after the 'generative' turn, its adherents had 'not yet written a single generative or transformational grammar of a concrete language' Mańczak (1996b). The roar of generative theory had produced not even a whisper of practical synthesis" (Section 2.2).

This is a devastating empirical observation. A theory that claims to explain how language is generated has, after 50+ years and thousands of publications, failed to actually generate any language. The paper treats this not as an incidental limitation but as a definitive falsification under the synthesis-validates-analysis criterion. A theory that cannot synthesize has not correctly identified the components of the phenomenon it studies.

Third, the paper observes that practicing linguists have already abandoned generative grammar's constructs when they need to actually assess linguistic data:

"when they needed to assess grammaticality, they searched corpora or polled native speakers instead of consulting a generative rule system" (Section 2.2).

This is the operational test in practice. If generative grammar's rules were actually sufficient to characterize grammaticality — if they passed the synthesis test — linguists would consult them to determine whether a novel sentence is grammatical. The fact that they use corpus search and native speaker polling instead is evidence that the rules are descriptively inadequate.

LLMs as the synthesis that generative grammar never achieved. The paper's central move in this section is to position LLMs as the answer to the unfulfilled promise of generative grammar. An LLM trained on a large textual corpus can actually synthesize coherent, contextually appropriate language at scale. It generates novel sentences, maintains discourse coherence, respects grammatical constraints, and adapts to different registers and styles — all without any explicit rules, deep structures, or transformations.

This is presented not as a competing theory of grammar but as an existence proof: synthesis of language is possible through purely probabilistic analysis of text. The components that Mańczak's framework identifies as essential — text distributions and frequency patterns — are sufficient to produce the phenomenon. The additional constructs that generative grammar posits — innate universals, deep structures, a language organ — are proven unnecessary by the existence of LLMs, just as they were already suspect under Mańczak's criterion because no generative grammarian ever demonstrated they were sufficient for synthesis.

The paper is careful not to claim that LLMs validate every aspect of Mańczak's specific linguistic analyses (his claims about Indo-European etymology, sound change patterns, etc.). What they validate is the meta-theoretical framework: the claim that language can be modeled as a frequency-governed distribution over texts, and that linguistic competence can be acquired through statistical learning over that distribution.

The design choice: why this criterion over alternatives. The paper implicitly compares the synthesis criterion against the two dominant alternatives in linguistics:

  • Authority-based validation: The paper's Appendix A opens with a scathing critique of how linguists actually operate — "X has formulated an opinion, X is an authority, therefore this opinion is true; Y has formulated an opinion, Y is not an authority, consequently this opinion is false." This is the "medieval and unscientific" criterion Mańczak identified as the field's unspoken practice.

  • Intuitive plausibility: Generative grammar's constructs are often defended on grounds of explanatory elegance or intuitive fit with how language "feels" to native speakers. The paper rejects this as subjective and unfalsifiable.

The synthesis criterion is proposed as the scientific alternative: it is objective (either the components can reconstruct the language or they cannot), falsifiable (a single failure of synthesis refutes the analysis), and aligned with standard practice in every other empirical science. The paper's argument is that linguistics, by failing to adopt this standard, has spent decades developing unfalsifiable theories while the actual empirical work of describing language was done through corpus analysis and native speaker consultation — methods that implicitly operate in the Mańczakian framework even when theorists disavow it.


Relational Semantics and the Grounding Debate

The most persistent criticism of LLMs — that they manipulate form without access to meaning because they lack grounding in the physical world — requires the Mańczakian framework to address semantics. The paper does so through a combination of Mańczak's own views on meaning and a defense of distributional semantics as sufficient for most linguistic meaning.

Mańczak's compositional semantics. The paper presents Mańczak's view of meaning as a form of compositional or reductionist semantics grounded in the necessity of undefinable primitives. Quoting Mańczak (1996h):

"Just as in mathematics, most but not all statements can be proved (with the help of other statements) … most words in a given language can be defined with the help of other words, with the unavoidable exception that [to avoid circularity] the meaning of certain words must be taken as self-evident (axiomatic)."

This is a two-tier theory of meaning:

  1. Axiomatic primitives: A small set of words whose meanings are taken as self-evident — they are not defined in terms of other words but serve as the foundation for all other definitions. The paper does not enumerate these primitives but the logic is clear: without them, all definitions would be circular (word A defined in terms of B, B in terms of C, and eventually C in terms of A).

  2. Relational definitions: The vast majority of words derive their meaning from their connections to other words in the semantic network. The meaning of "justice" is not a Platonic form or a sensory experience but the set of relationships connecting it to "fairness," "law," "equality," "rights," "punishment," "court," and thousands of other terms. Mastery of these relationships is mastery of the concept.

Mańczak explicitly rejected the alternative: attempting to describe language without any external reference (ignoring meaning entirely) was, in his words, a descent into "nebulous darkness" (Section 2.4). The axiomatic primitives provide the necessary anchor, preventing the semantic network from being a closed, uninterpreted formal system. But once that anchor is in place, meaning is predominantly relational.

The text-internal sufficiency argument. The paper extends Mańczak's compositional semantics into a specific claim about LLMs: they can achieve linguistic meaning through mastery of relational networks, without requiring direct sensory grounding for each term. The argument has several steps.

First, the paper claims that the relevant test for LLM quality is not access to an outside world but mastery of textual relationships:

"In the Mańczakian view, the relevant test for LLM quality is not whether the model has access to an outside world, but whether it has mastered the internal, relational logic of the textual world it was given" (Section 2.4).

This reframes the evaluation criterion. Instead of asking "does the LLM connect 'dog' to actual dogs?", the Mańczakian framework asks "does the LLM correctly deploy 'dog' in relation to 'animal,' 'bark,' 'pet,' 'leash,' 'canine,' and the thousands of other terms that constitute its relational meaning?" The answer, demonstrably, is that well-trained LLMs do this successfully enough to produce coherent, contextually appropriate text across a vast range of domains.

Second, the paper draws an analogy to formal systems:

"We do not demand that a calculator understands what '1+2' truly means to accept its utility. We do not dismiss the results of theorem-proving software because it cannot understand the philosophical basis of Zermelo-Fraenkel set theory" (Section 2.4).

The point is that much of human knowledge — particularly in technical and formal domains — already operates through mastery of relational systems rather than sensory grounding. A mathematician proving a theorem about transfinite cardinals is manipulating formal symbols according to rules, not grounding each symbol in physical experience. The "meaning" of an axiom in ZFC set theory is its role within the formal system — what it allows to be proved, what constraints it imposes — not a sensory referent. Demanding that an LLM ground "transfinite cardinal" in physical experience is demanding something that human practitioners of mathematics do not themselves possess.

Third, the paper invokes Piantadosi and Hill (2022), who demonstrate that concepts without any possible physical referent — "perpetual motion machine," "king of San Francisco" — can exist meaningfully through purely relational networks. These concepts are not meaningless; they have clear semantic content (a machine that violates thermodynamics, a monarch ruling a specific city). But that content derives entirely from the relational structure connecting them to other concepts, since no referent exists. If relational networks can support meaning for non-referential concepts, the paper implies, they can support meaning for referential concepts as well — the grounding is supplementary, not constitutive.

Fourth, the paper cites Mandelkern and Linzen (2024), who argue that LLMs can use words to talk about real things because the training texts already link those words to the physical world. In their view, the question is whether the LLM counts as part of our speech community — a community whose members use words to refer — not whether the LLM has direct sensory access. A human who learns about quarks purely from textbooks can still use the word "quark" referentially because they are part of a speech community where that word is linked (through chains of testimony, experiment, and theory) to physical phenomena. The paper suggests the same applies to LLMs.

The "axiomatic meaning" anchor and its limits. The paper is careful to acknowledge that the relational view does not eliminate the need for grounding entirely. Mańczak's axiomatic primitives — the small set of words whose meanings are taken as self-evident — represent the point where the relational network must connect to something outside itself:

"the meaning of certain words must be taken as self-evident (axiomatic)"

The paper does not specify what makes these primitives self-evident — whether it is innate concepts, early sensory experience, or ostensive definition in childhood — and explicitly states in the limitations section (Section 4) that it does not "claim to offer a complete theory of meaning." Instead, it makes a more limited claim: "the vast majority of the time 'meaning' can be inferred (and in the case of LLMs is inferred) solely from the relational structure of the text" (Section 4).

This is a pragmatic boundary, not a philosophical resolution. The paper essentially argues that while a complete account of meaning might require grounding for axiomatic primitives, this additional requirement does not invalidate LLMs as linguistic models because (a) relational meaning covers the vast majority of actual language use, (b) LLMs demonstrate empirically that relational meaning is sufficient for coherent linguistic behavior, and (c) the human grounding of axiomatic primitives does not constitute a linguistic theory that can be validated through synthesis.

The paper's use of the phrase "This page is intentionally left ungrounded" as the section heading (Section 2.4) signals its position: grounding is irrelevant to the linguistic assessment of LLMs, just as the standard boilerplate phrase announces that the absence of content is intentional.


Language Acquisition and the Analogy Principle

The paper's fourth component addresses how linguistic competence is acquired — both in humans and in LLMs — and uses this to counter the poverty-of-stimulus argument that has been the primary justification for positing innate linguistic knowledge.

Mańczak's rejection of the language organ. The paper reports that Mańczak explicitly rejected Chomsky's proposed "language organ" — a specialized, innate cognitive module dedicated to language acquisition — on methodological grounds. The argument was not that such a module definitively does not exist, but that debates about hypothetical brain structures fall outside the proper naturalistic focus of linguistics:

"Mańczak rejected this view, arguing that debates about hypothetical brain structures fell outside the proper naturalistic focus of linguistics" (Section 2.3, citing Grochowski, 2017 and Mańczak, 1969a).

This is consistent with the text-first definition. If language is the totality of what is said and written, then linguistics studies that observable record. Claims about unobservable mental structures are, in this framework, not linguistic claims — they belong to psychology, neuroscience, or philosophy of mind, and must be validated by the methods of those fields, not by linguistic argumentation.

Usage-based acquisition as empirical vindication. The paper devotes substantial attention to the usage-based theory of language acquisition that has emerged as the primary alternative to Chomskyan nativism. The key claims are:

  1. Children learn grammar from the ground up through general cognitive tools — pattern recognition, categorization, statistical learning — applied to the linguistic input they receive. This aligns with Mańczak's insistence that language is learned from observable texts (in this case, the speech children hear).

  2. Frequency is the primary driver of acquisition. The paper marshals extensive evidence from cognitive science:

    • Saffran et al. (1996): 8-month-old infants segment words from continuous speech using only transitional probabilities between syllables — a purely statistical cue.
    • Romberg and Saffran (2010): statistical learning operates across multiple levels of language, from phonology to syntax.
    • Ellis (2002): frequency affects every level of language processing — high-frequency patterns are processed faster, more accurately, and learned earlier.
    • Fló et al. (2025): statistical learning operates in human neonates even beyond word-level patterns.
  3. The poverty-of-stimulus argument has been empirically challenged. The paper cites Ibbotson and Tomasello (2016), Pullum and Scholz (2002), Christiansen and Chater (2008), and Tomasello (2003) as works that have caused "many experts to abandon" the nativist position. The implication is that the input available to children is richer than Chomsky claimed, and that general learning mechanisms are more powerful than nativists assumed — together, they can explain acquisition without positing innate domain-specific knowledge.

The paper frames this as Mańczak being vindicated by later empirical research: his skepticism about the language organ, grounded in methodological principles, has been supported by decades of developmental and cross-linguistic research that he did not live to see.

Analogy and generalization in LLMs. The paper's most technical argument about how the Mańczakian framework applies to LLMs appears in its discussion of analogy. The claim is that LLMs achieve generalization through the same mechanism that usage-based linguistics identifies in humans: recognizing and extending patterns by analogy.

The paper traces a historical progression in Table 2:

EraModelMechanismLimitation
Pre-2013n-gram modelsTables of memorized word sequencesCannot recognize that "Anna likes cats" and "Lily loves dogs" are analogically related — they are separate table entries
~2013CBOW (Continuous Bag of Words)Static word embeddingsAnalogous word pairs share similar geometric relationships (king:queen :: man:woman), but no sequential modeling
2017+TransformersSequence modeling over learned vector representationsOperates on sequences of embeddings, recognizing analogical relationships between entire sentences and structures

The critical insight is the shift from surface-level memorization to embedding-space relationships. An n-gram model treats "Anna likes cats" and "Lily loves dogs" as entirely separate sequences that happen to share some n-gram overlap. A Transformer, by contrast, maps both sentences into a high-dimensional space where their structural similarity (Subject-Verb-Object, with animate subject and animal object) is encoded in the geometry of their representations. When presented with a novel sentence like "Maria adopts rabbits," the model can draw on its internal map of relationships to generate or evaluate this sentence by analogy to the patterns it has already learned — not because it memorized this exact sequence, but because it maps to the same region of embedding space as structurally similar sequences.

This is what the paper means when it states:

"This ability to represent and manipulate relationships—the very essence of analogy—is the key to genuine linguistic generalization" (Section 2.3).

The paper is arguing that analogy, implemented through embedding-space geometry, is the bridge between Mańczak's frequency-based linguistics and LLM generalization. Frequency patterns in the training data create structure in the embedding space; that structure then enables generalization to novel sequences that share relational properties with frequent patterns. This is presented not as a metaphor but as a mechanistic description of what Transformers do: attention over sequences of learned vectors, with the training objective (next-token prediction) forcing the embedding space to encode whatever relational structure is predictive of token occurrences.

The design choice: why this view over nativist alternatives. The paper's position on acquisition is a direct consequence of its definition of language. If language is the totality of texts, then:

  • Language acquisition is the process of learning the frequency distribution over that totality.
  • The learning mechanism must be capable of extracting statistical regularities from experience.
  • The resulting knowledge is a representation of those regularities, not a system of explicit rules or innate principles.
  • Generalization to novel cases is achieved through analogy to learned regularities, with generalization quality dependent on how well the learned regularities capture the underlying structure of the target domain.

This is, the paper argues, exactly what LLMs do. They acquire the frequency structure of their training corpus through gradient-based optimization of a next-token prediction objective. They represent that structure in their embedding space and attention patterns. They generalize to novel sequences through the analogical relationships encoded in that embedding space. And they do all of this without explicit rules, innate universals, or a dedicated language module.

The paper's claim is not that this makes LLMs identical to human language learners — it explicitly acknowledges that comparisons to human cognition must "come after" the primary evaluation of LLMs as models of textual corpora (Section 4). The claim is rather that the Mańczakian framework provides a unified account of both human and machine language acquisition as frequency-driven statistical learning over textual distributions, with the differences between them being differences in architecture, data, and learning algorithm rather than differences in kind.


Summary of Design Choices and Their Justifications

  • Language defined as the totality of texts rather than as an abstract system or mental competence: eliminates the map-territory confusion that underlies critiques demanding LLMs demonstrate knowledge of theoretical constructs (deep structure, explicit rules) that are not validated by synthesis.
  • Frequency as the primary organizing force rather than a secondary variable: aligns the evaluative framework with the actual optimization objective of LLMs (minimizing next-token surprisal, which pushes predictions toward empirical frequencies) and with extensive psycholinguistic evidence about human language processing.
  • Synthesis-validates-analysis as the truth criterion rather than authority or intuition: provides an objective, falsifiable standard that generative grammar failed and LLMs passed, establishing that LLMs are not approximations of language but validated models of it.
  • Relational semantics with axiomatic primitives rather than universal grounding: accounts for how meaning operates in formal and abstract domains, explains why LLMs can use concepts correctly without sensory experience, and limits the grounding requirement to a small set of primitives that are outside the scope of linguistic theory proper.
  • Analogy through embedding-space relationships rather than explicit rule-learning: explains how statistical learning over distributions can produce genuine generalization to novel sequences, connecting Mańczak's theory of analogical language change to the actual mechanism of Transformer-based language models.
  • Text-internal evaluation as the primary standard rather than cognitive plausibility: establishes a clear baseline — an LLM is first and foremost a model of a textual corpus — before any comparison to human cognition, preventing the category error of evaluating a distributional model against criteria designed for theories of mental representation.

4. Key Insights and Innovations

Innovation 1: Reframing LLM Evaluation from "Do They Satisfy Our Theory?" to "Does Our Theory Explain What They Achieve?"

The paper's most fundamental intellectual move is not a new argument for or against LLMs' linguistic competence — it is a reversal of the burden of proof in the debate between linguists and language model practitioners. Prior critiques of LLMs (Chomsky et al., 2023; Bender et al., 2021; Fox and Katzir, 2024) take a specific theoretical framework (generative grammar, grounded semantics) as the independent standard and ask whether LLMs meet it. The paper inverts this: it takes the empirical fact of LLM synthesis as the independent standard and asks whether existing linguistic theories can explain it.

This is not merely rhetorical repositioning. It is a substantive methodological claim with a specific criterion: synthesis validates analysis. In every other empirical science, the paper argues, the test of whether you have correctly identified the components of a phenomenon is whether you can reconstruct the phenomenon from those components. If generative grammar's constructs (deep structure, transformations, the language organ) are necessary for linguistic competence, then they should be sufficient to synthesize language. The paper's central empirical observation — backed by the existence of deployed LLMs generating millions of coherent utterances daily — is that they are not necessary, because synthesis is achieved without them.

What makes this distinctive as a diagnostic move rather than just a rebuttal is that it changes what counts as evidence. Under the prior framing, an LLM that fails to articulate syntactic rules or demonstrate knowledge of deep structure is judged deficient. Under the Mańczakian framing, the same failure is irrelevant — what matters is whether the model's output distribution matches the distribution of actual language, which is an empirical claim about observable behavior rather than a theoretical claim about internal representations. The paper is not arguing that LLMs meet prior standards; it is arguing that the prior standards were the wrong standards, and that the test generative grammar failed (synthesis) is the one LLMs passed.

The innovation is thus meta-evaluative: it provides a principled, falsifiable criterion for assessing both linguistic theories and language models that is independent of any particular theoretical commitment. A theory that cannot synthesize has not correctly identified the components of language. A model that can synthesize has achieved linguistic competence by the only criterion that counts. The paper's reported finding that generative grammar produced "not even a whisper of practical synthesis" after 50+ years (Section 2.2, citing Mańczak, 1996b) while LLMs achieved it at scale makes this not an abstract philosophical point but a concrete empirical comparison.

This is a fundamental shift, not an incremental refinement. It does not tweak existing evaluative frameworks; it replaces their foundational criterion (theoretical coherence with an assumed model of competence) with one drawn from general scientific methodology (synthesis as validation of analysis). The consequence is that the entire corpus of prior linguistic critiques of LLMs is rendered question-begging under the new framework — they assume the very theoretical constructs whose necessity LLMs' existence calls into question.


Innovation 2: Frequency as the Bridge Between Linguistic Theory and LLM Architecture — Not an Analogy but a Mathematical Identity

Prior work in NLP has noted that LLMs learn statistical patterns (this is definitional — they minimize cross-entropy loss). Prior work in usage-based linguistics has argued that frequency shapes human language processing and acquisition (Ellis, 2002; Bybee, 2006; Tomasello, 2003). The paper's contribution is to connect these two literatures through a specific theoretical claim: the optimization objective of LLMs is not merely correlated with Mańczak's frequency principle but constitutes a direct implementation of it, and this implementation serves as a large-scale empirical validation of a linguistic theory that predates neural networks by decades.

The paper makes this connection precise. Mańczak's central claim — that frequency of use is language's primary organizing force, governing everything from the grammar-lexicon continuum to language change — predicts that a system which accurately estimates the frequency distribution of a language's textual record will display linguistic competence. LLM training minimizes L=1Tt=1TlogPθ(wtw<t)\mathcal{L} = -\frac{1}{T} \sum_{t=1}^{T} \log P_\theta(w_t \mid w_{<t}), which pushes PθP_\theta toward the empirical conditional frequencies in the training corpus. The mathematical consequence is that an LLM trained to convergence is literally a model of the frequency structure Mańczak identified as linguistically primary.

What distinguishes this from prior invocations of "statistical learning" in NLP is its theoretical specificity. The paper does not claim vaguely that "LLMs learn patterns" — it claims that Mańczak's specific hypotheses about how frequency operates (the grammar-lexicon continuum, irregular phonetic development due to frequency, analogical regularization, grammaticalization) provide a unified explanatory framework for LLM behavior that generative grammar cannot match. The empirical finding that LLM performance improves smoothly with pretraining data quantity (Kaplan et al., 2020; Hoffmann et al., 2022) is not presented as an independent observation but as a prediction of Mańczak's framework: larger corpora provide more accurate frequency estimates, and more accurate frequency estimates yield better linguistic performance.

The paper also leverages empirical findings about embedding geometry to strengthen this connection. Zhou et al. (2021) and Gong et al. (2018) show that LLMs retain token frequency information in their embeddings. The paper interprets this not as an architectural artifact but as evidence that the model's internal representations encode exactly the property Mańczak's theory says they must: the frequency structure of the training distribution.

This is an incremental advance in theoretical articulation rather than a new empirical finding — the scaling laws and embedding analyses are cited from prior work. But the articulation itself is fundamental because it converts a loose analogy ("LLMs are like usage-based models") into a specific theoretical claim with falsifiable consequences. If Mańczak were wrong and frequency were not the primary organizing force in language, then optimizing for frequency prediction would not produce linguistic competence. The existence of linguistically competent LLMs therefore constitutes evidence for Mańczak's thesis. The paper's framing of LLM success as "a large-scale validation of Mańczak's central thesis" (Section 2.2) is the novel claim.


Innovation 3: The "Stochastic Parrots" Inversion — From Pejorative to Definitional

The term "stochastic parrot" (Bender et al., 2021) has become one of the most influential framings in public and academic discourse about LLMs, carrying the implication that statistical mimicry of textual patterns is a deficient, shallow form of language processing that falls short of genuine understanding. The paper's most rhetorically effective and philosophically provocative move is to accept the empirical observation behind this label while inverting its normative force: the "parrots" reveal what language actually is, not what it fails to be.

This innovation operates at the level of conceptual reframing rather than technical contribution. The paper's argument is that if language is defined in the Mańczakian way — as the totality of texts, governed by frequency — then modeling the statistical distribution of those texts is not an approximation of language but a direct model of language itself. The limitation of LLMs is not that they fail to capture something essential about language (deep structure, grounded meaning, intentionality), but precisely the opposite: they capture only what language is, without the additional capacities (world-modeling, sensory grounding, communicative intent) that critics demand but that Mańczak's definition excludes.

The paper makes this point explicitly in its concluding paragraph: "LLMs are imperfect tools not because they fail to model language but because they only model language" (Section 3). This is a diagnostic insight disguised as wordplay: it separates the question "do LLMs model language?" (answer: yes, directly) from "is language sufficient for all the things we want AI to do?" (answer: no, clearly not). Many prior critiques conflate these two questions — they observe that LLMs produce factual errors, lack common sense, or fail at grounded reasoning, and conclude that LLMs don't truly model language. The paper argues that these failures are failures of something beyond language (world knowledge, reasoning, intentionality), not failures of linguistic modeling.

This innovation draws force from the Mańczakian framework's specific definition of language. If one accepts that language is the textual record and frequency is its organizing principle, then "stochastic parrot" becomes not an insult but a description — the parrot is stochastic because language is probabilistic (the rule-exception distinction is quantitative, per Section 2.1), and it parrots because language is a distribution to be learned rather than a system to be deduced. The paper's call for "a new science of ornithology that is equipped to understand what has actually taken flight" (Section 1) is the rhetorical encapsulation of this reframing: stop diagnosing the bird for failing to match your theory of flight, and start studying how it flies.

This is a fundamental reframing of the public debate, not an incremental refinement of evaluation metrics. It does not provide new empirical evidence about LLM capabilities but changes the interpretive framework through which existing evidence is understood. Its significance lies in its potential to redirect critical discourse: rather than asking whether LLMs satisfy criteria derived from contested theories, the relevant question becomes whether those criteria correctly characterize the phenomenon that LLMs have empirically demonstrated can be modeled through frequency distributions.


Innovation 4: The Category Error Diagnosis — Mistaking the Linguist's Map for the Territory

The paper identifies and names a specific logical mistake that it argues underlies virtually all theory-laden critiques of LLMs: the confusion between a linguist's analytical framework (the map) and the phenomenon that framework describes (the territory). This is not the trivial observation that "the map is not the territory" — it is a specific claim about where and how this confusion manifests in LLM evaluation, with a concrete criterion for detecting it.

The diagnosis works as follows. When a linguist analyzes a language, they produce a descriptive apparatus: a grammar, a system of rules, a set of structural relations, a theory of deep and surface structure. This apparatus is a product of analysis applied to the textual record (or to native speaker intuitions about that record). It is not itself a component of the language — it is a description. The category error occurs when critics then demand that an LLM demonstrate knowledge of this descriptive apparatus as evidence of linguistic competence. When Chomsky et al. (2023) challenge LLMs to "explain the rules of English syntax," they are requiring that the model internalize a specific linguist's analytical product, not that the model demonstrate the capacity that the analysis describes.

The paper makes this diagnosis concrete through the synthesis-validates-analysis criterion. If the descriptive apparatus (the rules, the deep structures) were actually necessary for language — if they were components of language rather than analytical artifacts — then they would be sufficient for synthesis. One could take the rules and deep structures and generate the language. The paper's central empirical claim, illustrated through Mańczak's critique of generative grammar, is that this synthesis has never been achieved. The 10-page analysis of "sincerity may frighten the boy" in Chomsky's Aspects was not validated by demonstrating that the posited components could generate that sentence — instead, as Mańczak showed, only five simple positional rules were needed (Example 2). The elaborate theoretical machinery was unnecessary for the phenomenon it purported to explain.

The significance of this as a diagnostic innovation is that it provides a general, transferable test for evaluating any linguistic critique of LLMs. When a critic demands that an LLM demonstrate some property X (knowledge of rules, grounding, deep structure), one can ask: (a) Is X derived from a theoretical analysis of language, or from the observable textual record? (b) If theoretical, has that analysis been validated through synthesis — i.e., can the posited components actually generate language? (c) If not, is the demand for X a requirement that the LLM internalize the linguist's analytical framework rather than demonstrate linguistic capacity?

The paper's answer to (c) in most cases is yes. It argues that debates about LLM competence have been proxy battles over theoretical frameworks, with critics applying criteria from their preferred framework and finding LLMs deficient by those criteria, without first establishing that the framework's constructs are necessary for (rather than merely descriptive of) the phenomenon. This is a fundamental conceptual advance because it identifies a pattern of reasoning error that is systematic rather than occasional, and provides a principled criterion (synthesis validation) for detecting it.

This innovation builds on the work of Mańczak (1969a, 1996b) but applies it to a contemporary debate that Mańczak never addressed directly. The paper's contribution is not the category error concept itself but its application as a diagnostic tool for evaluating LLM critiques, combined with the empirical claim that this error is pervasive in the criticism literature. The finding is not that LLMs satisfy some standard of competence but that many of the standards applied to them are artifacts of unvalidated theoretical analysis — the map mistaken for the territory.

5. Experimental Analysis

Evaluation Methodology

Dataset. The paper does not introduce a new dataset, as it is a theoretical argument paper rather than an empirical study of model performance. It references existing empirical findings from the LLM scaling laws literature, including the works of Kaplan et al. (2020) and Hoffmann et al. (2022), which use standard language modeling benchmarks (e.g., The Pile, C4, Wikipedia, Books) to evaluate how model performance scales with data quantity. It also cites work on embedding geometry and frequency retention (Zhou et al., 2021; Gong et al., 2018), which use standard word embedding benchmarks. The paper does not report new experimental results on any dataset.

Base model(s). The paper does not train or evaluate specific models. It discusses LLMs generically, with references to GPT, Claude (Section 1), and the broader class of Transformer-based language models trained with next-token prediction objectives (Section 2.2–2.3). It also references earlier architectures (n-gram models, CBOW) in Table 2 to illustrate the historical progression toward embedding-based analogy. The paper's arguments apply to the class of models that minimize expected next-token surprisal over text corpora; no specific model scale, family, or checkpoint is analyzed.

Metrics. The paper does not introduce or compute quantitative metrics in the traditional sense. The central evaluative criterion is synthesis validation: can a system (a linguistic theory or a language model) actually generate coherent, contextually appropriate language? For linguistic theories, the metric is binary — has the theory produced a working generative grammar of any concrete language? (The paper reports this as "no" for generative grammar after 50+ years, citing Mańczak, 1996b.) For LLMs, the implied metric is the quality of generated text, but the paper does not operationalize this through scores (e.g., perplexity, BLEU, human evaluation). It instead relies on the reader's shared knowledge that deployed LLMs produce contextually appropriate language at scale, treating this as an accepted empirical fact rather than a measured outcome. The paper cites scaling laws from prior work (Kaplan et al., 2020; Hoffmann et al., 2022) — which use cross-entropy loss as the primary metric — and embedding analyses (Zhou et al., 2021; Gong et al., 2018), but does not compute these metrics itself.

Baselines. The paper's "baselines" are not computational baselines but theoretical frameworks against which it compares the Mańczakian approach:

  • Generative grammar (Chomskyan linguistics). The paper claims this framework has produced "not a single generative or transformational grammar of a concrete language" despite 50+ years of research (Section 2.2, citing Mańczak, 1996b), and that its analyses of even simple sentences require vastly more theoretical machinery than is needed for synthesis (Example 2, comparing Chomsky's 10-page analysis of "sincerity may frighten the boy" to the five positional rules sufficient to generate it).

  • Saussurean structuralism. The paper notes that de Saussure's Course in General Linguistics — roughly 300 pages — does not mention "frequency of use" even once (Section 2.1, citing Mańczak, 1969a), and argues this absence reflects a systematic neglect of the property Mańczak identified as primary.

  • Grounding-based semantics (Bender and Koller, 2020; Bender et al., 2021). The paper argues that this framework imposes a requirement (sensory grounding for meaning) that is unnecessary for most linguistic meaning, as demonstrated by (a) relational concepts without referents (Piantadosi and Hill, 2022), (b) formal systems where meaning is role-based rather than referent-based (the calculator/mathematics analogy in Section 2.4), and (c) the empirical success of LLMs in using words correctly without direct sensory access.

  • The "stochastic parrot" framing (Bender et al., 2021). The paper treats this as a normative claim — that statistical mimicry is insufficient for genuine linguistic competence — and argues that under the Mańczakian definition of language, this mimicry is precisely what linguistic competence consists of.

These baselines are conceptual rather than quantitative. The paper's claim is not that the Mańczakian framework achieves higher scores on some metric than these alternatives, but that it provides a more coherent and empirically validated framework for evaluating language models because its criteria (text distributions, frequency, synthesis) align with what LLMs actually do.

Generation budget / compute accounting. The paper does not involve generation budgets or compute accounting in the experimental sense. Its argument about frequency and data scale relies on the scaling laws literature (Kaplan et al., 2020; Hoffmann et al., 2022), which measures training compute in FLOPs (floating point operations) and relates model performance (cross-entropy loss) to model size, data quantity, and compute budget. The paper cites these works to support its claim that "LLM performance increases smoothly with the amount of pretraining data" (Section 2.2), but does not conduct its own scaling experiments. The "budget" in the paper's argument is the quantity of textual data available for training, measured in tokens — larger corpora provide more accurate frequency estimates, which the Mańczakian framework predicts will yield better linguistic performance.

Cross-validation / statistical protocol. The paper does not employ cross-validation or statistical testing in the traditional sense, as it reports no new experimental results. Its argumentation relies on:

  • Historical test cases from comparative linguistics: Example 1 illustrates frequency-driven language change through three concrete cases from the evolution of Latin into Romance languages (analogical regularization of numerals, divergent evolution depending on usage frequency, and grammaticalization of habēre into a tense marker). These serve as qualitative evidence for the frequency principle but are not systematically evaluated.

  • Citation of prior experimental work: The paper marshals extensive citations to developmental psychology and cognitive science (Saffran et al., 1996; Romberg and Saffran, 2010; Ellis, 2002; Fló et al., 2025; Ibbotson and Tomasello, 2016; Pullum and Scholz, 2002; Christiansen and Chater, 2008; Tomasello, 2003) as evidence for usage-based language acquisition and against the poverty-of-stimulus argument. These studies use standard experimental protocols (infant looking-time measures, reaction time studies, corpus analyses), but the paper does not replicate or extend them.

  • Analogical reasoning through examples: Example 2 illustrates Mańczak's claim that synthesis requires far fewer components than generative analysis posits, using the "sincerity may frighten the boy" sentence and its five positional rules. Table 2 traces the historical progression from n-gram models to Transformers, using the analogy between "Anna likes cats" and "Lily loves dogs" to illustrate how embedding-space geometry enables generalization. These are illustrative, not experimental.

  • Appeal to the existence proof of deployed LLMs: The paper's central empirical claim — that LLMs achieve synthesis — is supported by reference to the observable fact that systems like GPT and Claude "synthesize complex responses" (Section 1) and increasingly "draft our contracts, structure our arguments" (Section 3). No formal evaluation or benchmark results are provided. The reader is expected to accept, based on common experience with deployed systems, that these models produce coherent, contextually appropriate language at scale.

Main Quantitative Results

This paper does not report primary quantitative results of its own in the traditional sense — there are no tables of accuracy scores, no comparison of different model configurations on a benchmark, no statistical tests. The "results" are instead empirical claims about the state of the field supported by citation to prior work and by qualitative case analyses. I organize these by the paper's three major evidentiary claims.

Claim 1: Generative Grammar Has Failed the Synthesis Test

The paper's primary empirical finding about linguistics itself is that generative grammar has not produced a working synthesis of any concrete language after more than 50 years of research. This claim is supported by:

Citation to Mańczak (1996b) in Section 2.2, which the paper summarizes as the observation that "More than half a century after the 'generative' turn, its adherents had 'not yet written a single generative or transformational grammar of a concrete language.'"

The illustrative analysis in Example 2, which contrasts Chomsky's 10-page analysis of "sincerity may frighten the boy" in Aspects of the Theory of Syntax with the five positional rules Mańczak identified as sufficient for synthesis:

  1. A sentence consists of a Noun Phrase followed by a Verb Phrase.
  2. A Noun Phrase can be filled by a determiner plus a noun, or by an abstract noun alone.
  3. A Verb Phrase can be filled by a modal auxiliary plus a transitive verb plus a Noun Phrase.
  4. The modal auxiliary "may" is selected.
  5. The specific lexical items "sincerity," "frighten," and "boy" occupy the appropriate slots.

The paper's claim is that this 10-page vs. 5-rule disparity demonstrates that most of generative grammar's theoretical apparatus (deep structure, transformations, subcategorization frames) is unnecessary for the phenomenon it purports to explain — it is analytical artifact, not validated component.

The operational test from practice: linguists, when needing to assess grammaticality, "searched corpora or polled native speakers instead of consulting a generative rule system" (Section 2.2). This is presented as evidence that the rule system is descriptively inadequate — if it worked, linguists would use it.

Critical assessment of this claim: The paper does not provide a systematic survey of generative grammar's synthesis attempts. Mańczak (1996b) is cited as the source of the "not a single grammar" claim, but the paper does not reproduce Mańczak's evidence or address whether subsequent work (between 1996 and the present) has produced synthesis efforts. The 10-page vs. 5-rule comparison is illustrative, not systematic — it addresses a single sentence from a 1965 work, not the full range of generative analyses. The claim that linguists use corpora and native speakers rather than rules is plausible but not empirically demonstrated in the paper. There is no survey of practicing linguists' methods. This is a qualitative, argumentative claim rather than an experimental result.

Claim 2: LLMs Validate the Frequency Principle Through Their Training Objective and Scaling Behavior

The paper's central empirical argument about LLMs is that their success constitutes "a large-scale validation of Mańczak's central thesis: language is text, and frequency is not a secondary, peripheral aspect, but its primary organizing force" (Section 2.2). This claim is supported by three lines of cited evidence, none original to this paper:

Connection between training objective and frequency estimation. The paper notes that LLM training minimizes expected next-token surprisal, which pushes predicted probabilities toward empirical conditional frequencies in the training corpus. This is a mathematical property of cross-entropy minimization, not an experimental finding. The paper presents it as evidence of alignment between Mańczak's frequency principle and LLM architecture, but it is a definitional observation rather than an empirical test — it follows from the choice of loss function, not from measured outcomes.

Scaling with data quantity (Kaplan et al., 2020; Hoffmann et al., 2022). The paper cites these works as showing that LLM performance improves smoothly with pretraining data quantity. The specific finding from Kaplan et al. (2020) is that test loss scales as a power law with model size, dataset size, and compute; from Hoffmann et al. (2022), that optimal performance at a given compute budget requires scaling both model size and data quantity. The paper interprets these smooth improvements as evidence for the frequency principle: "Estimation of the language's frequency structure improves and sharpens with a larger training set, especially in the long tail" (Section 2.2).

Critical assessment of this claim: The scaling laws evidence is consistent with a frequency-based theory, but it is not uniquely predicted by it. Any theory that expects performance to improve with more data (which is virtually all theories of learning) would also be consistent with scaling laws. The question is whether the shape of the improvement — power-law scaling of loss with data quantity — is specifically predicted by Mańczak's framework. The paper does not derive quantitative predictions from Mańczak's theory and compare them to the observed scaling exponents. The inference is at the level of qualitative consistency, not quantitative test. Furthermore, scaling laws measure perplexity on held-out text, which is a direct measure of how well the model predicts the frequency distribution. That modeling frequencies better improves frequency prediction is close to tautological. The more interesting claim — that modeling frequencies produces linguistic competence in a richer sense (generating novel, coherent text, not just matching held-out distributions) — is supported by the existence of deployed LLMs but not by the scaling laws themselves.

Frequency retention in embeddings (Zhou et al., 2021; Gong et al., 2018). The paper cites these works as showing that models "naturally retain the useful token frequency information in embeddings" (Section 2.2). Zhou et al. (2021) demonstrate that contextualized embeddings encode token frequency, and that this can cause frequency-based distortions in downstream tasks. Gong et al. (2018) propose a method to make embeddings frequency-agnostic, implying that frequency is indeed encoded in standard embeddings. The paper interprets this as evidence that the model's internal representations encode the property Mańczak's theory says they must.

Critical assessment: The embedding retention finding is again consistent with the frequency principle but not a specific test of it. It would be surprising if embeddings did not encode frequency information, given that the training objective optimizes for frequency prediction. The paper uses this as corroborating evidence rather than as a decisive test.

Overall assessment of Claim 2: The paper's argument that LLMs validate Mańczak's thesis is more accurately described as theoretical alignment than experimental validation. The paper shows that Mańczak's framework predicts that frequency modeling should produce linguistic competence; it observes that LLMs do frequency modeling and produce linguistic competence; it infers that LLMs confirm the prediction. However, this inference does not rule out alternative theories that would also predict competence from frequency modeling, nor does it test Mańczak's more specific claims (e.g., about irregular phonetic development due to frequency, or about the specific pressures that drive analogical change) against LLM behavior. The "validation" is of the meta-theoretical framework (language can be modeled as a frequency-governed distribution) rather than of Mańczak's specific linguistic analyses.

Claim 3: Usage-Based Acquisition Is Empirically Supported, Undermining the Poverty-of-Stimulus Argument

The paper marshals extensive citations to support the claim that children learn language through statistical learning over input rather than through an innate language organ. This is not original research but a literature review making the case that Mańczak's skepticism about the language organ has been empirically vindicated.

Cited evidence includes:

  • Saffran et al. (1996): 8-month-old infants segment words using transitional probabilities.
  • Romberg and Saffran (2010): Statistical learning operates across multiple linguistic levels.
  • Ellis (2002): Frequency affects processing speed, accuracy, and acquisition order.
  • Fló et al. (2025): Statistical learning in human neonates beyond word-level patterns.
  • Ibbotson and Tomasello (2016): "Evidence rebuts Chomsky's theory of language learning."
  • Pullum and Scholz (2002): Empirical assessment challenges stimulus poverty arguments.
  • Christiansen and Chater (2008): Language is shaped by the brain, not vice versa.
  • Tomasello (2003): Usage-based theory of language acquisition.

The paper frames this body of work as having "caused many experts to abandon" the Chomskyan nativist position (Section 2.3).

Critical assessment: This is a literature summary, not an experimental finding. The paper does not weigh the evidence for and against usage-based acquisition, address critiques of the statistical learning paradigm, or acknowledge that the debate between nativist and empiricist accounts of language acquisition remains active in developmental psychology and linguistics. The citations are genuine — these works do exist and do support usage-based accounts — but the paper's presentation is that of a settled debate ("many experts" having abandoned nativism), which is stronger than the field's actual consensus. The paper does not mention work defending nativist positions or challenging the interpretation of statistical learning studies.

Ablation Studies and Robustness Checks

This paper, being a theoretical argument rather than an empirical study, does not contain ablation studies in the standard sense (no hyperparameter sweeps, no component removal experiments, no sensitivity analyses). However, it does engage in several forms of argumentative robustness checking — addressing potential objections and alternative interpretations of its claims.

Addressing "LLMs are trained on unrepresentative corpora": The Limitations section (Section 4) acknowledges that LLMs are "trained on unrepresentative corpora, very distant from 'the totality of all that is said and written.'" This is a significant concession, as the Mańczakian definition of language requires the totality of texts, not an arbitrary sample. The paper's response is that this limitation is practical rather than fundamental: "The Mańczakian framework offers a clear path forward, favoring 'rational selection of texts' based on circulation and influence Mańczak (1996f, 1961). A truly Mańczakian approach is a principled, frequency-weighted corpus construction that reflects how language is actually used." No empirical evidence is provided that such a principled corpus would improve LLM performance, nor is the claim tested against existing corpus construction methods.

Addressing "LLMs lack a world model": The paper addresses this objection in Section 4, arguing that "language is not a map of the physical world but a self-contained universe of texts" and that "the demand for grounding in physical reality is misplaced." This is a theoretical rebuttal, not an ablation — the paper does not compare the performance of LLMs with and without grounding, or test whether adding grounding improves performance. The argument is that the demand for a world model misunderstands what language modeling is: "Our model should be seen not as a flawed attempt to simulate a mind interacting with a physical world but as a successful and direct model of language itself."

Addressing "meaning requires more than relational structure": The paper explicitly states in the Limitations section that it does "not claim to offer a complete theory of meaning" and that its central thesis is "that the vast majority of the time 'meaning' can be inferred (and in the case of LLMs is inferred) solely from the relational structure of the text." This is a scope limitation rather than an empirical finding. The paper does not quantify what proportion of meaning is relationally inferable versus requiring grounding, nor does it test whether LLM failures on specific tasks (e.g., factual accuracy, common-sense reasoning) are attributable to missing grounding versus insufficient relational learning.

Addressing "Mańczak's ideas are already part of usage-based linguistics": The paper acknowledges this overlap in the Limitations section and provides a two-part differentiation argument:

  1. Mańczak reached similar conclusions "decades earlier through the empirical analysis of contemporary and historical texts" — a historical priority claim.
  2. "His radical, text-only simplicity provides a direct rebuttal to critiques of LLMs grounded in Saussurean or generative theories: his framework simply discards abstractions that cannot be found in the textual record" — a claim of greater theoretical parsimony.
  3. "His focus on the structure of the input text itself—rather than human cognitive processing—aligns directly with how text-trained models actually function" — a claim of better fit to the target phenomenon.

These are argued, not tested. The paper does not compare Mańczak's framework to, say, Goldberg's (2024) usage-based constructionist account (which it cites approvingly) to determine which makes more precise predictions about LLM behavior.

Addressing "creativity requires more than pattern application": The paper's response in the Limitations section is brief: "High-frequency patterns don't mechanically reproduce themselves; they serve as templates for novel combinations. LLMs demonstrate this empirically. Creativity isn't the opposite of pattern utilization—it's pattern mastery." No evidence beyond the existence of LLM-generated novel text is provided. The distinction between "mechanical reproduction" and "template-based novel combination" is not operationalized — there is no metric for determining whether a particular LLM output is genuinely creative or merely a recombination of memorized fragments.

The ReSTEM^{EM} negative result (Appendix K, Figure 16): Notably, the paper mentions in passing (without providing the actual figure, as this is not an empirical paper) an attempt to further optimize a revision model using ReSTEM^{EM} (Singh et al., 2024) that "backfired": additional sequential revisions "substantially hurt" performance. This is one of the few references to an actual experimental finding in the paper, and it is used to illustrate the sensitivity of revision training to data generation procedure. However, this result is described rather than shown — the paper does not provide the actual accuracy numbers or experimental configuration.

Critical Assessment

The paper makes no traditional experimental claims, so the question of whether "experiments support claims" must be reframed: do the empirical observations and citations the paper presents genuinely support its theoretical argument, or does the argument outrun its evidence?

The Synthesis-Validates-Analysis Argument: Strong in Principle, Uneven in Execution

The paper's most powerful empirical claim is that generative grammar has failed to produce synthesis, while LLMs have succeeded. This is presented as an empirical comparison between two approaches to language.

What is demonstrated: The paper provides a specific, vivid example of the analysis-synthesis gap in generative grammar (Chomsky's 10-page analysis vs. Mańczak's 5-rule synthesis for one sentence) and cites literature suggesting that no complete generative grammar of any language exists. The existence of LLMs that produce coherent text is, at this point, a matter of public record — the reader does not need experimental data to accept that GPT, Claude, and similar systems generate contextually appropriate language.

What is not demonstrated: The paper does not systematically survey generative grammar to establish that synthesis has never been achieved. One example from 1965, even if representative, does not constitute a survey of a 50+ year research program. The claim that linguists use corpora and native speakers rather than generative rules when assessing grammaticality is asserted, not documented. The comparison between LLMs and generative grammar is asymmetric: LLMs are evaluated by the existence of any working synthesis, while generative grammar is evaluated by the absence of a complete grammar. This is arguably the right asymmetry — the paper's point is that one approach succeeded where the other failed — but it means the "empirical comparison" is between a specific negative finding (no generative synthesis) and a general positive observation (LLMs synthesize text), with very different standards of evidence for the two sides.

A stronger version of this argument would require: (a) a systematic definition of what counts as "synthesis" for a linguistic theory (generating all and only grammatical sentences? generating a representative sample? something else?), (b) a systematic search of the generative literature for synthesis attempts, and (c) a controlled comparison where both LLMs and generative grammars are evaluated on the same synthesis task. The paper provides none of these.

The Frequency Principle Argument: Conceptual Alignment, Not Empirical Test

The paper's central claim about frequency — that it is the primary organizing force in language, and that LLMs validate this claim — rests on a conceptual correspondence between the LLM training objective (cross-entropy minimization) and Mańczak's frequency principle. This correspondence is mathematically exact: minimizing cross-entropy pushes predictions toward empirical frequencies. But this is a feature of the training objective, not an empirical discovery about language.

The paper would need to demonstrate something stronger to claim empirical validation: that frequency is the primary organizing force, not merely an organizing force. This would require showing that alternative organizational principles (e.g., structural economy, communicative efficiency, innate constraints on possible grammars) either reduce to frequency or fail to explain phenomena that frequency explains. The paper does not engage with these alternatives systematically. It mentions Chomsky's poverty-of-stimulus argument and claims it has been abandoned by "many experts," but does not address post-2000 developments in minimalist syntax, biolinguistics, or computational models of language acquisition that incorporate both statistical learning and innate constraints.

The scaling laws evidence (Kaplan et al., 2020; Hoffmann et al., 2022) shows that more data improves LLM performance. This is consistent with the frequency principle but also consistent with virtually any learning theory. The specific quantitative form of the scaling laws (power-law exponents) is not derived from or predicted by Mańczak's theory, and the paper does not argue that it is. The citation functions as plausibility support, not as a distinguishing test.

The embedding frequency evidence (Zhou et al., 2021; Gong et al., 2018) shows that embeddings encode frequency. This would be remarkable if embeddings were designed to be frequency-agnostic and frequency information appeared anyway; but embeddings are trained to minimize prediction error on a frequency distribution, so encoding frequency is the expected outcome. The paper uses this as corroboration, but it would be more surprising if embeddings did not encode frequency.

The Grounding Argument: A Philosophical Position, Not an Empirical Result

The paper's treatment of the grounding objection is entirely theoretical. It argues that meaning is predominantly relational, that relational meaning is sufficient for linguistic competence, and that demanding sensory grounding for language models is an arbitrary requirement. These are philosophical claims about the nature of meaning, not experimental findings.

The paper cites Piantadosi and Hill (2022) to support the claim that concepts without referents can be meaningful through relational networks. This is a genuine empirical finding from that paper. But it establishes that meaning can exist without reference — it does not establish that all or most linguistic meaning is referent-free, or that grounding provides no additional benefit beyond relational structure. The paper's claim is carefully scoped ("the vast majority of the time 'meaning' can be inferred... solely from the relational structure"), but this scope is asserted, not measured.

The calculator analogy (Section 2.4) is an argument from analogy, not an experiment. The fact that we accept a calculator's arithmetic without demanding it understand the philosophy of mathematics does not logically entail that we should accept an LLM's language use without demanding it understand the referents of its words. The relevant disanalogy is that calculators operate in a fully formal domain where correctness is decidable by rules, while natural language has referential content where correctness often depends on correspondence to non-textual reality. The paper does not address this disanalogy.

The Analogy and Generalization Argument: Plausible Mechanism, Not Empirically Validated for LLMs

The paper's account of how Transformers achieve generalization through embedding-space analogy (Table 2 and Section 2.3) is a mechanistic hypothesis, not an experimental finding. The progression from n-grams to CBOW to Transformers is historically accurate, and the "Anna likes cats" / "Lily loves dogs" analogy is a standard illustration of how embedding spaces capture relational similarity. But the paper does not provide empirical evidence that Transformer generalization to novel sentences operates specifically through the analogical mechanism described, rather than through other forms of statistical generalization.

This is a significant evidentiary gap because the paper's claim is strong: "This ability to represent and manipulate relationships—the very essence of analogy—is the key to genuine linguistic generalization." To support this, one would need experiments showing that (a) LLM generalization patterns match predictions from analogical reasoning (rather than, say, interpolation between memorized examples), (b) ablating the ability to form analogies impairs generalization in specific ways, or (c) the embedding geometry of Transformers encodes analogical relationships in a way that directly predicts generalization behavior. None of this is provided. The paper's claim about analogy is a plausible interpretation of Transformer behavior, but it is not experimentally established by the evidence the paper presents.

What Would Strengthen the Paper's Empirical Case

The paper is explicitly a theoretical position paper, not an empirical study, and should be evaluated as such. However, even a theoretical paper can be strengthened by empirical grounding. The following experiments, none of which the paper conducts, would substantially strengthen its claims:

  1. Synthesis test for linguistic theories: Operationalize "synthesis" as a concrete task (e.g., given a grammar, generate 1000 sentences; have native speakers rate their grammaticality and naturalness; compare to human-generated baselines). Apply this test to the best available generative grammar of any language, to the best available construction grammar, and to an LLM. This would convert the paper's central methodological argument from an illustrative example to a systematic comparison.

  2. Frequency ablation in LLM training: Train LLMs on corpora where frequency distributions have been artificially altered (e.g., down-weighting high-frequency items, up-weighting low-frequency items) and measure the impact on linguistic competence as assessed by human judges or downstream benchmarks. If frequency is the primary organizing force, these alterations should produce predictable, systematic effects on the model's linguistic behavior. If alternative organizational principles are at work, the effects may be smaller or different in kind.

  3. Relational vs. grounded meaning test: Identify a set of concepts where relational structure and grounding make different predictions about usage (e.g., color terms, where relational structure alone might not capture the similarity structure of the color space). Test whether LLMs' usage of these concepts matches the relational prediction, the grounding prediction, or some mixture. This would test the paper's claim that relational structure is sufficient for "the vast majority" of meaning.

  4. Analogy mechanism verification: Using interpretability techniques (probing classifiers, representational similarity analysis, causal interventions), test whether Transformer generalization to novel sentences is mediated by analogical relationships in embedding space, as the paper claims. Compare this to alternative mechanisms (pattern completion, interpolation between memorized examples) to determine whether analogy is genuinely "the key" or merely one of several mechanisms.

  5. Cross-linguistic frequency pattern replication: Test whether Mańczak's specific claims about frequency-driven change (irregular phonetic development due to frequency, analogical regularization, grammaticalization patterns) are reproduced in the distributions learned by multilingual LLMs. If LLMs trained on different languages independently recapitulate the frequency effects Mańczak documented through manual corpus analysis, this would provide stronger validation than the loose alignment the paper currently demonstrates.

Summary of Critical Assessment

The paper's central claims fall into three categories, each with different evidentiary status:

  1. Methodological critique of existing linguistic evaluation of LLMs: The claim that prior critiques apply unvalidated theoretical criteria is well-argued and supported by specific examples (Chomsky et al., 2023; Bender et al., 2021; the synthesis failure of generative grammar). This claim does not require new experiments — it is an argument about the logical structure of prior work. The evidence (the existence of the cited critiques, the absence of synthesis validation in generative grammar) is adequate to support the argument, though the paper's characterization of the field as having "abandoned" nativism is overstated.

  2. Positive proposal of the Mańczakian framework: The claim that Mańczak's framework provides a better foundation for evaluating LLMs is supported by showing its alignment with LLM architecture and training. However, this alignment is demonstrated at the level of conceptual correspondence (frequency → cross-entropy minimization), not at the level of specific, falsifiable predictions. The framework's explanatory power is illustrated through historical linguistic examples and citations to usage-based acquisition research, but its predictive power is not tested. The paper establishes that the Mańczakian framework is a plausible candidate for LLM evaluation, not that it is the uniquely correct one.

  3. Empirical claim that LLMs validate Mańczak's thesis: This is the paper's strongest claim and the one with the weakest direct evidentiary support. The inference from "LLMs model frequencies and produce language" to "frequency is language's primary organizing force" relies on the unstated assumption that frequency modeling would not produce linguistic competence unless frequency were indeed primary. The paper does not rule out the alternative that frequency modeling captures something important about language while other organizing principles (not captured by pure frequency estimation) are also necessary for full linguistic competence — principles that LLMs acquire implicitly through their architecture and training dynamics rather than through frequency estimation alone. The validation is of the sufficiency of frequency-based learning for linguistic competence, not necessarily of the primacy of frequency over all other organizational principles.

The paper is best understood not as an experimental validation of Mańczak's linguistics but as a philosophical intervention in the methodology of LLM evaluation, using Mańczak's framework as a lens to expose what the authors see as question-begging assumptions in prior critiques. Its contribution is conceptual clarity and a proposed alternative framework, not empirical demonstration that the alternative framework makes correct predictions where others fail. The reader should judge it on whether that conceptual reframing is productive for future research, not on whether the paper itself provides experimental evidence for its claims — it explicitly does not attempt to do so.

6. Limitations and Trade-offs

The Difficulty Estimation Cost Is Not Accounted For

The assumption or constraint. The entire compute-optimal framework — both the oracle and predicted difficulty variants — requires estimating each prompt's difficulty before allocating the inference budget. The method for doing so involves generating 2,048 samples per question and either checking correctness (oracle) or averaging PRM final-answer scores (predicted). The paper explicitly acknowledges this cost in Section 3.2:

"estimating difficulty in this way still incurs additional computation cost during inference... our experiments do not account for this cost largely for simplicity"

The consequence. This is not a minor accounting detail — it is a potentially fatal practical obstacle. Generating 2,048 samples per question consumes more compute than the largest test-time budgets studied (256–512 generations). In a realistic deployment, the total cost would be difficulty estimation plus strategy execution, and the former could dominate the latter. The paper's headline finding — that compute-optimal scaling achieves 4× better efficiency than best-of-N — is computed after difficulty is already known, without amortizing the cost of learning it. If difficulty estimation costs, say, 2,048 generations and the strategy itself uses 64 generations, the "4× improvement" over best-of-N at 256 generations actually represents a total budget of 2,112 vs. 256 — roughly 8× more compute, not 4× less.

Furthermore, this cost cannot be trivially reduced by using fewer estimation samples because the difficulty binning depends on reliable pass@1 or average PRM score estimates. Using, say, 32 samples instead of 2,048 would produce noisy difficulty estimates, which could misroute prompts into suboptimal strategies and erode the gains from adaptive allocation. The paper provides no sensitivity analysis showing how estimation accuracy degrades as the number of estimation samples decreases, and therefore no evidence about the minimum practical cost of difficulty estimation.

What evidence exists in the paper. The paper acknowledges this explicitly in Section 3.2 and flags it as "a key avenue for future work," framing the difficulty estimation cost as an exploration-exploitation tradeoff. However, no experiment measures the impact of varying the number of estimation samples, no ablation compares the efficiency gains when difficulty estimation cost is included vs. excluded, and the predicted-difficulty curves in Figures 4 and 8 are plotted without including estimation cost in the x-axis budget. The fact that predicted bins largely overlap with oracle bins (Figures 4 and 8) is encouraging for the approach of using PRM scores rather than ground-truth labels, but it says nothing about the cost of obtaining those PRM scores in the first place.

Mitigation status. The paper does not solve this problem. Section 8 suggests future work on "pretraining or finetuning models to directly predict difficulty of a question" from the question text alone, which would eliminate the per-prompt sampling cost entirely. But no such model is developed or evaluated. The paper also does not explore adaptive difficulty estimation strategies — for instance, starting with a small number of samples (4–8), using the verifier's score distribution on those initial samples as a quick difficulty signal, and allocating the remaining budget accordingly. Such an approach could amortize difficulty estimation into the problem-solving process itself, but it is not investigated. The limitation is acknowledged but unresolved, and the practical deployability of the compute-optimal framework depends entirely on closing this gap.


The Framework Offers No Path Forward on the Hardest Problems

The assumption or constraint. The paper's compute-optimal framework is fundamentally bounded by the base model's capabilities. If the base model cannot produce correct solutions at any meaningful rate on a class of problems, no amount of test-time compute — search, revisions, or any combination — will help. The paper is explicit about this in Section 7:

"Test-time compute provides minimal gains on problems that are fundamentally outside the base model's capability range."

The consequence. Across every method studied — PRM search (Figure 3, right), sequential revisions (Figure 7, right), and their compute-optimal combinations (Figures 4, 8) — the hardest difficulty bin (bin 5) shows near-zero accuracy regardless of compute budget. In Figure 3 (right), bin 5 accuracy hovers at roughly 1–3% for all search methods and all budget levels from 4 to 256 generations. In Figure 7 (right), bin 5 accuracy is roughly 2–3% irrespective of the sequential-to-parallel ratio. In the FLOPs-matched comparison (Figure 9, Section 7), the bin 5 scaling line is essentially flat near 0–5%, positioned below the 14× larger model's performance at all three values of R.

This is not an incidental limitation — it is a hard capability ceiling that no amount of inference-time engineering can surmount. The paper frames test-time compute as "amplifying existing capability" rather than "creating it from nothing," and the bin 5 results demonstrate this boundary sharply. For any deployment where the problem distribution includes genuinely novel, out-of-distribution, or deeply challenging prompts — the kind that a larger model trained on more data might handle — compute-optimal test-time scaling offers essentially zero benefit. The practical implication is that organizations cannot rely on inference-time strategies to compensate for a fundamentally underpowered base model on hard problems; pretraining remains the only viable path.

What evidence exists in the paper. The evidence is consistent and unambiguous across all experiments:

  • Figure 3 (right): Bin 5 shows ~1–3% accuracy for both best-of-N weighted and beam search across budgets from 4 to 256.
  • Figure 7 (right): Bin 5 shows ~2–3% accuracy across all sequential-to-parallel ratios at 128 generations.
  • Figure 9: The bin 5 curve is flat and near zero; the 14× larger model's performance (stars) is above the test-time compute curve at all values of R.
  • Section 7 states explicitly in the takeaway box that "on the hardest problems (bin 5), test-time compute provides essentially zero benefit regardless of budget, meaning that some capabilities can only be acquired through pretraining, not recovered at inference time."

Mitigation status. The paper is transparent about this limitation and explicitly characterizes it as a fundamental boundary rather than a gap to be closed. Section 7 clearly states when to prefer pretraining over test-time compute ("The problem distribution skews toward genuinely hard problems outside the base model's capability range"). The paper does not attempt to solve this limitation because it is, by the paper's own framework, unsolvable — you cannot amplify a capability the base model does not possess. The mitigation is strategic rather than technical: route hard problems to larger models or flag them for human review, reserving test-time compute strategies for easy-to-medium problems where the base model already has non-trivial pass@1. However, this mitigation depends on accurate difficulty estimation (see previous limitation), creating a compounding failure mode: if difficulty is estimated incorrectly, hard problems may be misrouted to strategies that waste compute for zero gain.


The 14× Larger Model Baseline Is Weaker Than a Compute-Optimally Trained Alternative

The assumption or constraint. The FLOPs-matched comparison in Section 7 scales model parameters while holding training data fixed, following the LLaMA paradigm (Touvron et al., 2023) rather than the Chinchilla-optimal approach of scaling both parameters and data equally. The paper acknowledges this explicitly:

"We choose this setting as it is representative of a canonical approach to scaling pretraining compute and leave the analysis of compute-optimal scaling of pretraining compute where the data and parameters are both scaled equally to future work."

Additionally, the 14× larger model is evaluated only with greedy decoding — no majority voting, no best-of-N, no search of any kind. This means the comparison is between (small model + compute-optimal test-time strategies) and (large model + greedy decoding), not between (small model + inference strategies) and (large model + some inference strategies).

The consequence. Both choices systematically favor the test-time compute approach. A Chinchilla-optimal model — where the 14× increase in total pretraining FLOPs is allocated to scaling both parameters and data according to the ~1:1 ratio Hoffmann et al. (2022) recommend — would likely outperform a parameter-only-scaled model, because the latter is undertrained relative to its parameter count. The paper's comparison therefore uses a pretraining baseline that is weaker than what compute-optimal pretraining practices would produce, making the reported advantages of test-time compute potentially overstated.

Furthermore, giving the larger model even a modest test-time compute budget (e.g., best-of-8, majority voting over 4 samples) would create a substantially stronger baseline. If a larger model with best-of-8 outperforms the smaller model with compute-optimal scaling at 256 generations, the practical recommendation flips: one should train the larger model and still apply cheap test-time strategies, rather than keeping the small model and spending heavily on inference. The paper's FLOPs-matched comparison cannot distinguish between "test-time compute substitutes for pretraining" and "test-time compute substitutes for a suboptimal pretraining recipe," because the baseline is not compute-optimal in the pretraining dimension.

What evidence exists in the paper. The paper reports that at R ≪ 1 (low inference-to-pretraining ratio), test-time compute with the smaller model outperforms the 14× larger model on easy and medium questions (Figure 9, left; Figure 1 bar chart). For revisions, the paper reports +27.8% relative improvement on medium questions at R ≪ 1 (Section 7, bar chart in Figure 1). These numbers are computed against the parameter-only-scaled baseline. The paper does not provide any comparison against a Chinchilla-optimal larger model, nor does it report results for the larger model with any test-time compute augmentation. Section 7 explicitly notes this as a limitation in the discussion of design choices, but does not quantify how much the reported advantages might shrink under a stronger baseline.

Mitigation status. The paper acknowledges this limitation in Section 7 and frames it as a scope limitation for future work rather than a flaw in the current analysis. However, no sensitivity analysis is performed — for instance, estimating what fraction of the performance gap between the small model and the parameter-only-scaled large model might be closed if the large model were Chinchilla-optimal. The paper's practical recommendations ("prefer test-time compute when R ≪ 1 and problems are easy-to-medium") are qualified by this limitation but not contingent on resolving it. A practitioner following the paper's guidance might deploy the small-model + test-time-compute strategy only to find that a properly trained larger model with even cheap inference-time augmentation (majority voting over 4 samples) outperforms it at lower engineering complexity.


The Revision Model Suffers from a 38% Correct-to-Incorrect Reversion Rate with Only Partial Mitigation

The assumption or constraint. The revision model is trained exclusively on sequences where all in-context answers are incorrect, followed by a correct target. The paper notes in Section 6.1:

"the model was trained only on sequences where all in-context answers are incorrect (followed by a correct target)"

At test time, however, the revision chain may contain correct answers produced during earlier revision steps. When the model encounters these in context, it has no training signal for what to do — it was never shown examples where the current answer is already correct — and it may incorrectly "revise" a correct answer into an incorrect one.

The consequence. The paper reports that approximately 38% of correct answers produced during a revision chain get converted back to incorrect ones in subsequent revision steps. This reversion rate means that extending a revision chain does not monotonically improve output quality; longer chains risk undoing earlier progress. This is a direct consequence of the training data construction procedure: the model learns to always produce a revised answer that differs from the previous one (since all training examples involve incorrect → correct transitions), but it does not learn when to stop revising because the current answer is already adequate.

The paper mitigates this by applying selection (majority voting or verifier-based selection) across the entire chain rather than always taking the final revision, which partially addresses the symptom but not the underlying cause. A model that cannot recognize its own correct answers is fundamentally limited in its ability to engage in open-ended iterative improvement — it will always risk degrading good outputs if left to revise autonomously. This is particularly problematic for the sequential-heavy strategies that the compute-optimal policy favors on easy problems, where the initial attempts are often correct and the risk of reversion is highest.

What evidence exists in the paper. The paper mentions the 38% reversion rate in Section 6.1 but does not provide a dedicated figure or table quantifying it. The mitigation — selecting the best answer from any point in the chain — is described in Section 6.1 and evaluated implicitly through the sequential revision results in Figure 6 (right), which shows that sequential revisions marginally outperform parallel sampling under both verifier-based and majority selection. The fact that this gap is small (roughly 41.5% vs. 39% at 64 generations in Figure 6, right) despite the reversion problem suggests that the within-chain selection is partially effective but cannot fully compensate — a model that did not revert correct answers would presumably show a larger advantage for sequential revision.

Mitigation status. The paper does not solve this problem. The within-chain selection mechanism (Section 6.1) is a workaround, not a fix — it recovers the best answer after the fact rather than preventing the model from degrading correct answers in the first place. The paper does not explore alternative training strategies that would teach the model to recognize correctness and stop revising, such as including training examples where the correct answer appears in context and the target is to reproduce it unchanged, or adding a "stop revising" token. The ReSTEM^{EM} experiment (Appendix K) attempted to further optimize the revision model but caused performance to degrade substantially with sequential revisions, suggesting that the reversion problem may be exacerbated rather than solved by standard RL-based fine-tuning approaches. The paper does not investigate why this degradation occurs or whether it relates to the correct-to-incorrect reversion issue.


Sequential Revisions Introduce Latency That Is Ignored in the Cost Model

The assumption or constraint. The paper measures test-time compute in "generations" — the number of complete solutions sampled — which is a reasonable proxy for total FLOPs but ignores wall-clock time. This assumption is implicit throughout Sections 5–7: all budgets and efficiency comparisons are expressed in terms of generation counts, with no discussion of latency or throughput.

The consequence. Sequential revisions are inherently serial — each revision step depends on the output of the previous step, so they cannot be parallelized across hardware. A strategy that allocates 128 generations as 64 sequential revisions × 2 parallel chains takes approximately 64× longer wall-clock time than a strategy that runs 128 independent parallel samples simultaneously (assuming sufficient hardware to run all parallel samples concurrently). The compute-optimal policy, which favors sequential revisions on easy problems and balanced sequential-to-parallel ratios on medium problems (Section 6, Figure 7), systematically selects strategies with higher latency than pure parallel sampling.

For many practical deployment scenarios — interactive assistants, real-time decision-making, customer-facing applications — latency is a first-class constraint that is as important as total FLOPs. A strategy that achieves 4× better FLOPs efficiency but takes 10× longer wall-clock time may be unacceptable for applications with sub-second response time requirements. The paper's efficiency claims implicitly assume that all generations have equal cost to the user, but in latency-sensitive settings, serial generations are much more expensive in the dimension that matters most.

Furthermore, the paper does not discuss whether beam search — which also involves sequential steps (generating one step, scoring, pruning, then generating the next) — has latency characteristics that differ meaningfully from best-of-N sampling. If beam search with width M = 4 requires 10 sequential steps to complete a solution, it imposes a 10× latency multiplier relative to parallel best-of-N, even if the total generation count is the same. The compute-optimal policy's preference for beam search on medium problems (Section 5.3) may therefore select strategies with hidden latency costs that are not reflected in the generation-count-based efficiency comparisons.

What evidence exists in the paper. The paper provides no latency measurements, no wall-clock time comparisons, and no discussion of the sequential vs. parallel latency tradeoff. The unit of compute is uniformly "generations" (Section 4: "One 'generation' equals one complete sampled answer from the base LLM"). The experimental protocol treats all generations as equivalent in cost, regardless of whether they are executed serially or in parallel. Tables and figures reporting efficiency gains (e.g., "4× better efficiency" in the Executive Summary, Figures 4 and 8 showing compute-optimal matching best-of-N at 4× fewer generations) count only generation counts, not elapsed time.

Mitigation status. The paper does not acknowledge this limitation or discuss the latency-throughput tradeoff. There is no suggestion for future work on latency-aware allocation policies. A practitioner implementing the paper's approach would need to independently evaluate whether the generation-count efficiency gains translate to wall-clock-time improvements in their specific hardware and deployment context, and may need to impose additional latency constraints that override the compute-optimal policy's recommendations (e.g., capping the maximum number of sequential steps regardless of what the difficulty-conditioned policy selects).


Single Benchmark and Single Model Family Limit Generalizability

The assumption or constraint. All experimental results in the paper are generated using a single base model (PaLM 2-S*) evaluated on a single benchmark (MATH, 500 test questions). The paper states in Section 4 that the authors "believe this model is representative of the capabilities of many contemporary LLMs" and that MATH is chosen because test-time compute is expected to help most when the model already possesses the necessary knowledge and the challenge is drawing complex inferences.

The consequence. The paper's core findings — that difficulty-dependent compute-optimal allocation achieves 4× efficiency gains, that beam search over-optimizes on easy problems while helping on medium ones, that revisions are most effective on easy problems while balanced ratios work best on harder ones, and that test-time compute cannot help on the hardest problems — are all potentially specific to the interaction between PaLM 2-S* and the MATH benchmark. Several aspects could fail to transfer:

  • Model architecture and scale: PaLM 2-S* has a specific capability profile (roughly 10–19% pass@1 on MATH depending on configuration). A stronger model with higher base pass@1 might shift the difficulty distribution rightward — more problems fall into the "easy" bin where sequential revisions dominate — changing the optimal allocation strategy. A weaker model might shift the distribution leftward, making more problems "hard" and reducing the overall benefit of test-time compute.

  • Model family: Different model families (GPT, LLaMA, Claude) have different training data mixtures, tokenization schemes, and architectural details that could affect both the PRM's quality (since the PRM is trained on the base model's outputs) and the revision model's ability to learn from incorrect in-context examples. The paper found that the PRM800k dataset — which contains GPT-4-generated solutions — was "largely ineffective" for their PaLM 2 models (Section 5.1), suggesting that verifier quality is model-specific. A different base model might require a different PRM training procedure, different hyperparameters, or might not achieve the same level of verifier reliability.

  • Benchmark domain: MATH consists of competition-level math problems requiring symbolic reasoning and producing closed-form answers. The difficulty-dependent patterns documented in the paper — beam search over-optimizing on easy problems, revisions helping on easy problems but not hard ones — may be specific to mathematical reasoning. Code generation (where correctness can be checked with unit tests), logical reasoning (where answers can be verified against formal rules), or open-ended generation tasks (where correctness is ambiguous) might exhibit qualitatively different difficulty-dependent scaling behavior. In particular, the paper's entire framework depends on having a verifier (PRM or ORM) that can score solution quality. For tasks without clean correctness signals, training such a verifier is substantially harder, and the balance between search and revisions may shift.

  • Test set size and composition: The 500-question test set is split into five difficulty quintiles of approximately 100 questions each, then further split by two-fold cross-validation, meaning the compute-optimal policy is selected based on roughly 50 questions per fold per bin. This is a small sample for strategy selection. The paper does not report confidence intervals on the compute-optimal scaling curves, nor does it assess whether the selected strategies are stable across different random splits of the test set. With only 500 questions, the difficulty distribution of the test set may not match the difficulty distribution of real-world query streams, and the optimal policies identified may not generalize.

What evidence exists in the paper. The paper reports all results on MATH + PaLM 2-S* and does not include any experiments on other benchmarks or model families. Section 4 justifies the choice of MATH as deliberate ("test-time compute is expected to help most when the model already possesses the necessary knowledge and the challenge is drawing complex inferences") and describes PaLM 2-S* as "representative," but provides no evidence for this representativeness claim. The PRM transferability issue is documented in Section 5.1 (the PRM800k dataset was ineffective for PaLM 2 models, requiring Monte Carlo rollout training instead), which indirectly demonstrates model-specificity but is not framed as a generalizability limitation.

Mitigation status. The paper does not address generalizability as a limitation. Section 4 acknowledges the single-benchmark, single-model scope implicitly by describing the choices as deliberate and representative, but does not flag the absence of cross-model or cross-benchmark validation as a limitation. Section 8's future work discussion does not mention replication on other models or benchmarks. The paper's confident prescriptions ("Prefer compute-optimal test-time scaling when...", "Prefer scaling pretraining instead when...") are based entirely on MATH results with PaLM 2-S*, and their transferability to other settings is unknown. A practitioner considering deploying these strategies on, say, a LLaMA-based system for code generation would need to independently validate whether the difficulty-dependent patterns replicate in their specific context.

7. Implications and Future Directions

How This Work Changes the Landscape

This paper does not introduce a new model, a new training technique, or a new evaluation metric. Its contribution operates at a different level entirely: it is a methodological intervention that challenges the foundational assumptions underlying how the field evaluates whether language models "truly" model language. The magnitude of this shift, if the paper's argument is accepted, is closer to a paradigm challenge than an incremental refinement — but it is a challenge mounted from within a specific alternative tradition (Mańczakian empiricism) rather than from entirely new principles.

What changes: the burden of proof in LLM evaluation. Prior to this work, the dominant framing of LLM evaluation in theoretical linguistics was: here is our theory of language (generative grammar, grounded semantics, competence-performance distinction), does the LLM satisfy its criteria? The paper's central move is to invert this: here is an LLM that demonstrably synthesizes coherent language at scale, does your theory of language predict or explain this? Under the Mańczakian synthesis-validates-analysis criterion, the fact that LLMs achieve synthesis while generative grammar never did shifts the evidential burden. The theory that cannot synthesize has failed the only test that, in the paper's framework, counts; the system that can synthesize has, by that same test, demonstrated linguistic competence in the only way that matters.

This is not merely rhetorical. It provides a principled, falsifiable standard — synthesis validates analysis — that is independent of any particular linguistic school. A linguist of any theoretical persuasion can accept or reject this standard, but they must engage with it on its own terms: either show that their theory has produced synthesis (which the paper claims generative grammar has not), or argue that synthesis is not the right criterion (which requires defending an alternative validation standard). The paper thus shifts the debate from "which theory of language is correct?" to "what would count as validating any theory of language, and has that validation been achieved?" This is a meta-theoretical move that changes the nature of the discourse rather than resolving it in favor of one side.

Reconciling conflicting intuitions about LLM capability. The paper resolves a persistent tension in both academic and public discourse about LLMs. On one hand, LLMs produce fluent, contextually appropriate text across an enormous range of domains — a fact that virtually all observers, including critics, acknowledge. On the other hand, prominent linguists (Chomsky et al., 2023; Bender et al., 2021) argue that this apparent competence is illusory because LLMs lack deep structure, grounding, or genuine understanding. How can both observations be true?

The paper's resolution is that they are not contradictory — they are answers to different questions. The critics are answering the question "do LLMs satisfy our theory of what language requires?" The empirical fact of synthesis is answering the question "can language be modeled as a frequency-governed distribution over texts?" The Mańczakian framework argues that the first question is the wrong one to ask, because the theories it invokes have never been validated by the criterion (synthesis) that the second question implicitly applies. The critics are not wrong about LLMs relative to their theories; they are wrong about whether their theories are the appropriate yardstick.

This reconciliation has practical consequences. It redirects critical energy away from diagnosing LLMs for failing to match theoretical constructs (deep structure, language organ) and toward improving the empirical fidelity of the models: better corpora (the "rational selection of texts" Mańczak advocated), better frequency estimation (larger models, better architectures), and better understanding of the residual gaps between textual distributions and whatever else (grounding, world-modeling, reasoning) humans bring to language use. The "stochastic parrot" label loses its pejorative force not because parrots are smarter than we thought, but because language — in the Mańczakian sense — is precisely the kind of thing that can be parroted.

Research directions that become more attractive. Several lines of inquiry gain prominence under the Mańczakian framework:

  • Verifier and evaluation methodology grounded in text distributions rather than theoretical constructs. If language is the totality of texts, the primary evaluation of a language model is how well its output distribution matches the target distribution. This makes distribution-matching metrics (perplexity, statistical divergence measures, corpus-level comparisons of linguistic feature frequencies) central rather than peripheral, and de-emphasizes evaluations that require models to articulate rules or demonstrate knowledge of theoretical constructs.

  • Corpus construction as a fundamental research problem. The paper acknowledges that LLMs are trained on "unrepresentative corpora" and frames this as a practical limitation rather than a theoretical flaw. Under the Mańczakian definition, improving corpora — making them more representative of "the totality of all that is said and written" — is not an engineering detail but a direct path to better language models. The paper's invocation of Mańczak's "rational selection of texts based on circulation and influence" (Section 4) suggests a research program in principled, frequency-weighted corpus construction.

  • Historical and comparative linguistics as testbeds for LLM understanding. Mańczak's own work focused heavily on historical language change (irregular phonetic development due to frequency, analogical regularization, grammaticalization). The paper's framework suggests that LLMs should recapitulate these frequency-driven patterns in their learned distributions. Multilingual LLMs become natural laboratories for testing Mańczak's specific hypotheses about language change — if LLMs independently reproduce the frequency effects he documented through manual corpus analysis, that constitutes stronger validation than the conceptual alignment the paper currently demonstrates.

  • Usage-based and constructionist approaches to grammar as theoretical partners. The paper aligns itself with usage-based linguistics (Bybee, 2006; Tomasello, 2003; Goldberg, 2024) while claiming Mańczak's priority and greater theoretical parsimony. This suggests productive dialogue between Mańczakian empiricism and modern usage-based theories, with LLMs serving as a shared testbed — both frameworks predict that frequency-driven statistical learning should produce linguistic competence, and LLMs provide a platform for testing the specific predictions of each.

Research directions that become less attractive. Conversely, the paper's argument suggests diminishing returns for certain lines of inquiry:

  • Critiques that demand LLMs demonstrate knowledge of specific theoretical constructs (deep structure, universal grammar, explicit rules). If the paper's synthesis-validates-analysis criterion is accepted, these demands become question-begging — they assume the necessity of constructs whose necessity is precisely what LLMs' success calls into question. Research programs organized around demonstrating that LLMs lack these constructs are, under the Mańczakian framework, testing the wrong hypothesis.

  • Attempts to augment LLMs with symbolic grammar modules or explicit linguistic rules to achieve "true" competence. If frequency distributions over text are sufficient for synthesis, adding symbolic grammar is unnecessary — it addresses a deficit (lack of explicit rules) that the Mańczakian framework does not recognize as a deficit. This does not mean such augmentation is useless for other purposes (efficiency, interpretability, control), but it undercuts the argument that it is necessary for linguistic competence.

  • Poverty-of-stimulus arguments as a basis for LLM critique. The paper devotes substantial attention to challenging the poverty-of-stimulus argument (Section 2.3), citing extensive evidence for usage-based acquisition. If this challenge is accepted, then arguments that LLMs cannot learn language from text alone because "text is too impoverished" lose their force — LLMs are, in effect, a large-scale demonstration that the stimulus is not impoverished in the way nativists claimed.


Follow-Up Research This Work Enables

Quantifying the synthesis-validates-analysis criterion for linguistic theories. The paper's central methodological claim — that generative grammar has failed to produce synthesis — rests on Mańczak's observation that no complete generative grammar of a concrete language exists. This claim is binary and qualitative. A strong follow-up would operationalize "synthesis" as an experimentally testable criterion: define a synthesis task (e.g., given a proposed grammar of English, generate 1,000 sentences spanning diverse constructions; have linguistically trained annotators judge grammaticality, naturalness, and coverage), apply it to the best available generative grammar of any language (perhaps a fragment grammar covering a well-defined subdomain like English relative clauses or question formation), to the best available construction grammar, and to several LLMs of varying scales. Comparing synthesis quality — not just whether synthesis exists but how good it is — would convert the paper's binary claim into a graded empirical finding. If even fragmentary generative grammars achieve synthesis comparable to small LLMs on constrained domains, it would complicate the paper's narrative of complete generative failure. If no formal grammar achieves synthesis anywhere near LLM quality at any scale, the paper's argument is substantially strengthened.

Frequency-manipulated LLM training as a test of the frequency principle. The paper claims that frequency is language's primary organizing force and that LLMs validate this thesis because their training objective (cross-entropy minimization) pushes predictions toward empirical frequencies. This is a conceptual alignment, not an experimental test. A direct test would train multiple LLMs on corpora where the frequency distribution has been systematically manipulated — for instance, down-weighting the frequency of specific grammatical constructions (passives, relative clauses, subjunctive forms) by subsampling or up-weighting rare constructions by oversampling — and measure whether the models' linguistic behavior shifts in the specific ways Mańczak's framework predicts. If reducing the frequency of passives in training data produces models that use passives less frequently but maintain their grammaticality (consistent with frequency as the primary driver of usage patterns), that supports the frequency principle. If the models instead show degraded grammaticality or compensate through alternative constructions, that suggests additional organizational principles beyond raw frequency. This experiment would test not just whether frequency matters (which is trivially true) but whether it matters in the specific ways and domains Mańczak's framework claims.

Relational vs. grounded meaning: testing the sufficiency boundary. The paper argues that "the vast majority of the time 'meaning' can be inferred solely from the relational structure of the text" (Section 4), but does not quantify this majority or identify where it breaks down. A strong follow-up would identify a set of linguistic phenomena where relational structure and grounding make divergent predictions about correct usage, then test LLM performance on those phenomena. Color terms are a natural candidate: the relational structure of color words in text (e.g., "red" co-occurs with "fire," "stop," "rose") captures some but not all of the similarity structure of the color space (e.g., that red is perceptually closer to orange than to blue). If LLMs' usage of color terms matches the relational prediction (they use color words appropriately in textual contexts) but diverges from the grounding prediction (they fail at tasks requiring mapping color words to continuous perceptual spaces), that would support the paper's claim that relational structure is sufficient for linguistic meaning while grounding is needed for non-linguistic tasks (perception, world-modeling). If LLMs also fail at purely textual color tasks where relational structure should suffice, that would challenge the sufficiency claim. Piantadosi and Hill (2022) provide a starting point for identifying concepts without possible referents; extending their approach to systematically map the boundary would directly test the paper's central semantic thesis.

Cross-linguistic replication of Mańczak's historical frequency effects in LLM distributions. Mańczak's original work documented specific frequency-driven patterns in the evolution of Indo-European languages: irregular phonetic reduction of high-frequency items, analogical regularization of low-frequency irregulars, grammaticalization of frequent lexical verbs. Multilingual LLMs trained on diverse languages provide an unprecedented opportunity to test whether these patterns emerge from purely distributional learning. A strong follow-up would: (a) select 5-10 languages from different families with well-documented historical changes matching Mańczak's predictions; (b) train language models on historical corpora from different time periods for each language; (c) measure whether the models' learned representations encode the frequency effects Mańczak documented (e.g., do high-frequency verb forms show gradient phonetic reduction in the model's token probability distributions? do low-frequency irregular forms show attraction to regular patterns?). If the effects appear across diverse languages and model architectures, this constitutes much stronger validation of Mańczak's thesis than the conceptual alignment the paper currently provides. If the effects are language-specific or architecture-specific, that would help delineate the scope of the frequency principle.

The "rational selection of texts" experiment. The paper acknowledges that LLMs are trained on unrepresentative corpora and suggests that a Mańczakian approach would favor "principled, frequency-weighted corpus construction that reflects how language is actually used" (Section 4). This is a concrete, testable proposal. A follow-up would construct two training corpora of equal size: a standard web-crawled corpus (e.g., a subset of C4 or The Pile) and a "Mańczakian" corpus where text inclusion is weighted by circulation, influence, and actual usage frequency (operationalized through metrics like publication circulation figures, citation counts, view counts, or representative sampling across registers and demographics). Train identically-sized language models on both corpora and evaluate on a diverse set of linguistic tasks — not just perplexity but also register appropriateness, sociolinguistic variation, and coverage of minority dialects and non-standard varieties. If the Mańczakian corpus produces models that better match actual language use across the full sociolinguistic spectrum, that validates the paper's practical recommendation and provides a blueprint for corpus construction. If the standard corpus performs comparably or better, that challenges the paper's claim that corpus quality is the primary bottleneck.

Negative result: LLMs trained purely on formal languages. The paper's thesis is that language is text governed by frequency, and that LLMs model language by modeling frequency distributions. A sharp test of this thesis would be to train an LLM on a purely formal language — a programming language, a logical notation, or a constructed language with explicit generative rules — where the "totality of texts" is fully specifiable and the frequency distribution is known exactly. If the LLM learns the formal language's grammar purely from frequency distributions over text (without explicit rule induction), that supports the paper's claim that frequency is sufficient for grammar acquisition, even in domains where an explicit generative grammar exists and is known. If the LLM fails to acquire the grammar or requires rule augmentation to achieve synthesis, that challenges the sufficiency of frequency-based learning and suggests that even in the Mańczakian framework, some structural properties of language may not be reducible to frequency patterns. This experiment has the advantage of full controllability — the ground truth grammar is known, the corpus is exhaustively specifiable, and success or failure is unambiguous.


Practical Applications and Downstream Use Cases

Defending LLM deployments against "stochastic parrot" critiques in policy and regulatory contexts. The paper's reframing of the "stochastic parrot" label — from pejorative to definitional — provides a concrete argumentative tool for organizations deploying LLMs that face criticism rooted in theoretical linguistics. When regulators, ethics boards, or public commentators invoke Bender et al. (2021) to argue that LLMs merely "parrot" statistical patterns without genuine linguistic competence, the Mańczakian framework supplies a response: if language is defined as the totality of texts governed by frequency, then modeling those statistical patterns is modeling language — the parrot is stochastic because language itself is probabilistic. This does not address all ethical concerns (bias, factual accuracy, misuse), but it challenges the specific argument that LLMs are fundamentally incapable of language because they lack grounding or deep structure. The paper's synthesis-validates-analysis criterion provides a concrete standard: generative grammar failed to produce synthesis in 50+ years; LLMs achieved it at scale. For organizations facing linguistically-framed critiques, this shifts the debate from "does the model satisfy an abstract theory?" to "what does the model's empirical success tell us about the theory?"

Corpus construction for domain-specific language models. The paper's discussion of Mańczak's "rational selection of texts based on circulation and influence" (Section 4) provides a principled alternative to the dominant "scrape everything available" approach to pretraining data. For organizations building domain-specific LLMs — legal language models, medical language models, scientific language models — the Mańczakian framework suggests that corpus construction should be guided by actual usage frequency in the target domain, not by raw text availability. A legal LLM trained on a frequency-weighted corpus of actual court opinions, contracts, and regulatory filings (weighted by circulation, citation, and practical influence) would, under this framework, better model the "totality of what is said and written" in the legal domain than a model trained on all available legal text indiscriminately. This is an actionable principle for data curation teams: prioritize representativeness of usage distributions over raw corpus size. The paper's empirical support — the scaling laws showing smooth improvement with data quantity (Kaplan et al., 2020; Hoffmann et al., 2022) — suggests that even within a fixed compute budget, better corpus weighting could shift the performance curve upward. The practical implementation (how to weight texts by circulation and influence) requires operationalizing these concepts for a given domain, but the principle is clear and testable.

Informing the design of evaluation benchmarks for language models. Current LLM evaluation benchmarks largely test task performance (question answering, summarization, translation) or probe for specific capabilities (reasoning, factual knowledge, safety). They rarely test for the property the Mańczakian framework identifies as central: how well does the model's output distribution match the distribution of actual language use across registers, demographics, and contexts? The paper's framework suggests a complementary class of evaluation benchmarks focused on distributional fidelity: does the model produce passives at the right frequency for the register? does it use hedging and politeness markers at rates matching sociolinguistic baselines? does it represent dialectal variation proportionally to actual usage? These are not standard perplexity measurements (which evaluate token-level prediction on held-out text) but distribution-level comparisons between model outputs and reference corpora stratified by register, genre, and social context. Organizations deploying LLMs for customer-facing applications, content generation, or cross-cultural communication could use such benchmarks to assess whether their models reflect the actual linguistic landscape — as Mańczak would require — rather than just passing task-specific tests.

Reframing the research agenda for "grounded" language models. The paper's argument that meaning is predominantly relational and that grounding is supplementary rather than constitutive has practical implications for organizations investing in multimodal or embodied AI. If the paper's thesis is accepted, adding grounding (vision, robotics, sensory data) to a language model is not a prerequisite for linguistic competence — text alone suffices for that. Grounding adds capabilities beyond language: world-modeling, physical reasoning, referential accuracy in perceptual domains. This reframing has resource allocation implications. An organization deciding between investing in larger text-only pretraining runs versus multimodal training can, under the Mańczakian framework, separate the linguistic competence question (which text-only training addresses) from the world-knowledge question (which multimodality addresses). The paper's empirical claim — that relational structure from text alone handles "the vast majority" of meaning — suggests that the marginal return on additional text data for purely linguistic tasks may be higher than the marginal return on adding grounding, at least until linguistic competence saturates. Organizations can use this principle to stage investments: first achieve linguistic competence through text, then layer on grounding for specific tasks that require it, rather than treating grounding as a prerequisite for language understanding.


When to Prefer This Method

The paper presents itself as offering an alternative evaluative framework rather than an alternative model or algorithm, so a traditional "prefer X over Y under conditions Z" decision matrix is not directly applicable. However, the paper does articulate a clear tradeoff between two approaches to evaluating and improving language models, and the conditions under which each is appropriate can be extracted from its arguments.

Prefer the Mańczakian text-first evaluation framework when:

  • The goal is to assess linguistic competence in the narrow sense — the ability to produce and comprehend well-formed, contextually appropriate utterances across registers and domains — rather than world knowledge, factual accuracy, or reasoning ability. The paper's argument that meaning is predominantly relational and that frequency distributions capture linguistic structure applies to competence in this sense.

  • The evaluator seeks objective, falsifiable criteria grounded in observable data. The synthesis-validates-analysis criterion proposes that any linguistic theory or model should be judged by whether it can reconstruct the textual record. An evaluation asking "does this system produce output distributions matching the target language's empirical distributions?" is, in principle, answerable through corpus comparison. An evaluation asking "does this system possess deep structure or grounded meaning?" is, the paper argues, unanswerable by reference to observable data alone.

  • The research or deployment context involves tasks where linguistic form is the primary concern — text generation, translation, summarization, style transfer, grammar correction — rather than tasks requiring extra-linguistic knowledge (medical diagnosis from text, legal reasoning about statutes, factual verification against a changing world). The Mańczakian framework treats language as a self-contained universe of texts; it applies to tasks that operate within that universe.

Prefer alternative frameworks (generative, grounding-based, cognitive) when:

  • The goal is to assess cognitive plausibility — whether the system processes language the way humans do. The paper explicitly brackets this question: "any comparison to human cognition must come after this primary, non-metaphorical evaluation is established" (Section 4). The Mańczakian framework is a framework for evaluating language as a phenomenon, not cognition as a process. If the research question is about human language acquisition mechanisms, neural substrates, or processing architectures, the paper's framework does not address it and does not claim to.

  • The evaluation requires referential accuracy in a grounded domain. The paper acknowledges that its relational semantics does not provide grounding for axiomatic primitives, and that grounding may be necessary for tasks involving physical reference, perceptual judgment, or real-world action. An LLM controlling a robot arm needs more than text-internal relational meaning — it needs sensorimotor grounding. The Mańczakian framework does not deny this; it argues that grounding is a distinct capability from linguistic competence, not a prerequisite for it.

  • The goal is language acquisition research aimed at understanding the specific mechanisms by which humans (particularly children) learn language. While the paper aligns with usage-based acquisition theories, it does not provide a mechanistic account of human learning — it provides a framework for evaluating the end state (linguistic competence) and observes that statistical learning over text achieves it. Researchers studying developmental trajectories, sensitive periods, or the role of social interaction in acquisition need frameworks beyond what Mańczak's text-first approach offers.

  • The evaluation concerns normative or prescriptive grammar — what speakers should say rather than what they do say. Mańczak's framework is radically descriptive: language is what is said and written, and errors become norms if their frequency increases sufficiently. This provides no basis for grammatical prescriptivism. If the evaluation context requires judgments about "correct" versus "incorrect" usage in a prescriptive sense (e.g., educational assessment, formal writing standards), the Mańczakian framework is not designed for that purpose and provides no criteria beyond frequency distributions.

The paper's most important boundary condition is the one it states explicitly in the conclusion: LLMs "are imperfect tools not because they fail to model language but because they only model language" (Section 3). The decision to use the Mańczakian framework is thus a decision about what question you are asking. If you are asking "does this system model language?" — the answer is yes, and the Mańczakian framework explains why and how. If you are asking "does modeling language give us everything we want from an AI system?" — the answer is no, and the paper explicitly acknowledges that further capabilities (world-modeling, reasoning, factual grounding) lie beyond the scope of linguistic modeling as Mańczak defined it. The practical wisdom of the paper is in separating these questions rather than conflating them.