Draft paper argues some intended meaning cannot be recovered from text alone
An arXiv draft by Emily Cheng and Ryan Cotterell proposes information-theoretic limits on recovering intended meaning from text-derived representations. The authors argue that some ambiguity can only be resolved with extralinguistic context, and they report supporting experiments. The record establishes neither peer review nor enough methodological detail to judge the theory’s practical scope. [1]
This edition passed Imananq's enhanced publication checks. Some material claims remain explicitly attributed to official or company sources because no independent source is currently bound to this edition. The engine continues checking approved sources and will add corroboration only through a new edition that passes the full gate.

A new arXiv draft argues that text alone can place a formal ceiling on how reliably a system recovers a speaker’s intended meaning. Emily Cheng and Ryan Cotterell say the remaining ambiguity has both an irreducible component and a component that requires extralinguistic context. The claim is theoretical and remains unreviewed in the supplied record. [1]
01
What we know now
- 01
[1] arXiv, “A Formal Limitation on Learning Human Language From Textual Corpora,” Emily Cheng and Ryan Cotterell, version 1, submitted 28 August 2026. https://arxiv.org/abs/2608.28560v1.
- 02
The supplied primary record is complete for the arXiv abstract page, but it does not include the paper’s full methods or results.
02
Research status
arXiv identifies this submission as a draft and invites comments; the supplied record does not establish peer review.Input considered
The authors examine recovery of intended meaning from a representation of an utterance.Proposed missing signal
The authors say some uncertainty can be resolved by extralinguistic context rather than the utterance alone.A source-near outline of the distinction made in the abstract, not a measurement of any model's capability.
03
What the authors claim
The preprint models language use through a joint distribution of meanings, contexts, and utterances. Its authors say they derive information-theoretic upper bounds on the chance that a decoder can recover intended meaning from a representation of an utterance. They state that the result applies to text featurizers, including hidden states of contemporary large language models, and to discrete or continuous meaning spaces. [1]
In the authors’ account, linguistic form leaves uncertainty about meaning. They divide that uncertainty into a part they describe as irreducible and a part that extralinguistic context can resolve, while the utterance alone cannot. This is a claim about the framework presented in the draft, rather than verified evidence that all text-based systems fail in the same way or by the same amount. [1]
04
Evidence reported, and what is still unknown
The authors report experiments in three areas: artificial languages, Mandarin zero-pronoun resolution, and color reference. They say these experiments provide empirical evidence supporting their theory. [1]
However, the available record is an abstract page. It does not supply methods, quantitative outcomes, datasets, or comparison baselines, so readers cannot assess the reported experiments’ strength from this record alone. [1]
- Artificial languages
- Mandarin zero-pronoun resolution
- Color reference
05
Why the distinction matters
The draft offers a cautious way to separate two questions that are often conflated: whether a text representation has been improved, and whether the needed information was present in the utterance at all. If the authors’ formal account holds under its assumptions, adding text-derived representation or supervision would not resolve meaning that depends on outside context. [1]
For readers assessing language-interpretation claims, the practical takeaway is limited but useful: text-only results should not automatically be read as evidence of access to a speaker’s full intended meaning. The paper does not, from the supplied record, quantify this limitation for a specific model or application. [1]
06
07
How to read the preprint
Treat the paper as a proposed theoretical framework, not as a settled finding about every language system.
- 01
Use its central distinction when evaluating text-only interpretation: ask what information may depend on context outside the utterance.
- 02
Consult the linked arXiv record for the draft and any later versions before relying on the work in research or product decisions.
- 03
Do not infer benchmark performance, model-specific limits, or practical effect sizes from the abstract alone.
08
Limits of this edition
The arXiv record labels the work as a draft; peer review, acceptance, revision history beyond version 1, and independent replication are not established. [1]
The supplied abstract does not provide the formal assumptions, proof details, datasets, baselines, experimental procedures, or numerical results needed to assess the scope of the claims. [1]
The abstract reports evidence in support of the theory, but that report is the authors’ own account and does not independently establish the theory’s correctness or its practical effect on a particular system. [1]
SRC
Source desk
Direct links to the material behind this selection. Seeing the source matters as much as reading the synthesis.


