Preprint questions whether readable reasoning traces reveal step importance
An arXiv preprint reports that the text of a chain-of-thought step only partly reveals its functional importance under the authors' reward-based measure. The authors say capable LLM judges beat a prevalence baseline but remained below a noise ceiling. They also report strong improvement from a fine-tuned step-level critic on incorrect responses, while performance on correct responses remained distant from the ceiling. The supplied evidence is limited to a preprint record and abstract. [1]
This edition passed Imananq's enhanced publication checks. Some material claims remain explicitly attributed to official or company sources because no independent source is currently bound to this edition. The engine continues checking approved sources and will add corroboration only through a new edition that passes the full gate.

An arXiv preprint examines whether the text of a chain-of-thought step reveals its contribution to a model's expected reward. Its authors report that judges can identify some high-importance steps, but that trace text only partly reveals step importance under their measure. [1]
01
What we know now
- 01
[1] arXiv, “Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning,” version 1: https://arxiv.org/abs/2609.04194v1.
- 02
The supplied primary evidence is the arXiv record and abstract. It supports attribution of the authors' method and reported findings, not independent validation.
02
Record status
arXiv records version 1 as submitted on 3 September 2026.Paper's importance measure
The authors define a step's importance as advantage: the change in expected reward from including it, estimated with Monte Carlo rollouts.LLM judge result
The authors report that capable judges beat a prevalence baseline for high-advantage steps, but remained below a noise ceiling.Fine-tuned critic
The authors report strong improvement for incorrect responses. For correct responses, they say performance remained distant from the ceiling.The reported methods, results, and interpretation are attributed to the preprint's authors and have not been independently verified in the supplied evidence. [1]
03
What the paper measures
Reasoning traces are often read as an accessible account of how a model reached an answer. The preprint tests a narrower question: whether the text of an individual step carries information about that step's importance under the authors' reward-based definition. [1]
04
Reported experimental results
The authors interpret these results as evidence that step importance is only partly recoverable from the text of a reasoning trace. That is the paper's interpretation of its experiments, not an independently verified conclusion. [1]
- The authors report that sufficiently capable LLM judges outperformed a prevalence baseline at identifying high-advantage steps. [1]
- They also report that the judges fell well short of a noise ceiling. [1]
- They report strong improvement from fine-tuning a model as a step-level critic for incorrect responses. [1]
- For correct responses, the authors say the fine-tuned critic remained distant from the ceiling. [1]
05
Record and source
arXiv records the item as version 1, submitted on 3 September 2026. The available source is sufficient to identify the paper and attribute its abstract's claims, but not to independently assess the experiments. [1]
06
Practical takeaway
Treat visible reasoning steps as material for evaluation, not as a direct measure of which steps are functionally important or change expected reward under the paper's measure. [1]
- 01
Where feasible, compare step-level assessments by LLM judges or process reward models with outcome-linked measures rather than relying on wording alone. [1]
- 02
Do not assume that a clear-looking reasoning step is important simply because it appears plausible in a trace. [1]
- 03
Consult the full paper before applying its method: the supplied evidence does not include detailed methods, datasets, model identities, numerical results, confidence intervals, or replication materials. [1]
07
Limits of this edition
This is an arXiv preprint. The supplied evidence does not establish peer review, independent replication, or independent validation of its findings. [1]
The supplied record does not provide full methods, datasets, model identities, numerical results, confidence intervals, or replication materials. [1]
arXiv records a comment that the work was published at COLM 2026, but the supplied evidence does not independently verify publication details. [1]
SRC
Source desk
Direct links to the material behind this selection. Seeing the source matters as much as reading the synthesis.


