Preprint reports a two-stage path from summary errors to LLM ratings
A preprint reports that two fine-tuned summary evaluators use early attention layers to compare and route error signals, followed by later MLP layers that integrate those signals into ratings. A Llama-3-8B base-model control reportedly retained routing and crystallization but not the same stage separation; the authors attribute this difference to specific fine-tuning effects. These findings are model-specific author claims without independent verification in the supplied evidence.
This edition passed Imananq's enhanced publication checks. Some material claims remain explicitly attributed to official or company sources because no independent source is currently bound to this edition. The engine continues checking approved sources and will add corroboration only through a new edition that passes the full gate.

An arXiv preprint reports that two fine-tuned language-model evaluators convert detected summary errors into ratings through two internal stages: early attention layers compare local errors and route a signal, while later MLP layers integrate that signal and write the rating. The proposed stage separation was reported for Themis and Prometheus; the authors also tested a Llama-3-8B base-model control to examine routing, crystallization and what fine-tuning may change. The findings are author-reported, limited to summary Readability and Adequacy in the studied setting, and not independently verified in the supplied evidence. [1]
01
What we know now
- 01
arXiv records the v1 preprint by Himil Vasava and Ming Jiang as submitted on 1 September 2026. [1]
- 02
The authors report controlled-perturbation experiments on Themis and Prometheus for summary Readability and Adequacy, plus a Llama-3-8B base-model control. [1]
- 03
The authors report error comparison and routing below layer 15, followed by MLP-based integration and rating generation at higher layers. [1]
- 04
The authors report control-model routing and crystallization without stage separation, and attribute two specific changes to fine-tuning. [1]
02
Fine-tuned evaluator models with the reported two-stage mechanism
The authors report the mechanism for Themis, based on Llama-3-8B, and Prometheus, based on Mistral-7B.Reported early processing stage
Below layer 15, attention reportedly compares local errors and routes a resulting signal to the final input position.Reported decision crystallization layers
The authors place crystallization at layer 26 in Themis and layer 25 in Prometheus.Additional model examined as a control
A Llama-3-8B base-model control was used to examine routing, crystallization and the reported separation of stages.This visual summarizes author-reported findings from one preprint. The supplied evidence contains no independent replication.
03
What the record establishes
arXiv records “Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation” by Himil Vasava and Ming Jiang as submitted on 1 September 2026. The record says the paper was accepted at the EMNLP 2026 Main Conference, but no conference page or review material is included in the supplied evidence to independently confirm that statement. [1]
The authors report experiments on Themis, described as Llama-3-8B-based, and Prometheus, described as Mistral-7B-based. They evaluated summary Readability and Adequacy using controlled perturbations and report applying causal tracing, logit-lens vocabulary projection and attention-head knockout. [1]
04
The authors’ proposed mechanism
The paper attributes a two-stage pipeline to both fine-tuned evaluator models. Below layer 15, attention reportedly performs local error comparison and routes the result to the final input position. At higher layers, MLP layers reportedly integrate the signal and write the rating. This is the authors’ interpretation of the tested models, not an established general explanation for LLM evaluators. [1]
The authors report that the decision signal crystallizes at layer 26 in Themis and layer 25 in Prometheus. Their same-scale Llama-3-8B base-model control reportedly reproduced routing and crystallization, but not the reported separation between early attention work and later MLP work. The authors further interpret the comparison as showing two fine-tuning changes: suppression of below-layer-15 MLP contribution at the final position, and crystallization occurring two layers earlier. [1]
05
Why the distinction matters
The reported control result means the paper does not describe fine-tuning as creating every part of the pathway from scratch. Instead, the authors argue that fine-tuning modifies a pre-existing routing and crystallization pattern. That interpretation remains a claim from this preprint and has not been independently corroborated in the supplied material. [1]
For readers assessing automated summary scoring, the work suggests a bounded audit question: where does an evaluator move error-related information, and at what layer does a rating become fixed? It does not establish that this question, method or reported layer pattern will apply to other models or evaluation tasks. The primary record is https://arxiv.org/abs/2609.01604v1. [1]
06
How to use this result
Treat the paper as a model-specific research report, not as a general account of how automated evaluators work.
- 01
Read the primary arXiv record and separate the authors’ mechanistic interpretations from independently replicated evidence.
- 02
When using an LLM to score summaries, test the evaluator, task and quality dimensions in use rather than assuming the reported pattern transfers.
- 03
Verify the authors’ code and data materials directly before relying on them, because the fully expanded repository address is not available in the supplied arXiv text. [1]
07
Limits of this edition
This is a preprint record, not independent verification of its findings. [1]
The supplied evidence includes no independent replication, review materials or separate conference source confirming the findings or the record’s stated acceptance status. [1]
It is unknown whether the reported mechanism extends beyond Themis, Prometheus, the Llama-3-8B control, Readability and Adequacy, or summarization evaluation. [1]
The arXiv record says that source code and data are released, but the supplied text does not expose a fully expanded repository URL or release dates. [1]
SRC
Source desk
Direct links to the material behind this selection. Seeing the source matters as much as reading the synthesis.


