Preprint reports a model-harness mismatch in full-trajectory imitation
The authors report that full expert-trajectory imitation harmed weaker models when used with harnesses evolved around those models, while a method that corrects only a failing turn in the weaker model’s own rollout may preserve model-harness fit. The claims remain unreviewed and lack task-level and reproducibility detail in the supplied record. [1]
This edition passed Imananq's enhanced publication checks. Some material claims remain explicitly attributed to official or company sources because no independent source is currently bound to this edition. The engine continues checking approved sources and will add corroboration only through a new edition that passes the full gate.

A newly submitted arXiv preprint reports that fine-tuning a weaker AI model on a stronger model’s complete task trajectories can reduce performance when the surrounding agent harness was first evolved around the weaker model. Its authors propose a narrower, turn-level expert-correction approach, but the supplied record does not establish independent validation or reproducibility. [1]
01
What we know now
- 01
[1] arXiv record and abstract for “Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails”: https://arxiv.org/abs/2609.09134v1 (submitted 8 September 2026; primary source; preprint). [1]
- 02
The authors report results across seven enterprise agent tasks, including a 4-to-30-point regression after complete expert-trajectory fine-tuning under an evolved harness, and describe a turn-level on-policy correction pipeline. [1]
02
Study setting reported by the authors
The abstract describes seven enterprise agent tasks and tests involving Qwen3-Coder and Gemma 4.Reported regression after full imitation
Under an evolved harness, the authors report declines of 4 to 30 points on every tested task.Proposed intervention
The expert rewrites a failing turn found in the weaker model’s own rollout, rather than supplying a full trajectory.These are author-reported preprint findings, not independently verified results. [1]
03
What the preprint reports
The paper, “Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails,” was submitted to arXiv’s Artificial Intelligence category by Zhou Yu and ten co-authors. The arXiv record is the primary source for the authors’ reported results. [1]
The central warning is conditional: complete-trajectory imitation may not combine safely with a scaffold tailored to a weaker model. The record does not provide enough detail to determine which task types, evaluation measures, or harness features drive the reported effect. [1]
- A harness is described as the prompt, tools, execution hooks, and context-management scaffolding around a model.
- The authors say they first evolved a harness using a weaker model, then observed that a stronger expert often used that harness more effectively.
- They report that training the weaker model on complete expert trajectories under that evolved harness reduced performance on all seven tested enterprise tasks by 4 to 30 points across Qwen3-Coder and Gemma 4. They also say the same training procedure helped with an unevolved harness.
04
Why complete expert demonstrations may fail here
The paper frames the issue as an interaction between model weights and scaffolding. A harness that has evolved around a model’s native planning behaviour may cease to fit after the model is trained to follow a stronger expert’s approach. [1]
That account should not be generalized beyond the reported study. The supplied evidence contains neither a causal ablation record nor external evaluations that would confirm the explanation across other models or agent systems. [1]
- The authors attribute the regression to disrupted model-harness fit.
- In their account, imitation transfers some knowledge and increases use of the scaffold, but the weaker model adopts the expert’s planning strategy without having enough competence to carry it out.
- This is an author-proposed explanation, not an independently established cause.
05
The narrower correction method
The authors call their alternative an on-policy expert-correction pipeline. They say the method preserves the weaker model’s planning style while adapting it at a locally identified failure point, and they report that it combines gains from harness evolution and model adaptation. [1]
The abstract does not specify how failing turns are detected, how corrections are selected or trained on, what expert model is used, or the comparative measurements behind the claimed combined gains. Those omissions mean the method cannot be assessed from this record as a reproducible recipe. [1]
- Start from the weaker model’s own rollout.
- Locate a failing turn, using a process the authors say is automated by a meta-level MLE agent.
- Ask the expert to rewrite only that turn rather than producing a complete replacement trajectory.
06
Practical reading of the result
For teams evaluating agent systems, the useful takeaway is to measure the model and its harness as a coupled system. The study’s reported regression suggests that a stronger model’s demonstrations can alter a weaker model in ways that do not suit the environment optimized for the weaker model’s original behaviour. [1]
This is a hypothesis and evaluation consideration, not a proven general rule. The reported scope is seven unspecified enterprise agent tasks, and the primary record offers no independent replication, public materials, or task-level evidence in the supplied packet. [1]
- Do not treat a successful fine-tuning result under one scaffold as evidence it will work under a different or evolved scaffold.
- Evaluate end-to-end task outcomes, not only imitation quality or the rate at which a model uses tools and scaffold components.
- Keep full-trajectory imitation and localized correction as separate experimental conditions when testing an evolved harness.
07
What to check before applying the idea
The preprint supports a narrow evaluation lesson rather than a ready-to-deploy procedure.
- 01
Test model fine-tuning together with the exact harness it will use, rather than assuming gains transfer between unevolved and evolved scaffolding.
- 02
Compare complete expert-trajectory imitation with interventions that modify only identified failure points in the weaker model’s own rollout.
- 03
Treat the reported correction pipeline as a research claim until task definitions, measurements, implementation details, and independent results are available. [1]
08
Limits of this edition
This is an arXiv preprint submitted on 8 September 2026, not peer-reviewed evidence. [1]
The supplied material does not identify the seven tasks, their metrics, baselines, task-level results, or implementation details. [1]
No code, data, checkpoints, harness configurations, or independent replication are established by the supplied record. [1]
SRC
Source desk
Direct links to the material behind this selection. Seeing the source matters as much as reading the synthesis.


