StudentBench preprint reports comparable AI and human GRE tutoring gains
The StudentBench authors report that AI tutoring was statistically equivalent to expert human tutoring for GRE learning gains in their study. The supplied evidence does not establish peer review or independent replication, and the abstract leaves key methodological and cost details unknown.
This edition passed Imananq's enhanced publication checks. Some material claims remain explicitly attributed to official or company sources because no independent source is currently bound to this edition. The engine continues checking approved sources and will add corroboration only through a new edition that passes the full gate.
A September 23, 2026 arXiv preprint reports that AI tutoring produced learning gains statistically equivalent to expert human tutoring in a study using GRE questions. Peer review and independent replication are not established by the supplied evidence. [1]
01
02
Participants in the reported tutoring comparison
The authors report 2,383 participants receiving AI tutoring, human tutoring, or no tutoring on GRE questions.GRE domains where the best AI averaged above the human tutor
The authors report this average comparison in five of seven domains. The supplied abstract does not provide enough detail to assess the systems or conditions.Reported cost difference per percentage point gained
The authors report USD 0.0052 for one AI tutor and USD 4.81 for human tutoring. The abstract does not provide the full cost assumptions.These figures are author-reported findings in an arXiv preprint; the supplied evidence does not independently validate them.
03
What was published
The work was posted as an arXiv preprint. The supplied record confirms the paper’s title, authorship, and submission date, but does not establish whether it has undergone peer review. [1]
- The arXiv record lists “StudentBench: AI and human tutoring yield equivalent GRE learning gains” by Curtis Northcutt and five coauthors as version 1, submitted on September 23, 2026.
- The authors describe StudentBench as an AI-teaching evaluation suite and public platform, and report more than 175,000 student-AI messages. The supplied evidence does not independently verify the platform, message total, or current availability.
04
What the authors report
These are findings reported by the preprint’s authors, rather than independently established results. The supplied abstract does not provide enough methodological detail to assess the comparison or determine the underlying meaning and basis of statistical equivalence. [1]
- The authors report 2,383 participants receiving AI tutoring, human tutoring, or no tutoring on Quantitative and Verbal GRE questions.
- They report that AI tutoring was statistically equivalent to expert human tutoring for learning gains, with p = .015.
- They report that their best-performing AI tutor averaged above the human tutor in five of seven GRE domains.
- They also report a second study with 2,028 pairwise rubric evaluations by expert human tutors of LLM-generated lesson plans and practice problems.
05
The cost claim has important gaps
The figures are a reported comparison, not an independently established measure of the real-world price or value of AI tutoring. The supplied abstract does not set out the full cost assumptions or show whether all relevant operational costs were included. [1]
- For one AI tutor, the authors report learning gains equivalent to human tutoring, with p = .044.
- They report a cost of USD 0.0052 per percentage point gained for that AI tutor and USD 4.81 for human tutoring, which they describe as a 918-fold difference.
06
Why the boundary matters
The preprint presents a specific result about GRE-question tutoring, not proof that AI tutoring generally matches expert human tutoring. Readers should separate the authors’ reported study results from broader conclusions about tutoring quality, cost, or effectiveness. [1]
- Primary source: https://arxiv.org/abs/2609.28470v1
- The supplied record includes the paper’s title, authors, submission date, and abstract, but not the detail needed for an independent assessment of the experiment.
07
How to read the result
Treat this as a narrowly described research report and review the primary record before relying on it for a tutoring decision.
- 01
Read the arXiv record and look for the full methods, including participant procedures, outcome measures, tested systems, and cost assumptions.
- 02
Do not extend the reported GRE-specific comparison to other subjects or tutoring settings without further evidence.
- 03
Compare tutoring options using the learner’s needs, coverage, feedback, accessibility, privacy terms, and total cost.
08
Limits of this edition
The supplied evidence is an arXiv preprint record and abstract. It does not establish peer review, independent replication, or independent validation. [1]
The abstract does not provide enough detail to assess participant procedures, learning measures, equivalence criteria, statistical methods, tested AI systems, or human-tutoring conditions. [1]
The reported cost comparison does not include the full underlying assumptions, so the supplied evidence cannot establish whether all relevant operational costs were counted. [1]
The reported findings concern the GRE questions and participants described by the authors. Their applicability to other subjects, learners, or tutoring settings is not established. [1]
SRC
Source desk
Direct links to the material behind this selection. Seeing the source matters as much as reading the synthesis.



