A paper posted to arXiv on 3 August ran four autonomous research frameworks through automated peer review by three frontier language models . The best any system scored was 2.47 out of 5 . That is below the midpoint on a 1-to-5 scale. The study claims to establish the first quantitative benchmark for AI scientist systems, but the number that should stop you is not who won. It is that the winner is still failing, and one of the three AI reviewers appears to be grading on a completely different rubric.

My read: This is the first systematic attempt I have seen to grade AI-generated science using AI reviewers, and the results cut both ways. Gemini and Claude agree with each other almost perfectly (Spearman correlation 0.907), which suggests automated review could become a real screening tool for research labs drowning in AI-generated submissions. But the absolute scores tell a harder story: even the best system averages 2.47 out of 5. I do not buy the framing that this proves FARS is the superior system, because FARS supplied both the proposals and the benchmark papers. That is a structural advantage the abstract does not address. And GPT-5.4's near-zero correlation with the other two reviewers (approximately 0.32) is the detail I would watch most closely.

How the test was built

The authors ran four AI scientist frameworks, Sakana AI (versions 1 and 2), CycleResearcher, and Data-to-Paper, on a fixed set of 15 research proposals . The proposals came from FARS, a commercial autonomous AI scientist company. Each framework produced a paper from each proposal, generating 60 AI-written papers in total. Those were evaluated alongside 15 papers produced by FARS itself as a benchmark .

Three frontier language models acted as reviewers: GPT-5.4, Gemini, and Claude . Each reviewer scored every paper on four dimensions: originality, scientific rigor, clarity, and significance . According to the study's GitHub repository, the experiment matched several competing automated research systems against a hand-selected set of AI-generated reference papers on an AI safety subject, then submitted the entire collection to LLM-based review P⁴.

The scores nobody is celebrating

FARS benchmark papers scored between 2.14 and 2.47 on average across the three reviewers . Every other framework scored between 1.00 and 1.87 . On Gemini and Claude evaluations specifically, FARS scored more than twice as high as the next-best system .

But read those numbers again. The best score, 2.47 out of 5, sits below the midpoint. On a grading scale where 3 would be adequate, the strongest AI-generated science in this study is still inadequate by its own automated reviewers' standards. The worst framework averaged 1.00, which is the floor.

This matters because AI scientist systems are already being deployed. Sakana AI's framework has been publicly available since 2024. If the first systematic benchmark shows that even the leading system produces papers rated below average by automated reviewers, the gap between what these tools promise and what they deliver is wide.

When AI reviewers disagree

The agreement between Gemini and Claude was remarkably strong, with a Spearman correlation of 0.907 (p < 0.001) . Both also correlated extremely strongly with the synthesis score, the combined metric across all reviewers, at 0.961 (p < 0.001) .

How strongly AI reviewers agree with each other

GPT-5.4 told a different story. Its correlation with the other reviewers was approximately 0.32, which the authors interpret as evidence that it evaluates papers using different criteria . In practical terms, a paper that Gemini and Claude both rated highly could receive a completely different score from GPT-5.4, with no clear explanation of why.

The designation GPT-5.4 is unusual. It does not match any widely recognised public model name as of this writing, and the paper's abstract does not clarify the version or provider. This creates a verification risk: if the model identity cannot be independently confirmed, its divergent scores are hard to interpret.

The FARS question

The study's structure raises a conflict-of-interest concern that the abstract does not address. FARS, a commercial company, supplied the 15 research proposals that all four competing frameworks worked from. FARS also supplied the 15 benchmark papers against which those frameworks were compared . The authors do not state their affiliations in the abstract, so it is unclear whether they are independent of FARS or connected to it.

Even without a formal conflict, the setup gives FARS a structural advantage. Its benchmark papers were produced by its own system, tuned to its own proposals, and evaluated by criteria that its own workflow may have been optimised to satisfy. The competing frameworks were working from FARS-designed prompts, not their own native research agendas.

The authors state that their results establish the first quantitative benchmark for AI scientist systems and that multi-model LLM evaluation provides a scalable, consistent framework for assessing autonomous research quality . Both claims may prove correct. But the benchmark's first run shows a system winning a race it helped design.

What to do about it

For a research lab manager considering AI scientist tools, this paper offers a useful protocol even if its conclusions are premature. The four-dimension rubric (originality, rigor, clarity, significance) is a reasonable starting point for evaluating any AI-generated research output. The multi-model approach, using two or more frontier LLMs as reviewers, provides a consistency check that a single reviewer cannot.

Consider a pharmacology lab that wants to screen AI-generated literature reviews before a human researcher reads them. Running each paper through two independent LLM reviewers and flagging any paper where the reviewers disagree by more than one point would catch the GPT-5.4 problem: when two models see the same paper differently, that paper needs human eyes before it enters a research pipeline.

One practical step this week: if your team uses any AI research assistant, pick five recent outputs and run them through two different frontier models with a simple 1-to-5 rubric. If the scores diverge by more than a point on any paper, you have found the limit of automated quality control for your current workflow.

What we don't know yet

The paper has not been peer-reviewed. It is an arXiv preprint, and its methodology, data, and conclusions will need independent verification .

Several critical gaps remain. The study does not compare its automated reviewers' scores against human expert judgment. Without that comparison, we cannot know whether a 2.47 from Gemini and Claude means the same thing as a 2.47 from a human reviewer. The absolute scores, all below the midpoint, could reflect genuinely poor output, an overly harsh rubric, or a systematic bias in how language models rate AI-generated text.

The GPT-5.4 model identity needs clarification. If it is an internal or unreleased version, its divergent behaviour may not generalise to publicly available models.

And the FARS relationship needs disclosure. The study's GitHub repository was created on 3 April 2026 P⁴, four months before the paper appeared. The authors' affiliations, when they surface in a full version, will determine whether this is an independent benchmark or a company-validated product test.

The next signal: a peer-reviewed version of this paper, or an independent replication using non-FARS proposals and human reviewer scores. If you want to follow that thread with us, the subscribe button is below.


Sources: S1 — Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Rese · P2 — Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Rese · P3 — tongjingqi/AI-Can-Learn-Scientific-Taste · P4 — vaibhavalakshmiravideshik/AI-Research-Automation-Study · P5 — Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchm

More from Not A Tech Guy


Generated from an audited evidence pack with primary-source research. Social-media items are discussion signals, not verified facts. Nothing here is financial, legal or medical advice.