Hanyu Wang, a Penn State researcher who did this work as an intern at Honda Research Institute USA, and colleagues posted a paper on arXiv on October 9 that tackles a dead-end in training AI to reason S¹P². When a language model attempts a math problem and every try fails, reinforcement learning has no signal to learn from. The obvious fix, giving the model more of the answer, helps in one way and hurts in another, and the paper shows why S¹.

My read: This is the clearest treatment I've seen of the all-failure group problem in GRPO. Anyone training reasoning models with RL has hit it, but few have broken down why the naive fix, longer prefixes, doesn't monotonically help. The KL divergence derivation is the real contribution. I don't buy the pass@12 ranking yet, because we have no specific benchmark numbers, no peer review, and the results are entirely self-reported on a single model family. I'd want to see this reproduced on something outside Qwen3 before calling it settled.

When every answer is wrong

GRPO, Group Relative Policy Optimization, samples several attempts at the same problem and compares them. If some succeed and some fail, the model learns to prefer what worked. But when every attempt in a group fails, there is no comparison to make. The group produces no learning signal. The model stays stuck.

This is not a rare edge case. As reasoning problems get harder, the fraction of all-failure groups rises. A model that solves 30 percent of training problems on its own produces all-failure groups for the other 70 percent, and those are the problems where it needs to learn most. The problem compounds: the harder the task, the less the model can learn, precisely because it fails too often to get feedback. Reasoning models are increasingly the target of both improvement efforts and security research, as we found when backdoors hid attacks inside logical reasoning chains.

A head start that shrinks at the right rate

Wang and colleagues study prefix continuation S¹. The model starts from part of a known-correct solution, a prefix, and generates the rest. If the completed trajectory passes verification (produces the right answer), the model keeps it. If not, it falls back to the reference solution S¹.

The non-obvious part is how long that prefix should be. The paper derives the KL divergence, a measure of how far the model's output distribution drifts from the reference, for a single continuation in closed form S¹. Up to a bounded term, this divergence shrinks with the product of two quantities: the probability that the model generates a correct but different trajectory, and the reference surprisal, which measures how surprising the reference solution is to the model S¹.

A longer prefix raises the first term, because the model is more likely to produce something correct when given more of the answer. But it lowers the second, because the reference solution becomes less surprising when the model has already seen most of it S¹. The two effects pull in opposite directions. Continuation success alone does not tell you the right prefix length.

From this analysis, the authors learn a prefix selector, shared across all training questions, that picks prefix lengths based on continuation outcomes S¹. The selector needs no success-probability estimates and no extra generation passes.

What the experiments show, and what they don't

The resulting method, Adaptive Reference Guidance, builds correct trajectories within a fixed generation budget and applies specifically to all-failure groups in GRPO S¹. The authors tested it on Qwen3-4B and Qwen3-8B across five mathematical reasoning benchmarks S¹.

They report that ARG achieves the highest aggregate pass@12 among the evaluated methods, with competitive average sampled accuracy S¹. Pass@12 measures whether at least one of 12 sampled attempts reaches the right answer. The paper does not publish specific per-benchmark numbers in its abstract, and the results have not been independently verified.

The paper is an arXiv preprint and has not been peer-reviewed S¹. Author affiliations span Penn State and Honda Research Institute USA, according to the paper's HTML version P². No code link appears in the abstract. A separate effort from a different research group, ReBalance, appeared at ICLR 2026 and also targets efficient reasoning, with 130 stars on GitHub under an MIT licence P³. The broader problem of making reasoning training more efficient is attracting multiple groups, much as retrieval methods lifted code generation on CoderEval in a related vein.

For a team training reasoning models with GRPO, the practical first step is to measure what fraction of training groups are all-failure. That number determines whether ARG would help. If most groups already produce at least one correct attempt, rescue methods add little. But when most groups fail entirely, as happens with harder problem sets, that gap is where ARG operates.

The paper is available on arXiv as 2610.11128, posted October 9, 2026 S¹. The next checkpoint for anyone evaluating it is independent reproduction on a model outside the Qwen3 family, which the paper does not provide.


Sources: S1 — Balancing Reference Guidance and Free Generation in Trajectory Rollout · P2 — Balancing Reference Guidance and Free Generation in Trajectory Rollout · P3 — yu-lin-li/ReBalance · P4 — FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoni · P5 — aiyouthalliance/Free-Image-Generation · Hugging Face


Written from 5 sourced items, 4 of them primary.