arXiv published a preprint on 25 September 2026 whose authors argue that GRPO, the dominant training method for reasoning models, is systematically biased when scoring individual steps in multi-step agent tasks S¹. The preprint proposes GRAFT, a framework that stitches agent trajectories into a shared graph to estimate step-level value, and claims consistent gains over GRPO across multi-turn benchmarks S¹.
Those gains are author-reported in the abstract with no disclosed metrics, benchmarks, or statistical tests S¹. No error bars are provided, no code has been released, and the figures are the authors' own S¹.
My read: This is the clearest attack I have seen on the step-level credit assignment problem in agentic RL. The graph construction is elegant: by merging trajectories that pass through the same state, GRAFT gets multiple samples at intermediate nodes without the prohibitive cost of explicitly sampling from each one. But I do not buy the superiority claims yet. The abstract names no benchmarks, no metrics, no baselines beyond GRPO, and no error bars. The code is promised at a GitHub repository that was not live at the time of writing. Until someone independent reproduces these results, GRAFT is a well-motivated idea, not a proven method.
GRPO's step-level scoring shows systematic bias
The authors argue that GRPO succeeds in single-turn scenarios by drawing several responses from one prompt and averaging their rewards, which provides a reliable estimate of the initial state's value S¹.
Multi-step agents disrupt this assumption. When an agent browses the web, calls APIs, or writes and debugs code, it passes through numerous intermediate states. Accurately estimating the value of each state would require sampling multiple actions from every single one, a process the authors deem too expensive S¹. As a result, practitioners rely on trajectory-level scoring, where the entire sequence receives a single reward that every step inherits. This means a failed run might penalize valuable steps, while a successful run might reward wasteful ones, creating what the authors describe as a systematic bias S¹.
GRAFT merges trajectories to estimate step value
GRAFT combines all rollout trajectories from a training batch into one directed graph. If two trajectories arrive at the same state — sharing the same observation and context, that state becomes a shared node with multiple outgoing edges. The authors then apply Bellman iteration to the graph, moving backward from terminal rewards to estimate each node's value S¹. A single step's advantage is calculated as the difference between the node's value before and after that step.
The authors also introduce Graph GAE, adapting Generalized Advantage Estimation — a standard variance reduction technique in RL training, to the trajectory graph. This adaptation mitigates the effects of inaccurate value estimates at individual nodes S¹. They claim their estimated step-level advantage aligns with the fundamental definition of advantage in reinforcement learning S¹.
GRAFT is not the only graph-based approach to this problem. A separate arXiv paper, "Beyond Trajectory-Level Attribution" by Xin Cheng, Shuo He, and colleagues P⁴, tackles the same problem with a similar graph construction for agentic RL. The broader field is moving quickly: THUDM's AgentRL P⁵, a multi-turn agentic RL training framework with 337 stars and 27 forks on GitHub, shows the demand for better step-level training, and an ACL 2026 Findings paper on step-level advantage selection P² adds to the research momentum.
Gains remain unverified without code or metrics
The paper reports "consistent gains over GRPO and superior performance compared to recent agentic RL algorithms" S¹. The abstract does not name the benchmarks, disclose specific metrics, or report statistical significance. The paper is not peer-reviewed. No author affiliations or funding disclosures appear in the source. Code is promised at github.com/xcyao00/GRAFT but was not available at the time of writing.
For an RL engineer training a web-browsing agent, GRAFT's graph approach would mean shifting from trajectory-level rewards to step-level credit, allowing them to pinpoint which specific API calls failed rather than penalizing the entire session S¹. A team training a coding agent that writes, tests, and debugs in sequence would benefit most from this approach. Instead of labelling an entire failed debugging session as negative, the trainer could identify which specific steps, the initial code generation or a wrong fix, dragged the outcome down. That finer signal means each training run teaches more.
The next checkpoint is whether the code lands at the stated repository and whether the experimental details survive independent review.
Sources: S1 — Back to the Definition: Estimating Step-Level Advantages via Trajector · P2 — HanNight/SAS · P3 — gorkaydemir/FTGN · P4 — Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for · P5 — THUDM/AgentRL
Written from 5 sourced items, 4 of them primary.