> ## Content Index
> Fetch the complete content index at: https://www.notatechguy.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# GRPO step-level bias: GRAFT graph method claims gains on agent tasks
- URL: https://www.notatechguy.com/grpo-step-level-bias-graft-graph-method-claims-gains-on-agent-tasks/
- Published: 2026-09-27T20:13:14.000Z
- Updated: 2026-09-27T20:13:14.000Z
- Description: An arXiv preprint proposes GRAFT, a graph framework for step-level credit in agentic RL, reporting gains over GRPO on multi-turn benchmarks.
- Author: Marcello Babbili
- Tags: Technology & AI, AI Agents

arXiv published a preprint on 25 September 2026 whose authors argue that GRPO, the dominant training method for reasoning models, is systematically biased when scoring individual steps in multi-step agent tasks [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com). The preprint proposes GRAFT, a framework that stitches agent trajectories into a shared graph to estimate step-level value, and claims consistent gains over GRPO across multi-turn benchmarks [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com).

Those gains are author-reported in the abstract with no disclosed metrics, benchmarks, or statistical tests [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com). No error bars are provided, no code has been released, and the figures are the authors' own [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com).

**My read:** This is the clearest attack I have seen on the step-level credit assignment problem in agentic RL. The graph construction is elegant: by merging trajectories that pass through the same state, GRAFT gets multiple samples at intermediate nodes without the prohibitive cost of explicitly sampling from each one. But I do not buy the superiority claims yet. The abstract names no benchmarks, no metrics, no baselines beyond GRPO, and no error bars. The code is promised at a GitHub repository that was not live at the time of writing. Until someone independent reproduces these results, GRAFT is a well-motivated idea, not a proven method.

### GRPO's step-level scoring shows systematic bias

The authors argue that GRPO succeeds in single-turn scenarios by drawing several responses from one prompt and averaging their rewards, which provides a reliable estimate of the initial state's value [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com).

Multi-step agents disrupt this assumption. When an agent browses the web, calls APIs, or writes and debugs code, it passes through numerous intermediate states. Accurately estimating the value of each state would require sampling multiple actions from every single one, a process the authors deem too expensive [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com). As a result, practitioners rely on trajectory-level scoring, where the entire sequence receives a single reward that every step inherits. This means a failed run might penalize valuable steps, while a successful run might reward wasteful ones, creating what the authors describe as a systematic bias [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com).

### GRAFT merges trajectories to estimate step value

GRAFT combines all rollout trajectories from a training batch into one directed graph. If two trajectories arrive at the same state — sharing the same observation and context, that state becomes a shared node with multiple outgoing edges. The authors then apply Bellman iteration to the graph, moving backward from terminal rewards to estimate each node's value [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com). A single step's advantage is calculated as the difference between the node's value before and after that step.

The authors also introduce Graph GAE, adapting Generalized Advantage Estimation — a standard variance reduction technique in RL training, to the trajectory graph. This adaptation mitigates the effects of inaccurate value estimates at individual nodes [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com). They claim their estimated step-level advantage aligns with the fundamental definition of advantage in reinforcement learning [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com).

GRAFT is not the only graph-based approach to this problem. A separate arXiv paper, "Beyond Trajectory-Level Attribution" by Xin Cheng, Shuo He, and colleagues [P⁴](https://arxiv.org/html/2605.26684?ref=notatechguy.com), tackles the same problem with a similar graph construction for agentic RL. The broader field is moving quickly: THUDM's AgentRL [P⁵](https://github.com/THUDM/AgentRL/?ref=notatechguy.com), a multi-turn agentic RL training framework with 337 stars and 27 forks on GitHub, shows the demand for better step-level training, and an ACL 2026 Findings paper on step-level advantage selection [P²](https://github.com/HanNight/SAS?ref=notatechguy.com) adds to the research momentum.

### Gains remain unverified without code or metrics

The paper reports "consistent gains over GRPO and superior performance compared to recent agentic RL algorithms" [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com). The abstract does not name the benchmarks, disclose specific metrics, or report statistical significance. The paper is not peer-reviewed. No author affiliations or funding disclosures appear in the source. Code is promised at github.com/xcyao00/GRAFT but was not available at the time of writing.

For an RL engineer training a web-browsing agent, GRAFT's graph approach would mean shifting from trajectory-level rewards to step-level credit, allowing them to pinpoint which specific API calls failed rather than penalizing the entire session [S¹](https://arxiv.org/abs/2609.28963?ref=notatechguy.com). A team training a coding agent that writes, tests, and debugs in sequence would benefit most from this approach. Instead of labelling an entire failed debugging session as negative, the trainer could identify which specific steps, the initial code generation or a wrong fix, dragged the outcome down. That finer signal means each training run teaches more.

The next checkpoint is whether the code lands at the stated repository and whether the experimental details survive independent review.

---

*Sources: [S1 — Back to the Definition: Estimating Step-Level Advantages via Trajector](https://arxiv.org/abs/2609.28963?ref=notatechguy.com) · [P2 — HanNight/SAS](https://github.com/HanNight/SAS?ref=notatechguy.com) · [P3 — gorkaydemir/FTGN](https://github.com/gorkaydemir/ftgn?ref=notatechguy.com) · [P4 — Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for](https://arxiv.org/html/2605.26684?ref=notatechguy.com) · [P5 — THUDM/AgentRL](https://github.com/THUDM/AgentRL/?ref=notatechguy.com)*

---

*Written from 5 sourced items, 4 of them primary.*

## More from Not A Tech Guy

- [Chinese AI text humanizer trends with 18,000 GitHub stars](https://www.notatechguy.com/chinese-ai-text-humanizer-trends-with-18-000-github-stars/)
- [Rivet Actors hits GitHub trending with 20ms cold-start claim](https://www.notatechguy.com/rivet-actors-hits-github-trending-with-20ms-cold-start-claim/)
- [Pistis multimodal models blend distillation and RL training](https://www.notatechguy.com/pistis-multimodal-models-blend-distillation-and-rl-training/)