> ## Content Index
> Fetch the complete content index at: https://www.notatechguy.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# LLM pipelines lose 40 accuracy points at interfaces
- URL: https://www.notatechguy.com/llm-pipelines-lose-40-accuracy-points-at-interfaces/
- Published: 2026-09-29T02:35:41.000Z
- Updated: 2026-09-29T02:35:41.000Z
- Description: Pipeline interfaces destroy reasoning accuracy across 21 open-weight models, with a re-grounding fix holding on all three newest models tested.
- Author: Marcello Babbili
- Tags: Technology & AI, AI Models

Researchers testing 21 open-weight models found that a four-stage LLM pipeline can surrender 40.5 accuracy points at its own interfaces, the largest loss in their primary test family [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). The received view says breaking a hard problem into steps helps the model reason. The team, posting on arXiv in cs.AI and cs.LG on 26 September 2026, reports that the joints between those steps are where the reasoning dies [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). The 40.5-point figure is from a preprint that has not been peer-reviewed, and the decomposition tax has not been tested outside math benchmarks [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com).

**My read:** This is the first systematic isolation of interface loss I have seen that holds model, problem, prompts and budget fixed and varies only what each stage can see. The 40.5-point figure is the worst case, not the average, and it is on math benchmarks only. But the placebo result, where a stage carries 60% of the extra tokens and recovers nothing, tells me the loss is about missing relationships, not missing words. I would watch whether this holds on non-math tasks before calling it universal.

### Holding everything fixed isolates the handoff cost

The "decomposition tax" measures what accuracy vanishes when everything else — model, problem, stages, stage prompts, completion budget, stays constant and the only thing that changes is whether each stage can still see the original problem [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). Strip that visibility and you measure what the handoff costs.

Across 21 open-weight models from nine organisations, the researchers ran 200 paired items per cell on two math benchmarks, GSM-Hard and MATH-500 [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). Of 118 primary-family tests, 70 survive Benjamini-Hochberg correction and 54 survive the stricter Holm correction [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). The tax is real in roughly half the tests under the strictest standard.

Not universal, but widespread enough to take seriously. Gemma-4-12B gives up 37.0 points at its interfaces [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). Gemma-3-12B hits the 40.5-point ceiling on MATH-500, with a Holm-corrected p-value of 1.66e-19 [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com).

The failure is often not in the model's reasoning but in what it can see at the moment it needs to reason. A separate preprint by Jeonghun Yoon at KAIST and Dongchan Kim at NAVER describes how patching one module in a multi-module LLM agent can break adjacent modules through linguistic co-adaptation [P²](https://arxiv.org/pdf/2605.21958.pdf?ref=notatechguy.com). The decomposition tax is the quieter version of the same problem: the patch is not malicious, it is just a summary.

### One rewritten instruction, a 32-point swing

Rewriting one stage's instruction moved gemma-3-12B's tax from 4.5 to 36.5 points [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). The same model, the same problem, the same pipeline shape. A change in wording produced an eightfold difference in what the interface costs.

![gemma-3-12B decomposition tax: effect of rewriting one stage instruction](https://storage.ghost.io/c/6e/89/6e896869-22ef-4281-a213-b4c462c17cff/content/images/2026/09/chart_e723c442fd4bcf43ec18.png)

A placebo test on GSM-Hard sharpens the picture. When a stage absorbed at least 60% of the surplus tokens yet retained no more than a single word of the original problem, it recovered zero accuracy [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). The loss is not about token volume. It is about the relationships between quantities, the connective tissue that a summary strips out.

The damage happens at the boundary between components, not inside them. A related preprint by Kaituo Zhang at the University of Houston and colleagues documents a "tool-use tax" on LLM agents, a separate accuracy cost imposed by the act of calling external tools [P⁴](https://arxiv.org/html/2605.00136?ref=notatechguy.com). The decomposition tax is the internal version: you do not need an external tool to lose information, just another stage of your own pipeline.

### Re-ground after the break, not before

The researchers prescribe two repairs. First, re-ground the stage after the lossy interface, not before it. With one lossy interface, re-grounding after the break beat re-grounding before it on 7 of 7 models on both benchmarks [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). On MATH-500, the earlier repair was worse than doing nothing at all on 7 of 7 models [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com).

Fixing the wrong stage actively harms performance.

Second, if a stage must list numerical quantities, instruct it to keep the relationships. Adding the phrase "every relationship stated between them" to a stage that lists quantities lowered the tax on 9 of 9 models on MATH-500 [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). The repair holds on all three of the newest models tested [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com).

A sealed held-out test refuted a stronger rule the researchers had registered, which predicted the paying stage from the interface and receiver types [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com). They could not predict which stage would suffer, only where to patch after the fact.

For a legal-tech engineer running a multi-stage contract-review pipeline, the change on the desk is concrete: before adding any new extraction stage, verify that downstream stages still receive the original contract text. If they do not, insert a re-grounding step immediately after the lossy handoff, and if any stage extracts figures or clauses, instruct it to preserve the relationships between them.

The preprint is open on arXiv for community review, and the findings are limited to open-weight models on math benchmarks. Sixty-four of 118 primary-family tests did not survive Holm correction [S¹](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com), and nobody has tested whether the decomposition tax appears on non-mathematical tasks yet.

---

*Sources: [S1 — The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at ](https://arxiv.org/abs/2609.32825v1?ref=notatechguy.com) · [P2 — Diagnosis Is Not Prescription: Linguistic Co-Adaptation Explains Patch](https://arxiv.org/pdf/2605.21958.pdf?ref=notatechguy.com) · [P3 — Pipeline Parallelism with Controllable Memory](https://arxiv.org/html/2405.15362?ref=notatechguy.com) · [P4 — Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents](https://arxiv.org/html/2605.00136?ref=notatechguy.com) · [P5 — WePOINTS/WePOINTS](https://github.com/WePOINTS/WePOINTS?ref=notatechguy.com)*

---

*Written from 5 sourced items, 4 of them primary.*

## More from Not A Tech Guy

- [Holo4 27B scores 61.7% on OSWorld 2.0, trails Opus 5.5](https://www.notatechguy.com/holo4-27b-scores-61-7-on-osworld-2-0-trails-opus-5-5/)
- [LLVM trends on GitHub as AI agents target compiler code](https://www.notatechguy.com/llvm-trends-on-github-as-ai-agents-target-compiler-code/)
- [AI security tool reverse-skill hits 38,000 GitHub stars](https://www.notatechguy.com/ai-security-tool-reverse-skill-hits-38-000-github-stars/)