> ## Content Index
> Fetch the complete content index at: https://www.notatechguy.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# RECAST derives AI evidence, beating strongest baseline 15.9%
- URL: https://www.notatechguy.com/recast-derives-ai-evidence-beating-strongest-baseline-15-9/
- Published: 2026-10-10T13:53:18.000Z
- Updated: 2026-10-10T13:53:18.000Z
- Description: RECAST's RouterLM learns to derive evidence through computation and retrieval, hitting 75.6% success across six benchmark families.
- Author: Marcello Babbili
- Tags: Technology & AI, Google, OpenAI, AI Models

RECAST's authors posted the framework to arXiv on October 7, reporting a 75.6% mean success rate across six benchmark families [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com). That result outperforms the strongest large-model baseline they tested by 15.9 percentage points, though the baseline is not named in the available paper excerpts [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com).

The preprint has not been peer-reviewed, and all performance figures are self-reported [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com).

Conventional retrieval-augmented generation, or RAG, fetches text chunks that resemble a query and hands them to a language model [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com). When the answer requires filtering a spreadsheet or computing a ratio from two joined tables, similarity search returns the wrong thing. Agentic variants can adapt queries and tool use but remain largely retrieval-centric [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com). RECAST's bet is that evidence should be actively derived through computation rather than retrieved from a store [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com).

**My read:** This is the first framework I've seen that treats evidence gathering as a sequential decision problem where the model computes what it needs rather than hoping the right chunk sits in a vector database. The 15.9% margin over an unnamed baseline is striking, but I don't buy the zero-shot generalization claim yet, because the held-out benchmarks were selected by the authors and nobody outside the team has verified the results. The comparison between a trained Qwen3.5-9B and a training-free Gemini 3.5 Flash is also apples-to-oranges: training always helps, and the 5.0% gap says more about the value of fine-tuning than about either model's raw capacity.

### Three models, one pipeline

RECAST splits the work across three language models. A lightweight RouterLM iteratively selects and formulates operations, either choosing from primitives or specifying custom ones [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com). A frozen CompilerLM translates those specifications into executable code [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com). Once the RouterLM judges the evidence sufficient, it passes everything to a frozen AnswerLM that produces the final answer [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com).

The RouterLM is the only component that gets trained. The team used supervised fine-tuning followed by group relative policy optimization, or GRPO, a reinforcement learning method that rewards the model for choosing operations that lead to correct answers [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com).

Keeping the CompilerLM and AnswerLM frozen makes the system modular, and teams can swap in different base models for those roles.

This separation matters because it decouples the decision of what to compute from the computation itself. A financial analytics team, for instance, could train the RouterLM to recognise when a question requires pulling data from an SEC filing and computing a ratio, while keeping the CompilerLM and AnswerLM as off-the-shelf models. The same pattern echoes the broader trend of smaller, task-specific models matching larger ones, as we found when [an NYU 7.4B model matched GPT-3 13B with 20x less compute](https://www.notatechguy.com/nyu-7-4b-model-matches-gpt-3-13b-with-20-less-compute/).

### What the numbers show, and what they don't

Across six benchmark families, RECAST achieved a 75.6% mean success rate [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com). The strongest large-model baseline scored 15.9 percentage points lower, though the paper does not identify that baseline by name [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com). On three held-out benchmarks, RECAST improved over the strongest baseline by 15.0% on average [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com).

![RECAST's reported margins over baselines](https://storage.ghost.io/c/6e/89/6e896869-22ef-4281-a213-b4c462c17cff/content/images/2026/10/chart_57fdcc7fc4990fd9bc3e.png)

The authors frame these held-out results as evidence of strong zero-shot generalisation across tasks and heterogeneous source representations [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com).

That claim rests on benchmarks the authors chose, and the paper reports no independent verification [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com).

The training comparison is narrower than it sounds. A trained Qwen3.5-9B RouterLM outperformed a training-free Gemini 3.5 Flash RouterLM by 5.0%, which confirms that task-specific training helps but does not establish that Qwen3.5-9B is the stronger base model [S¹](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com).

The push for better context handling is not unique to this team. Google's work on [Gemini Omni 1.1 Flash processing 10x more scene context](https://www.notatechguy.com/google-gemini-omni-1-1-flash-10x-more-scene-context/) tackles a related problem from the model-capacity side. RECAST approaches it from the evidence-gathering side: instead of making the model's window bigger, it makes the evidence smarter.

### Who would use this first

The architecture is most relevant to teams building question-answering systems over heterogeneous data: financial filings that mix text and tables, legal databases that require cross-referencing, scientific repositories where answers need computation over raw data. A legal tech startup could train a RouterLM to recognise when a query needs statute text cross-referenced with case outcomes, then let the CompilerLM generate the code to execute that lookup.

No public code repository for RECAST appears in the evidence pack, so teams cannot yet run the framework themselves. The October 7 preprint is available on arXiv for scrutiny, and the results await independent reproduction.

---

*Sources: [S1 — RECAST: Learning to Compute the Right Context through Adaptive Evidenc](https://arxiv.org/abs/2610.10507v1?ref=notatechguy.com) · [P2 — deandevz/king-context](https://github.com/deandevz/king-context?ref=notatechguy.com) · [P3 — Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging](https://arxiv.org/html/2609.30751?ref=notatechguy.com) · [P4 — dieuroi/Routing-Evidence](https://github.com/dieuroi/Routing-Evidence?ref=notatechguy.com)*

## Related reading

- [NYU 7.4B model matches GPT-3 13B with 20× less compute](https://www.notatechguy.com/nyu-7-4b-model-matches-gpt-3-13b-with-20-less-compute/) — our technology desk, 2026-09-18
- [Google Gemini Omni 1.1 Flash: 10x more scene context](https://www.notatechguy.com/google-gemini-omni-1-1-flash-10x-more-scene-context/) — our technology desk, 2026-08-27
- [OpenAI Academy adds role-based learning paths for five audiences](https://www.notatechguy.com/openai-academy-adds-role-based-learning-paths-for-five-audiences/) — our technology desk, 2026-09-21

---

*Written from 4 sourced items, 3 of them primary.*