RECAST's authors posted the framework to arXiv on October 7, reporting a 75.6% mean success rate across six benchmark families S¹. That result outperforms the strongest large-model baseline they tested by 15.9 percentage points, though the baseline is not named in the available paper excerpts S¹.

The preprint has not been peer-reviewed, and all performance figures are self-reported S¹.

Conventional retrieval-augmented generation, or RAG, fetches text chunks that resemble a query and hands them to a language model S¹. When the answer requires filtering a spreadsheet or computing a ratio from two joined tables, similarity search returns the wrong thing. Agentic variants can adapt queries and tool use but remain largely retrieval-centric S¹. RECAST's bet is that evidence should be actively derived through computation rather than retrieved from a store S¹.

My read: This is the first framework I've seen that treats evidence gathering as a sequential decision problem where the model computes what it needs rather than hoping the right chunk sits in a vector database. The 15.9% margin over an unnamed baseline is striking, but I don't buy the zero-shot generalization claim yet, because the held-out benchmarks were selected by the authors and nobody outside the team has verified the results. The comparison between a trained Qwen3.5-9B and a training-free Gemini 3.5 Flash is also apples-to-oranges: training always helps, and the 5.0% gap says more about the value of fine-tuning than about either model's raw capacity.

Three models, one pipeline

RECAST splits the work across three language models. A lightweight RouterLM iteratively selects and formulates operations, either choosing from primitives or specifying custom ones S¹. A frozen CompilerLM translates those specifications into executable code S¹. Once the RouterLM judges the evidence sufficient, it passes everything to a frozen AnswerLM that produces the final answer S¹.

The RouterLM is the only component that gets trained. The team used supervised fine-tuning followed by group relative policy optimization, or GRPO, a reinforcement learning method that rewards the model for choosing operations that lead to correct answers S¹.

Keeping the CompilerLM and AnswerLM frozen makes the system modular, and teams can swap in different base models for those roles.

This separation matters because it decouples the decision of what to compute from the computation itself. A financial analytics team, for instance, could train the RouterLM to recognise when a question requires pulling data from an SEC filing and computing a ratio, while keeping the CompilerLM and AnswerLM as off-the-shelf models. The same pattern echoes the broader trend of smaller, task-specific models matching larger ones, as we found when an NYU 7.4B model matched GPT-3 13B with 20x less compute.

What the numbers show, and what they don't

Across six benchmark families, RECAST achieved a 75.6% mean success rate S¹. The strongest large-model baseline scored 15.9 percentage points lower, though the paper does not identify that baseline by name S¹. On three held-out benchmarks, RECAST improved over the strongest baseline by 15.0% on average S¹.

RECAST's reported margins over baselines

The authors frame these held-out results as evidence of strong zero-shot generalisation across tasks and heterogeneous source representations S¹.

That claim rests on benchmarks the authors chose, and the paper reports no independent verification S¹.

The training comparison is narrower than it sounds. A trained Qwen3.5-9B RouterLM outperformed a training-free Gemini 3.5 Flash RouterLM by 5.0%, which confirms that task-specific training helps but does not establish that Qwen3.5-9B is the stronger base model S¹.

The push for better context handling is not unique to this team. Google's work on Gemini Omni 1.1 Flash processing 10x more scene context tackles a related problem from the model-capacity side. RECAST approaches it from the evidence-gathering side: instead of making the model's window bigger, it makes the evidence smarter.

Who would use this first

The architecture is most relevant to teams building question-answering systems over heterogeneous data: financial filings that mix text and tables, legal databases that require cross-referencing, scientific repositories where answers need computation over raw data. A legal tech startup could train a RouterLM to recognise when a query needs statute text cross-referenced with case outcomes, then let the CompilerLM generate the code to execute that lookup.

No public code repository for RECAST appears in the evidence pack, so teams cannot yet run the framework themselves. The October 7 preprint is available on arXiv for scrutiny, and the results await independent reproduction.


Sources: S1 — RECAST: Learning to Compute the Right Context through Adaptive Evidenc · P2 — deandevz/king-context · P3 — Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging · P4 — dieuroi/Routing-Evidence


Written from 4 sourced items, 3 of them primary.