> ## Content Index
> Fetch the complete content index at: https://www.notatechguy.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI agent for chip verification hits 95% accuracy
- URL: https://www.notatechguy.com/ai-agent-for-chip-verification-hits-95-accuracy/
- Published: 2026-10-06T13:43:52.000Z
- Updated: 2026-10-06T13:43:54.000Z
- Description: An arXiv preprint turns chip simulation dumps into a searchable database for verification engineers, hitting 95.33% on 150 queries.
- Author: Marcello Babbili
- Tags: Technology & AI, AI Agents, AI Models

arXiv hosted an October 5 preprint reporting an AI framework that answers chip-verification questions with 95.33% accuracy across 150 queries [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com). The framework targets a part of the chip design process that most AI research has skipped: roughly 74.6% of existing LLM studies in electronic design automation focus on writing Register-Transfer Level code, and the harder work of debugging chips after simulation remains almost untouched [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com).

**My read:** This is the first verification framework I've seen that treats simulation dumps as a database problem rather than a code-generation problem. The 95.33% execution accuracy sounds strong, but the metric is undefined in the available material, the benchmark is self-reported, and nobody outside the authors has checked it. The SQLite angle is clever: waveform data is notoriously unstructured, and relational storage could make it queryable in ways that matter. But until the code or benchmark is public, this is a claim, not a tool.

## The gap in chip verification

Modern AI runs on chips with billions of transistors [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com), the kind of complexity that [Qualcomm says can run a 30B-parameter model on a phone](https://www.notatechguy.com/qualcomm-says-new-chip-runs-30b-ai-model-on-a-phone/). Yet the workflows that verify those chips still depend on engineers reading simulation waveforms by hand [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com).

Large language models have moved into electronic design automation, the software toolchain that turns chip designs into silicon. But the preprint's authors surveyed the field and found that about 74.6% of studies target static RTL code generation, the step where engineers write the hardware description language that defines a chip's logic [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com). Post-simulation verification and interactive waveform debugging remain largely untouched by LLM research [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com).

Code generation is the easy part.

Finding a timing violation buried in a massive simulation dump, tracing it back to the specific line of RTL that caused it, and confirming the fix is where verification teams spend most of their time.

## What the framework does

The preprint introduces Back-to-the-Future, or BTTF, an end-to-end agentic framework for chip design verification [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com).

It works in two parts.

First, BTTF distills massive, unstructured simulation dumps into a normalized relational SQLite database [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com). Simulation dumps are the raw output of chip verification runs, often gigabytes of signal data showing every transistor state at every clock cycle. SQLite is a lightweight, file-based database engine that stores data in tables and lets you query it with SQL.

Second, BTTF couples that database with a collaborative multi-agent orchestration engine that translates natural-language verification queries into schema-aware SQL [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com). An engineer could type a question about which signals changed between specific clock cycles on a particular block, and the engine would generate the SQL query to pull that answer from the database. BTTF also correlates signal anomalies with versioned RTL repositories. It links a bug found in simulation back to the specific code change that introduced it [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com).

The approach echoes a broader movement in agentic EDA. FluxEDA, a framework from Zhejiang University researchers including Zhengrui Chen and Cheng Zhuo, proposes a unified execution infrastructure for stateful agentic EDA [P²](https://arxiv.org/html/2603.25243?ref=notatechguy.com). AgentDV, a separate preprint by Navya Goli and colleagues, explores closed-loop agentic AI for hardware design verification [P⁵](https://arxiv.org/abs/2608.27148?ref=notatechguy.com). On GitHub, the OSCC-Project's AiEDA repository, with 86 stars, offers an RTL-to-Vector-to-GDS library under the Mulan Permissive Software License [P³](https://github.com/oscc-project/AiEDA?ref=notatechguy.com). BTTF differs by focusing on post-simulation debugging rather than code generation or full-flow automation.

## What 95.33% means for verification teams

The authors report that BTTF attained 95.33% execution accuracy across a 150-query benchmark [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com). That means the framework correctly answered roughly 143 of 150 verification questions, if "execution accuracy" means what it sounds like: the generated SQL query returned the correct answer when run against the database.

Several things are unclear. The preprint does not define "execution accuracy" in the available material, so the metric could mean something narrower or broader than it sounds. The 150-query benchmark is self-reported, with no external validation. The paper is an arXiv preprint, not peer-reviewed [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com). The framework does not appear to be publicly available, so no one outside the authors has run the benchmark independently. The 74.6% figure for RTL-focused studies reflects the authors' own literature-review methodology and boundaries, not an industry-standard survey [S¹](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com).

If BTTF works as described, the first users would be verification engineers at semiconductor firms who spend their days debugging simulation failures. A team running thousands of regression tests could ask BTTF to find which signals diverged from the expected waveform and trace the divergence to a specific RTL commit, the same kind of [multi-billion-transistor chip complexity that Qualcomm demonstrated last month](https://www.notatechguy.com/qualcomm-says-new-chip-runs-30b-ai-model-on-a-phone/).

That workflow is familiar to every verification team. But the preprint describes a research prototype, not production-ready EDA tooling, and the gap between a 150-query benchmark and a real chip project with millions of signals is large. The next checkpoint is whether the authors release the code and benchmark for external testing, or whether a semiconductor firm reports results on a real design.

---

*Sources: [S1 — Back to the Future: Rethinking EDA Infrastructure for Agentic Systems ](https://arxiv.org/abs/2610.06790v1?ref=notatechguy.com) · [P2 — FluxEDA: A Unified Execution Infrastructure for Stateful Agentic EDA](https://arxiv.org/html/2603.25243?ref=notatechguy.com) · [P3 — OSCC-Project/AiEDA](https://github.com/oscc-project/AiEDA?ref=notatechguy.com) · [P4 — v2-io/agentic-systems](https://github.com/v2-io/agentic-systems?ref=notatechguy.com) · [P5 — \[2608.27148\] AgentDV: Closed-Loop Agentic AI for Hardware Design Verif](https://arxiv.org/abs/2608.27148?ref=notatechguy.com)*

## Related reading

- [Qualcomm says new chip runs 30B AI model on a phone](https://www.notatechguy.com/qualcomm-says-new-chip-runs-30b-ai-model-on-a-phone/) — our technology desk, 2026-09-22

---

*Written from 5 sourced items, 4 of them primary.*