VeriChat: Multi-agent AI claims 87.73% faithfulness in chip-security verification
An unreviewed arXiv paper shows how three specialised AI agents catch chip-security flaws that general chatbots miss.
A newsletter breaking down AI research, technology, and Australian property in plain English
An unreviewed arXiv paper shows how three specialised AI agents catch chip-security flaws that general chatbots miss.
Researchers have distilled 457 papers into the first unified map of the human, tech and organisational factors that make or break a breach response.
New CNeVA framework dials per-channel behaviour in traffic sims — no frozen cars, no reward hacking.
Old-school engineering constraints let a tiny AI reviewer catch nine in ten backdoors, using fewer tokens than agentic scaffolding.
Researchers can spot stripped safety layers in open-weight AI with 95% accuracy, but the fix has limits.
AgentFlow proposes framework-agnostic dependency graphs for auditing agent code, though its 238 flagged risks and superiority claims remain unverified by peer review.
The 0.8B-parameter constitutional classifier reportedly outperforms 27B rivals on seven benchmarks, though the paper awaits peer review
An unpeer-reviewed arXiv preprint finds that limited coherent memory makes stabilizer state testing as costly as learning, with implications for certification bottlenecks.
An unreviewed arXiv paper argues that hardware trojans could hide in the building blocks of chip design, shifting supply-chain risk for fabless firms and their customers.