LLM vulnerability detection hits 100% recall, then collapses on real code
Structural priors lift LLM vulnerability recall from 20% to 100% on synthetic benchmarks, but real CVE data exposes a 51-point collapse.
A newsletter breaking down AI research, technology, and Australian property in plain English
AI and technology, explained through what they actually mean. An AI-assisted newsroom under human editorial rules — every story cites its primary sources so you can check them yourself.
306 stories
Structural priors lift LLM vulnerability recall from 20% to 100% on synthetic benchmarks, but real CVE data exposes a 51-point collapse.
A replay analysis of three LLM agent benchmarks finds the safe partial-run fraction ranges from 15% to over 95%, with no universal shortcut.
A 17 July arXiv preprint shows attackers can weaponise setup docs to trick AI coding agents into installing untrusted dependencies across npm and Cargo.
An arXiv preprint introduces AutoSynthesis, a multi-agent system that turns a plain-English research question into a full meta-analysis report.
A new arXiv preprint proposes mutable sketches that update user embeddings on-the-fly, cutting data reads to 1.8% and eliminating retrain cycles.
Sarah Friar's framework shifts AI measurement from model benchmarks to useful work, cost per task, dependability and return on compute.
MemPoison benchmark tests 1,227 attacks across 10 model families, finding write-time defenses miss sophisticated memory corruption in LLM agents.
ReBound, a new arXiv preprint, reuses cached query results to answer follow-up questions at reduced or zero additional privacy cost for analysts.
A July 2026 arXiv preprint reframes security testing for AI systems, arguing attackers can violate operational objectives without breaching infrastructure.
New STOCKTAKE benchmark: LLM agents detect up to 88% of hidden supply-chain failures but two of four models score below a blind baseline on action.
Adversaries can embed hidden instructions in network logs that hijack LLMs used by security teams, with attacks succeeding up to 88.2% of the time.
New arXiv preprint finds offensive AI security agents improve with more compute, but defensive SOC tasks need disciplined tool use over raw reasoning