GPT-4o audit finds 66% of solved problems have flawed reasoning
A black-box audit swaps predicates in chain-of-thought to expose when GPT-4o reaches correct answers through reasoning that ignores its stated premises.
A newsletter breaking down AI research, technology, and Australian property in plain English
A black-box audit swaps predicates in chain-of-thought to expose when GPT-4o reaches correct answers through reasoning that ignores its stated premises.
arXiv paper explains why vulnerability scanners give different results for the same software, tracing inconsistency to a distributed ecosystem.
A new arXiv preprint shows LLMs can read threat reports, generate attack playbooks, and self-correct failures, reaching 84% with Claude Sonnet 4.5.
An arXiv preprint shows an RL agent using a Lyapunov exponent reward found the Kapitza pendulum and a new upright state, with implications for robotics.
Structural priors lift LLM vulnerability recall from 20% to 100% on synthetic benchmarks, but real CVE data exposes a 51-point collapse.
A replay analysis of three LLM agent benchmarks finds the safe partial-run fraction ranges from 15% to over 95%, with no universal shortcut.
A 17 July arXiv preprint shows attackers can weaponise setup docs to trick AI coding agents into installing untrusted dependencies across npm and Cargo.
An arXiv preprint introduces AutoSynthesis, a multi-agent system that turns a plain-English research question into a full meta-analysis report.
A new arXiv preprint proposes mutable sketches that update user embeddings on-the-fly, cutting data reads to 1.8% and eliminating retrain cycles.