SWE-bench needs 90% of tasks for reliable agent benchmark results
A replay analysis of three LLM agent benchmarks finds the safe partial-run fraction ranges from 15% to over 95%, with no universal shortcut.
A newsletter breaking down AI research, technology, and Australian property in plain English
A replay analysis of three LLM agent benchmarks finds the safe partial-run fraction ranges from 15% to over 95%, with no universal shortcut.
A 17 July arXiv preprint shows attackers can weaponise setup docs to trick AI coding agents into installing untrusted dependencies across npm and Cargo.
An arXiv preprint introduces AutoSynthesis, a multi-agent system that turns a plain-English research question into a full meta-analysis report.
A new arXiv preprint proposes mutable sketches that update user embeddings on-the-fly, cutting data reads to 1.8% and eliminating retrain cycles.
MemPoison benchmark tests 1,227 attacks across 10 model families, finding write-time defenses miss sophisticated memory corruption in LLM agents.
ReBound, a new arXiv preprint, reuses cached query results to answer follow-up questions at reduced or zero additional privacy cost for analysts.
New STOCKTAKE benchmark: LLM agents detect up to 88% of hidden supply-chain failures but two of four models score below a blind baseline on action.
Adversaries can embed hidden instructions in network logs that hijack LLMs used by security teams, with attacks succeeding up to 88.2% of the time.
New arXiv preprint finds offensive AI security agents improve with more compute, but defensive SOC tasks need disciplined tool use over raw reasoning
A July 2026 arXiv preprint tests PAT, a RAG-based system that feeds whole-document context to LLMs for English-to-Spanish translation, with mixed results.
Oracle agent memory hits 93.8% on LongMemEval with 10.7x fewer tokens than flat-history baselines, per a new arXiv preprint on long-horizon AI agents.
A new arXiv survey formalises how AI agents update their own prompts, memory and tools with minimal human input, and what breaks when they do.