LLM agents fail 70% of skill safety tests in SLBench study
A new 86-case benchmark finds Codex and Claude Code agents violate logical constraints in skill files up to 70% of the time, causing privacy leaks and
A newsletter breaking down AI research, technology, and Australian property in plain English
A new 86-case benchmark finds Codex and Claude Code agents violate logical constraints in skill files up to 70% of the time, causing privacy leaks and
New arXiv preprint proposes auction-based routing for LLM agents, sending reasoning steps to the most capable solver rather than the most overconfident.
A preprint shows backdoors in feedforward neural networks evade every statistical test, even with full access to all weights, breaking model trust.
Training AI models directly on visual documents consistently outperforms text-only pretraining, challenging a core assumption in how foundation models
A new arXiv paper proposes the Hypothesis Evolution Protocol, making AI agents' scientific reasoning explicit and auditable instead of buried in logs.
New arXiv preprint 4DR360 treats 3D scene occupancy as a persistent state, reshaping radar-camera fusion for autonomous driving.
A new arXiv survey maps how LLMs could move front-end chip design from isolated tasks to autonomous agents, but offers no benchmarks to prove it works.
A new arXiv preprint proposes a training-free method that helps language models actually use evidence already sitting inside their 128K-token context
LLM-as-a-Verifier treats checking AI answers as a scaling axis, hitting 86.5% on Terminal-Bench V2 and 78.2% on SWE-Bench Verified without extra training.
VAORA, a new reward design on arXiv, targets hallucinated reasoning and reasoning-action misalignment in vision-language models on physical tasks.