AI agent benchmark tests tool-switching when reliability shifts
A new arXiv paper uses cognitive psychology's set-shifting concept to test whether LLM agents adapt when reliable tools silently change mid-session.
A newsletter breaking down AI research, technology, and Australian property in plain English
AI and technology, explained through what they actually mean. An AI-assisted newsroom under human editorial rules — every story cites its primary sources so you can check them yourself.
304 stories
A new arXiv paper uses cognitive psychology's set-shifting concept to test whether LLM agents adapt when reliable tools silently change mid-session.
A new clustering-based feature selection method matches standard accuracy for wind and solar forecasting while cutting compute cost 21%, with open-source
A new arXiv preprint introduces an Information Flow Graph monitor that stops AI coding agents secretly weakening security before deployment.
A new arXiv framework cuts transportation management AI costs 97% by mixing open-source and closed APIs, with a greedy heuristic that picks the cheapest
A lightweight probe on the Evo 2 DNA foundation model flags antimicrobial resistance in raw metagenomic data, skipping costly genome assembly steps.
A new two-agent framework finds code bugs across whole repositories at 16-48x lower cost, but over half its correct answers rely on flawed reasoning.
A July 2026 study of nine endoscopy AI systems finds high answer accuracy doesn't guarantee faithful clinical reasoning, affecting hospital AI buyers.
A new arXiv paper from ShanghaiTech replaces binary safe/unsafe guards with a three-way EXECUTE-ASK-REFUSE routing that cuts alert fatigue for anyone
A Hybrid Oscillator Arbiter PUF generates crypto keys at 2.7 microwatts, pointing to hardware security for battery-constrained wearables and IoT devices.
Graph neural network hints lift small language model molecular property prediction by up to 74% on Tox21, according to a new arXiv preprint.
A July 2026 arXiv preprint reframes how AI systems measure their own uncertainty, deriving competing measures from one mathematical framework. Here is what
New Mila preprint combines self-supervised and supervised learning to improve neural decoding across species when labelled data is scarce.