Agent-ready websites nearly double AI shopping agent success
A new arXiv framework lifts AI browser-agent task completion from 49% to 89% by restructuring pages for machine reading, hitting every e-commerce site
A newsletter breaking down AI research, technology, and Australian property in plain English
A new arXiv framework lifts AI browser-agent task completion from 49% to 89% by restructuring pages for machine reading, hitting every e-commerce site
An arXiv preprint details a five-stage AI upskilling framework. Three learners passed NVIDIA's Agentic AI exam, but evidence is thin and unpeer-reviewed.
A new arXiv preprint proposes a training-free detector that reads LLM hidden states to catch prompt injection attacks on purpose-specific agents, reporting
IBM Research found Claude Sonnet cost half as much as GPT-4.1 across 417 agent tasks, because caching matters more than sticker price for AI routing.
New E3 framework cuts AI agent cost 85% by estimating task difficulty before acting, matching 100% success on a 121-edit benchmark with code released.
A non-peer-reviewed arXiv preprint proposes everyday AI agent prompting loops quietly strengthen impatience and self-criticism through neuroplasticity.
First meta-evaluation of LLM-generated rubrics for paper reproduction finds AI graders are overly generous and too detailed, but augmentation helps.
MCP security scanners flag 96.89% of servers as risky, but a study of 64,611 servers finds fewer than half of alerts are true positives.
New arXiv preprint Prezta replaces application gateways at critical infrastructure edges with zero-knowledge proofs generated on client devices.
New arXiv preprint introduces AHA, a system using one AI agent to auto-discover reusable vulnerabilities in production agents like Claude Code and Codex.
A new open benchmark with 500+ tools reveals even frontier AI models can't reliably read an image and act on it, with most failures traced to seeing, not
A new Lean 4 library formalises three quantum coding theorems and builds reusable infrastructure for AI-assisted proof search in quantum information