AI models guess your real intent only 22-32% of the time
New arXiv research finds AI models recover users' intended tasks just 22-32% of the time when instructions are ambiguous, versus 48% for humans.
A newsletter breaking down AI research, technology, and Australian property in plain English
New arXiv research finds AI models recover users' intended tasks just 22-32% of the time when instructions are ambiguous, versus 48% for humans.
A July 2026 arXiv preprint shows AI agents following every consensus rule can still certify wrong answers, and shared model weights make them fail
SciForge, an open-source AI workbench posted to arXiv on July 20, lets scientists keep judgment while agents handle search, parsing, plotting and writing.
New arXiv framework auto-builds massive training environments from 400 Model Context Protocols to teach AI agents long-horizon tool use.
OpenAI's July 2026 safety post details new risks in long-horizon AI models and iterative safeguards, relevant to any business deploying autonomous agents.
A new arXiv study of OECD-catalogued AI trust tools finds the field heavy on fairness and transparency but thin on security, explainability and
NVIDIA's SIGGRAPH keynote unveils MCP connections letting AI agents work inside creative apps, plus Cosmos 3 Edge for local physical AI.
An arXiv preprint analysing 14 years of Qubes Security Bulletins finds most flaws originate in Xen and CPU components, not Qubes itself.
A new arXiv preprint introduces Mycelium, a shared workspace connecting researchers and AI agents, routing hypotheses to whoever can act on them.
New arXiv preprint shifts MCP security from guessing at tool risk to watching real execution, reporting 523 findings across 326 live servers.
A new arXiv paper uses cognitive psychology's set-shifting concept to test whether LLM agents adapt when reliable tools silently change mid-session.
A new arXiv preprint introduces an Information Flow Graph monitor that stops AI coding agents secretly weakening security before deployment.