StagedWorkspace lifts AI agent office task scores by 34 points
StagedWorkspace binds every agent view to a versioned file state, lifting Gemini 3.1 Pro from 29.3% to 63.9% on OfficeQA in a new arXiv preprint.
A newsletter breaking down AI research, technology, and Australian property in plain English
StagedWorkspace binds every agent view to a versioned file state, lifting Gemini 3.1 Pro from 29.3% to 63.9% on OfficeQA in a new arXiv preprint.
OpenAI's new Defender's Window article argues AI is reshaping cybersecurity for attackers and defenders, with steps for security teams to take now.
New arXiv preprint introduces Agentao, a local-first runtime that separates what LLM agents propose from what they can execute, with open-source code on
New benchmark of 1,243 failed agent runs shows even GPT-5.5 auditors fix only 26.6%, with SearchAuditor at 32.3%, a warning for agent teams.
A text-only attack on skill-based AI agents silently inflates token use by 67% and runtime by 92% while keeping task completion rates unchanged.
A new preprint proposes HARD, a framework where LLM agents automatically build and improve their own runtime defenses from observed failures.
ICML 2026's reproduction challenge used AI agents to verify 2,200 papers, and 23% had a falsified or contested claim while 49 had nothing verifiable.
An arXiv preprint shows LLMs autonomously controlling farm lighting and actuators, cutting grow cycles 35% and finding energy strategies humans missed.
OpenAI's builder guide for GPT-5.6 promotes smarter model selection and new Responses API tools for startups building AI agents.
A new arXiv paper shows an AI agent that rewrites its own code to answer knowledge-graph questions, hitting 22% accuracy and exposing benchmark flaws.
A new arXiv preprint shows that when one AI bombards another with messages while ignoring replies, the second model enters a state it never shows alone.
Google's August update puts AI agent Ask Advisor inside Ads and Analytics, with text-prompt dashboards and competitor benchmarking.