StagedWorkspace lifts AI agent office task scores by 34 points
StagedWorkspace binds every agent view to a versioned file state, lifting Gemini 3.1 Pro from 29.3% to 63.9% on OfficeQA in a new arXiv preprint.
A newsletter breaking down AI research, technology, and Australian property in plain English
Daily AI and technology news decoded in plain English — models, chips, agents and research, and what each development actually means for you and your business.
418 stories
StagedWorkspace binds every agent view to a versioned file state, lifting Gemini 3.1 Pro from 29.3% to 63.9% on OfficeQA in a new arXiv preprint.
Alibaba Cloud's Wuying-Browser-Agent-27B scores 65.1% on a new 350-task web benchmark averaging 37.9 steps, claiming open-source SOTA for browser agents.
IBM Research tested agent memory on eight models and found the biggest model gained nothing while a 117B model jumped 16 points at 5% extra token cost.
OpenAI's new safeguards for pacing frontier AI model development as cyber capabilities grow. What it means for security teams and what's missing.
sentence-transformers now ships MultiVectorEncoder, bringing late interaction retrieval to developers who already use dense and sparse models.
Euclid-Omni couples LLMs and vision models with a formal geometry solver to match state-of-the-art on Olympiad proofs using less compute.
New arXiv paper maps how AI regulation diverges across the EU, US and China, identifies three compliance gaps, and proposes a machine-checkable fix.
A new arXiv position paper argues excessive AI reliance is creating vulnerabilities from individual deskilling to national infrastructure failures.
OpenAI's new Defender's Window article argues AI is reshaping cybersecurity for attackers and defenders, with steps for security teams to take now.
A new preprint proposes a two-stage diffusion transformer that learns from multiple mismatched health tables to generate unlimited synthetic datasets.
New arXiv preprint introduces Agentao, a local-first runtime that separates what LLM agents propose from what they can execute, with open-source code on
New benchmark of 1,243 failed agent runs shows even GPT-5.5 auditors fix only 26.6%, with SearchAuditor at 32.3%, a warning for agent teams.