Neurosymbolic world model transfers tasks without retraining
An August 2026 arXiv preprint splits RL world models so reward prediction uses only symbolic state, letting agents switch tasks without further training.
A newsletter breaking down AI research, technology, and Australian property in plain English
An August 2026 arXiv preprint splits RL world models so reward prediction uses only symbolic state, letting agents switch tasks without further training.
A new arXiv paper proposes DeAR, a framework where AI agents coordinate reasoning without a central orchestrator, tested across nine benchmarks.
Cambridge researchers find looped language models improve at multi-step API tool calling, with adaptive inference offering the best compute-performance
A new arXiv preprint surveys agentic AI from working principles to adoption factors, as developer frameworks signal explosive real-world traction.
A new arXiv preprint introduces LeakGauge, a detector that flags when LLMs leak confidential context, tested across 11 models with AUROC to 0.996.
NVIDIA teams use OpenAI's ChatGPT Work agent to cut manual tasks and scale workflows globally, according to a new OpenAI case study published August 18.
StagedWorkspace binds every agent view to a versioned file state, lifting Gemini 3.1 Pro from 29.3% to 63.9% on OfficeQA in a new arXiv preprint.
Alibaba Cloud's Wuying-Browser-Agent-27B scores 65.1% on a new 350-task web benchmark averaging 37.9 steps, claiming open-source SOTA for browser agents.
IBM Research tested agent memory on eight models and found the biggest model gained nothing while a 117B model jumped 16 points at 5% extra token cost.