Agent memory boosts gpt-oss 16 points but does nothing for GLM-5
IBM Research tested agent memory on eight models and found the biggest model gained nothing while a 117B model jumped 16 points at 5% extra token cost.
A newsletter breaking down AI research, technology, and Australian property in plain English
IBM Research tested agent memory on eight models and found the biggest model gained nothing while a 117B model jumped 16 points at 5% extra token cost.
New arXiv paper maps how AI regulation diverges across the EU, US and China, identifies three compliance gaps, and proposes a machine-checkable fix.
A new arXiv position paper argues excessive AI reliance is creating vulnerabilities from individual deskilling to national infrastructure failures.
OpenAI's new Defender's Window article argues AI is reshaping cybersecurity for attackers and defenders, with steps for security teams to take now.
A new preprint proposes a two-stage diffusion transformer that learns from multiple mismatched health tables to generate unlimited synthetic datasets.
New arXiv preprint introduces Agentao, a local-first runtime that separates what LLM agents propose from what they can execute, with open-source code on
New benchmark of 1,243 failed agent runs shows even GPT-5.5 auditors fix only 26.6%, with SearchAuditor at 32.3%, a warning for agent teams.
Federated learning keeps data on-device, but gradient updates leak it. New preprint from Google and USC researchers cuts the error cost of adding privacy.
An arXiv preprint maps security threats across chiplet systems and LLM-driven chip design tools, directly affecting fabless semiconductor teams adopting