Voice agents drop 9.7% when instructions are merely implied
DSB-IFEval tests 1,038 cases across eight assistant roles and finds voice agents drop up to 9.7% when instructions are implied, not stated.
A newsletter breaking down AI research, technology, and Australian property in plain English
Daily AI and technology news decoded in plain English — models, chips, agents and research, and what each development actually means for you and your business.
482 stories
DSB-IFEval tests 1,038 cases across eight assistant roles and finds voice agents drop up to 9.7% when instructions are implied, not stated.
New 'epistemic warrant' framework grades individual LLM recommendations from unstable to broadly supported, capturing what confidence scores miss.
SMC pairs a small drafter model with a large actor model to speculatively execute tool calls, reducing wall time by up to 45% on AppWorld.
GPT-6 Astra is OpenAI's most capable deployed model and first to reach Critical cybersecurity capability, able to find and exploit unknown flaws.
NousResearch's Hermes Agent has 240,000 GitHub stars for an open-source AI agent that claims to learn from every conversation and create new skills.
Gilbert + Tobin is scaling ChatGPT Enterprise and Codex across the firm with CEO-led governance, targeting operational work rather than legal advice
SIR, a self-improving prompt injection attack, raised its success rate from 0% to 28% on Gemini 3.5 Flash while the agent's legitimate task still
AIMC, a visual analytics framework, lets researchers track quality, themes and weaknesses across papers produced by autonomous AI scientist FARS.
New research shows ToolSiphon can extract most source records from LLM agent tools using only queries, affecting anyone building agent-based services.
A new arXiv paper adapts the blameless M&M conference for AI errors, with two clinicians agreeing on all 20 classifications across five cases.
Firecracker, AWS's Rust microVM engine behind Lambda and Fargate, hit 36,443 GitHub stars this week. Here is what it means for serverless teams.
Synthetic data is widely used in privacy-sensitive settings without threat models or falsifiable claims, and rare and minority records face the greatest