Voice agents drop 9.7% when instructions are merely implied
DSB-IFEval tests 1,038 cases across eight assistant roles and finds voice agents drop up to 9.7% when instructions are implied, not stated.
A newsletter breaking down AI research, technology, and Australian property in plain English
DSB-IFEval tests 1,038 cases across eight assistant roles and finds voice agents drop up to 9.7% when instructions are implied, not stated.
New 'epistemic warrant' framework grades individual LLM recommendations from unstable to broadly supported, capturing what confidence scores miss.
SMC pairs a small drafter model with a large actor model to speculatively execute tool calls, reducing wall time by up to 45% on AppWorld.
SIR, a self-improving prompt injection attack, raised its success rate from 0% to 28% on Gemini 3.5 Flash while the agent's legitimate task still
New research shows ToolSiphon can extract most source records from LLM agent tools using only queries, affecting anyone building agent-based services.
DiaSentinel, a new arXiv preprint from a Taiwan hospital team, describes a fully on-premise multi-agent system for type 2 diabetes risk screening from
MMJailBench, a new factorized benchmark, evaluated 16 multimodal LLMs and found prompt framing, not visual tricks, drives most jailbreak vulnerabilities.
NVIDIA's $3.5 billion convertible bond investment in MediaTek locks in a partner spanning AI factories, consumer PCs and autonomous vehicles.
A systematic review of LLM-based security agents from 2023 to 2026 finds the field can act but cannot yet bound authority or audit behaviour.
Osmantic ODS bundles Ollama, Open WebUI, n8n and ComfyUI into a one-command local AI server. Here's what the README claims and what's unverified.
The multi-agent LLM trading framework has 100,000 GitHub stars after a release that quietly fixed a classic quant backtesting error.
New preprint STAR catches AI translations that silently drop or invent sentences, and a method that may let small models beat GPT-4o.