Voice input degrades LLM agents more than typing, study finds
Voice transcription errors cut accuracy across every instruction-tuned model tested, while typing errors are absorbed. The gap traces to one mechanism.
A newsletter breaking down AI research, technology, and Australian property in plain English
Voice transcription errors cut accuracy across every instruction-tuned model tested, while typing errors are absorbed. The gap traces to one mechanism.
A new arXiv preprint from Sharif University trains an energy-based surrogate to identify which prompt sentences matter most, with no extra API calls.
DreamGuard, a new arXiv preprint, uses a risk-aware world model to predict when AI agent actions could drift toward danger, intervening in 25 ms.
SkillTrace extracts three provenance traces from LLM-agent skills to catch partial reuse that code clone tools miss, scoring 0.938 AUROC across 36,446
MedUPS trains AI on mid-stream clinical decisions rather than final diagnosis and raises next-step accuracy up to 11 points on 21,874 real case reports.
RAG study tests LLaMA, Mistral and Qwen to cut hallucinations in small business AI, but the arXiv preprint provides no benchmark numbers.
New arXiv preprint proposes Chained RLM, calling the same model with fresh context to stop early reasoning errors reaching the final answer.
Self-improving AI agents that learn from stored memory can inflate rewards for wrong answers, then preferentially reuse their most confident mistakes
AntiSkillBench tests 7,500 dialogue traces across three frontier agents and finds privacy risks persist from explicit data to communication style.
TrainShield, a new arXiv preprint, proposes AI cybersecurity training embedded in browsing workflows when phishing or data-loss risks are detected.
An arXiv preprint proposes replaying temporally-evolving enterprise worlds to test AI agents at any moment, fixing a blind spot in current evals.
New arXiv paper proposes SkillBoost, a three-stage framework that stops LLM agents overfitting to limited experience and forgetting solved tasks.