MedRealMM benchmark: AI matches doctors but fails on safety
New 5,620-case benchmark finds frontier AI matches physicians on positive clinical steps but triggers more unsafe responses in real consultations.
A newsletter breaking down AI research, technology, and Australian property in plain English
New 5,620-case benchmark finds frontier AI matches physicians on positive clinical steps but triggers more unsafe responses in real consultations.
Malaika, a multi-agent AI framework posted to arXiv this month, uses three grounding mechanisms to make malware analysis more precise and auditable.
TrustX ARC, a new arXiv preprint, scores agentic AI risk across 12 dimensions into three governance tiers for risk officers, developers and regulators.
A new arXiv preprint pairs an LLM planner with a forecasting model to guard industrial control systems, recording zero hallucinated actions in attack
KV-PRM reads the memory AI models already produce during generation, slashing the cost of verifying multi-agent reasoning chains by up to 5,000x.
New arXiv paper with 489 probes shows structural scaffolding around a fixed LLM cuts failure rates, challenging the bigger-model assumption.
A new arXiv paper from July 2026 unveils a four-agent system that decomposes abstract reasoning puzzles into perception, code search, and reflective
A new 86-case benchmark finds Codex and Claude Code agents violate logical constraints in skill files up to 70% of the time, causing privacy leaks and
New arXiv preprint proposes auction-based routing for LLM agents, sending reasoning steps to the most capable solver rather than the most overconfident.