OpenAI CFO introduces four-metric AI ROI scorecard
Sarah Friar's framework shifts AI measurement from model benchmarks to useful work, cost per task, dependability and return on compute.
A newsletter breaking down AI research, technology, and Australian property in plain English
Sarah Friar's framework shifts AI measurement from model benchmarks to useful work, cost per task, dependability and return on compute.
New STOCKTAKE benchmark: LLM agents detect up to 88% of hidden supply-chain failures but two of four models score below a blind baseline on action.
Adversaries can embed hidden instructions in network logs that hijack LLMs used by security teams, with attacks succeeding up to 88.2% of the time.
New arXiv preprint finds offensive AI security agents improve with more compute, but defensive SOC tasks need disciplined tool use over raw reasoning
NVIDIA's Vera Rubin platform reframes AI economics around intelligence per dollar, making continuous post-training the central workload for agentic AI.
A study of 2,361 popular GitHub repos finds most generate just one or two AI-agent pull requests per quarter, with heavy use in small teams.
Google DeepMind and Isomorphic Labs unveil a bioresilience strategy covering AI model safety, outbreak detection, and vaccine design for governments.
A July 2026 arXiv preprint tests PAT, a RAG-based system that feeds whole-document context to LLMs for English-to-Spanish translation, with mixed results.
Oracle agent memory hits 93.8% on LongMemEval with 10.7x fewer tokens than flat-history baselines, per a new arXiv preprint on long-horizon AI agents.