OpenAI CFO introduces four-metric AI ROI scorecard
Sarah Friar's framework shifts AI measurement from model benchmarks to useful work, cost per task, dependability and return on compute.
A newsletter breaking down AI research, technology, and Australian property in plain English
Sarah Friar's framework shifts AI measurement from model benchmarks to useful work, cost per task, dependability and return on compute.
Australia's PM wants big AI data centres to underwrite new power and put back as much energy as they draw, but stalled renewables decide if it works.
New arXiv paper pairs world models with human preferences and justifications to train safe AI agents without risky trial-and-error deployment.
arXiv preprint reveals exact token cost of attributing AI text to a user, plus a window where generated text is provably machine-made but unattributable.
IBM Research found Claude Sonnet cost half as much as GPT-4.1 across 417 agent tasks, because caching matters more than sticker price for AI routing.
First meta-evaluation of LLM-generated rubrics for paper reproduction finds AI graders are overly generous and too detailed, but augmentation helps.
OpenAI's automated red teaming system GPT-Red uses self-play to find model weaknesses like prompt injection gaps affecting every AI user.
OpenAI's July 15 proposal wants state AI laws to build toward a national safety framework, shaping how every American AI developer gets regulated.
Reinforcement fine-tuning with 30 prompts cut an open-weight model's building emissions to 61.2 kg-CO2, near the 60.8 optimum, a preprint shows.
New arXiv preprint shows Transformer training on inductive reasoning can be confined to a low-dimensional manifold for automatic circuit detection.
AdvancedMathBench tests AI on graduate and doctoral math proofs, finding frontier models struggle with both writing and verifying advanced mathematics.
Reasoning scaffold lifts GPT-4.1-mini by 0.21 but degrades GPT-5-mini by 0.63, arXiv study finds, exposing an architecture-dependent split.