LLMs flip answers 23% of the time when you rephrase the question
LLM answers change on over 23% of questions when wording shifts, and standard accuracy scores hide the reliability gap from anyone deploying AI.
A newsletter breaking down AI research, technology, and Australian property in plain English
AI and technology, explained through what they actually mean. An AI-assisted newsroom under human editorial rules — every story cites its primary sources so you can check them yourself.
498 stories
LLM answers change on over 23% of questions when wording shifts, and standard accuracy scores hide the reliability gap from anyone deploying AI.
The June 2026 quarter CPI landed on 29 July with two structural changes: earlier monthly releases from February 2027 and no mid-year weight update
OpenAI's July 28 field report says AI coding agents speed up genomics discovery, but offers no benchmarks or named scientists.
Gemini API managed agents now run 3.6 Flash by default and let developers intercept tool calls with custom hooks, changing how agentic workflows run on
An MIT-linked arXiv preprint argues autonomous research systems are graded on final results but ignore how much compute they burn getting there, and
An arXiv thesis from Adelaide proposes decoy placement and admin feedback loops for AD security, but admits the underlying maths is intractable.
New arXiv study finds reasoning models fail tasks they can solve due to premature self-doubt. The fix may not require bigger, more expensive models.
ISPCloak projects AI images through a simulated camera pipeline, imprinting sensor noise that deepfake detectors cannot tell from genuine photographs.
OpenAI's latest research says ChatGPT users are taking on tasks across roles, reshaping job boundaries. Here's what that means for workers.
FineServe, a new dataset from a commercial LLM marketplace, reveals serving traffic varies fundamentally by model architecture and task type.
OPTScientist uses four AI agents to discover optimizer algorithms for transformer pretraining, including a new reduced-state matrix optimizer called RS-MR.
An arXiv preprint proposes pairing data into quadruplets to cut gradient variance in unsupervised domain adaptation, improving target accuracy on three