LLM confidence scores fail basic coherence test
New arXiv paper finds calibration, the standard test for AI confidence, misses deeper incoherence in how models estimate their own certainty.
A newsletter breaking down AI research, technology, and Australian property in plain English
New arXiv paper finds calibration, the standard test for AI confidence, misses deeper incoherence in how models estimate their own certainty.
NVIDIA partner Wistron opened a 324,000-square-foot Texas plant making GB300 chips, with 500 jobs and a $700M commitment.
PEARL puts a math solver inside the LLM loop to test, debug and revise optimization models, beating models 170× larger on verified solve rates.
DBMol pairs structure prediction models to design target-specific drug molecules, but every affinity result is a computational prediction, not a lab test
AI safety discourse focuses on dramatic harms, but a new preprint argues the real danger is in quiet failures normalised by everyday workflows.
Australia's Wage Price Index rose 3.3% in the year to March 2026, but private sector pay slowed to 3.2% as new home loans fell 6.2%.
A July 2026 arXiv preprint introduces TPIPS, a text-prompted image similarity metric showing frontier vision-language models lag human perception.
New arXiv framework uses twelve reasoning angles to separate jokes from hate in memes, hitting 80.3% humor and 75.9% hate detection on two benchmarks.
Six LLMs tested across navigation, triage and finance tasks showed stable risk preferences, a hidden trait anyone deploying AI agents needs to understand.