LLM confidence scores fail basic coherence test
New arXiv paper finds calibration, the standard test for AI confidence, misses deeper incoherence in how models estimate their own certainty.
The AI knowledge 99.9% of people never discover
Daily AI and technology coverage — models, chips, agents and research, and what each development actually means for you and your business.
551 stories
New arXiv paper finds calibration, the standard test for AI confidence, misses deeper incoherence in how models estimate their own certainty.
A new arXiv preprint claims small language models trained on synthetic data outperform prompt-based LLM safety checks for hallucination and topic drift.
New arXiv preprint proposes Probabilistic Concept-Aware Steering, fixing incoherent behaviour in LLM steering vectors for open-weight model teams.
An arXiv preprint traced 232,270 AI supply chains, finding permissive licences survive 95% while obligation-bearing ones drop below 7%.
NVIDIA partner Wistron opened a 324,000-square-foot Texas plant making GB300 chips, with 500 jobs and a $700M commitment.
PEARL puts a math solver inside the LLM loop to test, debug and revise optimization models, beating models 170× larger on verified solve rates.
DBMol pairs structure prediction models to design target-specific drug molecules, but every affinity result is a computational prediction, not a lab test
AI safety discourse focuses on dramatic harms, but a new preprint argues the real danger is in quiet failures normalised by everyday workflows.
A July 2026 arXiv preprint introduces TPIPS, a text-prompted image similarity metric showing frontier vision-language models lag human perception.
New arXiv framework uses twelve reasoning angles to separate jokes from hate in memes, hitting 80.3% humor and 75.9% hate detection on two benchmarks.
Six LLMs tested across navigation, triage and finance tasks showed stable risk preferences, a hidden trait anyone deploying AI agents needs to understand.
BSB jailbreak tricks text-to-video models like Veo and Sora by hiding harmful intent between two safe frames, beating rivals by 18.6% in attack success.