AI reasoning models fail from self-doubt, not capability gaps
New arXiv study finds reasoning models fail tasks they can solve due to premature self-doubt. The fix may not require bigger, more expensive models.
A newsletter breaking down AI research, technology, and Australian property in plain English
New arXiv study finds reasoning models fail tasks they can solve due to premature self-doubt. The fix may not require bigger, more expensive models.
ISPCloak projects AI images through a simulated camera pipeline, imprinting sensor noise that deepfake detectors cannot tell from genuine photographs.
FineServe, a new dataset from a commercial LLM marketplace, reveals serving traffic varies fundamentally by model architecture and task type.
OPTScientist uses four AI agents to discover optimizer algorithms for transformer pretraining, including a new reduced-state matrix optimizer called RS-MR.
Adversarial Frontiers, a new arXiv preprint, argues single-budget AI robustness rankings are unstable and proposes a frontier-based evaluation framework.
Language models build shared rules instead of storing individual facts, KAIST AI finds, and the overgeneralisation affects teams fine-tuning on nested
HyGRL, a new arXiv paper from Beijing Institute of Technology, blends text with knowledge graphs to answer multi-entity questions that defeat standard RAG.
A new arXiv preprint from Purdue and Princeton shows LLM query routing can work without training data, using model self-agreement as a signal.
New OrderBench benchmark runs 2,400 calls across four open models and finds schema validity alone doesn't guarantee semantic correctness.
Agentic Real2Sim uses vision-language agents to convert real robot recordings into simulatable twins, aiming to cut the labor cost of training robots in
JAXBench, the first TPU benchmark for AI kernel generation, shows curated docs lift correctness from 5.8% to 37.3% and reach 1.6x speedup over XLA.
New arXiv paper finds calibration, the standard test for AI confidence, misses deeper incoherence in how models estimate their own certainty.