AI reasoning models fail from self-doubt, not capability gaps
New arXiv study finds reasoning models fail tasks they can solve due to premature self-doubt. The fix may not require bigger, more expensive models.
A newsletter breaking down AI research, technology, and Australian property in plain English
New arXiv study finds reasoning models fail tasks they can solve due to premature self-doubt. The fix may not require bigger, more expensive models.
OPTScientist uses four AI agents to discover optimizer algorithms for transformer pretraining, including a new reduced-state matrix optimizer called RS-MR.
An arXiv preprint proposes pairing data into quadruplets to cut gradient variance in unsupervised domain adaptation, improving target accuracy on three
New preprint extends Ex-Fuzzy with interpretable regression, hitting R² of 0.86 across 10 benchmarks using 10 to 15 human-readable rules.
Adversarial Frontiers, a new arXiv preprint, argues single-budget AI robustness rankings are unstable and proposes a frontier-based evaluation framework.
Language models build shared rules instead of storing individual facts, KAIST AI finds, and the overgeneralisation affects teams fine-tuning on nested
HyGRL, a new arXiv paper from Beijing Institute of Technology, blends text with knowledge graphs to answer multi-entity questions that defeat standard RAG.
A new arXiv preprint from Purdue and Princeton shows LLM query routing can work without training data, using model self-agreement as a signal.
New OrderBench benchmark runs 2,400 calls across four open models and finds schema validity alone doesn't guarantee semantic correctness.