AI models guess your real intent only 22-32% of the time
New arXiv research finds AI models recover users' intended tasks just 22-32% of the time when instructions are ambiguous, versus 48% for humans.
A newsletter breaking down AI research, technology, and Australian property in plain English
New arXiv research finds AI models recover users' intended tasks just 22-32% of the time when instructions are ambiguous, versus 48% for humans.
AI trust gap: new arXiv preprint argues firms can't prove safety claims, leaving buyers and regulators unable to distinguish safe systems from imitations.
A new arXiv survey argues logic plus modern optimization makes rule-based AI practical, offering explainable systems where neural networks fall short.
A July 2026 preprint proposes harmonizing the capability thresholds frontier AI labs publish, which differ so much that no one can verify them.
An arXiv preprint describes Gasp, an L2 rollup DEX using EigenLayer restaking for gas-free cross-chain swaps without traditional bridges.
Prompt syntax changes whether open-source LLMs generate secure or vulnerable code, a new preprint finds, with direct impact on self-hosting teams.
A July 2026 arXiv preprint shows AI agents following every consensus rule can still certify wrong answers, and shared model weights make them fail
arXiv preprint CRAFT turns grading rubrics into capability diagnoses, generating targeted fine-tuning data that beats EvalTree on four models.
SciForge, an open-source AI workbench posted to arXiv on July 20, lets scientists keep judgment while agents handle search, parsing, plotting and writing.