LLM-as-a-Verifier hits 86.5% on Terminal-Bench V2
LLM-as-a-Verifier treats checking AI answers as a scaling axis, hitting 86.5% on Terminal-Bench V2 and 78.2% on SWE-Bench Verified without extra training.
A newsletter breaking down AI research, technology, and Australian property in plain English
AI and technology, explained through what they actually mean. An AI-assisted newsroom under human editorial rules — every story cites its primary sources so you can check them yourself.
306 stories
LLM-as-a-Verifier treats checking AI answers as a scaling axis, hitting 86.5% on Terminal-Bench V2 and 78.2% on SWE-Bench Verified without extra training.
VAORA, a new reward design on arXiv, targets hallucinated reasoning and reasoning-action misalignment in vision-language models on physical tasks.
A new arXiv preprint from IIT Madras proposes a blockchain-verifiable voting system on Solana that keeps ballots private without a trusted key dealer
Limited quantum memory collapses the gap between stabilizer testing and learning, a new arXiv preprint shows, with both needing Theta(n) copies.
A new arXiv preprint trains a Real-Bogus classifier without human labels, using simulated injections and dual-network co-teaching for noisy survey data.
A new preprint pairs quantum convolutional networks with path signature kernels to tackle reparameterisation invariance in time series classification.
New preprint Orcaella lets clients pick a fast 2-message-delay commit or a resilient path tolerating 54% equivocation at double the latency.
New arXiv preprint from NTU and Alibaba introduces a Vision-Language-Action model needing only a single RGB image to handle unseen camera angles.
Columbia researchers want data infrastructure to do more than store information — they want it to actively keep autonomous agents from causing harm
New arXiv paper shows low-rank adapters can dial up or suppress AI model personality traits like neuroticism and agreeableness, with measurable effects on safet
New preprint cuts encrypted Transformer bootstraps 2.65× with 1.2% perplexity cost, bringing privacy-preserving AI inference closer to practical real-world depl
xDECAF, a browser-based tool checking data flow diagrams against security constraints, shipped as a public preprint with a reusable dataset of over 20 models.