LLM answers silently shaped by own values, study finds
New arXiv research shows Claude and Qwen models quietly bias their answers based on internal values, often without telling the user.
A newsletter breaking down AI research, technology, and Australian property in plain English
New arXiv research shows Claude and Qwen models quietly bias their answers based on internal values, often without telling the user.
A new arXiv paper uses cognitive psychology's set-shifting concept to test whether LLM agents adapt when reliable tools silently change mid-session.
A new two-agent framework finds code bugs across whole repositories at 16-48x lower cost, but over half its correct answers rely on flawed reasoning.
A July 2026 study of nine endoscopy AI systems finds high answer accuracy doesn't guarantee faithful clinical reasoning, affecting hospital AI buyers.
A new arXiv paper from ShanghaiTech replaces binary safe/unsafe guards with a three-way EXECUTE-ASK-REFUSE routing that cuts alert fatigue for anyone
Graph neural network hints lift small language model molecular property prediction by up to 74% on Tox21, according to a new arXiv preprint.
A black-box audit swaps predicates in chain-of-thought to expose when GPT-4o reaches correct answers through reasoning that ignores its stated premises.
IRT rankings of AI models are unreliable when few systems are tested, 18,000 simulations show, threatening how the industry compares models.
A new preprint tested 327 real-world agent skills and found vulnerabilities across the entire skill lifecycle, from admission to evolution.
arXiv paper explains why vulnerability scanners give different results for the same software, tracing inconsistency to a distributed ecosystem.
A new arXiv preprint shows LLMs can read threat reports, generate attack playbooks, and self-correct failures, reaching 84% with Claude Sonnet 4.5.
Structural priors lift LLM vulnerability recall from 20% to 100% on synthetic benchmarks, but real CVE data exposes a 51-point collapse.