LLM refusal neurons: causal audit finds no unique mechanism
A preprint audit across five LLMs finds refusal lives in a redundant subspace — and trusted rank-stable methods can be the least causally valid.
A newsletter breaking down AI research, technology, and Australian property in plain English
AI and technology, explained through what they actually mean. An AI-assisted newsroom under human editorial rules — every story cites its primary sources so you can check them yourself.
506 stories
A preprint audit across five LLMs finds refusal lives in a redundant subspace — and trusted rank-stable methods can be the least causally valid.
MedPMC's 11M curated medical image-text pairs trained a CLIP model beating larger datasets by 7.1 points — 95% of its images are actually medical.
LLM hidden states hold richer confidence than models verbalise; a new method extracts it before generation finishes, affecting any AI system needing to know whe
Researchers propose SPELLSMITH, which rewrites MCP server tool descriptions to steer LLM agents away from taint-style vulnerabilities without code changes.
A new arXiv study finds ChatGPT sends users to external sites in just 5.2% of sessions, cutting Google search use 9.4% and reshaping the web's referral economy.
SciReasoner, a multimodal model for structural reasoning across proteins, molecules and crystals, matches frontier LLMs in 98% of expert reviews.
A new arXiv preprint reframes neural network training as optimal control, inserting layers where error is highest — but evidence is limited to scientific datase
OpenAI's new analysis exposes reliability issues in SWE-bench Pro — the very benchmark it recommended to replace the contaminated SWE-bench Verified — leaving A
OpenAI Academy and the Walton Family Foundation are launching hands-on AI Skills Jams to help K-12 educators build practical classroom AI skills.
Anthropic researchers found a small broadcast hub of silent concepts inside Claude — the J-space — that drives its deliberate reasoning, and built a lens that lets them watch it think.
An ETH Zurich preprint proves four discrete diffusion methods optimize the same object — and reveals why one popular parameterization diverges at initialization
GaP, a multi-agent coding framework on arXiv, builds computation graphs for robot tasks and rehearses them in simulation to tackle variable automation's reliabi