Technology & AI
LLM-as-a-Verifier hits 86.5% on Terminal-Bench V2
LLM-as-a-Verifier treats checking AI answers as a scaling axis, hitting 86.5% on Terminal-Bench V2 and 78.2% on SWE-Bench Verified without extra training.
The AI knowledge 99.9% of people never discover
Anthropic and Claude coverage — model releases, research papers and policy moves, decoded for readers who use AI rather than build it.
25 stories
LLM-as-a-Verifier treats checking AI answers as a scaling axis, hitting 86.5% on Terminal-Bench V2 and 78.2% on SWE-Bench Verified without extra training.