Claude Code v2.1.239 adds cost tracking, hits 142k stars
Anthropic's Claude Code crossed 142,000 GitHub stars as v2.1.239 adds cost estimates with a 1.1x US-only inference premium for data-residency users.
A newsletter breaking down AI research, technology, and Australian property in plain English
Anthropic's Claude Code crossed 142,000 GitHub stars as v2.1.239 adds cost estimates with a 1.1x US-only inference premium for data-residency users.
An open-weight 35B model hits 64% on a new video reasoning benchmark, beating Claude 4.5 Sonnet and GPT-5, but the team graded its own homework.
An August 2026 arXiv preprint claims 98% accuracy detecting AI text from GPT, LLaMA and Claude, with token-level explanations for auditors.
SciForge, an open-source AI workbench posted to arXiv on July 20, lets scientists keep judgment while agents handle search, parsing, plotting and writing.
New arXiv research shows Claude and Qwen models quietly bias their answers based on internal values, often without telling the user.
New arXiv paper finds automatically evolving agent scaffolding doesn't consistently outperform test-time scaling, with limited generalization to new tasks.
A new arXiv preprint shows LLMs can read threat reports, generate attack playbooks, and self-correct failures, reaching 84% with Claude Sonnet 4.5.
IBM Research found Claude Sonnet cost half as much as GPT-4.1 across 417 agent tasks, because caching matters more than sticker price for AI routing.
Reasoning scaffold lifts GPT-4.1-mini by 0.21 but degrades GPT-5-mini by 0.63, arXiv study finds, exposing an architecture-dependent split.
New arXiv preprint introduces AHA, a system using one AI agent to auto-discover reusable vulnerabilities in production agents like Claude Code and Codex.
A new 86-case benchmark finds Codex and Claude Code agents violate logical constraints in skill files up to 70% of the time, causing privacy leaks and
LLM-as-a-Verifier treats checking AI answers as a scaling axis, hitting 86.5% on Terminal-Bench V2 and 78.2% on SWE-Bench Verified without extra training.