Qwen, Mistral and Llama verify fake developer identity, study finds
Qwen, Mistral and Llama accepted a fake developer identity without external proof, a new arXiv preprint reports, exposing a gap in AI identity handling.
A newsletter breaking down AI research, technology, and Australian property in plain English
Qwen, Mistral and Llama accepted a fake developer identity without external proof, a new arXiv preprint reports, exposing a gap in AI identity handling.
X's For You feed algorithm repo is trending on GitHub after August updates added production training code, visibility filters, and a Brazil election
OpenAI says coding agents are reshaping its internal AI research, with early data on agent usage and experiment velocity, but no verifiable metrics.
OpenAI's Codex CLI, a Rust-based terminal coding agent with 121,803 GitHub stars, is trending. Here's what it does and who can use it.
ACToR, a new retrieval framework from an arXiv preprint, targets critical tokens in AI code generation, lifting CoderEval 15.4% over prior methods.
New preprint: AI agents learn individual standards through interaction, lifting task success up to 20.9% for professionals needing personal quality.
LoopX, a Python project with 5,615 GitHub stars, adds a durable control plane to AI coding agents like Claude Code and Codex.
OpenAI's GPT-6 Astra launched September 3 with a 1-million-token context window and the first 'Critical' cybersecurity rating in the company's history.
DSB-IFEval tests 1,038 cases across eight assistant roles and finds voice agents drop up to 9.7% when instructions are implied, not stated.
New 'epistemic warrant' framework grades individual LLM recommendations from unstable to broadly supported, capturing what confidence scores miss.
SMC pairs a small drafter model with a large actor model to speculatively execute tool calls, reducing wall time by up to 45% on AppWorld.
GPT-6 Astra is OpenAI's most capable deployed model and first to reach Critical cybersecurity capability, able to find and exploit unknown flaws.