OpenAI says coding agents reshape its own AI research
OpenAI says coding agents are reshaping its internal AI research, with early data on agent usage and experiment velocity, but no verifiable metrics.
The AI knowledge 99.9% of people never discover
Daily AI and technology coverage — models, chips, agents and research, and what each development actually means for you and your business.
551 stories
OpenAI says coding agents are reshaping its internal AI research, with early data on agent usage and experiment velocity, but no verifiable metrics.
OpenAI's Codex CLI, a Rust-based terminal coding agent with 121,803 GitHub stars, is trending. Here's what it does and who can use it.
ACToR, a new retrieval framework from an arXiv preprint, targets critical tokens in AI code generation, lifting CoderEval 15.4% over prior methods.
New preprint: AI agents learn individual standards through interaction, lifting task success up to 20.9% for professionals needing personal quality.
LoopX, a Python project with 5,615 GitHub stars, adds a durable control plane to AI coding agents like Claude Code and Codex.
OpenAI's GPT-6 Astra launched September 3 with a 1-million-token context window and the first 'Critical' cybersecurity rating in the company's history.
Google's Fairwind Program pairs Gemini 3.8 Flash Cyber with CodeMender, letting 650-plus partners autonomously find and fix code vulnerabilities in
New arXiv preprint: terminal AI agents gain 14 to 18 benchmark points by training against progressively harder environments. Results are unverified.
Vercel Labs' agent-browser, a Rust CLI that lets AI agents drive Chrome without Node.js or Playwright at runtime, is trending on GitHub with 41,800 stars.
DSB-IFEval tests 1,038 cases across eight assistant roles and finds voice agents drop up to 9.7% when instructions are implied, not stated.
New 'epistemic warrant' framework grades individual LLM recommendations from unstable to broadly supported, capturing what confidence scores miss.
SMC pairs a small drafter model with a large actor model to speculatively execute tool calls, reducing wall time by up to 45% on AppWorld.