ACToR retrieval lifts AI code generation 15% on CoderEval
ACToR, a new retrieval framework from an arXiv preprint, targets critical tokens in AI code generation, lifting CoderEval 15.4% over prior methods.
A newsletter breaking down AI research, technology, and Australian property in plain English
AI and technology, explained through what they actually mean. An AI-assisted newsroom under human editorial rules — every story cites its primary sources so you can check them yourself.
493 stories
ACToR, a new retrieval framework from an arXiv preprint, targets critical tokens in AI code generation, lifting CoderEval 15.4% over prior methods.
New preprint: AI agents learn individual standards through interaction, lifting task success up to 20.9% for professionals needing personal quality.
LoopX, a Python project with 5,615 GitHub stars, adds a durable control plane to AI coding agents like Claude Code and Codex.
OpenAI's GPT-6 Astra launched September 3 with a 1-million-token context window and the first 'Critical' cybersecurity rating in the company's history.
Google's Fairwind Program pairs Gemini 3.8 Flash Cyber with CodeMender, letting 650-plus partners autonomously find and fix code vulnerabilities in
New arXiv preprint: terminal AI agents gain 14 to 18 benchmark points by training against progressively harder environments. Results are unverified.
Vercel Labs' agent-browser, a Rust CLI that lets AI agents drive Chrome without Node.js or Playwright at runtime, is trending on GitHub with 41,800 stars.
DSB-IFEval tests 1,038 cases across eight assistant roles and finds voice agents drop up to 9.7% when instructions are implied, not stated.
New 'epistemic warrant' framework grades individual LLM recommendations from unstable to broadly supported, capturing what confidence scores miss.
SMC pairs a small drafter model with a large actor model to speculatively execute tool calls, reducing wall time by up to 45% on AppWorld.
GPT-6 Astra is OpenAI's most capable deployed model and first to reach Critical cybersecurity capability, able to find and exploit unknown flaws.
NousResearch's Hermes Agent has 240,000 GitHub stars for an open-source AI agent that claims to learn from every conversation and create new skills.