ProjAgent hits 41% on REPOCOD with procedural code retrieval
ProjAgent adds a new retrieval signal that finds code doing the same job under different names, hitting 41.14% Pass@1 on the REPOCOD benchmark for repository-le
A newsletter breaking down AI research, technology, and Australian property in plain English
Daily AI and technology news decoded in plain English — models, chips, agents and research, and what each development actually means for you and your business.
418 stories
ProjAgent adds a new retrieval signal that finds code doing the same job under different names, hitting 41.14% Pass@1 on the REPOCOD benchmark for repository-le
A July 2026 preprint finds models compressed to lower precision diverge in which questions they get right, even when overall accuracy holds steady — exposing a
A new arXiv preprint called SLORR adds under 1% training overhead in LLM pretraining while making models more compressible, without SVDs or architecture changes
OpenAI's new ChatGPT Work agent promises to turn goals into finished work by acting across your apps and files for hours — but key details remain undisclosed.
OpenAI's GPT-5.6 is now the preferred model powering Microsoft 365 Copilot across Word, Excel and PowerPoint, affecting millions of daily office users worldwide
OpenAI's GPT-5.5 Bio Bug Bounty invites researchers to test universal jailbreaks for biological risks in a model built for autonomous, multi-tool work.
OpenAI's GPT-5.6 launches a three-model lineup — Sol, Terra and Luna — promising token efficiency and frontend gains, but with no independent benchmarks yet.
Stronger privacy cuts AI generalisation error in high-noise regimes, yet the robustness-privacy tension returns when noise is low, according to new arXiv prepri
ALER-TI, a new arXiv preprint, retrieves cached historical patterns to reconstruct missing time series values, tested across six real-world datasets under varyi
CNeVA lets developers steer virtual driver aggression and caution on the Waymo benchmark, offering per-channel control that higher-ranked imitation models lack.
ActionCache cuts inference latency for vision-language-action robot models by up to 34× without retraining, a potential breakthrough for real-time robotic deplo
G-RRM uses neural reasoning models to guide classical SAT solvers, cutting Sudoku backtracking 33× — but the speedup vanishes if the solver can't reject bad hin