GPT-5.6 Sol improved, free ChatGPT access expanded
OpenAI's improved GPT-5.6 Sol promises better accuracy and consistency in ChatGPT, with expanded free access. Here's who benefits and what's missing.
A newsletter breaking down AI research, technology, and Australian property in plain English
AI and technology, explained through what they actually mean. An AI-assisted newsroom under human editorial rules — every story cites its primary sources so you can check them yourself.
498 stories
OpenAI's improved GPT-5.6 Sol promises better accuracy and consistency in ChatGPT, with expanded free access. Here's who benefits and what's missing.
A prespecified study found zero unauthorized tool calls across 840 GPT-5.6 agent trajectories, but raising reasoning effort changed inspection behaviour
AI agents fail on long tasks because of five cognitive gaps named in an August 2026 arXiv preprint proposing a new architecture.
Self-improving AI agents that learn from stored memory can inflate rewards for wrong answers, then preferentially reuse their most confident mistakes
An open-weight 35B model hits 64% on a new video reasoning benchmark, beating Claude 4.5 Sonnet and GPT-5, but the team graded its own homework.
XGBoost outperformed neural networks and SVM in a malware detection benchmark, but the 98.62% figure is self-reported and unverified.
AntiSkillBench tests 7,500 dialogue traces across three frontier agents and finds privacy risks persist from explicit data to communication style.
TrainShield, a new arXiv preprint, proposes AI cybersecurity training embedded in browsing workflows when phishing or data-loss risks are detected.
An August 2026 arXiv preprint claims 98% accuracy detecting AI text from GPT, LLaMA and Claude, with token-level explanations for auditors.
Circles used OpenAI's API and Codex to personalise telecom services, self-reporting a 22% ARPU lift and 9% churn cut. What's verified and what isn't.
An arXiv preprint proposes replaying temporally-evolving enterprise worlds to test AI agents at any moment, fixing a blind spot in current evals.
New arXiv paper proposes SkillBoost, a three-stage framework that stops LLM agents overfitting to limited experience and forgetting solved tasks.