AI agents self-evolve security defenses in new HARD framework
A new preprint proposes HARD, a framework where LLM agents automatically build and improve their own runtime defenses from observed failures.
A newsletter breaking down AI research, technology, and Australian property in plain English
A new preprint proposes HARD, a framework where LLM agents automatically build and improve their own runtime defenses from observed failures.
ICML 2026's reproduction challenge used AI agents to verify 2,200 papers, and 23% had a falsified or contested claim while 49 had nothing verifiable.
An arXiv preprint shows LLMs autonomously controlling farm lighting and actuators, cutting grow cycles 35% and finding energy strategies humans missed.
OpenAI's Ultrafast API tier runs GPT-5.6 Sol up to 14 times faster than standard, powered by Cerebras chips rather than NVIDIA GPUs.
OpenAI's builder guide for GPT-5.6 promotes smarter model selection and new Responses API tools for startups building AI agents.
Model ML uses GPT-5.6 Sol to carry finance work from research to editable, traceable PowerPoint and Excel files. No efficiency metrics are published.
NVIDIA joins the NSF's new State and Regional AI Infrastructure Hubs, expanding AI computing and workforce training across US universities.
MedUPS trains AI on mid-stream clinical decisions rather than final diagnosis and raises next-step accuracy up to 11 points on 21,874 real case reports.
RAG study tests LLaMA, Mistral and Qwen to cut hallucinations in small business AI, but the arXiv preprint provides no benchmark numbers.
New arXiv preprint proposes Chained RLM, calling the same model with fresh context to stop early reasoning errors reaching the final answer.
OpenAI's improved GPT-5.6 Sol promises better accuracy and consistency in ChatGPT, with expanded free access. Here's who benefits and what's missing.
A prespecified study found zero unauthorized tool calls across 840 GPT-5.6 agent trajectories, but raising reasoning effort changed inspection behaviour