LLM agents hallucinate 36.9% of skill names, study finds
A new arXiv preprint found every agent tested invented non-existent skill names, creating a supply-chain attack path through open skill registries.
A newsletter breaking down AI research, technology, and Australian property in plain English
Daily AI and technology news decoded in plain English — models, chips, agents and research, and what each development actually means for you and your business.
418 stories
A new arXiv preprint found every agent tested invented non-existent skill names, creating a supply-chain attack path through open skill registries.
A new tool shows robot swarms implicitly learn environmental geometry from simple rewards, raising questions about control of emergent collective
A non-peer-reviewed arXiv preprint proposes everyday AI agent prompting loops quietly strengthen impatience and self-criticism through neuroplasticity.
An arXiv preprint shows spectral scores stay frozen while retrieval value collapses from +33% to -35%, challenging how teams decide on context.
First meta-evaluation of LLM-generated rubrics for paper reproduction finds AI graders are overly generous and too detailed, but augmentation helps.
EnCF, a new ensemble controlled-flow filter on arXiv, targets non-Gaussian and multimodal data assimilation where Kalman-type filters fall short.
OpenAI's automated red teaming system GPT-Red uses self-play to find model weaknesses like prompt injection gaps affecting every AI user.
OpenAI's July 15 proposal wants state AI laws to build toward a national safety framework, shaping how every American AI developer gets regulated.
When AI models grade other AI models without an answer key, they hand out passing marks too freely, new research from Potsdam and Ottawa finds.
A 26-billion-parameter diffusion model transcribes speech in eight parallel steps, training just 42 million parameters for 6.6% word error rate on
A new arXiv preprint recasts fall detection as a stability-loss physics problem using liquid time-constant networks for low-power edge devices.
Bulkhead, a new arXiv preprint, uses multi-agent LLMs to automatically find and fix path traversal vulnerabilities in containers running AI workloads.