LLM agents spot 88% of supply-chain failures but can't act
New STOCKTAKE benchmark: LLM agents detect up to 88% of hidden supply-chain failures but two of four models score below a blind baseline on action.
The AI knowledge 99.9% of people never discover
Daily AI and technology coverage — models, chips, agents and research, and what each development actually means for you and your business.
552 stories
New STOCKTAKE benchmark: LLM agents detect up to 88% of hidden supply-chain failures but two of four models score below a blind baseline on action.
Adversaries can embed hidden instructions in network logs that hijack LLMs used by security teams, with attacks succeeding up to 88.2% of the time.
New arXiv preprint finds offensive AI security agents improve with more compute, but defensive SOC tasks need disciplined tool use over raw reasoning
NVIDIA's Vera Rubin platform reframes AI economics around intelligence per dollar, making continuous post-training the central workload for agentic AI.
A study of 2,361 popular GitHub repos finds most generate just one or two AI-agent pull requests per quarter, with heavy use in small teams.
Google DeepMind and Isomorphic Labs unveil a bioresilience strategy covering AI model safety, outbreak detection, and vaccine design for governments.
A July 2026 arXiv preprint tests PAT, a RAG-based system that feeds whole-document context to LLMs for English-to-Spanish translation, with mixed results.
Oracle agent memory hits 93.8% on LongMemEval with 10.7x fewer tokens than flat-history baselines, per a new arXiv preprint on long-horizon AI agents.
Australia's PM wants big AI data centres to underwrite new power and put back as much energy as they draw, but stalled renewables decide if it works.
A new arXiv survey formalises how AI agents update their own prompts, memory and tools with minimal human input, and what breaks when they do.
PriEval-Protect, a new arXiv preprint, uses a fine-tuned legal LLM and technical data analysis to automate GDPR and HIPAA privacy risk scoring for
A Google Public Sector preprint proposes splitting geospatial AI between big-compute pretraining and expert fine-tuning, with LLMs as orchestrators.