AI coding agent defense cuts malware severity 83%
SkillShield bakes security skills into the system prompt, dropping malware severity from 3.37 to 0.58 across six models while refusing just 0.14% of benign
A newsletter breaking down AI research, technology, and Australian property in plain English
SkillShield bakes security skills into the system prompt, dropping malware severity from 3.37 to 0.58 across six models while refusing just 0.14% of benign
StepGuard from Shanghai AI Lab checks AI agent tool calls before execution, cutting attack success 77% while losing just 2.8 points of utility.
An arXiv preprint found adding security requirements to AI coding prompts cut confirmed flaws from 51 to 24 across six web apps, with zero critical issues.
Google's Gemini CLI offers 1,000 free AI requests daily in the terminal, with a 1M token context window and Gemini 3 models. Now 106,000 stars.
NVIDIA's Groq 3 LPX inference accelerator is now in full production, delivering a record 3,400 tokens/second for agentic AI workloads.
The TypeScript AI agent platform landed on GitHub's daily trending list on 22 August, pitching multi-tenant team features the bigger names lack.
Tencent Zhuque Lab's open-source tool scans AI agents, MCP servers and LLM jailbreaks. It has 5,200+ stars and a README that explicitly asks for them.
Anthropic's Claude Code crossed 142,000 GitHub stars as v2.1.239 adds cost estimates with a 1.1x US-only inference premium for data-residency users.
GxP-Agent encodes regulatory steps as a graph, turning 0% failure into 100% structural match on a new FDA-pilot clinical trial benchmark.
A new arXiv paper proposes DeAR, a framework where AI agents coordinate reasoning without a central orchestrator, tested across nine benchmarks.
Cambridge researchers find looped language models improve at multi-step API tool calling, with adaptive inference offering the best compute-performance
A new arXiv preprint surveys agentic AI from working principles to adoption factors, as developer frameworks signal explosive real-world traction.