AI coding agent defense cuts malware severity 83%
SkillShield bakes security skills into the system prompt, dropping malware severity from 3.37 to 0.58 across six models while refusing just 0.14% of benign
A newsletter breaking down AI research, technology, and Australian property in plain English
SkillShield bakes security skills into the system prompt, dropping malware severity from 3.37 to 0.58 across six models while refusing just 0.14% of benign
A preprint called WebMCP-Phalanx cuts browser agent attacks from 100% to 0% and blocks 80 of 80 prompt injection attempts, with task utility unchanged.
StepGuard from Shanghai AI Lab checks AI agent tool calls before execution, cutting attack success 77% while losing just 2.8 points of utility.
A multi-agent framework from the University of Stuttgart lets LLM agents design, run, and interpret experiments on pharmaceutical simulation models.
TradingAgents, an open-source multi-agent LLM framework modelling a trading desk with arguing AI agents, approaches 100,000 GitHub stars.
An arXiv preprint found adding security requirements to AI coding prompts cut confirmed flaws from 51 to 24 across six web apps, with zero critical issues.
Purdue researchers found stricter EU regulatory formatting makes LLMs hallucinate more, while vaguer rules need heavier prompts for consistent output.
SRPO turns a model's completed reasoning into per-token training signals without external critics and hits 73.3% on AIME'24 with Qwen3-8B at 8% of the
A new distillation recipe shrinks LLM safety guards to run on commodity CPUs in 24ms, matching 8-billion-parameter teachers on adversarial prompts.
Docling, the open-source Python tool parsing PDFs, video and charts into structured data for AI agents, is trending on GitHub with 65,346 stars.
BERT-LER, trained on 75 million de-identified patient records, matches benchmark models on clinical prediction tasks while explaining its own reasoning.
arXiv preprint argues fixing LLM errors is an operations problem, not a tooling problem, and governance for persisted corrections doesn't exist yet.