OpenAI GPT-Red automates red teaming with self-play
OpenAI's automated red teaming system GPT-Red uses self-play to find model weaknesses like prompt injection gaps affecting every AI user.
A newsletter breaking down AI research, technology, and Australian property in plain English
OpenAI's automated red teaming system GPT-Red uses self-play to find model weaknesses like prompt injection gaps affecting every AI user.
A new arXiv preprint recasts fall detection as a stability-loss physics problem using liquid time-constant networks for low-power edge devices.
Bulkhead, a new arXiv preprint, uses multi-agent LLMs to automatically find and fix path traversal vulnerabilities in containers running AI workloads.
A new arXiv study finds LLM-suggested password replacements are more secure than human-created ones and equally memorable after one week.
Reinforcement fine-tuning with 30 prompts cut an open-weight model's building emissions to 61.2 kg-CO2, near the 60.8 optimum, a preprint shows.
Researchers find LLM judge bias lives in a low-dimensional subspace inside model hidden states, and steering along it can switch bias on and off.
MCP security scanners flag 96.89% of servers as risky, but a study of 64,611 servers finds fewer than half of alerts are true positives.
New arXiv preprint Prezta replaces application gateways at critical infrastructure edges with zero-knowledge proofs generated on client devices.
Researchers propose Q-DIBA, a backdoor attack generating a unique trigger for each input to quantum neural networks, evading three tested defenses.