PriEval-Protect scores hospital privacy risk with legal LLM
PriEval-Protect, a new arXiv preprint, uses a fine-tuned legal LLM and technical data analysis to automate GDPR and HIPAA privacy risk scoring for
A newsletter breaking down AI research, technology, and Australian property in plain English
PriEval-Protect, a new arXiv preprint, uses a fine-tuned legal LLM and technical data analysis to automate GDPR and HIPAA privacy risk scoring for
A Google Public Sector preprint proposes splitting geospatial AI between big-compute pretraining and expert fine-tuning, with LLMs as orchestrators.
New arXiv preprint from Microsoft and Walmart researchers finds encoded prompts bypass AI safety filters most, raising the stakes for enterprise AI
A new arXiv preprint proposes a training-free detector that reads LLM hidden states to catch prompt injection attacks on purpose-specific agents, reporting
New E3 framework cuts AI agent cost 85% by estimating task difficulty before acting, matching 100% success on a 121-edit benchmark with code released.
A new arXiv preprint found every agent tested invented non-existent skill names, creating a supply-chain attack path through open skill registries.
First meta-evaluation of LLM-generated rubrics for paper reproduction finds AI graders are overly generous and too detailed, but augmentation helps.
OpenAI's automated red teaming system GPT-Red uses self-play to find model weaknesses like prompt injection gaps affecting every AI user.
When AI models grade other AI models without an answer key, they hand out passing marks too freely, new research from Potsdam and Ottawa finds.
A 26-billion-parameter diffusion model transcribes speech in eight parallel steps, training just 42 million parameters for 6.6% word error rate on
Bulkhead, a new arXiv preprint, uses multi-agent LLMs to automatically find and fix path traversal vulnerabilities in containers running AI workloads.
A new arXiv study finds LLM-suggested password replacements are more secure than human-created ones and equally memorable after one week.