AI persona agents leak private traits, defenses fail
AntiSkillBench tests 7,500 dialogue traces across three frontier agents and finds privacy risks persist from explicit data to communication style.
A newsletter breaking down AI research, technology, and Australian property in plain English
AntiSkillBench tests 7,500 dialogue traces across three frontier agents and finds privacy risks persist from explicit data to communication style.
An arXiv preprint proposes replaying temporally-evolving enterprise worlds to test AI agents at any moment, fixing a blind spot in current evals.
New arXiv paper proposes SkillBoost, a three-stage framework that stops LLM agents overfitting to limited experience and forgetting solved tasks.
EvoPINN uses an LLM agent to automatically discover and validate algorithms for physics-informed neural networks, replacing manual tuning.
TAPR uses reinforcement learning to turn vague user prompts into task-optimized instructions, lifting accuracy on Natural Questions and GSM8K benchmarks.
AI-generated research papers scored below the midpoint in the first automated multi-model peer review, and one AI reviewer disagreed with the others
GLASS, a new arXiv preprint, uses sparse autoencoders to extract a user's writing style and inject it at inference, no fine-tuning or retrieval needed.
An arXiv preprint shows organizations can collaboratively predict equipment failure using federated survival analysis without sharing raw sensor data.
More than half of Australians surveyed by the RBA think higher interest rates push inflation up, not down, and it changes how monetary policy actually