Scaling improves LLM social simulation — but not human biases
A Stanford preprint testing 85 LLMs finds bigger models better predict opinions and behaviour, but fail to capture cognitive biases like risk aversion.
A newsletter breaking down AI research, technology, and Australian property in plain English
AI and technology, explained through what they actually mean. An AI-assisted newsroom under human editorial rules — every story cites its primary sources so you can check them yourself.
500 stories
A Stanford preprint testing 85 LLMs finds bigger models better predict opinions and behaviour, but fail to capture cognitive biases like risk aversion.
LeRobot v0.6.0 adds three world model policies that teach robots to imagine the future during training, then discard that imagination at inference — for free. H
A new multi-agent AI framework converts natural-language Operations Research problems into executable math, hitting state-of-the-art on 3 of 4 benchmarks — and
Indic AI researchers propose 'Culture Sensing,' a hermeneutic approach to building foundation models that serve India's low-resource languages without erasing c
New arXiv preprint proposes representing assets from any blockchain as ERC-20 tokens on one backend chain, enabling cross-chain lending and privacy apps for dev
A new arXiv preprint constrains model fine-tuning to a trusted adapter subspace, blocking backdoor attacks to 8% while matching full LoRA on clean data.
Researchers show human-readable text in a robot's view can trigger 6.96x latency in vision-language models, with physical prints still hitting 4.74x.
A July 2026 arXiv preprint introduces GPBACC, a coded-computing technique that tackles privacy leakage and malicious workers simultaneously across federated and
A new study of 85 Chrome extension wallets finds routine operations silently connect users' addresses, enabling cross-site tracking and possible deanonymization
A July 2026 arXiv paper formalises the interval-width trade-off in polynomial activation replacement for encrypted AI inference, affecting privacy-preserving ML
Unpeer-reviewed July 2026 arXiv preprint proposes treating LLM workflows as persistent, inspectable knowledge objects. Here's what it means for builders.
UniClawBench, a new arXiv benchmark with 400 bilingual tasks, tests proactive AI agents in live containers with simulated human feedback — vital for anyone depl