MindReader: LLM tool makes replacement passwords harder to guess
A new arXiv study finds LLM-suggested password replacements are more secure than human-created ones and equally memorable after one week.
A newsletter breaking down AI research, technology, and Australian property in plain English
A new arXiv study finds LLM-suggested password replacements are more secure than human-created ones and equally memorable after one week.
Reinforcement fine-tuning with 30 prompts cut an open-weight model's building emissions to 61.2 kg-CO2, near the 60.8 optimum, a preprint shows.
Researchers find LLM judge bias lives in a low-dimensional subspace inside model hidden states, and steering along it can switch bias on and off.
A new arXiv preprint introduces StoryTeller, a training-free system that keeps film audio descriptions coherent for blind and low-vision viewers.
New arXiv preprint shows Transformer training on inductive reasoning can be confined to a low-dimensional manifold for automatic circuit detection.
An arXiv preprint finds sampling temperature on retrieval-augmented LLMs controls how strongly ideological framing from source documents bleeds into
Reasoning scaffold lifts GPT-4.1-mini by 0.21 but degrades GPT-5-mini by 0.63, arXiv study finds, exposing an architecture-dependent split.
FactorDiff decomposes diffusion samples into pixel-level factors and routes each to the best expert, beating global weighting on ARC-AGI reasoning tasks.
AMT-X multi-turn red-teaming framework hit up to 100% attack success on six frontier AI models, exposing critical gaps in current LLM safety testing.
PromptGraph models LLM prompts as graphs to catch privacy leaks that existing sanitisers miss, affecting anyone sending sensitive data to cloud AI.
A new open benchmark with 500+ tools reveals even frontier AI models can't reliably read an image and act on it, with most failures traced to seeing, not
Vision-language models internally encode the right answer but misread it, a new arXiv study finds. A simple fix lifts counting accuracy 15.6 points.