Small AI models outperform LLM prompt guardrails, Fence says
A new arXiv preprint claims small language models trained on synthetic data outperform prompt-based LLM safety checks for hallucination and topic drift.
A newsletter breaking down AI research, technology, and Australian property in plain English
A new arXiv preprint claims small language models trained on synthetic data outperform prompt-based LLM safety checks for hallucination and topic drift.
New arXiv preprint proposes Probabilistic Concept-Aware Steering, fixing incoherent behaviour in LLM steering vectors for open-weight model teams.
PEARL puts a math solver inside the LLM loop to test, debug and revise optimization models, beating models 170× larger on verified solve rates.
A July 2026 arXiv preprint introduces TPIPS, a text-prompted image similarity metric showing frontier vision-language models lag human perception.
New arXiv framework uses twelve reasoning angles to separate jokes from hate in memes, hitting 80.3% humor and 75.9% hate detection on two benchmarks.
Six LLMs tested across navigation, triage and finance tasks showed stable risk preferences, a hidden trait anyone deploying AI agents needs to understand.
MOSAIC, posted on arXiv on July 21, lifts long-conversation accuracy 27 points and catches 66% of factual conflicts before they corrupt an agent's memory.
New arXiv preprint shows AI agents can be redirected by reordering true facts, with 83.3% success across GPT, Claude, Gemini, DeepSeek and Qwen.
A new arXiv preprint maps 43 operations that corrupt self-hosted AI agents via legitimate OS calls, finding a residual surface no defense can distinguish.
LaCache caches unchanged tokens during diffusion LLM denoising, cutting inference 1.3X standalone and up to 40X stacked with other acceleration methods.
A July 2026 arXiv paper proposes that human raters' emotional states leak into preference labels used to train AI, with a five-prediction audit plan.
A new arXiv framework uses an LLM-based simulator to predict code outcomes, cutting agent training 14x and search-based inference 3-6x.