EvoPINN: AI agent invents new neural network for physics
EvoPINN uses an LLM agent to automatically discover and validate algorithms for physics-informed neural networks, replacing manual tuning.
A newsletter breaking down AI research, technology, and Australian property in plain English
EvoPINN uses an LLM agent to automatically discover and validate algorithms for physics-informed neural networks, replacing manual tuning.
TAPR uses reinforcement learning to turn vague user prompts into task-optimized instructions, lifting accuracy on Natural Questions and GSM8K benchmarks.
AI-generated research papers scored below the midpoint in the first automated multi-model peer review, and one AI reviewer disagreed with the others
OmegaUse-OfficeVal pairs 100 real office tasks with labour costs, showing AI agents are cheaper and faster but still fall short on quality.
OpenAI's July 31 post describes a full-stack push for cheaper, more capable AI, arguing falling intelligence costs change what work is worth doing for
An arXiv preprint proposes MTGuard, a hybrid analysis framework that catches malicious MCP tool use in AI agents while preserving normal task performance.
RSMeM adds a knowledge-enhanced memory system to remote sensing AI agents, improving DeepSeek-V3.2 accuracy 6% with under 1% extra tokens, per arXiv.
A new arXiv preprint tracks every repair attempt in LLM code generation, giving developers a replayable audit trail for AI-written code.
A preprint proposes a hybrid AI architecture for e-commerce search, extending discovery from 60% to 80% of queries at 30% of teacher model cost.
SecRespond, a new arXiv benchmark, tested 23 frontier LLMs on real-world post-compromise incident response across 10 cyber ranges. Zero passed.
A new open-source tool from University of Michigan researchers lets any PyTorch model get automatic CUDA kernel speedups without manual GPU programming.
LLM answers change on over 23% of questions when wording shifts, and standard accuracy scores hide the reliability gap from anyone deploying AI.