PagedWeight cuts MoE serving memory 72% with FP16 accuracy
A new preprint claims PagedWeight dynamically quantizes MoE model weights at runtime, saving 72% GPU memory and lifting throughput 1.94× for AI inference.
A newsletter breaking down AI research, technology, and Australian property in plain English
Daily AI and technology news decoded in plain English — models, chips, agents and research, and what each development actually means for you and your business.
418 stories
A new preprint claims PagedWeight dynamically quantizes MoE model weights at runtime, saving 72% GPU memory and lifting throughput 1.94× for AI inference.
A July 2026 arXiv preprint proposes learning interpretable trustworthiness levels and monitoring them across an AI system's entire lifecycle.
A new arXiv study of OECD-catalogued AI trust tools finds the field heavy on fairness and transparency but thin on security, explainability and
NVIDIA's SIGGRAPH keynote unveils MCP connections letting AI agents work inside creative apps, plus Cosmos 3 Edge for local physical AI.
A new arXiv preprint claims to solve a hard problem in probability, running on a desktop CPU without neural networks. Scientists and engineers take note.
A new arXiv preprint introduces a method to measure whether frontier AI helps non-experts plan CBRN attacks, with uplift limited to radiological threats.
An arXiv preprint analysing 14 years of Qubes Security Bulletins finds most flaws originate in Xen and CPU components, not Qubes itself.
A new arXiv preprint introduces Mycelium, a shared workspace connecting researchers and AI agents, routing hypotheses to whoever can act on them.
A new arXiv preprint strips named entities from training data, and the resulting models recall fewer facts but answer questions from context up to 25%
OpenAI's July 2026 article outlines age-appropriate protections, learning tools and parental controls for teen ChatGPT users, but key questions remain.
A new arXiv position paper argues AI must formalize entire mathematical theories, not isolated statements, to build verifiable knowledge bases.
Uniform mixing matches Transformer spatial attention across six traffic benchmarks at 0.14% MAE gap, while cutting complexity from O(N²) to O(N).