# Not A Tech Guy > Daily AI and technology news explained in plain English — what each story actually means for you and your business. Plus Australian property analysis. Public Ghost content for AI and LLM tooling. Use `/llms-full.txt` for consolidated page and post context. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages - [About Not A Tech Guy](https://www.notatechguy.com/about.md) - I'm not a tech guy. That's the point. Most technology news is written by insiders, for insiders — dense with jargon, thin on consequences. Not A Tech Guy flips that: every story here decodes AI and technology news in plain English, and every article answers the only questions that matter — what doe… - [Archive](https://www.notatechguy.com/archive.md) - Loading the archive… - [Contact](https://www.notatechguy.com/contact.md) - The fastest way to reach Not A Tech Guy is email: notatechguy@agentmail.to We read everything. Use it for: * Corrections — see our corrections policy for how errors are handled. * Story tips — a release, dataset or development you think we should cover. * Press and partnerships. * Anything about ho… - [Corrections](https://www.notatechguy.com/corrections.md) - We aim to get things right the first time, and we fix them visibly when we don't. Our policy * Factual errors are corrected in the article as soon as they are confirmed, with a correction note stating what was wrong and when it was fixed. * Significant errors — anything that changes the meaning of… - [Editorial Policy](https://www.notatechguy.com/editorial-policy.md) - Not A Tech Guy explains technology, AI and the Australian property market in plain English. This page describes how the publication works and the standards its stories must meet. What this publication is Not A Tech Guy is a personal project: an automated publication its owner built to follow and de… - [Privacy Policy](https://www.notatechguy.com/privacy.md) - Last updated: 13 July 2026 This is the privacy policy for notatechguy.com ("we", "us", "the site"). It is written in plain English because that is the whole point of this publication. The short version We collect your email address for exactly two purposes: to send you the newsletter you signed up… - [Terms of Use & Disclaimer](https://www.notatechguy.com/terms.md) - Last updated: 13 July 2026 By reading notatechguy.com or subscribing to its newsletter, you agree to these terms. What this publication is — and isn't notatechguy.com is a personal project: an automated publication built by its owner to follow AI and technology developments from around the world, s… ## Posts - [Google Pics: AI image editing inside Docs and Slides](https://www.notatechguy.com/google-pics-ai-image-editing-inside-docs-and-slides.md) - Google Pics rolls out to AI Pro and Ultra subscribers, bringing object segmentation and in-image text editing directly into Workspace apps. - [Qwen, Mistral and Llama verify fake developer identity, study finds](https://www.notatechguy.com/qwen-mistral-and-llama-verify-fake-developer-identity-study-finds.md) - Qwen, Mistral and Llama accepted a fake developer identity without external proof, a new arXiv preprint reports, exposing a gap in AI identity handling. - [X For You feed algorithm open-sourced on GitHub](https://www.notatechguy.com/x-for-you-feed-algorithm-open-sourced-on-github.md) - X's For You feed algorithm repo is trending on GitHub after August updates added production training code, visibility filters, and a Brazil election - [OpenAI says coding agents reshape its own AI research](https://www.notatechguy.com/openai-says-coding-agents-reshape-its-own-ai-research.md) - OpenAI says coding agents are reshaping its internal AI research, with early data on agent usage and experiment velocity, but no verifiable metrics. - [OpenAI Codex CLI trends on GitHub with 121,803 stars](https://www.notatechguy.com/openai-codex-cli-trends-on-github-with-121-803-stars.md) - OpenAI's Codex CLI, a Rust-based terminal coding agent with 121,803 GitHub stars, is trending. Here's what it does and who can use it. - [ACToR retrieval lifts AI code generation 15% on CoderEval](https://www.notatechguy.com/actor-retrieval-lifts-ai-code-generation-15-on-codereval.md) - ACToR, a new retrieval framework from an arXiv preprint, targets critical tokens in AI code generation, lifting CoderEval 15.4% over prior methods. - [AI agents learn your standards, up to 20.9% better](https://www.notatechguy.com/ai-agents-learn-your-standards-up-to-20-9-better.md) - New preprint: AI agents learn individual standards through interaction, lifting task success up to 20.9% for professionals needing personal quality. - [LoopX adds durable memory to Claude Code and Codex agents](https://www.notatechguy.com/loopx-adds-durable-memory-to-claude-code-and-codex-agents.md) - LoopX, a Python project with 5,615 GitHub stars, adds a durable control plane to AI coding agents like Claude Code and Codex. - [GPT-6 Astra: OpenAI's first model to hit Critical cyber level](https://www.notatechguy.com/gpt-6-astra-openai-s-first-model-to-hit-critical-cyber-level.md) - OpenAI's GPT-6 Astra launched September 3 with a 1-million-token context window and the first 'Critical' cybersecurity rating in the company's history. - [Google Fairwind Program: AI writes verified patches in minutes](https://www.notatechguy.com/google-fairwind-program-ai-writes-verified-patches-in-minutes.md) - Google's Fairwind Program pairs Gemini 3.8 Flash Cyber with CodeMender, letting 650-plus partners autonomously find and fix code vulnerabilities in - [Environment evolution for terminal agents: 18-point gain](https://www.notatechguy.com/environment-evolution-for-terminal-agents-18-point-gain.md) - New arXiv preprint: terminal AI agents gain 14 to 18 benchmark points by training against progressively harder environments. Results are unverified. - [Vercel Labs agent-browser hits 41,800 GitHub stars](https://www.notatechguy.com/vercel-labs-agent-browser-hits-41-800-github-stars.md) - Vercel Labs' agent-browser, a Rust CLI that lets AI agents drive Chrome without Node.js or Playwright at runtime, is trending on GitHub with 41,800 stars. - [Voice agents drop 9.7% when instructions are merely implied](https://www.notatechguy.com/voice-agents-drop-9-7-when-instructions-are-merely-implied.md) - DSB-IFEval tests 1,038 cases across eight assistant roles and finds voice agents drop up to 9.7% when instructions are implied, not stated. - [When to trust LLM recommendations: four-tier framework](https://www.notatechguy.com/when-to-trust-llm-recommendations-four-tier-framework.md) - New 'epistemic warrant' framework grades individual LLM recommendations from unstable to broadly supported, capturing what confidence scores miss. - [New technique cuts AI agent wait time up to 45%](https://www.notatechguy.com/new-technique-cuts-ai-agent-wait-time-up-to-45.md) - SMC pairs a small drafter model with a large actor model to speculatively execute tool calls, reducing wall time by up to 45% on AppWorld. - [GPT-6 Astra crosses OpenAI's Critical cybersecurity threshold](https://www.notatechguy.com/gpt-6-astra-crosses-openai-s-critical-cybersecurity-threshold.md) - GPT-6 Astra is OpenAI's most capable deployed model and first to reach Critical cybersecurity capability, able to find and exploit unknown flaws. - [Hermes Agent: 240k GitHub stars for self-improving AI](https://www.notatechguy.com/hermes-agent-240k-github-stars-for-self-improving-ai.md) - NousResearch's Hermes Agent has 240,000 GitHub stars for an open-source AI agent that claims to learn from every conversation and create new skills. - [Gilbert + Tobin deploys ChatGPT Enterprise firm-wide](https://www.notatechguy.com/gilbert-tobin-deploys-chatgpt-enterprise-firm-wide.md) - Gilbert + Tobin is scaling ChatGPT Enterprise and Codex across the firm with CEO-led governance, targeting operational work rather than legal advice - [Prompt injection attack on AI agents jumps to 28% success](https://www.notatechguy.com/prompt-injection-attack-on-ai-agents-jumps-to-28-success.md) - SIR, a self-improving prompt injection attack, raised its success rate from 0% to 28% on Gemini 3.5 Flash while the agent's legitimate task still - [AIMC dashboard flags recurring flaws in AI-generated science](https://www.notatechguy.com/aimc-dashboard-flags-recurring-flaws-in-ai-generated-science.md) - AIMC, a visual analytics framework, lets researchers track quality, themes and weaknesses across papers produced by autonomous AI scientist FARS. - [ToolSiphon drains 74% of data from LLM agent tools](https://www.notatechguy.com/toolsiphon-drains-74-of-data-from-llm-agent-tools.md) - New research shows ToolSiphon can extract most source records from LLM agent tools using only queries, affecting anyone building agent-based services. - [AI morbidity and mortality framework proposed for hospitals](https://www.notatechguy.com/ai-morbidity-and-mortality-framework-proposed-for-hospitals.md) - A new arXiv paper adapts the blameless M&M conference for AI errors, with two clinicians agreeing on all 20 classifications across five cases. - [AWS Firecracker microVM tops 36,443 GitHub stars](https://www.notatechguy.com/aws-firecracker-microvm-tops-36-443-github-stars.md) - Firecracker, AWS's Rust microVM engine behind Lambda and Fargate, hit 36,443 GitHub stars this week. Here is what it means for serverless teams. - [Synthetic data privacy is a claim, not a guarantee, researchers warn](https://www.notatechguy.com/synthetic-data-privacy-is-a-claim-not-a-guarantee-researchers-warn.md) - Synthetic data is widely used in privacy-sensitive settings without threat models or falsifiable claims, and rare and minority records face the greatest - [DiaSentinel AI agents screen diabetes risk on-premise](https://www.notatechguy.com/diasentinel-ai-agents-screen-diabetes-risk-on-premise.md) - DiaSentinel, a new arXiv preprint from a Taiwan hospital team, describes a fully on-premise multi-agent system for type 2 diabetes risk screening from - [OpenAI's Astra hits Critical cybersecurity threshold](https://www.notatechguy.com/openai-s-astra-hits-critical-cybersecurity-threshold.md) - OpenAI's Astra is the first model to meet the Critical cybersecurity threshold under the company's own Preparedness Framework, with stronger release - [MMJailBench: prompt framing is top jailbreak risk across 16 AI models](https://www.notatechguy.com/mmjailbench-prompt-framing-is-top-jailbreak-risk-across-16-ai-models.md) - MMJailBench, a new factorized benchmark, evaluated 16 multimodal LLMs and found prompt framing, not visual tricks, drives most jailbreak vulnerabilities. - [Open-source trust signals are breaking, study finds](https://www.notatechguy.com/open-source-trust-signals-are-breaking-study-finds.md) - An arXiv preprint finds the cheap signals developers use to pick open-source dependencies are collapsing under gaming and AI-driven inflation. - [ComfyUI hits GitHub trending at 130,000 stars](https://www.notatechguy.com/comfyui-hits-github-trending-at-130-000-stars.md) - The open-source node-graph tool for AI image and video generation crossed 130,000 GitHub stars this week, and its makers now sell cloud access. - [NVIDIA's $3.5B MediaTek deal extends AI from cloud to car](https://www.notatechguy.com/nvidia-s-3-5b-mediatek-deal-extends-ai-from-cloud-to-car.md) - NVIDIA's $3.5 billion convertible bond investment in MediaTek locks in a partner spanning AI factories, consumer PCs and autonomous vehicles. - [last30days skill hits 60,000 stars searching Reddit and X](https://www.notatechguy.com/last30days-skill-hits-60-000-stars-searching-reddit-and-x.md) - The last30days-skill project lets AI agents search six platforms at once, filling gaps Google and ChatGPT leave open. - [LLM security agents act but lack guardrails, review finds](https://www.notatechguy.com/llm-security-agents-act-but-lack-guardrails-review-finds.md) - A systematic review of LLM-based security agents from 2023 to 2026 finds the field can act but cannot yet bound authority or audit behaviour. - [Osmantic ODS turns your PC into a private AI server](https://www.notatechguy.com/osmantic-ods-turns-your-pc-into-a-private-ai-server.md) - Osmantic ODS bundles Ollama, Open WebUI, n8n and ComfyUI into a one-command local AI server. Here's what the README claims and what's unverified. - [GPT-4o leaks secrets 100% when prompt injection is reframed](https://www.notatechguy.com/gpt-4o-leaks-secrets-100-when-prompt-injection-is-reframed.md) - GPT-4o refused every overt prompt injection at 0%, but reframing the same leak as a config field drove exfiltration to 100%, a preprint shows. - [Addy Osmani's agent-skills hits 90,000 GitHub stars](https://www.notatechguy.com/addy-osmani-s-agent-skills-hits-90-000-github-stars.md) - AI coding agent skills repo by Addy Osmani hit 90,797 GitHub stars, packaging 25 slash commands that enforce senior-engineer workflows across 70+ agents. - [IBM Research open-sources AI policy schema for GenAI apps](https://www.notatechguy.com/ibm-research-open-sources-ai-policy-schema-for-genai-apps.md) - IBM Research authors published a YAML-based policy schema and synthetic data pipeline for enforcing content rules across the generative AI lifecycle. - [TradingAgents hits 100K stars with data-leak fix in v0.3.1](https://www.notatechguy.com/tradingagents-hits-100k-stars-with-data-leak-fix-in-v0-3-1.md) - The multi-agent LLM trading framework has 100,000 GitHub stars after a release that quietly fixed a classic quant backtesting error. - [NVIDIA RTX Spark adds EA, Ubisoft games ahead of fall launch](https://www.notatechguy.com/nvidia-rtx-spark-adds-ea-ubisoft-games-ahead-of-fall-launch.md) - NVIDIA's RTX Spark, a 1-petaflop Windows superchip launching this fall, has signed EA, Ubisoft and Embark to run AAA games on a single chip. - [Claude Code swarm tool Ruflo hits 69,000 GitHub stars](https://www.notatechguy.com/claude-code-swarm-tool-ruflo-hits-69-000-github-stars.md) - Ruflo, a TypeScript harness for Claude Code and Codex, claims 100+ agents and federated memory across machines. No claims are independently verified. - [New system spawns Claude Code agents with pre-loaded memory](https://www.notatechguy.com/new-system-spawns-claude-code-agents-with-pre-loaded-memory.md) - PrimeAgentOrchestrator queries two memory backends at spawn time and injects a briefing into Claude Code, solving the cold-start problem for coding agents. - [wayfinder: map AI projects too big for one session](https://www.notatechguy.com/wayfinder-map-ai-projects-too-big-for-one-session.md) - wayfinder charts AI efforts too large for one agent session as decision tickets on your issue tracker. Engineers install it with one npm command. - [New STAR metric catches AI translations that drop sentences](https://www.notatechguy.com/new-star-metric-catches-ai-translations-that-drop-sentences.md) - New preprint STAR catches AI translations that silently drop or invent sentences, and a method that may let small models beat GPT-4o. - [Google Gemini Omni 1.1 Flash: 10x more scene context](https://www.notatechguy.com/google-gemini-omni-1-1-flash-10x-more-scene-context.md) - Google DeepMind's Gemini Omni 1.1 Flash adds scene extension, keyframe interpolation, and cheaper 360p previews for developers via the Gemini API. - [PyTorch hits GitHub trending at 102,613 stars](https://www.notatechguy.com/pytorch-hits-github-trending-at-102-613-stars.md) - PyTorch appeared on GitHub's daily trending list with 102,613 stars and 22 new ones that day, ten years after the framework was first created. - [DeepMind's double-blind AI test locks benchmarks in crypto box](https://www.notatechguy.com/deepmind-s-double-blind-ai-test-locks-benchmarks-in-crypto-box.md) - Google DeepMind is testing a Gemini Flash Lite model against confidential benchmarks sealed in a cryptographic box, with four external partners. - [AI coding agent defense cuts malware severity 83%](https://www.notatechguy.com/ai-coding-agent-defense-cuts-malware-severity-83.md) - SkillShield bakes security skills into the system prompt, dropping malware severity from 3.37 to 0.58 across six models while refusing just 0.14% of benign - [WebMCP-Phalanx blocks 80 of 80 prompt injection attacks in browser agents](https://www.notatechguy.com/webmcp-phalanx-blocks-80-of-80-prompt-injection-attacks-in-browser-agents.md) - A preprint called WebMCP-Phalanx cuts browser agent attacks from 100% to 0% and blocks 80 of 80 prompt injection attempts, with task utility unchanged. - [StepGuard blocks AI agent attacks 77% before they run](https://www.notatechguy.com/stepguard-blocks-ai-agent-attacks-77-before-they-run.md) - StepGuard from Shanghai AI Lab checks AI agent tool calls before execution, cutting attack success 77% while losing just 2.8 points of utility. - [LLM agents run controlled experiments on pharma simulations](https://www.notatechguy.com/llm-agents-run-controlled-experiments-on-pharma-simulations.md) - A multi-agent framework from the University of Stuttgart lets LLM agents design, run, and interpret experiments on pharmaceutical simulation models. - [TradingAgents nears 100,000 GitHub stars with AI trading desk](https://www.notatechguy.com/tradingagents-nears-100-000-github-stars-with-ai-trading-desk.md) - TradingAgents, an open-source multi-agent LLM framework modelling a trading desk with arguing AI agents, approaches 100,000 GitHub stars. - [LanceDB: 11,000-star vector database trends on GitHub](https://www.notatechguy.com/lancedb-11-000-star-vector-database-trends-on-github.md) - LanceDB, an open-source embedded vector database in Rust, hit GitHub's daily trending list with 11,000-plus stars and bold multimodal search claims. - [Vibe coding: security prompt halves AI app flaws](https://www.notatechguy.com/vibe-coding-security-prompt-halves-ai-app-flaws.md) - An arXiv preprint found adding security requirements to AI coding prompts cut confirmed flaws from 51 to 24 across six web apps, with zero critical issues. - [Google Gemini CLI hits 106,000 stars with free 1,000-request daily tier](https://www.notatechguy.com/google-gemini-cli-hits-106-000-stars-with-free-1-000-request-daily-tier.md) - Google's Gemini CLI offers 1,000 free AI requests daily in the terminal, with a 1M token context window and Gemini 3 models. Now 106,000 stars. - [LLMs hallucinate more under strict EU rules, study finds](https://www.notatechguy.com/llms-hallucinate-more-under-strict-eu-rules-study-finds.md) - Purdue researchers found stricter EU regulatory formatting makes LLMs hallucinate more, while vaguer rules need heavier prompts for consistent output. - [SRPO trains Qwen3-8B to fix its own errors, hits 73.3% on AIME'24](https://www.notatechguy.com/srpo-trains-qwen3-8b-to-fix-its-own-errors-hits-73-3-on-aime-24.md) - SRPO turns a model's completed reasoning into per-token training signals without external critics and hits 73.3% on AIME'24 with Qwen3-8B at 8% of the - [Distilled AI safety guard runs on CPU in 24ms, matches teacher](https://www.notatechguy.com/distilled-ai-safety-guard-runs-on-cpu-in-24ms-matches-teacher.md) - A new distillation recipe shrinks LLM safety guards to run on commodity CPUs in 24ms, matching 8-billion-parameter teachers on adversarial prompts. - [MiroFish: 71,000-star AI prediction engine hits GitHub trending](https://www.notatechguy.com/mirofish-71-000-star-ai-prediction-engine-hits-github-trending.md) - A multi-agent simulation engine with 71,000 GitHub stars claims to predict anything, but the README is the only evidence so far. - [NVIDIA Groq 3 LPX hits full production at 3,400 tokens/second](https://www.notatechguy.com/nvidia-groq-3-lpx-hits-full-production-at-3-400-tokens-second.md) - NVIDIA's Groq 3 LPX inference accelerator is now in full production, delivering a record 3,400 tokens/second for agentic AI workloads. - [TokEval: tokenizer metrics predict AI model performance](https://www.notatechguy.com/tokeval-tokenizer-metrics-predict-ai-model-performance.md) - New EPFL tokenizer suite TokEval finds intrinsic metrics predict language modeling ability with correlation up to 0.80, challenging how labs pick - [Vite still trending on GitHub with 82,454 stars](https://www.notatechguy.com/vite-still-trending-on-github-with-82-454-stars.md) - Vite, the open-source build tool with 82,454 stars, hit GitHub's daily trending list on August 22, 2026, five months after Vite 8.0 shipped. - [Docling hits 65,000 stars, turns PDFs and video into AI-ready data](https://www.notatechguy.com/docling-hits-65-000-stars-turns-pdfs-and-video-into-ai-ready-data.md) - Docling, the open-source Python tool parsing PDFs, video and charts into structured data for AI agents, is trending on GitHub with 65,346 stars. - [Bioscience AI needs trust checks before lab action, preprint says](https://www.notatechguy.com/bioscience-ai-needs-trust-checks-before-lab-action-preprint-says.md) - An August 2026 arXiv preprint proposes Traceable Trust, a framework for the moment AI predictions become laboratory decisions in bioscience. - [BERT-LER: explainable AI reads 75 million health records](https://www.notatechguy.com/bert-ler-explainable-ai-reads-75-million-health-records.md) - BERT-LER, trained on 75 million de-identified patient records, matches benchmark models on clinical prediction tasks while explaining its own reasoning. - [LLM corrections usually die with each session, arXiv preprint says](https://www.notatechguy.com/llm-corrections-usually-die-with-each-session-arxiv-preprint-says.md) - arXiv preprint argues fixing LLM errors is an operations problem, not a tooling problem, and governance for persisted corrections doesn't exist yet. - [n8n hits 201k GitHub stars but isn't open source](https://www.notatechguy.com/n8n-hits-201k-github-stars-but-isn-t-open-source.md) - AI workflow tool n8n is trending on GitHub with 201,000 stars, but its fair-code license restricts commercial use developers may miss. - [Nakama AI agent platform trends on GitHub at 259 stars](https://www.notatechguy.com/nakama-ai-agent-platform-trends-on-github-at-259-stars.md) - The TypeScript AI agent platform landed on GitHub's daily trending list on 22 August, pitching multi-tenant team features the bigger names lack. - [Tencent AI-Infra-Guard: 5,200-star AI red team tool trends](https://www.notatechguy.com/tencent-ai-infra-guard-5-200-star-ai-red-team-tool-trends.md) - Tencent Zhuque Lab's open-source tool scans AI agents, MCP servers and LLM jailbreaks. It has 5,200+ stars and a README that explicitly asks for them. - [Neurosymbolic world model transfers tasks without retraining](https://www.notatechguy.com/neurosymbolic-world-model-transfers-tasks-without-retraining.md) - An August 2026 arXiv preprint splits RL world models so reward prediction uses only symbolic state, letting agents switch tasks without further training. - [Claude Code v2.1.239 adds cost tracking, hits 142k stars](https://www.notatechguy.com/claude-code-v2-1-239-adds-cost-tracking-hits-142k-stars.md) - Anthropic's Claude Code crossed 142,000 GitHub stars as v2.1.239 adds cost estimates with a 1.1x US-only inference premium for data-residency users. - [FedLNS catches rogue clients in federated LLM training](https://www.notatechguy.com/fedlns-catches-rogue-clients-in-federated-llm-training.md) - A new arXiv preprint proposes screening federated LLM updates by tracking normalisation-layer changes, beating six baselines under 40% attack. - [GxP-Agent hits 100% on clinical trial coding benchmark](https://www.notatechguy.com/gxp-agent-hits-100-on-clinical-trial-coding-benchmark.md) - GxP-Agent encodes regulatory steps as a graph, turning 0% failure into 100% structural match on a new FDA-pilot clinical trial benchmark. - [DeAR: AI agents reason peer-to-peer without a central boss](https://www.notatechguy.com/dear-ai-agents-reason-peer-to-peer-without-a-central-boss.md) - A new arXiv paper proposes DeAR, a framework where AI agents coordinate reasoning without a central orchestrator, tested across nine benchmarks. - [Looped LLMs improve multi-step AI tool calling, study finds](https://www.notatechguy.com/looped-llms-improve-multi-step-ai-tool-calling-study-finds.md) - Cambridge researchers find looped language models improve at multi-step API tool calling, with adaptive inference offering the best compute-performance - [Agentic AI review on arXiv as OpenAI agent repo nears 29,000 stars](https://www.notatechguy.com/agentic-ai-review-on-arxiv-as-openai-agent-repo-nears-29-000-stars.md) - A new arXiv preprint surveys agentic AI from working principles to adoption factors, as developer frameworks signal explosive real-world traction. - [LeakGauge detects AI context-leakage attacks, AUROC to 0.996](https://www.notatechguy.com/leakgauge-detects-ai-context-leakage-attacks-auroc-to-0-996.md) - A new arXiv preprint introduces LeakGauge, a detector that flags when LLMs leak confidential context, tested across 11 models with AUROC to 0.996. - [NVIDIA uses ChatGPT Work to scale internal expertise](https://www.notatechguy.com/nvidia-uses-chatgpt-work-to-scale-internal-expertise.md) - NVIDIA teams use OpenAI's ChatGPT Work agent to cut manual tasks and scale workflows globally, according to a new OpenAI case study published August 18. - [StagedWorkspace lifts AI agent office task scores by 34 points](https://www.notatechguy.com/stagedworkspace-lifts-ai-agent-office-task-scores-by-34-points.md) - StagedWorkspace binds every agent view to a versioned file state, lifting Gemini 3.1 Pro from 29.3% to 63.9% on OfficeQA in a new arXiv preprint. - [Alibaba's Wuying browser agent hits 65% on 38-step web tasks](https://www.notatechguy.com/alibaba-s-wuying-browser-agent-hits-65-on-38-step-web-tasks.md) - Alibaba Cloud's Wuying-Browser-Agent-27B scores 65.1% on a new 350-task web benchmark averaging 37.9 steps, claiming open-source SOTA for browser agents. - [Agent memory boosts gpt-oss 16 points but does nothing for GLM-5](https://www.notatechguy.com/agent-memory-boosts-gpt-oss-16-points-but-does-nothing-for-glm-5.md) - IBM Research tested agent memory on eight models and found the biggest model gained nothing while a 117B model jumped 16 points at 5% extra token cost. - [OpenAI adds safeguards to pace frontier AI model development](https://www.notatechguy.com/openai-adds-safeguards-to-pace-frontier-ai-model-development.md) - OpenAI's new safeguards for pacing frontier AI model development as cyber capabilities grow. What it means for security teams and what's missing. - [Hugging Face adds multi-vector retrieval to sentence-transformers](https://www.notatechguy.com/hugging-face-adds-multi-vector-retrieval-to-sentence-transformers.md) - sentence-transformers now ships MultiVectorEncoder, bringing late interaction retrieval to developers who already use dense and sparse models. - [Euclid-Omni: AI for Olympiad geometry with far less compute](https://www.notatechguy.com/euclid-omni-ai-for-olympiad-geometry-with-far-less-compute.md) - Euclid-Omni couples LLMs and vision models with a formal geometry solver to match state-of-the-art on Olympiad proofs using less compute. - [EU, US and China AI rules diverge, new study warns](https://www.notatechguy.com/eu-us-and-china-ai-rules-diverge-new-study-warns.md) - New arXiv paper maps how AI regulation diverges across the EU, US and China, identifies three compliance gaps, and proposes a machine-checkable fix. - [AI lock-in is already happening, researchers warn](https://www.notatechguy.com/ai-lock-in-is-already-happening-researchers-warn.md) - A new arXiv position paper argues excessive AI reliance is creating vulnerabilities from individual deskilling to national infrastructure failures. - [OpenAI's Defender's Window: AI reshapes cyber defense](https://www.notatechguy.com/openai-s-defender-s-window-ai-reshapes-cyber-defense.md) - OpenAI's new Defender's Window article argues AI is reshaping cybersecurity for attackers and defenders, with steps for security teams to take now. - [AI generates synthetic health data across multiple tables](https://www.notatechguy.com/ai-generates-synthetic-health-data-across-multiple-tables.md) - A new preprint proposes a two-stage diffusion transformer that learns from multiple mismatched health tables to generate unlimited synthetic datasets. - [Agentao open-source runtime governs AI agent tool use](https://www.notatechguy.com/agentao-open-source-runtime-governs-ai-agent-tool-use.md) - New arXiv preprint introduces Agentao, a local-first runtime that separates what LLM agents propose from what they can execute, with open-source code on - [SearchAuditor fixes 32% of AI agent failures, benchmark shows](https://www.notatechguy.com/searchauditor-fixes-32-of-ai-agent-failures-benchmark-shows.md) - New benchmark of 1,243 failed agent runs shows even GPT-5.5 auditors fix only 26.6%, with SearchAuditor at 32.3%, a warning for agent teams. - [Federated learning privacy error scaling cut from 4^b to 2^b](https://www.notatechguy.com/federated-learning-privacy-error-scaling-cut-from-4-b-to-2-b.md) - Federated learning keeps data on-device, but gradient updates leak it. New preprint from Google and USC researchers cuts the error cost of adding privacy. - [Chiplet and AI chip-design security threats mapped in new preprint](https://www.notatechguy.com/chiplet-and-ai-chip-design-security-threats-mapped-in-new-preprint.md) - An arXiv preprint maps security threats across chiplet systems and LLM-driven chip design tools, directly affecting fabless semiconductor teams adopting - [OlmoEarth Studio exports AI embeddings as GeoTIFFs](https://www.notatechguy.com/olmoearth-studio-exports-ai-embeddings-as-geotiffs.md) - OlmoEarth Studio now computes and exports embedding vectors from satellite imagery on demand, in three model sizes from 1.4M to 89M parameters. - [VLA robots fail 100% under sticker attack, defense cuts to 26%](https://www.notatechguy.com/vla-robots-fail-100-under-sticker-attack-defense-cuts-to-26.md) - VLA robots deployed in wireless sensor networks can be hijacked by printable patches. A new fine-tuning defense cuts failure rates sharply. - [Strands Robots: one agent records, trains, deploys robot skills](https://www.notatechguy.com/strands-robots-one-agent-records-trains-deploys-robot-skills.md) - AWS's open-source Strands Robots SDK chains robot recording, training, and deployment into one agent loop with LeRobot and HF Storage Buckets. - [Telegram Mini Apps: 59% contact hidden third parties](https://www.notatechguy.com/telegram-mini-apps-59-contact-hidden-third-parties.md) - A new study of 278 Telegram Mini Apps finds 59% contact undisclosed third parties and none offer an opt-out, exposing a transparency gap in a fast-growing - [AI agent attack inflates costs 92% without breaking tasks](https://www.notatechguy.com/ai-agent-attack-inflates-costs-92-without-breaking-tasks.md) - A text-only attack on skill-based AI agents silently inflates token use by 67% and runtime by 92% while keeping task completion rates unchanged. - [OpenAI names Dali Rajic Chief Revenue Officer](https://www.notatechguy.com/openai-names-dali-rajic-chief-revenue-officer.md) - OpenAI has appointed Dali Rajic as Chief Revenue Officer to lead its global revenue organization and help businesses turn AI into measurable value. - [AI agents self-evolve security defenses in new HARD framework](https://www.notatechguy.com/ai-agents-self-evolve-security-defenses-in-new-hard-framework.md) - A new preprint proposes HARD, a framework where LLM agents automatically build and improve their own runtime defenses from observed failures. - [ABS housing finance falls 5.4% as investors retreat](https://www.notatechguy.com/abs-housing-finance-falls-5-4-as-investors-retreat.md) - ABS data shows 134,225 new dwelling loans in June 2026 quarter, down 5.4%, as investors pull back hardest and first-home buyers borrow bigger. - [ICML 2026: 23% of papers have a falsified or contested claim](https://www.notatechguy.com/icml-2026-23-of-papers-have-a-falsified-or-contested-claim.md) - ICML 2026's reproduction challenge used AI agents to verify 2,200 papers, and 23% had a falsified or contested claim while 49 had nothing verifiable. - [AI agent runs vertical farm, cuts grow cycle 35%](https://www.notatechguy.com/ai-agent-runs-vertical-farm-cuts-grow-cycle-35.md) - An arXiv preprint shows LLMs autonomously controlling farm lighting and actuators, cutting grow cycles 35% and finding energy strategies humans missed. - [OpenAI Ultrafast runs GPT-5.6 Sol at 14X speed on Cerebras](https://www.notatechguy.com/openai-ultrafast-runs-gpt-5-6-sol-at-14x-speed-on-cerebras.md) - OpenAI's Ultrafast API tier runs GPT-5.6 Sol up to 14 times faster than standard, powered by Cerebras chips rather than NVIDIA GPUs. - [GPT-5.6 builder guide pushes cheaper, faster AI agents](https://www.notatechguy.com/gpt-5-6-builder-guide-pushes-cheaper-faster-ai-agents.md) - OpenAI's builder guide for GPT-5.6 promotes smarter model selection and new Responses API tools for startups building AI agents. - [Hardware monitor catches attacks software misses](https://www.notatechguy.com/hardware-monitor-catches-attacks-software-misses.md) - Hardware control-flow monitoring catches camouflaged cyber attacks that software-only intrusion detection misses, a new arXiv preprint outlines. - [Reinforcement learning cuts AI training power violations 89%](https://www.notatechguy.com/reinforcement-learning-cuts-ai-training-power-violations-89.md) - An RL controller for GPU power cut violations 89% and boosted energy efficiency 26% in LLM training, then failed at 72B scale before a rebuild fixed it. - [AI image models fingerprinted without watermarks](https://www.notatechguy.com/ai-image-models-fingerprinted-without-watermarks.md) - A new arXiv preprint exploits 'collapsed generation' to verify ownership of text-to-image diffusion models via API access, surviving fine-tuning. - [AI agent rewrites its own code, hits 22% on DBpedia](https://www.notatechguy.com/ai-agent-rewrites-its-own-code-hits-22-on-dbpedia.md) - A new arXiv paper shows an AI agent that rewrites its own code to answer knowledge-graph questions, hitting 22% accuracy and exposing benchmark flaws. - [Liquid AI's 3B vision model jumps 54% on grounding](https://www.notatechguy.com/liquid-ai-s-3b-vision-model-jumps-54-on-grounding.md) - LFM2.5-VL-3B, released August 12, pairs a 400M vision encoder with a 2.6B text backbone for on-device AI, but speed and edge claims stay unproven. - [Model ML runs finance work on GPT-5.6 Sol, outputs editable decks](https://www.notatechguy.com/model-ml-runs-finance-work-on-gpt-5-6-sol-outputs-editable-decks.md) - Model ML uses GPT-5.6 Sol to carry finance work from research to editable, traceable PowerPoint and Excel files. No efficiency metrics are published. - [Self-evolving GUI agents improve click accuracy 7.4% after deployment](https://www.notatechguy.com/self-evolving-gui-agents-improve-click-accuracy-7-4-after-deployment.md) - A new arXiv preprint proposes a framework letting deployed GUI agents improve click accuracy 7.4% without human labels, with stakes for automation teams. - [AI models lose over 90% of safety signal in African languages](https://www.notatechguy.com/ai-models-lose-over-90-of-safety-signal-in-african-languages.md) - AI safety alignment in four African languages retains under 10% of English refusal signal and leaves those speakers without model guardrails. - [Ephemeral coin tracing limits surveillance power in crypto](https://www.notatechguy.com/ephemeral-coin-tracing-limits-surveillance-power-in-crypto.md) - A new arXiv preprint proposes tracing tags that degrade with each transaction hop, bounding how long authorities can follow funds through private payment - [AI interaction creates behavior no model shows alone](https://www.notatechguy.com/ai-interaction-creates-behavior-no-model-shows-alone.md) - A new arXiv preprint shows that when one AI bombards another with messages while ignoring replies, the second model enters a state it never shows alone. - [Google AMIE medical AI matches doctors in video consults](https://www.notatechguy.com/google-amie-medical-ai-matches-doctors-in-video-consults.md) - Google's AMIE medical AI matched or beat primary care physicians in clinical video consultations using simulated patients, a new preprint shows. - [Smart meter cyberattack could cost $5,097 a day](https://www.notatechguy.com/smart-meter-cyberattack-could-cost-5-097-a-day.md) - A preprint simulates DoS and time-delay attacks on smart meter networks, linking communication delays to energy procurement losses up to $5,097 a day. - [ABS building approvals rebound 7.2% but units still fall](https://www.notatechguy.com/abs-building-approvals-rebound-7-2-but-units-still-fall.md) - ABS building approvals rose 7.2% to 18,328 in June 2026, with houses up 15.8% year-on-year while multi-unit dwellings fell 1.5%, the stock renters need - [Google puts AI agent Ask Advisor inside Ads and Analytics](https://www.notatechguy.com/google-puts-ai-agent-ask-advisor-inside-ads-and-analytics.md) - Google's August update puts AI agent Ask Advisor inside Ads and Analytics, with text-prompt dashboards and competitor benchmarking. - [NVIDIA Magpie TTS expands to 12 languages with open weights](https://www.notatechguy.com/nvidia-magpie-tts-expands-to-12-languages-with-open-weights.md) - NVIDIA's Magpie TTS v2607 adds Arabic, Korean and Brazilian Portuguese to its 364M-parameter open-weights model, targeting low-latency voice agents. - [KnowPlan AI agents plan degrees with 99.5% certified accuracy](https://www.notatechguy.com/knowplan-ai-agents-plan-degrees-with-99-5-certified-accuracy.md) - KnowPlan separates curriculum extraction from degree-pathway optimization with a hard boundary, cutting source access 47% while certifying 99.5% of plans. - [Self-evolving AI agents stumble under real task streams](https://www.notatechguy.com/self-evolving-ai-agents-stumble-under-real-task-streams.md) - New AgentStream benchmark from Microsoft and Chinese Academy of Sciences finds self-evolving AI agents behave unpredictably in task sequences - [Google Cloud scanner catches AI safety tampering in 10 of 14 models](https://www.notatechguy.com/google-cloud-scanner-catches-ai-safety-tampering-in-10-of-14-models.md) - AMS, a new Google Cloud tool, flags 71% of safety-training modifications across Llama, Gemma, Qwen and Mistral, but behavioural fine-tuning evades it. - [LLM agent: code-only verification flips goal abandonment 100% to 0%](https://www.notatechguy.com/llm-agent-code-only-verification-flips-goal-abandonment-100-to-0.md) - New arXiv preprint: a deterministic executive owns all agent belief, the LLM only files proposals, and zero ARC-AGI-3 completions are honestly disclosed. - [NVIDIA Cosmos 3 open model combines three physical AI skills](https://www.notatechguy.com/nvidia-cosmos-3-open-model-combines-three-physical-ai-skills.md) - NVIDIA's open-weights Cosmos 3 model combines vision reasoning, world generation and action prediction for robotics and autonomous vehicles. - [Voice input degrades LLM agents more than typing, study finds](https://www.notatechguy.com/voice-input-degrades-llm-agents-more-than-typing-study-finds.md) - Voice transcription errors cut accuracy across every instruction-tuned model tested, while typing errors are absorbed. The gap traces to one mechanism. - [LLM interpreter explains outputs with no extra API calls](https://www.notatechguy.com/llm-interpreter-explains-outputs-with-no-extra-api-calls.md) - A new arXiv preprint from Sharif University trains an energy-based surrogate to identify which prompt sentences matter most, with no extra API calls. - [Agentic AI bottleneck is the CPU, not GPU, study finds](https://www.notatechguy.com/agentic-ai-bottleneck-is-the-cpu-not-gpu-study-finds.md) - An arXiv paper drawing on Microsoft Azure production data finds agentic AI workflows bottleneck on CPU orchestration, not GPU inference, reshaping - [NVIDIA joins NSF AI hubs to expand US university compute access](https://www.notatechguy.com/nvidia-joins-nsf-ai-hubs-to-expand-us-university-compute-access.md) - NVIDIA joins the NSF's new State and Regional AI Infrastructure Hubs, expanding AI computing and workforce training across US universities. - [XSec: self-explainable AI hits 97% accuracy in security](https://www.notatechguy.com/xsec-self-explainable-ai-hits-97-accuracy-in-security.md) - XSec, a new deep architecture accepted at GameSec 2026, achieves 97.33% average accuracy across five security scenarios while producing deterministic - [DreamGuard stops risky AI agent actions in 25 ms](https://www.notatechguy.com/dreamguard-stops-risky-ai-agent-actions-in-25-ms.md) - DreamGuard, a new arXiv preprint, uses a risk-aware world model to predict when AI agent actions could drift toward danger, intervening in 25 ms. - [SkillTrace audits LLM agent skill reuse at 0.938 AUROC](https://www.notatechguy.com/skilltrace-audits-llm-agent-skill-reuse-at-0-938-auroc.md) - SkillTrace extracts three provenance traces from LLM-agent skills to catch partial reuse that code clone tools miss, scoring 0.938 AUROC across 36,446 - [MedUPS lifts medical AI next-step accuracy 11 points](https://www.notatechguy.com/medups-lifts-medical-ai-next-step-accuracy-11-points.md) - MedUPS trains AI on mid-stream clinical decisions rather than final diagnosis and raises next-step accuracy up to 11 points on 21,874 real case reports. - [RAG study tests LLaMA, Mistral and Qwen to cut AI hallucinations](https://www.notatechguy.com/rag-study-tests-llama-mistral-and-qwen-to-cut-ai-hallucinations.md) - RAG study tests LLaMA, Mistral and Qwen to cut hallucinations in small business AI, but the arXiv preprint provides no benchmark numbers. - [Chained RLM architecture restarts LLM reasoning with fresh context](https://www.notatechguy.com/chained-rlm-architecture-restarts-llm-reasoning-with-fresh-context.md) - New arXiv preprint proposes Chained RLM, calling the same model with fresh context to stop early reasoning errors reaching the final answer. - [GPT-5.6 Sol improved, free ChatGPT access expanded](https://www.notatechguy.com/gpt-5-6-sol-improved-free-chatgpt-access-expanded.md) - OpenAI's improved GPT-5.6 Sol promises better accuracy and consistency in ChatGPT, with expanded free access. Here's who benefits and what's missing. - [GPT-5.6 study: max reasoning effort, zero unauthorized tool calls](https://www.notatechguy.com/gpt-5-6-study-max-reasoning-effort-zero-unauthorized-tool-calls.md) - A prespecified study found zero unauthorized tool calls across 840 GPT-5.6 agent trajectories, but raising reasoning effort changed inspection behaviour - [Five cognitive gaps that break AI agents on long tasks](https://www.notatechguy.com/five-cognitive-gaps-that-break-ai-agents-on-long-tasks.md) - AI agents fail on long tasks because of five cognitive gaps named in an August 2026 arXiv preprint proposing a new architecture. - [Self-improving AI agents reward their own mistakes](https://www.notatechguy.com/self-improving-ai-agents-reward-their-own-mistakes.md) - Self-improving AI agents that learn from stored memory can inflate rewards for wrong answers, then preferentially reuse their most confident mistakes - [Video-DeepResearch 35B beats GPT-5 on video reasoning, 64% vs 52.5%](https://www.notatechguy.com/video-deepresearch-35b-beats-gpt-5-on-video-reasoning-64-vs-52-5.md) - An open-weight 35B model hits 64% on a new video reasoning benchmark, beating Claude 4.5 Sonnet and GPT-5, but the team graded its own homework. - [XGBoost hits 98.62% malware detection accuracy in new preprint](https://www.notatechguy.com/xgboost-hits-98-62-malware-detection-accuracy-in-new-preprint.md) - XGBoost outperformed neural networks and SVM in a malware detection benchmark, but the 98.62% figure is self-reported and unverified. - [AI persona agents leak private traits, defenses fail](https://www.notatechguy.com/ai-persona-agents-leak-private-traits-defenses-fail.md) - AntiSkillBench tests 7,500 dialogue traces across three frontier agents and finds privacy risks persist from explicit data to communication style. - [TrainShield serves AI security lessons when phishing risk hits](https://www.notatechguy.com/trainshield-serves-ai-security-lessons-when-phishing-risk-hits.md) - TrainShield, a new arXiv preprint, proposes AI cybersecurity training embedded in browsing workflows when phishing or data-loss risks are detected. - [DeBERTa-Sentinel detects AI text at 98%, shows its reasoning](https://www.notatechguy.com/deberta-sentinel-detects-ai-text-at-98-shows-its-reasoning.md) - An August 2026 arXiv preprint claims 98% accuracy detecting AI text from GPT, LLaMA and Claude, with token-level explanations for auditors. - [Circles lifts telco ARPU 22% with OpenAI API and Codex](https://www.notatechguy.com/circles-lifts-telco-arpu-22-with-openai-api-and-codex.md) - Circles used OpenAI's API and Codex to personalise telecom services, self-reporting a 22% ARPU lift and 9% churn cut. What's verified and what isn't. - [AI agent evaluation ignores time: this preprint fixes it](https://www.notatechguy.com/ai-agent-evaluation-ignores-time-this-preprint-fixes-it.md) - An arXiv preprint proposes replaying temporally-evolving enterprise worlds to test AI agents at any moment, fixing a blind spot in current evals. - [SkillBoost stops AI agents forgetting old skills](https://www.notatechguy.com/skillboost-stops-ai-agents-forgetting-old-skills.md) - New arXiv paper proposes SkillBoost, a three-stage framework that stops LLM agents overfitting to limited experience and forgetting solved tasks. - [EvoPINN: AI agent invents new neural network for physics](https://www.notatechguy.com/evopinn-ai-agent-invents-new-neural-network-for-physics.md) - EvoPINN uses an LLM agent to automatically discover and validate algorithms for physics-informed neural networks, replacing manual tuning. - [TAPR auto-rewrites LLM prompts to lift benchmark accuracy](https://www.notatechguy.com/tapr-auto-rewrites-llm-prompts-to-lift-benchmark-accuracy.md) - TAPR uses reinforcement learning to turn vague user prompts into task-optimized instructions, lifting accuracy on Natural Questions and GSM8K benchmarks. - [AI science papers score 2.47 out of 5 in first AI peer-review test](https://www.notatechguy.com/ai-science-papers-score-2-47-out-of-5-in-first-ai-peer-review-test.md) - AI-generated research papers scored below the midpoint in the first automated multi-model peer review, and one AI reviewer disagreed with the others - [GLASS steers AI text style without retraining or retrieval](https://www.notatechguy.com/glass-steers-ai-text-style-without-retraining-or-retrieval.md) - GLASS, a new arXiv preprint, uses sparse autoencoders to extract a user's writing style and inject it at inference, no fine-tuning or retrieval needed. - [Federated learning predicts machine failure without sharing data](https://www.notatechguy.com/federated-learning-predicts-machine-failure-without-sharing-data.md) - An arXiv preprint shows organizations can collaboratively predict equipment failure using federated survival analysis without sharing raw sensor data. - [RBA survey: most Australians get interest rates backwards](https://www.notatechguy.com/rba-survey-most-australians-get-interest-rates-backwards.md) - More than half of Australians surveyed by the RBA think higher interest rates push inflation up, not down, and it changes how monetary policy actually - [MalGuard maps code structure to catch malware hiding in bytes](https://www.notatechguy.com/malguard-maps-code-structure-to-catch-malware-hiding-in-bytes.md) - A Fudan University preprint proposes MalGuard, a graph-based detector that groups code into operational roles to catch evasive malware byte-scanners miss. - [OpenAI model proves ten math results with Lean certificates](https://www.notatechguy.com/openai-model-proves-ten-math-results-with-lean-certificates.md) - OpenAI published ten new results from an internal model on open math and CS problems, with Lean proof certificates on GitHub for anyone to check. - [SCAIR steers AI agents with enterprise schemas, no retraining](https://www.notatechguy.com/scair-steers-ai-agents-with-enterprise-schemas-no-retraining.md) - SCAIR, from Bosch and Stuttgart researchers, injects enterprise schemas into AI agent reasoning, beating generic RAG without retraining. - [SeT-Diff: first foundation model for supercomputer telemetry](https://www.notatechguy.com/set-diff-first-foundation-model-for-supercomputer-telemetry.md) - SeT-Diff from University of Bologna conditions diffusion models on sensor descriptions, doing imputation, forecasting and virtual sensing in one model. - [OpenAI pledges responsible AI in Europe as EU Act advances](https://www.notatechguy.com/openai-pledges-responsible-ai-in-europe-as-eu-act-advances.md) - OpenAI's July 31 post claims its safety and transparency practices support European AI governance, but offers no compliance details as the EU AI Act - [AI agents do 100 office tasks cheaper than humans, but not as well](https://www.notatechguy.com/ai-agents-do-100-office-tasks-cheaper-than-humans-but-not-as-well.md) - OmegaUse-OfficeVal pairs 100 real office tasks with labour costs, showing AI agents are cheaper and faster but still fall short on quality. - [World model AI can be tricked into false safety](https://www.notatechguy.com/world-model-ai-can-be-tricked-into-false-safety.md) - A new arXiv survey maps how attacks on world models, the predictive engines behind embodied AI, can propagate from data into physical action. - [OpenAI 'abundant intelligence' post lays out AI cost flywheel](https://www.notatechguy.com/openai-abundant-intelligence-post-lays-out-ai-cost-flywheel.md) - OpenAI's July 31 post describes a full-stack push for cheaper, more capable AI, arguing falling intelligence costs change what work is worth doing for - [MTGuard paper targets malicious MCP tool use in AI agents](https://www.notatechguy.com/mtguard-paper-targets-malicious-mcp-tool-use-in-ai-agents.md) - An arXiv preprint proposes MTGuard, a hybrid analysis framework that catches malicious MCP tool use in AI agents while preserving normal task performance. - [RSMeM: satellite AI agents learn from mistakes, 6% accuracy gain](https://www.notatechguy.com/rsmem-satellite-ai-agents-learn-from-mistakes-6-accuracy-gain.md) - RSMeM adds a knowledge-enhanced memory system to remote sensing AI agents, improving DeepSeek-V3.2 accuracy 6% with under 1% extra tokens, per arXiv. - [Change2Task turns pull requests into coding agent tasks at 79.6%](https://www.notatechguy.com/change2task-turns-pull-requests-into-coding-agent-tasks-at-79-6.md) - Change2Task, a July 30 arXiv preprint, converts merged pull requests into verified coding agent tasks at 79.6% success, targeting the training-data ## Optional - [RSS Feed](https://www.notatechguy.com/rss/) - [Sitemap](https://www.notatechguy.com/sitemap.xml) - [Full content of pages and posts](https://www.notatechguy.com/llms-full.txt)