A September 2 arXiv paper proposes giving clinical AI failures the same treatment hospitals give human errors: a blameless review S¹. Two clinicians tested its four-part classification on five cases and agreed on all 20 judgements S¹. Whether that perfect agreement survives real hospital errors, different institutions, and actual pressure is the question the paper itself leaves open.
My read: This is the first structured failure-review protocol I've seen that treats AI errors the way medicine already treats human errors: as system failures to study, not bugs to patch. The 20-for-20 inter-rater agreement is striking, but it rests on two reviewers and five illustrative cases, not confirmed adverse events. I don't think this is ready for deployment, and the authors don't either. They explicitly call for prospective evaluation. What I do think is real is the gap they've found. None of the existing model drift detection frameworks reconstruct the chain of decisions that leads to a specific patient harm event. That reconstruction is the piece nobody has built yet.
Why existing safety tools miss the point
Hospitals already have two layers of AI safety monitoring. Aggregate model monitoring tracks whether a model's overall performance is drifting, whether accuracy is dropping or false positives are climbing S¹. Traditional patient safety reporting captures adverse events after they happen S¹.
Neither, the paper argues, is built to explain how risk actually emerges when an AI system, a clinician, a clinical workflow, and institutional controls all interact S¹. A model's accuracy can hold steady at 94% while the 6% of errors concentrate in exactly the patients who can least afford them. A safety report can document that a patient received the wrong dose, but it won't trace whether the AI recommended it, whether the clinician overrode their own judgement to follow it, or whether the dosing tool was fed stale lab values.
The gap is causal reconstruction. The aim is to answer a different question: how did the chain of events, decisions, and system interactions produce this specific harm?
How the four-part framework works
Each AI-related event gets classified across four linked dimensions S¹:
Trigger: the condition that exposed the vulnerability. A data feed going down, a patient population shift, an edge-case input the model never saw in training.
Mechanism: the process that produced risk. How the AI system, the clinician's interaction with it, and the surrounding workflow combined to create danger.
Clinical Pathway: the consequence for patient care. What actually happened to the patient, or how close it came.
Corrective Action: the remediation assigned. What changes, who owns it, and whether it prevents the same chain from recurring.
The framework also includes standardized case intake, evidence preservation, investigator-level reconstruction, tool-in-loop attribution, and corrective-action tracking S¹. The design separates what exposed the vulnerability from what created the risk, what it did to the patient, and what fixes it S¹.
Five cases, two reviewers, perfect agreement
The authors demonstrated the framework on five illustrative outpatient medication and clinical decision-support cases S¹. Two clinician reviewers independently applied all four classification axes to each case, producing 20 classifications in total, and agreed on every one S¹.
That perfect agreement is the paper's strongest evidence. It is also its thinnest. Two reviewers, five cases, and illustrative scenarios rather than confirmed adverse events mean the inter-rater reliability claim is a proof of concept, not a validation S¹. The paper is an arXiv preprint and has not been peer-reviewed.
The broader ecosystem is moving in the same direction. A GitHub project called medical-ai-failure-atlas, created in June 2026, describes itself as a clinician-built benchmark and live leaderboard for medical AI safety evaluation P². Another repo, medical_reasoning, built in March 2026, is a formal reasoning system for detecting deviations from evidence-based clinical guidelines P³. These are early, small projects with one star and zero stars respectively, but they signal that clinicians are starting to build the tooling an AI M&M framework would need.
What to do about it
The framework is a proposal, not a protocol. No hospital has adopted it, and the authors are explicit that it complements rather than replaces existing safety systems S¹.
But the idea is actionable now for any clinical team running AI tools. Consider a primary care clinic where an AI decision-support tool recommends medication adjustments based on lab values. The tool suggests increasing a patient's statin dose, but the patient recently started a conflicting prescription that the AI's data feed hasn't picked up yet. The recommendation lands in the clinician's inbox. The clinician, trusting the system, nearly signs off.
Today, if that near-miss is caught, it might become a line in a safety report. Under an AI M&M review, the team would classify the trigger (stale medication reconciliation data), the mechanism (the model didn't cross-reference new prescriptions because the feed lagged), the clinical pathway (the recommendation reached the clinician and was nearly actioned), and the corrective action (adding a real-time drug-interaction check before AI recommendations surface). The review would be blameless. No one person failed. But the system gap would be named, tracked, and fixed.
If your hospital or clinic uses clinical AI tools, the practical step this week is to ask your safety or quality team whether any existing review process specifically examines AI-related events. If the answer is no, the AI M&M paper's four-axis classification is a free template to start from.
What we don't know yet
The framework has not been tested in a live clinical environment S¹. The five cases are illustrative, not confirmed adverse events S¹. The inter-rater agreement study involved two reviewers at what appears to be a single site S¹. The paper has not been peer-reviewed S¹. And no hospital, regulator, or health system has adopted the framework S¹.
The authors call for prospective evaluation across institutions, AI systems, and clinical settings S¹. That is the test that matters. A multi-site trial with real adverse events, more reviewers, and independent institutions would show whether the four-axis classification holds up when the cases are messy, the stakes are high, and the reviewers disagree.
The next signal: watch for a follow-up study or institutional pilot from the authors. If a hospital system announces an AI M&M trial within the next 12 months, this framework will have legs. If it stays on arXiv with no real-world adoption, it joins a long shelf of well-designed protocols that never reached a patient. We'll track both. Subscribe to catch the first real-world AI M&M trial when it lands.
Sources: S1 — AI Morbidity and Mortality: A Framework for Clinical AI Failure Review · P2 — goktugozkanmd/medical-ai-failure-atlas · P3 — mishra191/medical_reasoning · P4 — Corrective Retrieval Augmented Generation
More from Not A Tech Guy
- AWS Firecracker microVM tops 36,443 GitHub stars
- Synthetic data privacy is a claim, not a guarantee, researchers warn
- DiaSentinel AI agents screen diabetes risk on-premise
Generated from an audited evidence pack with primary-source research. Social-media items are discussion signals, not verified facts. Nothing here is financial, legal or medical advice.
