A preprint posted to arXiv on July 20, 2026 lays out a methodology for harmonizing the capability thresholds that frontier AI companies publish, because those thresholds currently differ so much that no third party can verify whether one has been crossed [S1, S2]. The authors warn that without common minimum thresholds, safety standards could slide in a race to the bottom [S1, S2]. But the fix they propose splits into two fundamentally different approaches depending on the type of risk, and the third risk domain they cover is not named in the abstract. If the companies building these models cannot even agree on how to measure danger, how is anyone supposed to trust the safety claims?
The incompatible yardsticks
Frontier AI companies have each published their own capability thresholds: the points at which a model becomes dangerous enough to trigger extra safeguards, restrictions, or reporting [S1, S2]. The problem is not that these thresholds exist. It is that they differ substantially from one company to the next, making it nearly impossible for an outside party to check whether a threshold has actually been crossed or to compare what one company requires versus another [S1, S2].
Think of it like speed limits. If every car manufacturer set its own definition of "too fast," measured speed differently, and never published the number, you could not compare a Toyota's safety record to a Ford's. No traffic cop could write a ticket. That is roughly where AI safety thresholds sit today.
Google DeepMind published an update to its Frontier Safety Framework in February 2025, describing stronger security protocols on the path to AGI [P4]. OpenAI, separately, has outlined its own approach to AI safety through US state and federal action, in a post dated July 15, 2026 [P6]. Each company frames risk in its own terms, with its own metrics.
This is the gap the preprint tries to close.
Two risks, two methods
The authors develop a methodology for deriving harmonized thresholds across three risk domains [S1, S2]. They do not treat all risks the same way.
For misuse risks, specifically cyber and biological threats, they take "expected harm" as the key primitive. That means they model the actual damage a model could cause if misused, factoring in the channels through which harm could flow and the conditions under which a model is released [S1, S2]. A model available to anyone via an API carries different risk than one locked inside a lab.
For automated AI R&D, they switch approaches entirely. Instead of estimating expected harm, they base their proposed threshold on the observed rate of AI progress [S1, S2]. The logic: if models are improving fast enough to accelerate their own development, that speed itself is the warning signal, regardless of whether you can quantify the harm.
The third risk domain is mentioned but not named in the available text [S1, S2].
The authors note their analysis builds on prior work and points to existing empirical gaps and limitations [S1, S2]. The paper is a preprint and has not been peer-reviewed [S2].
What it means
The core problem this paper identifies is simple to state and hard to solve: when every AI lab sets its own safety bar, nobody outside the lab can check the measurement. A regulator, a researcher, or a business thinking about deploying a model has to take the company's word for it that the model has or has not crossed a danger line.
The proposed fix is a common methodology, not a binding standard or a regulation, but a shared way of deriving thresholds so that different companies' numbers are at least comparable [S1, S2]. If adopted, it would let an outside auditor look at two companies' thresholds and tell whether they are measuring the same thing in the same units.
The split between expected-harm modeling and progress-rate observation is the most interesting detail. It acknowledges that not all AI risks can be measured the same way. A biological misuse risk has a plausible harm scenario you can model. An AI system that accelerates its own research does not have a clean harm number attached; the danger is in the velocity of improvement itself.
What it means for business
For a two-person AI consultancy evaluating which frontier model to build on, the threshold mismatch is a practical problem, not an academic one. If Company A says its model is safe below threshold X and Company B says its model is safe below threshold Y, and X and Y are measured differently, the consultancy has no way to compare risk. They are choosing blind.
A suburban law firm using an AI tool to draft contracts cares about one thing: is the model they are using above or below the line where it could help someone do real harm? Today, that line is different depending on which vendor they ask.
For compliance teams, the absence of a common standard means every model deployment requires its own risk assessment framework. A company using models from three different providers cannot run one safety check. It has to run three, each calibrated to that provider's definitions.
The preprint does not solve this overnight. It is a methodology paper, not a regulation. But if harmonized thresholds were adopted, even voluntarily, a small business could ask a single question, "has this model crossed the harmonized threshold?", and get a comparable answer regardless of vendor.
What we don't know yet
The third risk domain. The abstract references three risk domains but names only two: misuse risks covering cyber and biological threats, and automated AI R&D [S1, S2]. The third is not specified in the available text.
Whether the methodology produces specific numbers. The preprint describes an approach for deriving harmonized thresholds, but the available evidence does not include specific numerical values for any proposed threshold [S1, S2].
Whether any frontier company will adopt it. The paper proposes a harmonization methodology, not binding rules [S1, S2]. Google DeepMind and OpenAI have each published their own frameworks [P4, P6], but neither has indicated they will align with an external standard.
Whether peer review will change the methodology. The paper is a preprint and has not undergone independent academic review [S2]. Its approach to modeling expected harm and measuring AI progress rates may be revised.
The next concrete event to watch: whether the full paper, once read in detail, names the third risk domain and provides the specific threshold values the methodology produces. That will determine whether this is a framework others can actually apply or a conceptual argument that stays on the page.
If you want to keep reading stories that decode what AI safety developments actually mean for the people building and deploying these systems, subscribe.
Sources
- [S1] Harmonizing AI Safety Thresholds — arXiv cs.AI new (official RSS) (attributed)
- [S2] Harmonizing AI Safety Thresholds — arXiv preprint (cs.AI, cs.LG) (attributed)
- [P3] [2603.14825v1] Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection — [2603.14825v1] Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection (attributed)
- [P4] Updating the Frontier Safety Framework — Google DeepMind — Updating the Frontier Safety Framework — Google DeepMind (primary)
- [P5] FrontierCS/Frontier-CS — FrontierCS/Frontier-CS (attributed)
- [P6] The US is advancing AI safety through state and federal action | OpenAI — The US is advancing AI safety through state and federal action | OpenAI (primary)
More from Not A Tech Guy
- Gasp paper proposes bridge-free cross-chain DEX rollup
- Prompt syntax shifts open LLM code security, preprint finds
- Byzantine fault tolerance breaks with AI agents, preprint warns
Generated from an audited evidence pack with primary-source research. Social-media items are discussion signals, not verified facts. Nothing here is financial, legal or medical advice.