Matheo-b16 — What Changed Between MMv5 and OOv2#
This page exists so that the paper does not have to carry its own history in its margins. Everything here is provenance: what moved, why, and where the audit trail lives. None of it is needed to read the paper. It is here for a reader who wants to know whether the argument has been handled carefully — and for the reader who read an earlier version and wants to know what to un-remember.
A note on the register, because it matters. Most of what follows is refinement: the same findings, measured with a better instrument. A smaller part is correction: things that were wrong and are now right. One item is a removal: a claim that was never a result of the model and should not have been in the prose. These are different, and collapsing them into “withdrawn” would be its own kind of inaccuracy — it invites a reader to discount work that was, in the main, functioning exactly as intended. Conceive, test, catch the error, correct, strengthen: the corrections below are evidence that the cycle is running, not evidence that the study failed.
1. Refined — same finding, finer instrument#
These are the bulk of the changes, and none of them overturn anything.
What |
MMv5 (floor) |
OOv2 |
Why it moved |
|---|---|---|---|
Median time to onset |
~19 yr (40-run sample) |
21.1 yr (exact) |
The simulated median is a sample statistic; the closed form gives the population value. The two differ by 0.43 standard errors at \(n=40\) — ordinary noise. The simulation finding stands. |
Blow-ups within year 1 |
~1 of 40 runs |
1.16 of 40 (expected) |
Same: the exact rate predicts 1.16 per forty runs, so observing about one is exactly right. The simulation finding stands. |
Annual probability |
~2.5% (1 in 40, sampled) |
2.90% (1 in 34, exact) |
Forty runs cannot resolve a ~3% probability — the standard error is ~2.5%, nearly the estimate itself. The closed form can. The finer number reads slightly worse. |
Mean time to onset |
not reported separately |
30.3 yr (exact) |
Mean and median are not the same for this distribution (ratio 1.44); the paper now reports both, in separate columns, because conflating them is a common way to mislead. |
Waiting-time horizons |
1 / 10 / 40 / 85 yr |
1 / 10 / 20 / 40 / 80 yr |
Round horizons a policymaker can act within. The 85-year figure survives in §5.2 only, where it is needed to compare like-for-like with Hellman’s “life expectancy of a child born today”. |
Also refined: the sensitivity analysis (§2.5a) now reports exact values rather than 40-run medians; the death fraction \(q\) is anchored to Xia et al. (2022) rather than asserted as a placeholder; and the paper gained a stress test (§2.6) that pushes the crisis rate, the death fraction, and the road-death baseline against its own conclusion simultaneously — the inequality survives that too.
2. Corrected — things that were wrong#
Danger
The headline was computed with an approximation that was labelled exact, and it erred in the paper’s own favour.
MMv5 and the OOv1 draft computed the one-year probability as \(1 - \exp(-t/T_R)\) — an exponential fitted to the correct mean — while describing the result as “the exact first-passage solution”. The first-passage time of this model is hypoexponential, not exponential: it is the sum of two exponentials, because a crisis takes ~40 days to resolve before onset is possible at all, which suppresses the distribution near the origin.
The error was confined to the one-year column, where it overstated every scenario by ~12 percent. The base case was 3.24% (1 in 31) and is 2.90% (1 in 34).
No conclusion changed. What changed is that a paper whose only request is
#AuditTheMath no longer contains an approximation wearing the word “exact”. The
derivation is now in §2.4a and can be checked by hand. Found by the authors before
submission; reported in §2.0a at the point where it happened rather than repaired
quietly.
What |
Was |
Now |
|---|---|---|
Car-crash baseline |
US road deaths, \(1.22\times10^{-4}\)/yr |
Global, \(1.49\times10^{-4}\)/yr (WHO). The claim is about most people, so the baseline must be too. The US rate is lower than the global average, so the old choice inflated the paper’s own multiplier: 80x becomes 58x. |
Worst-country stress figure |
Libya ~73 per 100,000 |
~34 per 100,000 (WHO country data, 2021). The 73 is real but is the 2013 report’s figure, describing ~2010. Using a stale figure that WHO has since revised would not survive a reviewer checking it. |
The \(q\) bound |
“inequality survives for any \(q > 0.004\)” |
\(q > 0.017\). The old threshold was computed at the base case while the paper leans on the optimistic corner, where the requirement is ~4x higher. |
Tertrais (§5.3) |
Both “his rate is our optimistic corner” and “his scope excludes our subject” |
Only the first, plus the reason: his scope is narrower, so his count is an under-count — which is why it belongs as a corner rather than a centre. The two claims could not both hold, and the paper was relying on both. |
Theology (§9) |
“load-bearing, not decorative” and “§§2–4 are self-contained” |
Only the second. The framework is why the author looked; it is not why the number is what it is. Nothing in §§2–4 depends on it. |
“Nobody told them” (§4.0a, §10) |
The calculation is absent because no one ran it or said it |
The risk is widely known and has been for 81 years; the waiting-time distribution is what has not been seen. Running the two together was wrong, and plainly so. |
Uncited claims (§2.10) |
Russia’s doctrinal change; the 2026 Iran–US–Israel escalation |
Removed rather than sourced. Neither was load-bearing, and neither could be checked to the standard the rest of the paper is held to. |
§3.3 |
Deterrence professionals compared to “a tobacco executive” |
The incentive-alignment argument, kept and connected to the super wicked problem literature (Lazarus 2009; Levin et al. 2012), where “those who cause the problem also seek to provide the solution” is a named structural feature rather than an insult. |
3. Removed — one claim, and it was never a result#
The uniform “1 in 40”
Earlier prose stated that roughly 1 in 40 runs reached catastrophe within a year in each scenario — pessimistic, base, and optimistic alike. That is false, and it was never something the model produced. The expected number of blow-ups per forty runs is:
Scenario |
Expected per 40 runs |
P(onset ≤ 1 yr) |
|---|---|---|
optimistic (0.03/yr) |
0.35 |
0.88% (1 in 113) |
base (0.10/yr) |
1.16 |
2.90% (1 in 34) |
pessimistic (0.30/yr) |
3.34 |
8.34% (1 in 12) |
A spread of nearly ten. There is no invariant here and there never was.
What this was, most likely. The site’s own long-standing formulation is “worse than 1 in 40” — a floor, not a constant. As a floor it is correct at the base rate and in the pessimistic scenario, and it over-states the risk at the optimistic corner, where the exact value is 1 in 113. The scenario-independent version appears in the MMv3/OOv1 drafting, which the audit trail attributes to machine drafting: a hedged human statement was tidied into a cleaner claim that read better and asserted more. The correction here is to the tidying, not to the underlying work.
It is recorded at this length because it is the clearest example in this paper of a failure mode worth naming: a simplification that improves readability while manufacturing a certainty the evidence never had.
4. Added#
§2.0 BEST Names (Table 0) — every quantity in Brief / Explicit / Summarizing / Technical form, so the prose never needs a bare symbol and cross-paper identities (Hellman’s \(\lambda_{IE}\) vs \(\lambda_{CMTC}\)) are explicit. Uses the project’s own published naming convention.
§2.4a — the full first-passage derivation (Laplace transform → roots → CDF).
§2.6 — the three-axis stress test, and the uninsurability finding: nuclear war is excluded from essentially every insurance policy by standard clause, so the risk is not unpriced — it was priced, at infinity, decades ago, and then not spoken of.
§2.10 — New START expired 2026m02d05 with no successor and no negotiations, cited. The first hard, dated evidence the pessimistic corner has ever had.
§4.3 — the 10 Nuclear Kings defined as an abstraction for sovereign (ten thrones, nine armed); Canute; ResearchCity’s two structural functions.
§4.3a — the escape window priced: 7 / 9 / 11 years costs 20.4% / 25.5% / 30.3% at the base rate, against 73% over forty years of doing nothing.
§6 — rewritten as an uptake report: three audiences, the two convening offices, and the finding that there is no channel.
A bibliography (18 references), and a competing-interests statement.
5. Where the audit trail lives#
Every decision above is recorded verbatim, including the prompts that produced it and the arguments that were rejected:
File (in HELL) |
What it holds |
|---|---|
The adversarial panel review that produced this revision: the arithmetic finding, the EDEN classification, the per-reviewer analysis, every prompt verbatim, and the compliance failures recorded against the reviewer rather than hidden. |
|
The plan for the OOv1 revision, its decision list, and its stated gaps. |
|
The MMv3 → OOv1 revision session: prior art, the Baum classification, the Michaelis–Menten correction. |
|
How the nuclear-risk bibliography was assembled and checked. |
|
OOv2r0p0 itself. |
The MMv1–MMv3 drafts are retained unedited under source/matheology/hell/mm/b/16/ as the
archive. Nothing has been rewritten in place.
Why this page exists. An earlier version of this draft carried its own change history inline — admonitions scattered through the sections announcing what had moved and why. That is honest but it is not readable: a first-time reader does not need to know what a previous draft said, and being told repeatedly is a distraction from the argument. The provenance is kept here, in full, and the paper keeps only those corrections that a reader needs in order to judge the current claims — principally the §2.0a note, because a reader comparing against a quoted 3.24% is entitled to know why it moved.