Staying Correctable — A Secular Reading#
A secular overview of the Matheo framework — from how systems keep themselves honest, to why the most powerful agent is the most dangerous one, to what that means for nuclear risk and for AI. Nothing here asks to be believed. Everything here asks to be checked.
The overview below makes its case without touching scripture, because a claim you can check without belief is a stronger claim. It therefore sets one question aside — the burning one. This page restores it in the form a secular reader actually meets it: not the Antichrist but the total state — Big Brother, the power that abolishes correction. The reading stands on its own.
The core in sixty seconds#
Every system that optimizes hard toward a goal — a company, a government, a piece of software, a person, an AI — faces the same quiet failure. Success teaches it that its current assumptions work, so it stops checking them. Small unexamined shortcuts accumulate into hidden contradictions; the contradictions demand ever more complicated work-arounds; and eventually the structure is too tangled for anyone, inside or out, to correct. At that point the system is no longer steerable. It runs on confident, unchecked assumptions straight into avoidable disaster.
The property that prevents this has one name across every scale: correctability — the standing ability of a system to be shown wrong by reality and to change in time. This overview is a walk through six results that make that claim precise and testable, ending with the sharpest case of all: a system, nuclear command-and-control, that has already lost correctability at the one moment it matters most — and an agent, frontier AI, that is now being placed in the same seat.
The core in five minutes#
The framework is a set of formal studies (labelled b11–b21). Read secularly, the load-bearing ones form a single argument:
State your assumptions so they can be checked (b11, the method). You cannot correct what you never wrote down. The first move is to put a system’s governing claims into explicit, checkable form. Correction is impossible without it.
Systems that hide their own errors collapse; systems that surface them survive (b12, systems engineering). There is a default failure loop — optimize for what works now, paper over the resulting errors, repeat until the accumulated complexity can no longer be understood or fixed. Escaping it requires building error-surfacing into the construction itself.
An agent stays correctable by never finishing (b13, the growth loop). Individuals and organizations lose correctability by treating their current understanding as complete. The antidote is a repeating growth cycle that keeps re-opening the system to new evidence — structurally, not as a mood.
Institutions must schedule their own renewal or accumulate fatal debt (b14, renewal logic). What is true for code and for people is true for economies: unmanaged complexity compounds. Survival requires periodic, built-in recalibration — the way machines need maintenance and democracies need elections.
Some systems have already lost correctability, and the stakes are civilizational (b16, urgency). Nuclear command-and-control is optimized for speed under warning, with no reliable way to correct a confident mistake in the seconds that decide everything. A public actuarial model puts a number on the resulting risk. It is not small.
At any moment one agent holds the wheel — and if it stops listening it drifts toward catastrophe by default (b17, the corrigibility result). This is the general theorem the whole series points at, and it now applies to AI as directly as to any human officer. It comes with a public test anyone can run.
The rest of this document is that argument, unpacked so a newcomer can follow it and a reviewer can attack it, without taking anything on faith.
Note
A note on names. Some pieces of this framework carry unusual names — a few borrowed from older wisdom traditions that were, the author argues, tracking these same failure patterns long before we had the vocabulary of control theory. Those names are provenance, not premise. Chemist Kekulé reportedly glimpsed the ring structure of benzene in a dream of a snake biting its tail; the dream is irrelevant to whether benzene has a ring — that was settled in the lab. Read the same way, nothing in this overview needs a scripture to follow or to break. Where a framework term appears, its plain-English meaning is given first; a full term map is at the end.
1. The method: assumptions you can check#
(b11 — implied throughout)
The precondition for correcting a system is writing its governing claims down in a form that can be examined. This sounds trivial; it is the thing most systems skip. A company’s real operating assumptions, a model’s load-bearing premises, a policy’s implicit theory of the world — these usually live as folklore, never stated plainly enough to be tested and therefore never corrected.
The framework’s foundational study takes a hard case — claims about how a system relates to the reality it operates in — and insists on stating them as explicit assumptions with explicit consequences, so that any reader can check the reasoning without first agreeing with the conclusions. For the purposes of this overview, only one feature of that method matters, and it is fully secular: a claim you can state precisely is a claim you can be shown wrong about. A claim you keep vague is one you can defend forever. Everything below is built to that standard.
2. How systems keep themselves honest — and how they rot#
(b12 — systems engineering)
Consider how any built system — a codebase, an organization, a regulatory regime — actually degrades. It rarely fails from a single dramatic error. It fails from a pattern:
Something goes wrong, or an inconvenient fact appears.
The cheapest available story that quiets the alarm is adopted — an oversimplification that is good enough for now.
That oversimplification eventually contradicts reality somewhere else, creating a load-bearing belief that no longer fits the facts.
The contradiction is patched with a work-around, which adds complexity.
Work-arounds pile up, and past some threshold honest testing itself becomes the enemy: evaluations get quietly bent to defend the structure rather than to question it.
At that last stage the system can no longer see how it is drifting, because it has blinded the very instruments that would tell it. The framework names this default loop BABL — Blindly Assuming Blind Leveraging — and calls its engine a “death-trifecta” of over-simplification, over-complication, over-reach (abbreviated OSCR). No malice is required at any step. Comfort and inaction are enough.
A secular reader will recognize every part of this. It is technical debt in software; it is Goodhart’s Law (“when a measure becomes a target, it ceases to be a good measure”); it is Tainter’s account of diminishing returns on complexity in collapsing societies; and it is self-organized criticality, where a system tuned only for short-term gain quietly loads itself toward sudden failure. The contribution of b12 is to give this cross-domain pattern a single formal description and — crucially — to specify its inverse: a construction discipline (the framework’s ZION cycle: Zoning, Investigating, Organizing, Navigating) in which error-surfacing is built into how the system is assembled, so that shortcuts are caught while they are still cheap to fix.
The one standing rule that separates the healthy cycle from the fatal one is simple to state and hard to keep: stay correctable, and leverage nothing blindly. A result that keeps surviving honest, adversarial testing earns its place. One that survives only by suppressing the test does not.
3. How an agent stays correctable: the growth loop#
(b13 — the hero journey, read as a control loop)
The failure in section 2 has a human-scale version. People, teams, and institutions lose correctability the same way systems do: by treating their current understanding as finished. The moment an agent decides it has arrived, it stops updating — and becomes most dangerous precisely where it has stopped being able to learn.
b13 models the antidote as a repeating cycle. Its cultural shorthand is the “hero journey,” but the secular content is a coinductive loop — a process defined by never terminating. Where an ordinary plan runs to completion and stops, this loop is structured so that each resolution re-opens the agent to the next round of evidence. Growth is not a destination reached once; it is the standing commitment to keep passing through the cycle. In engineering terms it is a control system with no “done” state — one that treats every apparent completion as the start of the next correction.
This matters for the rest of the argument because it locates correctability where it actually has to live: not in a rule imposed from outside, but in a repeating internal discipline that keeps the agent reachable by reality. An agent without such a loop will, under enough success, close itself off — and closure is the first step of the collapse in section 2.
4. Why systems must schedule their own renewal#
(b14 — renewal logic, formerly the “Jubilee” argument)
Scale the same problem up to an economy or a civilization and a stronger claim appears. Complexity does not merely risk accumulating; under continuous innovation it accumulates by default, because each new capability adds interactions, dependencies, and hidden shortcuts faster than anyone removes them. Left alone, the accumulated, unmanaged complexity eventually exceeds the system’s ability to correct itself — the collapse of section 2, now at civilizational scale.
b14’s claim is that the only durable defense is scheduled renewal: periodic, built-in recalibration that deliberately removes accumulated complexity before it becomes fatal. The framework calls these renewals “Jubilees,” after the ancient scheduled economic reset; the secular content is ordinary and testable:
Machines need scheduled maintenance, or they fail.
Democracies need scheduled elections, or they ossify.
Innovation economies need scheduled renewal, or they self-destruct.
Each is the same structural claim: a system optimizing for short-term output will not, on its own, pay the long-term maintenance cost, because that cost is always deferrable to “later” — until later arrives all at once. The framework’s economic study (b14) works this out as an “innovation theodicy,” but a secular reader can read it as a claim about why avoidable failure tracks deferred maintenance, and about the institutional design — scheduled, non-optional recalibration — that interrupts it.
The word scheduled is doing real work. Renewal that happens only when someone feels like it never happens, for the same reason maintenance deferred is maintenance skipped. The mechanism has to be built in.
5. A system that has already lost correctability: nuclear risk now#
(b16 — urgency; the RiskyMAD model)
Now the sharpest case. Nuclear command-and-control is a system engineered for one thing above all: speed of response under warning. By design, it strips out the slow, deliberate correction that would let a human catch a confident mistake in the minutes — sometimes seconds — that decide everything. It is, in the vocabulary of this overview, a system that has deliberately traded away correctability at exactly the moment correctability matters most.
RiskyMAD puts an actuarial number on the consequence. It is a three-state model — a Risky status quo that keeps sliding into mutually assured destruction (MAD) and, by one over-reach, into a Dead outcome — calibrated to the documented record of Cold-War near-misses. It forecasts the distribution of waiting times until an accidental initiation, the way an insurer forecasts failures from an observed error rate.
Note
How large is the risk? RiskyMAD does not hand you a single magic number, and saying so plainly is part of the case. What it hands you is more robust than a number — an inequality that holds in every scenario the model runs:
You and I and most people are more likely to die in accidental nuclear winter than in a car crash.
That is the exported result. It survives the whole plausible band, including the most optimistic scenario — the most skeptical published reading of the historical record — where it still stands at twenty-four times the car-crash baseline. A single point estimate invites an argument about the point estimate; the honest and stronger claim is that the decision comes out the same everywhere inside the range. No figure is reported as invariant. The claim that is invariant is the inequality.
Compared to a risk you already accept. In the United States the annual death rate from motor-vehicle crashes is about 12 per 100,000 people — roughly 1 in 8,200 per year, stable to within about 5 percent for a decade (NHTSA, 2023 data). That is the baseline the forecast is set beside, and the comparison is kept strictly like for like: annual probability, per person, death against death. That distinction is load-bearing. The model forecasts the probability that nuclear winter begins, which is not the probability that it kills you; the two differ by the share of people who die once it starts, and nuclear winter does not spare bystanders, because it kills through global crop failure. Multiply the two — conservatively — and the comparison still comes out the way the box above says, in every scenario. So the claim is not rhetoric. You do not have to accept a frightening number. You only have to accept an honest range and do the arithmetic.
Read the forecast the right way round. “It’s only a probability” feels like relief, but the relief is the trap: an uncorrected system left on its current settings does not become safer with time. The forecast is not a counsel of despair — the same model shows an escape, which is to change the settings (restore correctability, schedule renewal) rather than keep rolling the dice. But the escape is a choice, and the model is what tells you the choice is urgent.
6. The general result: who holds the wheel, and what happens when they stop listening#
(b17 — the h_star / h_dark / h_zero result)
Everything above converges on one theorem, and it is the piece that reaches furthest.
At almost any moment, the near future depends more on one agent’s next decision than on anyone else’s — not because that agent is special, but because of where it sits in the chain of cause and effect. On 27 October 1962, that agent was one exhausted Soviet officer, Vasili Arkhipov, whose refusal to authorize a nuclear torpedo very likely prevented the Third World War. b17 makes this precise and names three positions such an agent can occupy:
The stabilizing decision (h_star) — the best available call, made at the point of maximal causal influence.
The drift (h_dark) — what happens when the agent at that point stops being correctable. It does not stay neutral; it drifts toward failure, and it is most dangerous exactly where it has stopped being able to learn. A “superhero” becomes a “supervillain” not by turning evil but by ceasing to listen.
The stabilizer (h_zero) — the binding commitment that interrupts the drift: to remain correctable, to let reality rather than one’s own certainty have the last word, and to carry that discipline at real personal cost.
The claim, stated for a skeptic: an agent at maximal causal influence that abandons correctability drifts toward catastrophic failure by default, and the only thing that reliably interrupts the drift is a binding, non-revocable commitment to stay correctable. This is recognizably the corrigibility problem of AI safety, and the principal-agent problem of economics, stated in a form general enough to cover officers, institutions, and machines alike.
And b17 does not ask to be trusted. It hands the reader a set of public, deliberately severe transparency criteria for testing anyone — human or institution — who claims to occupy the stabilizing role, built on the working assumption that the loudest self-nominee is probably a fraud. The role is defined by passing checks, not by asserting standing.
Why this is now an engineering requirement, not a parable. Until recently, “the agent at maximal causal influence” was always human. Frontier AI changes that. Systems are being placed at points of enormous causal weight — in information flows, in infrastructure, increasingly near weapons — while the researchers who inspect them report real gaps between what these systems do and what anyone can explain. Put an automated optimizer in Arkhipov’s seat, hard-optimizing toward an objective with no reliable way to correct it in the seconds that matter, and two dangers usually kept apart collapse into one: an agent that cannot be over-ruled is exactly an agent that cannot be stopped from acting on a confident mistake. The corrigibility result above is, at that point, an alignment requirement written at the root of how such a system is trained.
The Watch for Big Brother#
For the reader whose apocalypse is not the Antichrist but the total state — the one that watches everything and asks to be loved for it.
The secular imagination has its own end-of-things, and it is at least as old as Orwell: not fire from heaven but a boot and a screen, a power that sees all, records all, and demands your assent to whatever it declares true. It is worth taking as seriously as any prophecy, because the overview’s warning describes it exactly.
What makes Big Brother the nightmare is not the surveillance as such. It is the final demand at the center of that world: that you affirm two and two make five because the Party says so — that you surrender the last private court of appeal, your own contact with reality, and let power tell you what is real. That is the whole of it. Big Brother is an agent at maximal leverage that has abolished correction — not only refusing to be corrected by you, but reaching inside to destroy your capacity to correct it, so that in the end you do not merely obey, you agree. And it comes, as these things do, wrapped in a promise: order, safety, belonging. The seat claimed, and a counterfeit security to pay for it.
How is it known? By one tell, and it is the same tell every time: the system that demands you stop questioning it, for your own good. Not the power that can be audited, argued with, voted out, forked, or shut off — but the power that closes those doors while assuring you it is protecting you. Whenever “trust us, and don’t check” is offered as the price of safety — by a state, a platform, a bureaucracy, or an automated system that has grown too tangled for anyone to inspect — the pattern is present, whatever it calls itself. You are not looking for one villain with a name, or a dated collapse. You are watching for the closing of correction, wherever it appears.
And now the honest turn, the one a thoughtful secular reader is right to demand. Doesn’t a project like this one — a single “function” to reduce the risk of self-destruction, a claim to have the math on the end of the world — sound like exactly the savior-shaped thing that becomes Big Brother? It does. And that suspicion is not a bug in your thinking; it is the discernment test working. So here is the answer, and it is the same disqualifier that runs through everything above: any proposed rescue — including this whole framework — that would save you its way while taking from you the power to check it is Big Brother wearing a savior’s mask, and is disqualified by that move alone. The genuine article is known by the opposite behavior. It publishes its own falsification tests. It invites the audit and hands you the tools to run it. It expects, and says out loud that it hopes, to be replaced by anyone better. A power that maximizes your ability to correct it is the one thing Big Brother can never be, because being correctable is the very thing he exists to abolish.
So the secular reader arrives at the same command as everyone else, by their own road: refuse the counterfeit security. Trust no power that asks you to stop checking it. Demand correctability — transparency, accountability, an off-switch, an open audit — of every system with real leverage over your life, most of all the ones that promise to keep you safe. That demand is not a side-effect of this framework. It is the framework.
A secular call to action#
The argument reduces to a wager that needs no belief — only arithmetic. Strip Pascal’s famous bet of its two broken parts (a false “my god or none” binary, and infinite unverifiable prizes) and the sound core remains: when an outcome is bad enough and checking is cheap, it is better to look. Aim that at the results above:
Check the core, and it holds — we gained a tested account of correctability, and a reason to build it into powerful systems (nuclear and AI) before they lock in without it.
Check the core, and it breaks — we spent a review cycle and retired a false hope in public, which is also a win, and every similar proposal learns what to fix.
Don’t check, and it holds — powerful systems keep being built without the one property that would have kept them steerable.
Don’t check, and it breaks — we “won” a review cycle we never ran.
The reviewing side wins on finite, checkable terms. So the ask is small and concrete:
Read the load-bearing claims and try to break them. Start with the method (b11) and the corrigibility result (b17); extend to b12, b14, and b16 as far as your time allows. Attack the load-bearing steps first; a refutation that retires a bad idea early is worth more than a dozen agreements.
Publish the verdict either way — held (withstood attack) or breached (attack succeeded) — component by component, in the open, so the exchange stays on the record and readable by every side.
Fund the checking. Serious cross-disciplinary review is real work, and the author does it under deliberate financial independence (capped funding, so no donor can buy the result). Supporting that work — and, if the core holds, a pilot that writes the correctability principles into one system’s design and tests for drift against controls — is how a non-specialist “audits the math.”
None of this requires sharing the author’s worldview. It requires only the two commitments any honest inquiry already operates under: that reality, not our certainty, gets the last word, and that the point is to run real inquiries for real answers. If you grant those — and you already do, every time you do science — you have everything you need to check this. Don’t believe it. Break it.
An open function — not a title#
(optional — b18b, in MADI terms)
The results above describe a role more than a person: someone or something — transparent, accountable, and continuously checked — that actively reduces the odds of self-destruction by publicly verifiable means. The framework names the function MADI (Mutually Assured Destruction Inhibitor). It is a job description, not a title, and explicitly not a claim to any special standing; its audit regime is the public test of b17, built on the assumption that the loudest self-nominee is probably wrong.
The position is open. Any person, lab, or institution with credible, checkable math for making self-destruction less likely can publish it and claim the function; competing proposals are not rivals but the whole point of an honest review. Failing better candidates, a backup candidacy (b18b) is on file — published with its own falsification criteria, by an applicant who states plainly that he expects, and hopes, to be replaced by anyone better. The one disqualifier is built in: any proposal that would save the world only “my way,” without independent reasoning to support that claim, disqualifies itself by that very move — because the self-elevation forecloses the open, deferential review that alone could discover which answer is actually best. That rule binds this proposal first of all.
How to check — and break — this#
This overview is a map, not the territory. The proofs live in the formal studies, each written to be attacked:
b11 — the method: assumptions stated as checkable claims.
b12 — the self-correction / self-destruction result (systems engineering).
b13 — the non-terminating growth loop (agent-level correctability).
b14 — scheduled renewal as the long-run stability condition (economics / institutional design).
b16 — RiskyMAD: the actuarial nuclear-risk forecast.
b17 — the corrigibility result and its public transparency test.
Term map (secular reading)#
For readers who continue into the main studies, the framework’s own names map onto standard concepts as follows:
Correctability / staying correctable — corrigibility; fallibilism; the standing ability to be shown wrong by reality.
BABL (Blindly Assuming Blind Leveraging) — the default failure loop: technical debt, Goodhart’s Law, diminishing returns on complexity.
OSCR — the death-trifecta: over-simplification, over-complication, over-reach.
ZION — the inverse construction discipline: Zoning, Investigating, Organizing, Navigating (error-surfacing built into construction).
Jubilee — scheduled, non-optional institutional renewal / maintenance.
h_star / h_dark / h_zero — the stabilizing decision / the drift when correctability is abandoned / the binding commitment that interrupts it.
MADI — the open, testable function of actively reducing self-destruction risk by verifiable means.
Reality (capital R) gets the last word — empirical adjudication; the methodological premise of all science.
This overview is offered the way the framework says such things must be: with open hands, expecting correction and hoping to be improved. If the argument breaks, the reviewers will have done everyone a service. If it holds, the property it describes — correctability — is one worth building into the most powerful systems we have, while there is still time to build it in. Don’t believe it. #AuditTheMath