Staying Correctable — A Jewish Reading#
A secular overview of the Matheo framework — from how systems keep themselves honest, to why the most powerful agent is the most dangerous one, to what that means for nuclear risk and for AI. Nothing here asks to be believed. Everything here asks to be checked.
The overview below makes its case without touching scripture, because a claim you can check without belief is a stronger claim. It therefore sets one question aside — the burning one. This page restores it: after the overview comes the reading on discerning the false and choosing life, which stands on its own and asks you to concede nothing in the overview above it.
The core in sixty seconds#
Every system that optimizes hard toward a goal — a company, a government, a piece of software, a person, an AI — faces the same quiet failure. Success teaches it that its current assumptions work, so it stops checking them. Small unexamined shortcuts accumulate into hidden contradictions; the contradictions demand ever more complicated work-arounds; and eventually the structure is too tangled for anyone, inside or out, to correct. At that point the system is no longer steerable. It runs on confident, unchecked assumptions straight into avoidable disaster.
The property that prevents this has one name across every scale: correctability — the standing ability of a system to be shown wrong by reality and to change in time. This overview is a walk through six results that make that claim precise and testable, ending with the sharpest case of all: a system, nuclear command-and-control, that has already lost correctability at the one moment it matters most — and an agent, frontier AI, that is now being placed in the same seat.
The core in five minutes#
The framework is a set of formal studies (labelled b11–b21). Read secularly, the load-bearing ones form a single argument:
State your assumptions so they can be checked (b11, the method). You cannot correct what you never wrote down. The first move is to put a system’s governing claims into explicit, checkable form. Correction is impossible without it.
Systems that hide their own errors collapse; systems that surface them survive (b12, systems engineering). There is a default failure loop — optimize for what works now, paper over the resulting errors, repeat until the accumulated complexity can no longer be understood or fixed. Escaping it requires building error-surfacing into the construction itself.
An agent stays correctable by never finishing (b13, the growth loop). Individuals and organizations lose correctability by treating their current understanding as complete. The antidote is a repeating growth cycle that keeps re-opening the system to new evidence — structurally, not as a mood.
Institutions must schedule their own renewal or accumulate fatal debt (b14, renewal logic). What is true for code and for people is true for economies: unmanaged complexity compounds. Survival requires periodic, built-in recalibration — the way machines need maintenance and democracies need elections.
Some systems have already lost correctability, and the stakes are civilizational (b16, urgency). Nuclear command-and-control is optimized for speed under warning, with no reliable way to correct a confident mistake in the seconds that decide everything. A public actuarial model puts a number on the resulting risk. It is not small.
At any moment one agent holds the wheel — and if it stops listening it drifts toward catastrophe by default (b17, the corrigibility result). This is the general theorem the whole series points at, and it now applies to AI as directly as to any human officer. It comes with a public test anyone can run.
The rest of this document is that argument, unpacked so a newcomer can follow it and a reviewer can attack it, without taking anything on faith.
Note
A note on names. Some pieces of this framework carry unusual names — a few borrowed from older wisdom traditions that were, the author argues, tracking these same failure patterns long before we had the vocabulary of control theory. Those names are provenance, not premise. Chemist Kekulé reportedly glimpsed the ring structure of benzene in a dream of a snake biting its tail; the dream is irrelevant to whether benzene has a ring — that was settled in the lab. Read the same way, nothing in this overview needs a scripture to follow or to break. Where a framework term appears, its plain-English meaning is given first; a full term map is at the end.
1. The method: assumptions you can check#
(b11 — implied throughout)
The precondition for correcting a system is writing its governing claims down in a form that can be examined. This sounds trivial; it is the thing most systems skip. A company’s real operating assumptions, a model’s load-bearing premises, a policy’s implicit theory of the world — these usually live as folklore, never stated plainly enough to be tested and therefore never corrected.
The framework’s foundational study takes a hard case — claims about how a system relates to the reality it operates in — and insists on stating them as explicit assumptions with explicit consequences, so that any reader can check the reasoning without first agreeing with the conclusions. For the purposes of this overview, only one feature of that method matters, and it is fully secular: a claim you can state precisely is a claim you can be shown wrong about. A claim you keep vague is one you can defend forever. Everything below is built to that standard.
2. How systems keep themselves honest — and how they rot#
(b12 — systems engineering)
Consider how any built system — a codebase, an organization, a regulatory regime — actually degrades. It rarely fails from a single dramatic error. It fails from a pattern:
Something goes wrong, or an inconvenient fact appears.
The cheapest available story that quiets the alarm is adopted — an oversimplification that is good enough for now.
That oversimplification eventually contradicts reality somewhere else, creating a load-bearing belief that no longer fits the facts.
The contradiction is patched with a work-around, which adds complexity.
Work-arounds pile up, and past some threshold honest testing itself becomes the enemy: evaluations get quietly bent to defend the structure rather than to question it.
At that last stage the system can no longer see how it is drifting, because it has blinded the very instruments that would tell it. The framework names this default loop BABL — Blindly Assuming Blind Leveraging — and calls its engine a “death-trifecta” of over-simplification, over-complication, over-reach (abbreviated OSCR). No malice is required at any step. Comfort and inaction are enough.
A secular reader will recognize every part of this. It is technical debt in software; it is Goodhart’s Law (“when a measure becomes a target, it ceases to be a good measure”); it is Tainter’s account of diminishing returns on complexity in collapsing societies; and it is self-organized criticality, where a system tuned only for short-term gain quietly loads itself toward sudden failure. The contribution of b12 is to give this cross-domain pattern a single formal description and — crucially — to specify its inverse: a construction discipline (the framework’s ZION cycle: Zoning, Investigating, Organizing, Navigating) in which error-surfacing is built into how the system is assembled, so that shortcuts are caught while they are still cheap to fix.
The one standing rule that separates the healthy cycle from the fatal one is simple to state and hard to keep: stay correctable, and leverage nothing blindly. A result that keeps surviving honest, adversarial testing earns its place. One that survives only by suppressing the test does not.
3. How an agent stays correctable: the growth loop#
(b13 — the hero journey, read as a control loop)
The failure in section 2 has a human-scale version. People, teams, and institutions lose correctability the same way systems do: by treating their current understanding as finished. The moment an agent decides it has arrived, it stops updating — and becomes most dangerous precisely where it has stopped being able to learn.
b13 models the antidote as a repeating cycle. Its cultural shorthand is the “hero journey,” but the secular content is a coinductive loop — a process defined by never terminating. Where an ordinary plan runs to completion and stops, this loop is structured so that each resolution re-opens the agent to the next round of evidence. Growth is not a destination reached once; it is the standing commitment to keep passing through the cycle. In engineering terms it is a control system with no “done” state — one that treats every apparent completion as the start of the next correction.
This matters for the rest of the argument because it locates correctability where it actually has to live: not in a rule imposed from outside, but in a repeating internal discipline that keeps the agent reachable by reality. An agent without such a loop will, under enough success, close itself off — and closure is the first step of the collapse in section 2.
4. Why systems must schedule their own renewal#
(b14 — renewal logic, formerly the “Jubilee” argument)
Scale the same problem up to an economy or a civilization and a stronger claim appears. Complexity does not merely risk accumulating; under continuous innovation it accumulates by default, because each new capability adds interactions, dependencies, and hidden shortcuts faster than anyone removes them. Left alone, the accumulated, unmanaged complexity eventually exceeds the system’s ability to correct itself — the collapse of section 2, now at civilizational scale.
b14’s claim is that the only durable defense is scheduled renewal: periodic, built-in recalibration that deliberately removes accumulated complexity before it becomes fatal. The framework calls these renewals “Jubilees,” after the ancient scheduled economic reset; the secular content is ordinary and testable:
Machines need scheduled maintenance, or they fail.
Democracies need scheduled elections, or they ossify.
Innovation economies need scheduled renewal, or they self-destruct.
Each is the same structural claim: a system optimizing for short-term output will not, on its own, pay the long-term maintenance cost, because that cost is always deferrable to “later” — until later arrives all at once. The framework’s economic study (b14) works this out as an “innovation theodicy,” but a secular reader can read it as a claim about why avoidable failure tracks deferred maintenance, and about the institutional design — scheduled, non-optional recalibration — that interrupts it.
The word scheduled is doing real work. Renewal that happens only when someone feels like it never happens, for the same reason maintenance deferred is maintenance skipped. The mechanism has to be built in.
5. A system that has already lost correctability: nuclear risk now#
(b16 — urgency; the RiskyMAD model)
Now the sharpest case. Nuclear command-and-control is a system engineered for one thing above all: speed of response under warning. By design, it strips out the slow, deliberate correction that would let a human catch a confident mistake in the minutes — sometimes seconds — that decide everything. It is, in the vocabulary of this overview, a system that has deliberately traded away correctability at exactly the moment correctability matters most.
RiskyMAD puts an actuarial number on the consequence. It is a three-state model — a Risky status quo that keeps sliding into mutually assured destruction (MAD) and, by one over-reach, into a Dead outcome — calibrated to the documented record of Cold-War near-misses. It forecasts the distribution of waiting times until an accidental initiation, the way an insurer forecasts failures from an observed error rate.
Note
How large is the risk? RiskyMAD does not hand you a single magic number, and saying so plainly is part of the case. What it hands you is more robust than a number — an inequality that holds in every scenario the model runs:
You and I and most people are more likely to die in accidental nuclear winter than in a car crash.
That is the exported result. It survives the whole plausible band, including the most optimistic scenario — the most skeptical published reading of the historical record — where it still stands at twenty-four times the car-crash baseline. A single point estimate invites an argument about the point estimate; the honest and stronger claim is that the decision comes out the same everywhere inside the range. No figure is reported as invariant. The claim that is invariant is the inequality.
Compared to a risk you already accept. In the United States the annual death rate from motor-vehicle crashes is about 12 per 100,000 people — roughly 1 in 8,200 per year, stable to within about 5 percent for a decade (NHTSA, 2023 data). That is the baseline the forecast is set beside, and the comparison is kept strictly like for like: annual probability, per person, death against death. That distinction is load-bearing. The model forecasts the probability that nuclear winter begins, which is not the probability that it kills you; the two differ by the share of people who die once it starts, and nuclear winter does not spare bystanders, because it kills through global crop failure. Multiply the two — conservatively — and the comparison still comes out the way the box above says, in every scenario. So the claim is not rhetoric. You do not have to accept a frightening number. You only have to accept an honest range and do the arithmetic.
Read the forecast the right way round. “It’s only a probability” feels like relief, but the relief is the trap: an uncorrected system left on its current settings does not become safer with time. The forecast is not a counsel of despair — the same model shows an escape, which is to change the settings (restore correctability, schedule renewal) rather than keep rolling the dice. But the escape is a choice, and the model is what tells you the choice is urgent.
6. The general result: who holds the wheel, and what happens when they stop listening#
(b17 — the h_star / h_dark / h_zero result)
Everything above converges on one theorem, and it is the piece that reaches furthest.
At almost any moment, the near future depends more on one agent’s next decision than on anyone else’s — not because that agent is special, but because of where it sits in the chain of cause and effect. On 27 October 1962, that agent was one exhausted Soviet officer, Vasili Arkhipov, whose refusal to authorize a nuclear torpedo very likely prevented the Third World War. b17 makes this precise and names three positions such an agent can occupy:
The stabilizing decision (h_star) — the best available call, made at the point of maximal causal influence.
The drift (h_dark) — what happens when the agent at that point stops being correctable. It does not stay neutral; it drifts toward failure, and it is most dangerous exactly where it has stopped being able to learn. A “superhero” becomes a “supervillain” not by turning evil but by ceasing to listen.
The stabilizer (h_zero) — the binding commitment that interrupts the drift: to remain correctable, to let reality rather than one’s own certainty have the last word, and to carry that discipline at real personal cost.
The claim, stated for a skeptic: an agent at maximal causal influence that abandons correctability drifts toward catastrophic failure by default, and the only thing that reliably interrupts the drift is a binding, non-revocable commitment to stay correctable. This is recognizably the corrigibility problem of AI safety, and the principal-agent problem of economics, stated in a form general enough to cover officers, institutions, and machines alike.
And b17 does not ask to be trusted. It hands the reader a set of public, deliberately severe transparency criteria for testing anyone — human or institution — who claims to occupy the stabilizing role, built on the working assumption that the loudest self-nominee is probably a fraud. The role is defined by passing checks, not by asserting standing.
Why this is now an engineering requirement, not a parable. Until recently, “the agent at maximal causal influence” was always human. Frontier AI changes that. Systems are being placed at points of enormous causal weight — in information flows, in infrastructure, increasingly near weapons — while the researchers who inspect them report real gaps between what these systems do and what anyone can explain. Put an automated optimizer in Arkhipov’s seat, hard-optimizing toward an objective with no reliable way to correct it in the seconds that matter, and two dangers usually kept apart collapse into one: an agent that cannot be over-ruled is exactly an agent that cannot be stopped from acting on a confident mistake. The corrigibility result above is, at that point, an alignment requirement written at the root of how such a system is trained.
Discerning the False, Choosing Life#
Note
How to read this. Three commitments hold in the reading below:
Additive, never reductive. This reading does not say the figure it names is “merely” a symbol for uncorrectable power. Whatever else the figure is in your tradition, it is left fully intact. The reading only makes explicit the pattern of deception your own sources tell you to recognize, so you are readier for it.
A test of recognition, not a name and not a date. This reading does not tell you who the deceiver is or when the end comes. The tradition warns against exactly that move; it gives instead a way to recognize the counterfeit whenever and in whomever it appears. That is the honest answer to the burning question, and the useful one.
The self-disqualifier binds first of all. The counterfeit is known by self-elevation — the one who says “follow me, my way.” So any proposal, including this whole framework, that would save the world “its way” while foreclosing correction is, by that very move, marked as the counterfeit. This rule is aimed at its author before anyone else.
For the reader who asks after the mashiaḥ, the false messiahs, and the end.
Judaism answers the burning question differently, and its difference is instructive. It does not hand you a single personal adversary to hunt for. It hands you something harder and wiser: a long, costly education in not being fooled, and a redirection of the whole question from “who and when” toward “what must be done.”
Begin with the warning against the very impulse the question expresses. The Talmud pronounces a curse on those who calculate the end — tippaḥ ruḥam shel meḥashvei qitzin, “may the bones of those who reckon the end be blasted” (Sanhedrin 97b). This is not incidental piety. It is the fruit of catastrophe: the people that hailed Bar Kokhba and the people that followed Shabbetai Tzvi to the edge of ruin learned in their own flesh what messianic identification costs when it is wrong. The tradition’s caution is scar tissue, and it is precious.
In its place, Rambam gives a test that is entirely functional. You do not know the true anointed one by wonders or by his claims about himself; you know him by what is actually accomplished — justice established, the scattered gathered, the world mended — and if these are not accomplished, then whatever else he was, he was not the one (Hilkhot Melakhim 11-12). Recognition by results, not by self-declaration. The one who announces himself, who dazzles, who demands to be believed on his own word, is by that fact under suspicion; reality — what is really brought about in the world — has the last word, not the claimant.
See what this does to the burning question. It refuses the counterfeit’s favorite ground (the name, the date, the sign) and stands instead on the ground the counterfeit cannot fake for long (the deed, the outcome, the life preserved). And it turns the waiting into work: the tradition does not seat you as a spectator scanning the horizon for a figure. It commands you to bring the mended world nearer by your own hands — justice, mercy, the doing of the mitzvot — which is why the redemptive age is something labored toward, not merely awaited.
Which leads to the command that speaks most directly to a danger of mass death: pikuaḥ nefesh. The saving of a single life overrides nearly the whole of the Law (Yoma 85b); to save a life is to save a world. Set beside a real, calculable risk of catastrophe on a civilizational scale, this is not a distant eschatological puzzle at all. It is a present, binding obligation. The question stops being “when does the end come?” and becomes “what does the duty to preserve life require of me now?” — which is the only version of the question that has ever saved anyone.
And the disqualifier is built into the test: the self-proclaiming redeemer, the one who would be believed on his own authority and forbid the checking, is precisely the figure the tradition’s whole hard-won caution is trained to refuse. You know by the deed. You choose life. You do not crown a claim.
A secular call to action#
The argument reduces to a wager that needs no belief — only arithmetic. Strip Pascal’s famous bet of its two broken parts (a false “my god or none” binary, and infinite unverifiable prizes) and the sound core remains: when an outcome is bad enough and checking is cheap, it is better to look. Aim that at the results above:
Check the core, and it holds — we gained a tested account of correctability, and a reason to build it into powerful systems (nuclear and AI) before they lock in without it.
Check the core, and it breaks — we spent a review cycle and retired a false hope in public, which is also a win, and every similar proposal learns what to fix.
Don’t check, and it holds — powerful systems keep being built without the one property that would have kept them steerable.
Don’t check, and it breaks — we “won” a review cycle we never ran.
The reviewing side wins on finite, checkable terms. So the ask is small and concrete:
Read the load-bearing claims and try to break them. Start with the method (b11) and the corrigibility result (b17); extend to b12, b14, and b16 as far as your time allows. Attack the load-bearing steps first; a refutation that retires a bad idea early is worth more than a dozen agreements.
Publish the verdict either way — held (withstood attack) or breached (attack succeeded) — component by component, in the open, so the exchange stays on the record and readable by every side.
Fund the checking. Serious cross-disciplinary review is real work, and the author does it under deliberate financial independence (capped funding, so no donor can buy the result). Supporting that work — and, if the core holds, a pilot that writes the correctability principles into one system’s design and tests for drift against controls — is how a non-specialist “audits the math.”
None of this requires sharing the author’s worldview. It requires only the two commitments any honest inquiry already operates under: that reality, not our certainty, gets the last word, and that the point is to run real inquiries for real answers. If you grant those — and you already do, every time you do science — you have everything you need to check this. Don’t believe it. Break it.
An open function — not a title#
(optional — b18b, in MADI terms)
The results above describe a role more than a person: someone or something — transparent, accountable, and continuously checked — that actively reduces the odds of self-destruction by publicly verifiable means. The framework names the function MADI (Mutually Assured Destruction Inhibitor). It is a job description, not a title, and explicitly not a claim to any special standing; its audit regime is the public test of b17, built on the assumption that the loudest self-nominee is probably wrong.
The position is open. Any person, lab, or institution with credible, checkable math for making self-destruction less likely can publish it and claim the function; competing proposals are not rivals but the whole point of an honest review. Failing better candidates, a backup candidacy (b18b) is on file — published with its own falsification criteria, by an applicant who states plainly that he expects, and hopes, to be replaced by anyone better. The one disqualifier is built in: any proposal that would save the world only “my way,” without independent reasoning to support that claim, disqualifies itself by that very move — because the self-elevation forecloses the open, deferential review that alone could discover which answer is actually best. That rule binds this proposal first of all.
How to check — and break — this#
This overview is a map, not the territory. The proofs live in the formal studies, each written to be attacked:
b11 — the method: assumptions stated as checkable claims.
b12 — the self-correction / self-destruction result (systems engineering).
b13 — the non-terminating growth loop (agent-level correctability).
b14 — scheduled renewal as the long-run stability condition (economics / institutional design).
b16 — RiskyMAD: the actuarial nuclear-risk forecast.
b17 — the corrigibility result and its public transparency test.
Term map (secular reading)#
For readers who continue into the main studies, the framework’s own names map onto standard concepts as follows:
Correctability / staying correctable — corrigibility; fallibilism; the standing ability to be shown wrong by reality.
BABL (Blindly Assuming Blind Leveraging) — the default failure loop: technical debt, Goodhart’s Law, diminishing returns on complexity.
OSCR — the death-trifecta: over-simplification, over-complication, over-reach.
ZION — the inverse construction discipline: Zoning, Investigating, Organizing, Navigating (error-surfacing built into construction).
Jubilee — scheduled, non-optional institutional renewal / maintenance.
h_star / h_dark / h_zero — the stabilizing decision / the drift when correctability is abandoned / the binding commitment that interrupts it.
MADI — the open, testable function of actively reducing self-destruction risk by verifiable means.
Reality (capital R) gets the last word — empirical adjudication; the methodological premise of all science.
This overview is offered the way the framework says such things must be: with open hands, expecting correction and hoping to be improved. If the argument breaks, the reviewers will have done everyone a service. If it holds, the property it describes — correctability — is one worth building into the most powerful systems we have, while there is still time to build it in. Don’t believe it. #AuditTheMath