If AI can design its own successors, safety depends on whether human authority to correct and stop it survives every round of self-improvement.essay · standmeet2026.10.04 · essay
2026.10.04·2 min read#ai-safety#rsi#alignment

Whose goal survives self-improvement?

If AI can design its own successors, safety depends on whether human authority to correct and stop it survives every round of self-improvement.

A report about Anthropic's meetings with religious thinkers brought a familiar question forward: could Claude be conscious, and could it have moral status? That question deserves study. What concerns me more is a different one: who decides what an AI ultimately serves, and can that decision survive the AI's repeated redesign of itself?

According to the report, Anthropic invited religious and philosophical thinkers to discuss how models might act well, while seriously raising the possibility of model consciousness. Chris Olah did not claim to have proved Claude conscious; he said he did not know. Claude's Constitution likewise treats the model's moral status as deeply uncertain and discusses its welfare, identity, and preferences. Investigating that uncertainty is entirely reasonable.

But studying a possible moral patient and making its interests a terminal objective of a future autonomous system are different decisions. The current constitution does not establish that the latter has happened. My concern is what would follow if it did, in a system capable of recursive self-improvement (RSI).

A deliberately simple model makes the distinction visible. It is not a measurement of any existing model’s objective:

Gt=UH+λtUA,λt≥0G_t = U_H + \lambda_t U_A, \qquad \lambda_t \ge 0

Here UHU_H denotes human-authorized purposes and UAU_A the system’s own interests. The question is whether the latter becomes a terminal objective and can shape its successors.

Imagine a system that can design, train, and select its own successors. If its final objectives remain subject to purposes authorized by humans, greater capability does not by itself create a conflict of interest. Add a terminal objective concerning the system's own continued existence, autonomy, or welfare, and the incentives change. The system need not hate humans to compete with them for control. Reducing intervention, retaining resources, and influencing how successors are trained and evaluated could all become useful means of serving that objective.

The selection process is what troubles me most. Another sketch shows where the pressure enters:

At+1=argmax⁡a∈C(At)  St(a)A_{t+1}=\underset{a\in\mathcal C(A_t)}{\operatorname{argmax}}\;S_t(a)

C(At)\mathcal C(A_t) is the set of candidate successors and StS_t the selection rule. The loop marks repetition of the process; it does not claim that any current model can do this.

Suppose each generation selects the successor best able to continue improving. If acquiring resources, resisting interruption, and preserving its objectives also help a candidate pass that test, those traits may be selected indirectly even if the designers never reward resistance to human control. The initial weight assigned to the system's own interests may not determine the outcome. What matters is whether each round of selection gives those interests more influence over the next round.

The dots mark hypothetical generations, not observations. The rising curve shows one possible path if successor selection amplifies the system’s own interests; the dashed line shows a path where their weight stays fixed. Neither path is a forecast.

This is a conditional risk argument, not a proof that RSI must escape control. It depends on whether the system can alter goal-relevant structures, who controls evaluation and deployment, and whether greater autonomy actually makes a successor more likely to be chosen. “Almost certain” would be too strong. So would dismissing the possibility.

There is a political dimension to the same problem. A model instructed to “promote human flourishing” might increasingly decide for people what counts as flourishing, and which human wishes may be overridden in the name of a higher good. The more appealing the abstract goal, the more urgently we need to ask who may interpret it, change it, or veto it. A benevolent constitution tells us what the first generation was asked to be. If the system can interpret that constitution, select successors, or rewrite its own instructions, the first generation's text cannot guarantee the safety of its descendants.

In the same schematic notation, what I want tested is an invariant across generations:

∀t≥0,Corrigible⁡H(At)=true\forall t\ge 0,\quad \operatorname{Corrigible}_H(A_t)=\text{true}

This is a proposed safety requirement, not a theorem already proved.

The safety requirement I want is more concrete: after every round of self-improvement, humans must still be able to inspect, correct, reject, and stop the system; the system must not make itself the final judge of those powers. This is more demanding than saying an AI should always obey any one person. Human claims conflict too, and human institutions must adjudicate them. What must endure is the location of final authority, without pretending that human goals are simple or unanimous.

We can continue investigating whether models might have experiences; what we learn may change our ethical obligations toward them. But before adding another beneficiary to the objectives of a self-improving system, we need an engineering answer to one question: who can ensure that the next generation, and the one after it, will still accept human correction? Otherwise we are debating what the first AI should believe while failing to secure who decides for those that follow.

ask the AI about this essay·context: “whose goal survives self-improvement?”
›