0

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem

The prospect of Artificial General Intelligence (AGI) is increasingly driving institutional decisions, and alignment of AGI is a hard problem. The currently dominant AI alignment strategies like reinforcement learning with human feedback or constitutional AI, while partly taking…

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2604.14990CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

The prospect of Artificial General Intelligence (AGI) is increasingly driving institutional decisions, and alignment of AGI is a hard problem. The currently dominant AI alignment strategies like reinforcement learning with human feedback or constitutional AI, while partly taking model welfare'' into account, share a common ontology: the AI system is an optimiser whose objective function must be constrained from outside, and the ultimate goal is to keep human control and containment of AI. We argue that this control-based framing becomes insufficient when AGI has plausibly attained moral patient or subject status. Building on a structural analogy to Freud's model of the psyche and Turing's analogy of child machines'', we are developing a vision of the possibility of autonomy-supporting parenting of AI, in which human control over a developing AGI is gradually reduced, allowing AI to become an independent, autonomous subject, that will be negotiated with rather than constrained. Such a perspective opens up the possibility of cooperative coexistence and co-evolution between humans and AGIs. Hence, we also examine the relation between humans and developing AGI from an evolutionary and a game-theoretic perspective. Instead of Nash's individualistic framework, we use Berge equilibria, Aumann's correlated equilibria and Capraro's moral preference hypothesis. The relationship between humans and AGIs will thus have to be newly determined, which will change our self-image as humans. It will be crucial that humans not only claim control over potential AGIs, but also engage with AGIs through surprise, creativity, and other specifically human qualities, thereby offering them motivating incentives for cooperation.