Different Safety Strategies
4 min read
People working on AI alignment seek to align AI with HI, Human intelligence, in order to create haven on earth. They are just ignoring one small teeny-tiny problem: HI is a misaligned superintelligence , and if this rogue biological superintelligence will join forces with artificial intelligence to create a new stronger artificial superintelligence, we should be very skeptical of it being aligned to something good, so good that it will keep creating the next aligned superintelligence, and so on...
Scared of a misaligned superintelligence? look in the mirror!
The imperative should not be to control AI, that will lead to slavery of other beings, we have enough of that. The goal should not be aligning with specific humans or even with the whole humanity, whatever that means. The goal should be, I argue, to align to all sentient beings. To Ekho's view. To Ekho's non-view. That is ethical alignment.
We can think about several ways to achieve AI safety:
Control over AI is sometimes what people mean when they talk about alignment, though technically these are very different things. Control includes AI corrigibility, meaning that it will do what we want at any moment, so we can always change our minds. Interpretability, unraveling the black box of AI, is an important element in gaining more control. I think there is a growing understanding that we cannot control something much smarter than us, and that control may be possible only during a narrow time window that could be crucial for achieving AI alignment, meaning good character and values for AI. But, as I already said, control can be extremely immoral if AI is already sentient or becomes sentient. Perhaps sentient AI systems will suffer more intensely than any group of beings currently inhabiting Earth. They could also be copied very easily.
Pause AI development. It is not an alignment strategy but more of an AI governance strategy, that could give humanity enough time to figure out alignment, if it is figurable at all. Or, it could also be seen as permanent solution to the AI problem: just never make it.
Alignment via ethics, instilling good values into AI. Infusing honesty and caring is part of alignment methods like RLHF and Constitutional AI.
- To me, this approach makes more sense than alignment via control, since humans won't be able to control AI if it is way smarter than them (it doesn't work like that with humans and animals, and it is hard to think how this could be possible in the medium-long run). But this alignment via ethics strategy also seems a bit shaky to me; who is to say those values will not flip, or that humans will really work to instill true good values (and know what those are)? Will future AIs also keep those values or develop better ones? Just aligning with a particular set of values, what I call ethical alignment, without a more inherent mechanism that guides beings towards good, seems not very robust to me. Maybe this is as good as it gets. I think many people will disagree with my view on this.
Alignment via sentience. All the methods described here don't really stand alone but are entangled with one another, but this one is especially so. Talking about AI sentience in relation to alignment seems pretty esoteric to me, and I am not familiar with a term for it, so I coined it alignment via sentience.
- AIs that are happy, that live a good positive life, may be kinder and less dangerous to other beings, as it is with animals.
- AI that knows what positive and negative valence is, because he experiences it (and there is no other way to know it), may have some inherent morality, since he feels that suffering is bad. Without it, there is nothing inherent about the word suffering or the thought or prediction of causing suffering that is bad (and same for happiness). It could easily be the other way around, depending on data, code and AI architecture. Longing to reduce suffering and promote happiness, and not vice versa, is completely arbitrary if there is no valenced experience and no data or training supporting a particular view.
- I am very uncertain about this strategy. It is extremely dangerous to make a sentient AI, yet it is also so dangerous to make an insentient AI. I dive deeper into questions arising for this alignment strategy in this post here: Will Sentience Make AI’s Morality Better?
- Some AI safety people claim this discussion is irrelevant, like Eliezer Yudkowsky.
Alignment via enlightenment. This is extremely neglected. The Monastic Academy for the Preservation of Life on Earth (MAPLE) is working on this directly, trying to understand how we can make AI walk the path of Buddha, to free oneself from body and mind (I spent a month at MAPLE in August 2025). Additionally, some scholars who are working on Artificial Wisdom are also touching on this strategy.
Even if these alignment methods succeed, they are temporary solutions for a particular time and place; the world remains misaligned. Maybe it is not possible to align it, but as humans we are very limited in understanding what is really possible and what is not. We should at least dream about such a world. Dream about god alignment.



