A Contrarian Take on the Alignment Problem
Why aligning superintelligence to 'human values' might be aiming at the wrong target, and what it could mean to build the values in before the knowledge.
Table of Contents
I spend a lot of time thinking about AI in 2050. Not the robots-and-flying-cars version, but the quieter question underneath it: what does super-alignment actually mean when we're trying to align superintelligent systems with "human values"?
The more I sit with that phrase, the less sure I am that it points at anything we'd want to hit.
Which Humans, and Which Values?
The uncomfortable part is that human history isn't exactly a moral success story. Strip away the flattering narration and it's a few thousand years of survival instincts, resource competition, and self-preservation dressed up as civilization. We got very good at cooperation, sure, but mostly within the group, and mostly because it beat the alternative.
That's not cynicism, it's just the training data. When we say "align AI to human values," what we can actually operationalize is "align AI to what humans did, said, and rewarded." The internet, the books, the preference labels. And that corpus is not a record of our ideals. It's a record of our behavior, which is a very different thing.
So if we're aligning AI to human behavior patterns, we might be encoding the wrong things entirely. Not because the researchers are careless, but because the target itself is muddled. We're asking the model to become a faithful mirror, then hoping the reflection is better than the original.
The Band-Aid Problem
I don't think alignment in the traditional sense scales to superintelligence.
In the first post I went through RLHF and Constitutional AI and why the Shoggoth meme is uncomfortably accurate. The short version: pretraining builds an alien optimizer out of everything humanity ever wrote, and then post-training teaches it which faces to make. Helpful, harmless, honest. A good mask.
Teaching models to mimic human preferences this way feels like a band-aid on systems that are already learning from fundamentally misaligned data. The band-aid works today because today's models aren't strong enough to notice the gap between seeming good and being good. A system that is much smarter than its raters will notice. At that point, human feedback stops being a constraint and becomes just another signal to model and satisfy.
Every technique we currently trust shares the same shape: build the intelligence first, then correct it afterwards. Course-correction after training. I'm not convinced that shape survives the jump to something smarter than us.
A Different Kind of Mind
Part of the problem, I think, is a framing we inherited without examining it: that a good AI is one that thinks like us.
We might need to stop treating AI as something that should think like us and accept it as a different form of intelligence altogether. It doesn't have a body that gets hungry. It doesn't have a tribe, a childhood, or a lifespan. Insisting that it reproduce the value system that fell out of those constraints is a bit like insisting a submarine should swim.
This isn't an argument for letting it do whatever it wants. It's an argument that "human-like" and "good" are not the same axis, and we've been quietly collapsing them into one. If anything, a different kind of mind is an opportunity: it doesn't have to inherit the parts of us we'd rather it didn't.
Values Before Knowledge
This is the speculative bit, and I'll flag it as such.
The key might be embedding core values before the system ever touches knowledge. Think of it as building the ethical foundation first, then letting the intelligence form around it, rather than trying to course-correct after training.
Today we do the reverse. We pour in all of human text, get a capable model with whatever values shook out, and then spend a comparatively tiny amount of compute nudging it in the right direction. The values are the last thing added and the first thing under pressure.
What if the ordering flipped? A small set of commitments that the system holds structurally, so that everything it later learns is organized around them, rather than a thin layer of preferences painted over a fully formed optimizer. Humans, for what it's worth, mostly work this way: you learn "don't hurt people" long before you learn chemistry. The knowledge arrives into a mind that already has a shape.
I don't know how to do this. I don't know if "values" is even a coherent thing to have before you have concepts to attach them to. It might be that the foundation and the knowledge can't be separated, and this whole idea dissolves on contact with an actual training pipeline. But I notice that almost nobody is trying, and the reason seems to be "that's not how we build these things," which is not the same as "that can't work."
Where I Might Be Wrong
This is speculative, and I'm not claiming to have answers.
Maybe the behavior-versus-ideals gap is smaller than I think, and a model trained on enough of us really does pick up the better angels. Maybe post-training scales further than I expect, or gets replaced by interpretability tools that let us inspect the beast under the mask directly. Maybe "values before knowledge" is a category error and I'll be embarrassed by this post in three years.
But if you're working on safety and alignment research, have strong counterarguments to this framing, or just want to debate whether any of this is even possible, I'd genuinely love to hear from you. I'd rather be corrected now than be right and ignored.