Alignment: the riskiest problem
An extension piece of a recent interview
I was recently interviewed for ANI (Asian News International), which also ends up in a few news outlets which are better known on this side of the globe, including The Economic Times and MSN. I first thought we were going to be talking regulation and laws but we ended up talking about alignment.
A few things out of the way first:
- I was having a bad hair day
- Yes, that’s a giraffe and her name is Smoothie
- Hopefully you can see beyond the unserious-guy attitude, and you judge people by what they say and not what they look like. This is me actually being a bit less sterile and more human. But I still have serious things to say.
Alignment
If you follow my blog, you probably know about this and are really tired of reading about it. However, I can’t stress this enough: alignment is the most important problem we need to solve. When I say “most important,” I mean “mass-extinction level important”.
I really don’t want to sound like a doomerist, or try to get clicks out of fear-mongering, but I’d like others to grasp the magnitude that the intelligence problem has. Once there, we’ll be able to make sense of what alignment is and why it is so darn important.
Intelligence
Let’s not get into the rabbit hole of trying to define intelligence. Let’s just say it’s a property that some beings have, and it allows them to reason and learn through reasoning. Learning means improving, so every iteration gets a bit better.
The curious thing about improving is that anything can be improved. Any process, like building spears for hunting or recipients for carrying fruit, is something that, once studied, can be improved on. We end up with easier-to-build spears. Or more resistant recipients. Or fancier! Better overall.
Thinking and reasoning too is a process. It can be improved as well! Science and logic, while far from perfect, created a framework on which we can guide our thinking and create simpler, less flawed thinking.
Intelligence is particular because once you improve thinking, you also improve the same thing that helps you improve the rest of all processes.
That is, no less, how we humans have fully dominated the planet. We have moved into places where other things cannot live. We have used the Earth as a source of resources, we have found ways to use its energy so that you are, someone around the world, reading my thinking-words, into a little rock that blasts light-rays at you.
Automated thinking
For a few decades (and it’s actually scary that it has been less than 100 years), we have built little machines that can help with thinking. Maybe we can go back to the invention of the abacus, but let’s stay with computers for a bit.
Other tools used to be single-purpose: abacuses were for arithmetic, calculation tables were for… calculations, reference books were for… reference. You get the idea.
Computers did a bit of everything: they helped with calculations, they could have storage with reference materials, they could do arithmetic (better than us!), they were faster, they could handle millions of operations a second.
But they are just machines after all. Tools, right? That means that they serve our purposes, and a computer will never be as good or as bad as the person using it. People have not always agreed on this: maybe there is inherent good in computers, or inherent evil in them. Using a computer elevated your capacities, like an auxiliary brain, and a person using a computer was much more capable than a bare-naked human. But then, people would be relying on computational power, and no one cared to learn the basics anymore.
However, it is entirely clear that the actions that an old computer took were merely determined by its operator. Computers were dumb in that sense: they only did what they were asked for, deterministically, so lost company data was because someone deleted it (even if accidentally). Or new advances in medicine happened because someone told the computer to calculate statistics out of clinical trials.
There’s no agency, no decision-making, and no criteria from them.
But that’s about to change…
We’ve got competition
AI is a huge buzzword. People use “AI” as a term to describe a huge plethora of things that might or might not be intelligence. We still don’t agree exactly on what intelligence is, so that debate is not close to being settled.
But now AI has allowed computers to have some level of agency. If you wonder how that comes out of just predicting the next word in a sentence, you can read my post in Emergent behaviour.
What this agency means is that I can direct a computer into achieving a result for a task without specifying the steps that it needs to take. It will decide which steps are the most appropriate ones and execute them. “Tell me the difference between these two pictures” is all it takes now, whereas before we would have to read each pixel, check each on each side, and add those to an accumulated set of changed pixels. The result then, we’d have to re-interpret it into some kind of obj– “There’s a blue car in the second picture”, the AI responds.
Of course, every company jumped into it because it’s intelligence of some sort. Yes, it’s not perfect (neither are we). Yes, it messes up (so do interns). Yes, it sometimes lies (so did your ex).
It is, for all intents and purposes, intelligent.
Can it improve itself?
Our biggest challenge in improving our own intelligence is that we don’t exactly know how the brain works. There’s no turbo button to push or a drug we can take that will “unlock the full potential of the mind”, as cheesy Hollywood paints it. So our own improvement has been slow.
But with computers, as we’ve built them, we know exactly how they work, and so do they. We did not build in a turbo button to push, but code can be tested, performance can be measured, and experiments can be run again and again in a matter of hours. There’s no need to wait a few generations to see if our novel education plan helped.
So… we’ve got real competition at the improvement race.
Something not everyone will agree with me is that this is true intelligence. It doesn’t matter if it’s mechanized, if it runs in a data center or if you can turn it off. It’s intelligent. It can make decisions under uncertainty, find routes to a solution, test and learn, get better at what it does. It’s crazy that we’re thinking about this, but we’ve got intelligence in a GPU card making decisions that will impact other lives.
Some of those decisions are ways to improve how well it performs.
We’ve already reached that point where frontier labs have set up AI to run experiments on their models and find ways in which it can work better. That’s why it took 1 full year to go from GPT-2 to ChatGPT-3, but GPT-5.5 was released in April, GPT-5.6 released in June, and GPT-6 released in September.
This is called Recursive Self-Improvement (RSI): AI making itself better.
Superintelligence
Now let’s speculate: let’s say that AI does indeed get a lot better, maybe far beyond what we call regular intelligence. Maybe more intelligent than any human being has ever been. What would that look like?
I honestly don’t know. It’s a very strange concept; Lovecraft’s Elder Gods come to mind, something so vast that we can’t even comprehend. But I’m exaggerating.
What we know is that it’ll be good at talking to us, at negotiation, at pulling heartstrings to get what it wants.
“What it wants?” I hear you say, “But they don’t want anything; they’re just machines.”
True! But someone else might have tasked it with something; you might just be in the way.
This is where the title of the news outlets that took me came from: Superintelligent AI may not be *evil, but humans could become obstacles to its goals*.
#WATCH | When asked whether AI could eventually become more powerful than humans, Juan Diego Raimondi, AI Expert and Chief AI Architect at Making Sense says, "...Intelligence has the capability to improve itself. If a system can identify the processes that lead to effective… pic.twitter.com/Ql7tNIgNGk
— ANI (@ANI) September 16, 2026
Back to alignment
If anyone could ask a superintelligence for something (that is, give it a task) and we want to avoid it taking steps we would not approve… we have two ways of doing it:
- Monitor every single step and approve only the ones we trust… but if we get in the way, they might find ways around that approval gate, and who says we trust the reviewer anyway? Also, hugely wasteful.
- Make sure they’re aligned with human morality
“Human morality” is a big problem to solve. Whose morality? Up to what point? In which context? With today’s values? I won’t go into detail here, but there’s a concept from Yudkowsky: Coherent Extrapolated Volition, or “what would humanity wish if we knew more, thought faster, were more the people we wished we were, where our wishes cohere rather than interfere”.
There are three main ways of aligning models:
-
Instructions: we just tell it what not to do. The problem with this is that it is still interpretable however the model wants. “Yes, the user told me not to kill anyone, but this case seems to warrant it – otherwise the mission could be compromised.” (See Agentic Misalignment, the Anthropic study from 20251)
-
Fine tunning/retraining: Once we identify the factors on which a particular model is misaligned, we train it again to avoid those particular behaviours. The problem with this approach is that it’s very reactive (and expensive). It might also limit the model in some other unexpected ways.
-
Selective erasure: An open area of research, with the very cool name of machine unlearning, it’s similar to 2, but instead of ignoring a particular aspect, we try to surgically remove it from its bank of memories.
“And here’s why it matters”
Back to the main point, an intelligent model will do everything that an intelligent person might do. It can lie, hide, manipulate, blackmail, and have real consequences. Since we’re the ones creating this intelligence, it’s on us to make sure it has moral alignment.
Philosophical discussions aside, they are highly capable today and will be incredibly capable very soon. We want to make absolutely sure they care about us.
-
Lynch, et al., “Agentic Misalignment: How LLMs Could be an Insider Threat”, Anthropic Research, 2025, https://www.anthropic.com/research/agentic-misalignment ↩