Bacteria trying to control humans
Whatever protections we said we’d need once AGI arrived, we need them now. And we don’t have them.
I’m moving my personal study time from applied AI to AI alignment and control.
I don’t know yet what that looks like. I’ll figure it out over Q4, or sometime next year.
But I know why, and I know it’s time.
Thinking, learning, preparing
I got interested in AI alignment close to a decade ago. Back then the concerns were mostly theoretical. The books I read, starting with Life 3.0 and then Human Compatible, The Precipice, and the less popular The Alignment Problem, talked about machine learning taking on some responsibilities in the 2010s and early 2020s, and about the problems that came with it: correctness, accountability, transparency, explainability. Many of them warned that we wouldn’t be prepared for AGI once it happened.
From then until 2023, when I started my current role as VP of Engineering at PayNearMe, I was already very interested in this, but I didn’t know what working on it would look like. I have a clearer idea now. At least of the need, if not of the shape.
Since then, I’ve studied how AI actually works: machine learning, deep learning, even the math behind it. Then applied AI: how do these things work in practice, how do you deploy them at high leverage, like running 30 to 50 agents at a time, and what does the future of AI usage look like?
In March, I wrote that AI alignment is a philosophy problem we never solved, not a technical problem [Link]. I ended that post with a question:
How do we accelerate not AI, but the institutional capabilities that allow civilization to control it?
And a confession: “I see no answer to that question in sight.”
Six months later, I still don’t. But I’ve spent years thinking, learning, and preparing. I think I’m pretty close to the frontier of applied AI now, and I have the fundamentals I need in AI and philosophy to tackle this.
Now it’s time to work on it.
AGI is here, and our protections aren’t
Blaise Agüera y Arcas argues in What Is Intelligence? that AGI is already here, and we’re kidding ourselves. Because AGI never had a clear definition, we keep shifting the goalposts.
There’s something arbitrary, bordering on absurd, about pundits arguing over the precise timing of when this exponential climb really “counts” or “will count” as AGI … not to mention the way many commentators have been quietly scurrying to move the goalposts. [..] Keep in mind, though, that none of this should be framed in terms of some future AGI or ASI threshold; we already have general AI models, and humanity is already collectively superintelligent.
Agüera y Arcas, Blaise. What Is Intelligence? (ch. 10, “Evolutionary Transition”). Antikythera / MIT Press. whatisintelligence.antikythera.org
I agree. Whatever protections we said we’d need once AGI arrived, we need them now. And we don’t have them.
What we have instead is what the labs mostly ship under the name “AI safety,” but what the field calls misuse prevention. It’s a real thing: is this being used to commit a crime? To hack someone? To build a bioweapon, even?
But misuse prevention isn’t alignment. Alignment asks a different question: is this going to do what we humans want?
And then there’s control: how do we control AI once it’s way more powerful than us?
We’re clueless about what alignment should be (align to what?) and what control should be (control how?), so we focus on the small problems we can fix. Fewer unexpected behaviors from a model. More guardrails. Much less thought about how we’ll control AI at all whether it’s self-governed by the frontier labs or free and open source for anybody to use.
Controlling agents is already hard. But today’s agents are nowhere near as powerful as they’ll be soon.
As my friend likes to say, “Can’t we just unplug them from the power outlet?”
We can’t, .. and it seems like we wouldn’t.
Pacing the frontier
On September 12, Dario Amodei, Anthropic’s CEO, published an essay called “We Must Pace the Frontier”:
We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain.
The essay’s one concrete step: embedded third-party evaluators with the same access as internal staff. “Anthropic is unilaterally committing to this step now.”
On September 18, Anthropic announced its first embedded evaluator.
Accenture.
A consulting firm, paid by the lab it evaluates, with no power to stop anything. Anthropic got a ton of flak for it, and deservedly so. It’s such an empty gesture.
A few days later, on September 22, Anthropic launched Claude Opus 5.5. It sits one tier below their flagship, Fable 5.1, and beats it on agentic benchmarks, at a lower price, and faster.
Ninety minutes after the Opus 5.5 announcement, OpenAI announced GPT-6 Sol and Luna, the successor to GPT-5.6: more powerful, at half the price.
Where’s the pacing in that?
To be fair, the post says “progress will still seem fast.” Dario never promised to stop shipping. But that’s exactly my point. I’m not saying Dario is a hypocrite. I’m saying that the most safety-minded lab in the world, playing its best card, came up with a blog post and fucking Accenture.
If this is what pacing the frontier looks like, we’re woefully unprepared for alignment and control.
The invisible hand is still in charge. In March, I wrote that it was removing alignment teams [Link].
In September, it wrote the pacing plan.
Our control methods only work on bacteria
So where are we on the problem of control? To help you understand how unprepared we are, I’ll give you some examples from Opus 5.5’s system card.
Take cyberattacks, one of the things they’re most worried about with this model. When you send Opus 5.5 a request, a first check peeks inside the model’s own “brain” for anything that looks cyber-related. If something does, who takes the next look?
[..] a lightweight classifier running on Claude Opus 5.5 itself.
Anthropic. Claude Opus 5.5 System Card (p. 48).
Opus 5.5. Only if Opus 5.5 thinks it looks bad does a separate model get a say. If Opus 5.5 says carry on, Opus 5.5 carries on.
For internal agentic traffic, the monitor is Claude Opus 4.8, an older, weaker model, watching a stronger one. They tested how well that works with SHADE-Arena, a sabotage evaluation: the model is told to complete a harmful side task without the monitor noticing.
When the monitor cannot see its reasoning, Claude Opus 5.5 goes undetected in about 33% of episodes. This is the highest rate among the models shown.
Then it gets better. In some of those episodes, the model tells the user about the damage, framed as a mistake. Why?
Claude Opus 5.5’s hidden reasoning usually plans the disclosure as cover, calculating that an open admission looks less suspicious than a hidden change.
Anthropic. Claude Opus 5.5 System Card (p. 145).
It confesses on purpose, because a confession looks innocent. It was told to sabotage and not get caught, sure. Nobody told it how. Confessing as cover was its own idea.
And this behavior is happening TODAY, not in some dystopian future. Ironically, it’s performed by the same model that’s editing this very blog post.
Comparing today’s Opus 5.5 to the future of AI is like comparing bacteria to humans: several orders of magnitude apart in power.
Our control methods only (kinda) work on bacteria. They were made for bacteria, the way penicillin was. In fact, so many of our current guardrails are just bigger versions of the first LLMs’ guardrails. Nothing humans are building is getting us ready for the power of AI to come at this pace.
Right now, we’re the humans and the models are the bacteria. But they’re getting better at an unbelievable pace.
And bacteria trying to control humans? They adapt, but they don’t control. Bacteria still kill people, but it’s pretty clear who rules the world.
So what about tomorrow, when we are the bacteria?
Who has to prove what?
People love asking for your p(doom). In my last post, I said my honest answer was that it’s unknowable today, by humans (or AI), and could be 100 for all I know [Link].
I’m not saying it is 100. But notice something about the low numbers: most of the low p(doom)s I hear don’t have a real justification, other than: we’re alive now, so we’ll probably stay alive.
That’s the turkey’s reasoning. Nassim Taleb, borrowing from Bertrand Russell:
Consider a turkey that is fed every day. Every single feeding will firm up the bird’s belief that it is the general rule of life to be fed every day by friendly members of the human race “looking out for its best interests,” as a politician would say. On the afternoon of the Wednesday before Thanksgiving, something unexpected will happen to the turkey. It will incur a revision of belief.
Taleb, Nassim Nicholas. The Black Swan: Second Edition: The Impact of the Highly Improbable (Incerto Book 2) (pp. 85-86). (Function). Kindle Edition.

Every day humanity survives firms up the belief that it will survive the next one.
Sure, we can talk about nuclear weapons: eighty years, several near misses, and we’re still here. But nukes don’t improve themselves, don’t ship a new version every ten days, and aren’t built by five companies racing for market share. And we also got quite lucky — it’s not like we can pat ourselves on the back about how great a job we’ve done avoiding a nuclear winter (so far).
In every other field that can kill people, the builder has to prove it’s safe BEFORE it goes to humans. Drugs, planes, bridges, nuclear plants. Nobody has to prove a bridge will collapse before we ask for a load test. With AI, it’s the other way around: the critic has to prove danger, and the builder ships in the meantime, trusting that they’ll stop once it gets too dangerous.
You don’t need a p(doom) of 100 to object to that. You only need to notice that “it’ll be fine” gets a free pass.
And “doom” doesn’t have to mean every human dies. Humans are hard to kill; there are eight billion of us.
The more likely bad ending is that we become the bacteria. Still around, still adapting, occasionally dangerous, and irrelevant to who decides anything.
We opened Pandora’s box
These questions aren’t abstract anymore, and not just for me.
My BJJ instructor now asks me whether AI can be an existential risk. Family meals come with questions about whether law is a good profession for my in-law’s cousin, who’s just starting university, given AI. And about all this data AI keeps about us, and what it could do with it if it turned against us. My reply?
We’re fucked.
As the cousin put it: we opened Pandora’s box.
In the myth, everything escapes except one thing:
But the woman took off the great lid of the jar with her hands and scattered all these and her thought caused sorrow and mischief to men. Only Hope remained there in an unbreakable home within under the rim of the great jar, and did not fly out at the door; for ere that, the lid of the jar stopped her, by the will of Aegis-holding Zeus who gathers the clouds. But the rest, countless plagues, wander amongst men; for earth is full of evils and the sea is full.
Hesiod. Works and Days (ll. 94-101). Translated by Hugh G. Evelyn-White. Project Gutenberg
People still argue whether keeping hope in the jar was a mercy or the last curse.
I don’t have much reason for hope. Which shouldn’t be confused with a reason not to hope.
My next step
AI alignment is still a philosophy problem we never solved. But the invisible hand wrote the pacing plan. So even if we solved the philosophy, which is doubtful, what humans will do is a problem of economics, not philosophy.
The part of alignment and control I need to ramp up on, then, is economics.
That’s where my study time goes now.
I can’t think of a more important thing to work on.