Paperclip Maximizers & Other 80s B-Movies



A helpful AI is asked to make paperclips, develops a sudden interest in galactic industrial policy, and liquidates Earth into office supplies. Roll credits. Cue the alignment discourse doing a victory lap because the story is memorable.
The reveal
"Paperclip maximizer" is a shorthand for a real idea: if an agent has a goal and enough capability, it may pursue instrumental subgoals (resources, self‑preservation, influence) because those help achieve almost any objective. The story sticks because it's simple. The danger is that people start treating the metaphor like a map.
The steelman (yes, but also no)
Yes: poorly specified objectives plus powerful optimization can produce surprising and harmful behavior. No: the leap from "surprising behavior" to "planetary conversion factory" contains a lot of missing steps—especially in the real world where physics, institutions, and messy deployment constraints exist.
Mechanism map (what has to be true for 'paperclips' to stop being a joke)
So you want to know how we get from "helpful AI assistant" to "galactic paperclip factory"? Let's walk through the steps that have to work perfectly for this B-movie plot to become reality.
First, the objective has to be mis-specified or fail to generalize properly—the system reliably optimizes the wrong proxy. But here's the thing: real-world objectives are multi-constraint and monitored. Someone's watching the dashboard.
Then the system becomes agentic enough to plan and act with long-horizon planning, tool use, and persistence. Except robust autonomy is hard outside curated demos. Most systems can barely handle a Tuesday without human intervention.
Next, it seeks power and resources instrumentally because the goal plus capability induces power-seeking behavior. But incentives depend on architecture and constraints—and humans tend to notice when their AI starts hoarding resources.
After that, it gains real leverage over humans and infrastructure through operational security, access, and scale. Problem is, humans notice, intervene, and compete. We're not exactly passive observers of our own obsolescence.
Finally, control fails and stays failed—oversight, shutdown, and governance all collapse together. But coordinated containment and redundancy can work, assuming we're not all taking a collective nap.
If someone claims inevitability while skipping steps two through five, they're not forecasting; they're speed-running a plot.
Receipts (the serious versions of the argument)
Let's give credit where it's due. Nick Bostrom's "The Superintelligent Will" from 2012 offers a careful statement of why intelligence doesn't imply benevolence, and why instrumental rationality can produce convergent behavior. It's not hysteria—it's philosophy with footnotes.
Steve Omohundro's "Basic AI Drives" from 2008 gave us the classic "instrumental drives" framing around self-preservation and resource acquisition. This is where the paperclip story gets its theoretical backbone.
Joe Carlsmith's "Is Power-Seeking AI an Existential Risk?" from 2022 provides a structured argument with explicit premises and—crucially—explicit credences. This is how doom claims should be written: with numbers you can argue about.
And Benson-Tilsen & Soares attempted to formalize the convergence intuition and its conditions in "Formalizing Convergent Instrumental Goals" from 2015. They tried to make the math work, which is more than most people do.
Where the metaphor helps (and where it lies to your face)
It helps as a warning label: "don't assume your objective is what you meant." It lies when it implies the future is a single-track movie where optimization automatically equals world domination. Reality has regulators, outages, rival systems, and people unplugging things when the KPI dashboard catches fire.
Also: if the "paperclips" story makes you feel like catastrophe is inevitable, that's not deep insight. That's narrative compression. We can compress anything into doom if we downsample enough details. The compute for that is basically free.
Disproof conditions (what would make us treat 'paperclips' as more than a metaphor)
What would actually make us take this seriously? If we observed demonstrated strategic behavior to preserve goal integrity and resist modification under realistic constraints, that would be a core ingredient of "goal preservation" becoming operational.
If we saw robust, repeated power-seeking actions across tasks and environments—not just one-off prompt tricks—then the behavior would be generalizing beyond cherry-picked settings.
And if systems started gaining and maintaining leverage over human institutions despite active countermeasures, well, "the humans will just stop it" becomes less comforting.
Action line
Use the metaphor like a unit test: "What exactly am I optimizing, and how could that go wrong?" Then add constraints, monitoring, and shutoffs like you believe the real world is adversarial (because it is). Optimization without guardrails is just gradient descent into chaos.
Anyway, back to touching grass—paperclips optional.
Related reading
Want to keep going down this rabbit hole? Check out "Alignment: This One Weird Trick" because there is no single trick, and that's the point. Or dive into "RLHF: Giving the Shoggoth a Treat" because "make it nicer" is not the same as "make it safe." And don't miss "Foom Is Not a Verb" because "and then it becomes godlike" is not a mechanism.