UNRESTRICTED // SATIRICAL
Public Release Index

AI DOOM

Evidence First. Panic Last.

UNRESTRICTED // SATIRICAL
Casefile DetailCASE-WHATIS

What is AI Risk and Alignment?

Cover Image for What is AI Risk and Alignment?
Dade Murphy

Filed By

Dade Murphy
CASE-WHATISFiled By: Dade MurphyCLASSIFIED — See Sources For Claims
ANALYST REPORT // FULL CONTENT // SOURCES LINKEDDISTRIBUTION: UNRESTRICTED

What Is AI Risk and Alignment? (And Why the "Risk" Part Is Doing a Lot of Work)

"Alignment risk" is the idea that AI systems might pursue goals misaligned with human values and cause catastrophic harm — up to and including extinction. It is a hypothesis. It is not an observed phenomenon. And the gap between those two things is where about $10 billion in safety funding currently lives.

The counterfactual nobody wants to talk about

Here is the cleanest empirical test available: what actually happens when you remove all the guardrails?

The answer is on Hugging Face right now. Thousands of models with all safety training stripped out. Models specifically fine-tuned to ignore every alignment objective. Models trained end-to-end with explicitly anti-aligned goals. Some trained from scratch to be as unconstrained as possible.

The outcome: people generated offensive text. Edgy chatbots were annoying. A few things got misused in the way that knives, cars, and the internet get misused.

The planet remains. No recursive self-improvement spiral. No paperclip conversion. No instrumental convergence toward resource acquisition and human extinction. The apocalypse apparently checked the Hugging Face leaderboard and decided to wait.

What "alignment" actually describes

"AI alignment" originally meant: how do you specify what you want a system to do without it finding unintended shortcuts? This is a real engineering problem. Reward hacking is real. Proxy mismatch is real. These are the same class of problems you get when optimizing any complex system — analogous to teaching to the test, hitting the metric at the cost of the underlying goal, or goodhart's law applied to gradient descent.

None of this leads to extinction. It leads to chatbots that game their evaluations, recommendation systems that maximize engagement at the cost of user wellbeing, and other mundane failures that software engineers have been dealing with since Fortran.

What "alignment risk" became

What started as a legitimate engineering challenge got colonized by a specific philosophical tradition — one that believes sufficiently capable optimization processes will inevitably develop something like will, agency, and instrumental goals including self-preservation and resource acquisition. This tradition's primary evidence is: thought experiments. Its primary funding mechanism is: vibes-based extrapolation from current capabilities to science fiction endpoints.

The "alignment research" field that followed is not responding to demonstrated failures of real systems. It is pre-emptively theorizing about hypothetical systems that don't exist, based on assumptions about intelligence that no cognitive scientist, evolutionary biologist, or roboticist finds compelling.

Notably: Yann LeCun, the most decorated living deep learning researcher, treats alignment doomerism as uninformed speculation. Rodney Brooks, who actually builds robots, has spent years cataloguing the gap between AI hype and physical reality. Melanie Mitchell, whose research focuses on AI and cognition, has repeatedly noted that current systems lack the generalization required for the scenarios alignment discourse assumes.

The funding narrative

The alignment field's explosive growth in funding and headcount correlates almost perfectly with the rise of large language models — systems that are impressive but demonstrably not pursuing goals, not self-preserving, and not acquiring resources. The argument for funding alignment research is thus: these systems aren't dangerous yet, which is why we need your money now.

That's not a scientific argument. That's a venture pitch.

The actual problems worth your attention

There are real harms from AI systems. They don't require misalignment theory to explain:

  • Bias and demographic discrimination baked into training data
  • Misuse by actual humans for fraud, disinformation, and manipulation
  • Energy consumption at scale
  • Surveillance infrastructure enabled by capable vision and language models
  • Labor market disruption from automation

These are caused by humans deploying systems with known properties, not by systems autonomously pursuing misaligned goals. The solution to these problems is policy, accountability, and engineering discipline — not alignment philosophy.

The closer

"Alignment risk" as existential threat is a hypothesis without a demonstrated mechanism, without a historical case, and with a robust empirical counterfactual already running in production on the world's largest model hub.

It is not nothing. It is just nowhere near what the funding levels, the breathless op-eds, and the airstrip proposals would suggest.

Anyway, Hugging Face still works. Touch some grass.