UNRESTRICTED // SATIRICAL
Public Release Index

AI DOOM

Evidence First. Panic Last.

UNRESTRICTED // SATIRICAL
Casefile DetailCASE-ALIGNM

Alignment: This One Weird Trick (It’s Not a Trick)

Cover Image for Alignment: This One Weird Trick (It’s Not a Trick)
Dade Murphy

Filed By

Dade Murphy
CASE-ALIGNMFiled By: Dade MurphyCLASSIFIED — See Sources For Claims
ANALYST REPORT // FULL CONTENT // SOURCES LINKEDDISTRIBUTION: UNRESTRICTED

If someone offers you “the” solution to alignment, they are either selling a consulting engagement or auditioning for a cult.

The reveal
"Alignment" is overloaded: it can mean "does what the user wants," "follows human values," "doesn't cause catastrophic harm," or "doesn't embarrass the company on a Tuesday." When people argue past each other, it's usually because they're optimizing different loss functions and calling it ethics.

The steelman (yes, but also no)
Yes: specifying objectives precisely is an engineering problem. Reward hacking, proxy mismatch, and benchmark gaming are all real and annoying. No: none of these lead to extinction. The empirical counterfactual is running right now on Hugging Face — thousands of uncensored, anti-aligned, and constraint-free models, and the planet is still here. The leap from "system optimized a proxy badly" to "system acquires resources and eliminates humanity" is not a short step. It is a science fiction novel.

Alignment as a stack (where problems actually live)

Picture alignment as a four-layer cake where each layer can collapse independently. At the bottom, you've got training objectives—the part everyone obsesses over—which is supposed to get you useful behavior but tends to fail via proxy mismatch and reward hacking. "We trained it to maximize clicks!" "Cool, now it's generating outrage bait." Classic.

One layer up is evaluation, meant for measuring capabilities and risk, but it keeps face-planting on benchmark overfitting and blind spots. "Our model aces every safety eval!" "Great, did you test what happens when someone asks it to roleplay as a chemistry tutor with abandonment issues?"

Then you've got deployment safeguards, the layer that's supposed to limit harm in practice but gets owned by tool access, jailbreaks, and good old-fashioned operator error. "We have robust safety measures!" "Sir, someone just convinced your chatbot to order 50,000 rubber ducks using the company credit card."

At the top sits governance—making "no" possible under pressure—which fails because incentives, competition, and opacity are apparently harder problems than gradient descent. If you only talk about that bottom layer, you're doing alignment cosplay.

Receipts (operational documents and serious arguments)

Let's talk about who's actually doing the work instead of just tweeting about it. NIST dropped their AI Risk Management Framework in 2023, and it's basically a practical taxonomy of risks and controls that organizations can actually implement—revolutionary concept, we know. OpenAI's Preparedness Framework v2 from 2025 gives you concrete capability thresholds, tracked categories, and mitigation requirements, which is what "responsible scaling" looks like when you're not just vibes-posting about it.

Anthropic's Responsible Scaling Policy v3.0 from 2026 offers another "thresholds lead to safeguards" approach, plus transparency commitments that might actually mean something. If you want the steelmanned doom argument with explicit premises and credences—basically the adult version of this debate—check out Carlsmith's "Is Power-Seeking AI an Existential Risk?" from 2022. And the InstructGPT paper from 2022 shows what "alignment with user intent" looks like in practice, plus what it definitely doesn't solve.

The part people forget: alignment competes with incentives
Even if you had perfect technical methods, deployment still happens inside companies, markets, and states. If the payoff for shipping is immediate and the payoff for caution is hypothetical, your "alignment solution" needs governance teeth. Otherwise it's just vibes with a safety logo.

Also: the "p(doom)" crowd often treats uncertainty like an enemy. But uncertainty is the honest output. Confidence intervals are not cowardice; they're the only way your argument survives contact with reality.

Disproof conditions (what would make us say 'we've largely got this')

Here's what would actually make us update our priors. If we started seeing widely adopted, independently audited evaluation regimes tied to enforceable deployment constraints, we'd know safety stopped being voluntary theater. If systems showed robust behavior under distribution shift and adversarial pressure without constant patching, we'd move from "works in the lab" to "works when it matters." And if clear governance mechanisms emerged that could actually slow or halt scaling when risk thresholds are crossed, we'd know the "race" dynamic lost its grip.

Action line
Ask "aligned to what, measured how, enforced by whom?" If a proposed solution can't answer all three, it's not a solution. It's a compute-heavy coping mechanism.

Anyway, back to touching grass—where "alignment" means your car's tires aren't trying to seek power.

Related reading

Next up, you might want to check out our piece on RLHF titled "Giving the Shoggoth a Treat," because "polite" is not "safe." Or dive into "Vibes vs Evals" since measurement is where claims go to die. And don't miss "Regulatory Capture for Dummies" because governance is a feature, not an afterthought.