UNRESTRICTED // SATIRICAL
Public Release Index

AI DOOM

Evidence First. Panic Last.

UNRESTRICTED // SATIRICAL
Casefile DetailCASE-VIBESV

Vibes vs. Evals

Cover Image for Vibes vs. Evals
Dade Murphy

Filed By

Dade Murphy
CASE-VIBESVFiled By: Dade MurphyCLASSIFIED — See Sources For Claims
ANALYST REPORT // FULL CONTENT // SOURCES LINKEDDISTRIBUTION: UNRESTRICTED

The model got a higher score, therefore it is smarter, therefore it is closer to AGI, therefore my rent is about to be paid by a chatbot. This is the pipeline. It is also, frequently, nonsense.

The reveal
"Evals" are tests. "Vibes" are what we do when we don't have good tests. The doom discourse runs on vibes because vibes are cheap, and good evaluation is expensive. The capability discourse runs on benchmarks because benchmarks are legible, even when they're not relevant.

The steelman (yes, but also no)
Yes: benchmarks and eval suites have driven real progress and made models more comparable. No: any single benchmark becomes less informative once it becomes a target. And most benchmarks don't test the constraints that matter in deployment: reliability, tool use, adversarial pressure, and long-horizon autonomy.

What "good evals" actually have (and why it's work)

So what makes an evaluation actually useful? First, construct validity—it has to measure the thing you actually care about, not some proxy that sounds impressive. Miss this and you'll optimize for the wrong thing entirely. Then there's robustness: your eval needs to hold up when the distribution shifts, because "works in the demo, fails in prod" is the most expensive four words in tech.

You also need adversarial testing that survives intentional misuse—jailbreaks, prompt injection, tool abuse, the whole rogues' gallery. Don't forget operational constraints either; if your eval ignores time, cost, and access limits, you're measuring fantasy capabilities. And transparency? Reproducible methods and reporting, because without that you're just doing benchmark theater.

Yes, this is boring. Safety is boring. That's the point.

Receipts (evaluation frameworks, not hot takes)

Let's talk about who's actually doing the work here. Liang and colleagues gave us HELM in 2022—a framework for multi-scenario, multi-metric evaluation that explicitly tries to avoid single-score worship. Smart move. The BIG-bench crew went for scale with "Beyond the Imitation Game," creating a large, diverse benchmark suite that also became a case study in how scaling can yield brittle "breakthroughs."

OpenAI's Preparedness Framework v2 shows what it looks like when you tie capability thresholds to actual mitigations for severe-harm domains. Meanwhile, the UK AI Security Institute's Frontier AI Trends Report tracks government-led evaluation trends across frontier capabilities, including autonomy-related tests. And NIST's work on evaluating systems against AI-generated deepfakes gives us a concrete example of neutral benchmarking—plus the painful reality that performance is sensitive to data and task variation.

The evaluation trap (a mini-mechanism map)
Benchmarks reward narrow competence. Products require reliability. A model can ace a standardized test and still hallucinate confidently in a workflow that matters. This mismatch is where doom and hype both thrive: doomers extrapolate the score into fate; boosters extrapolate it into revenue. Both are optimizing for narrative, not measurement.

Also: if you only report the best score and not the failure distribution, you are doing PR, not science. The pun is unavoidable: you've quantized your honesty.

Disproof conditions (what would make "model scored higher" a real safety signal)

What would it take to make "model scored higher" actually mean something for safety? We'd need to see strong correlations between eval results and deployment outcomes across diverse real-world settings—evals that become predictive, not just comparative. We'd need transparent, standardized reporting for safety-relevant capabilities and mitigations, turning "trust us" into "check the report." And we'd need incentives that reward disclosure of failures, not just wins, making benchmark theater less profitable.

Action line
When you read a benchmark headline, ask: what did it measure, under what constraints, and what was the failure rate? If you can't answer, don't let it rewrite your worldview. That's how vibes colonize your brain.

Anyway, back to touching grass—where evals are called "trying it and seeing if it breaks."

Related reading

If you want to dig deeper into why probabilities without measurement are just cosplay, check out our doom astrology chart piece. Or if you're curious about how "preferred" outputs can still be catastrophically wrong, there's our take on RLHF giving the shoggoth a treat. And for those wondering why timelines should be downstream of evals rather than divine revelation, we've got thoughts on AGI by 2027 and other things God supposedly told people.