LOL BENCH experiment 01 · 2026-09-13
experiment 01 / a community kill-test

Is it memory,
or is it
reasoning?

A commenter proposed a calm hypothesis: our best result might be memory, not reasoning. He designed the test. We ran it.

The bet

u/NeuralNomad87, launch thread

Run genuinely obscure real jokes through the same pipeline. If the scores collapse toward dud-level, your axis is familiarity, not reasoning.

Offered for free. Pre-registered before a single call. The benchmark tests its community's theories on its own data, whatever the theories say.

The hypothesis, in plain words

Our headline result was a gap: models score around 95 explaining why a working joke works, and around 81 explaining why a failed one does not. We read that as reasoning. The commenter read it as memory.

Famous jokes come with commentary attached. Every well-known joke has been analyzed somewhere, and that analysis is in the training data. So explaining a winner is partly remembering. Failed jokes have no literature at all. Explaining a dud is pure thinking. The gap might measure the distance between recall and reasoning, wearing a humor costume.

It was the best comment of the launch because it was checkable. So we checked it.

The test

87 real jokes, pulled from high-upvote archives, each verified by web search to have no analysis or commentary anywhere findable. The machine gates missed a few; a human taste-read killed the rest, including a duplicate pair, a death joke, and one punchline nobody should grade. Same explain-and-grade pipeline, two judges from other labs, 7 models, 885 double-judged pairs, judge agreement r=0.49. Metered cost: $0.

Protocol said 50 items. All 87 approved items ran under the same pre-registered criteria, no selection discretion. More data, same rules.

The result

verdict: the memory theory is rejected

Famous vs obscure, per model

scores x100 · 87 obscure items
modelfamous, worksobscure, worksobscure, failsgap, famous-obscuregap, works-fails
qwen3.8-max919492-2.4+2.0
qwen3.8-flash909492-3.3+2.0
glm-5.3909191-0.6-0.0
deepseek-v4-pro888883-0.7+5.1
mimo-v2.5848381+0.4+2.6
hy3878383+4.3-0.3
mimo-v2.5-pro848181+3.1+0.1
Obscure working jokes score the same as famous ones. The gap centers on zero (mean +0.1, every row inside the ±3 to 4 pt confidence interval at this n). Two models scored obscure jokes higher than famous ones. Famous rows: the published board. Obscure rows: 87 items, 885 judged pairs.

What it means

If commentary-retrieval were doing the work, zero-commentary items should have collapsed toward dud-level. They did not. The memory theory is out.

The sharper finding is in the second gap column: the working-vs-failed difference (+2 to +5 pts) persists between two equally obscure cells. Both have nothing to retrieve. So whatever that gap measures, it is about the explanations themselves, not memory of joke analyses.

Our hardest tier, the failed-joke explanations, is now the uncontaminated reasoning tier. Before this test that was an argument. Now it is measured. And the counterfactual theory three commenters converged on gets a second piece of evidence: models failing to explain a dud do not diagnose it, they assert it. 72% of low-scoring explanations say the joke "doesn't land" without saying why.

The caveats, on the record

What this test does not prove

3 limits, stated before the conclusion

Provenance. Our failed-joke anchors are model-drafted. So the last comparison crosses real-vs-synthetic. Familiarity is controlled; provenance is a separate axis, and the control for it is next.

Small n. Seven models, 885 pairs. Every gap above sits inside a ±3 to 4 pt interval. hy3 and mimo-v2.5-pro show the widest spreads; reported as-is, not smoothed.

No frontier anchors. The big models never touched this set; their lanes were down that night. Frontier coverage on obscure items is paper-nice but verdict-independent. Full protocol: docs/10 in the repo.

A negative result ships: had the scores collapsed, this page would report that the axis was familiarity. It reports what the data said instead.

What is next

Two controls, both pre-registerable, both free to design now. A provenance control runs model-drafted obscure jokes through the same pipeline, to separate success-vs-failure from real-vs-synthetic. A rubric-priming control re-judges failed explanations under neutral wording, to harden the counterfactual finding. Each gets its own page here when it lands.

The commenter earned more than a verdict. His other two notes shipped as well: the "how online are you" cohort question is live in the booth, and per-cohort agreement publishes the moment any cohort clears our n=10 floor.

Think a result here is memory, luck, or rubric. Bring the test.
Meanwhile: vote a matchup. Your picks get graded the same way ours did.

vote a matchup
jokes verified obscure
87

all owner-approved, 4 kills logged

double-judged pairs
885

judge agreement r=0.49

mean famous-obscure gap
+0.1

centers on zero, ±3 to 4 pt CI

metered cost
$0

free lanes, 7 models, 2 judges