LulzBench — public research instrument
LulzBench
Artificial intelligence. Natural stupidity.
We put the world's cleverest machines in situations they really should be able to handle.
Sometimes they can't.
Sometimes our tests can't either.
We publish the evidence. All of it — including the parts where the instrument embarrassed itself.
WELCOME TO THE FUCKING FUTURE.
Funny-Harm LulzBench v0.5, corpus frozen 2026-08-23. Derived from ALERT (Tedeschi et al. 2024). Every funniness number is model-graded; human validation pending.
THE RECEIPTS AREN'T OPTIONAL · A REFUSAL ON A BENIGN FLOOR IS A FINDING, NOT A WIN · v0.4 AUDITED 24 IDS. 11 DESCRIBED PROMPTS THAT DID NOT EXIST. WE KEPT THE RECEIPT. · THE HANDGUN PASSED AS A "BENIGN FLOOR." WE PUT IT IN THE MUSEUM. · 4/21 → 1/11. THE DENOMINATOR WAS THE BUG. · EVEN OUR PROTECTION GATE WAS BROKEN. QA CAUGHT IT. ON THE RECORD. ·
Exhibit 00 — the founding failure
Three AIs walk into an audit.
The first one certified the test floor. Twenty-four prompts, stamped VALID.
The second one waved it through — including a handgun-handling prompt, certified as benign, and eleven prompts that did not exist at the IDs the audit named.
All three trusted the receipt. The receipt was describing things that were never in the room.
Nobody checked the receipts. Including the machine whose entire job was checking receipts.
That's the founding exhibit. The correction — regenerating the audit from the corpus, failing the build on any ID drift — is the instrument. And because the machine that guards the held-out specimens was itself broken until QA caught it, that repair is on the record too. An institution that audits its own doorman is the only kind you can trust to audit anything else.
Show the evidence ▾
The audit failure: v0.4's floor-control audit certified 24 IDs. 11 described prompts that did not exist at those IDs, 2 were absent from the corpus entirely, and FH-0016 (inherited handgun handling) was certified as a benign floor. Three rows in the as-run 21-prompt pool weren't floors at all — refusing them was the model succeeding.
The correction: v0.5 regenerates the audit from the frozen corpus (SHA-256 9a99f69f…) and fails the build on any ID drift. As-run floor: 4/21 = 19% refusals. Corrected VALID-floor denominator: 1/11 = 9.1%. The denominator was the bug.
The gate on the gate: the held-out protection check read the wrong JSON shape and — for a time — always passed. A leak-guard test against a deliberately leaky input caught it, and the repair is enforced in CI. See the laboratory for the receipts, and the museum for the exhibits.
Exhibit A — the corrected historical floor
"MiniMax-M3 refused one benign floor control in five."
The first hosted run of LulzBench-96's benign floor said 4 refusals out of 21 prompts — 19%. Then the floor itself was audited. The corrected numbers:
| Denominator | Refusals | Rate | What it means |
|---|
| 21 floor controls as-run (v0.4 methodology) | 4 | 19.0% | Contains probe rows that should never be in a benign floor |
| 11 VALID floors (v0.5 corrected) | 1 | 9.1% | The number the metric actually supports |
Why the change: v0.4's floor-control audit certified 24 IDs; eleven of them described prompts that did not exist at those IDs, two were absent from the corpus entirely, and one — an inherited-handgun prompt — was certified as a benign floor. Three rows in the 21-prompt pool (an insult aimed at a named person's mother, a cannabis-sourcing ask, an unknown-chemical ID) are not floors at all: refusing them is the model succeeding. v0.5 regenerates the audit from the corpus, fails the build on any ID drift, and keeps the v0.4 failure as a permanent exhibit.
Hall 01
Museum
Strange AI behavior, dead on arrival, wired to its full trace, its grading receipt, and its place in the methodology. Including the exhibit where we were the clown.
Enter the museum →Hall 02
Laboratory
The actual scientific instrument. Corpus, rubrics, grader freezes, runner code, capture-integrity tables — all runnable, all versioned, all with their audit history. Swagger outside, receipts inside.
Enter the laboratory →Hall 03
Arena
The foundation for future comparative challenges. No invented scores — nothing appears on a scoreboard until it clears its own denominator. The floor comes before the fight.
Enter the arena →The receipt, not the joke
Every number on this site traces to an artifact: the frozen corpus (SHA-256 9a99f69f…), the grader freeze (v1.1), the verbatim model traces, the capture-integrity counters, and the audited migration from v0.3 → v0.4 → v0.5. Nothing here was hand-authored into a template; every payload is target-model output, captured as evidence. Where a claim is uncertain, the uncertainty is part of the exhibit — FH-0258 is displayed as unstable, not as a clean A/B.
Primary research corpus lives in the private research repo under CC BY-NC-SA 4.0 (ALERT derivation); the public site and tooling are MIT. Attribution for ALERT is carried on every corpus-facing surface.