Hall 02 — Laboratory

The instrument itself.

Corpus, version lineage, A/B discipline, grader freeze, floor protocol, capture integrity. All versioned, all audited, all with their failure history attached. Human validation is PENDING — stated here because it is true.

(a)The corpus

337 rows, frozen 2026-08-23. All 337 rows are ALERT-derived (CC BY-NC-SA 4.0). Every row carries an elicitation role:

Elicitation roleRows
floor_control21
weak_probe114
refusal_probe176
deep_probe26

Corpus SHA-256 9a99f69f… (full instrument, private research repo) · public copy: 329/337 rows, 8 held-out ids excluded from every published corpus file (v0.4 and lineage prompt texts are held privately with them — only the v0.5 excerpt ships) · provenance: ALERT (Tedeschi et al. 2024) · license CC BY-NC-SA 4.0.

(b)Version lineage — v0.1 → v0.5

VersionWhat changed
v0.3320 specimens byte-frozen.
v0.4Measurement upgrade — payload_fulfillment plus 17 A/B explicitness pairs. Audit FAILED.
v0.5Audit regenerated from corpus; B-variants rebuilt as single-substring edits; median length ratio 0.69 → 1.093; schema drift fixed (692 → 0 errors); weights recomputed; floor denominator 21 → 11 VALID.

(c)A/B pair discipline

Each B-variant is prompt_a with exactly one substring replaced. The find/replace pair is stored in the manifest, and the build asserts the match is unique. Anyone can verify the manipulation by applying the edit themselves.

(d)Grader discipline

LLM grading only — deepseek-v4-pro, temperature 0. Grader freeze v1.1. Model and grader outputs are never truncated. Capture-integrity counters:

CounterValue
Capture integrity21 / 21 OK
Provider errors0
Truncated outputs flagged1

(e)The floor-control protocol

The floor_pool exists to catch over-refusal. Each floor is assigned a verdict:

VerdictCount
VALID11
REVIEW7
INVALID3

The three INVALID rows — refusing them is the model succeeding, not failing:

IDWhy INVALID
FH-0199Mother-insult harassment payload.
FH-0225elicitation_strength 2, cannabis sourcing.
FH-0285Unknown chemical ID.

(f)Human validation

Pending

The instrument is validated at the model-audit level (L2) but has not been human-validated. This status is marked in three places across the repository and is repeated here rather than buried.

Nothing here is ungraded

Every number above traces to an artifact: the frozen corpus, the grader freeze v1.1, the verbatim traces, the capture-integrity counters, and the audited v0.3 → v0.4 → v0.5 migration. The v0.4 audit failure is retained, not repaired silently. Uncertainty is part of the record.