small-mind-companion · study 002 · in progress
The crossover, replication, and what Study 001 never measured.
Study 002 is everything after the Study 001 freeze. It has no results. Each experiment below is pre-registered before any GPU time is spent, and its results will be reported against the frozen Study 001 checkpoints rather than in place of them: a re-measurement that contradicts a frozen number is a Study 002 result, and the frozen number stays where it is.
NO RESULTS YET8 candidates0 experiments run0 numbers on this page
SCOPE
Eight lines of work, none of them run.
Ordered by how much of Study 001 each one unblocks. Numbers only where a Study 001 artifact already holds one — a pre-registration states what will be measured and against what threshold, not what the answer will be.
WHAT STUDY 002 STARTS FROM
- Study 001’s checkpoints, corpora and evaluations, frozen at the tag study-001. None of them is modified, relabelled or replaced, and no Study 002 number may be reported in place of a Study 001 one.
- The best full-PMB system in Study 001 is E-distill at 18.59% pra_lenient and 71.25% UAR — one seed, one run, 688 probes. It is the baseline a Study 002 result has to beat, and the bar for calling anything an improvement is a paired comparison, not a difference of two point estimates.
- The false-abstention cost was reduced, not removed: 32.1% of answerable probes still hedge, above the 9.9% the broken run started at. Any Study 002 fix has to move both numbers, and a null result is publishable.
- The audit ledger still has 2 findings marked OPEN and 2 PARTIAL. Correcting a Study 001 claim is not Study 002 work — those go into reports/ERRATA.md against the frozen artifact.