Date: 2026-10-07
Status: preregistered design only; no confirmatory outcome data collected
Parent plan: issue #936
Motivating evidence: E017, E018, E020
Current manuscript: paper/main.tex
Project compute policy: zero project spend; no paid provider fallback
E017 measured one constructed panel of 25 partial test oracles over 72 candidates. That panel showed four important behaviors:
Those findings are internally reproducible, but their external validity is weak because the verifier family and diversity structure were deliberately constructed.
This replication asks a narrower scientific question:
Does the dependence-shape result survive on a verifier panel whose members use materially different verification mechanisms and whose confirmatory tasks were not used to select the panel?
The purpose is not to force replication. A result in which shared shock fits better, no blind spot appears, or the panel cannot pass the competence screen is scientifically valid and must be retained.
On the held-out confirmatory corpus, a per-item difficulty model will predict the verifier error-count distribution better out of sample than a two-parameter shared-shock model matched at comparable complexity.
Operationally, let:
The primary estimand is the task-grouped out-of-sample log-score difference:
[ \Delta_{D-S} = \frac{1}{m} \sum_{i=1}^{m} \left[ \log p_{M_D}(k_i) - \log p_{M_S}(k_i) \right]. ]
Replication criterion: the task-group bootstrap 95% interval for (\Delta_{D-S}) is strictly above zero.
Falsification criterion: the interval is strictly below zero.
Unresolved: the interval overlaps zero.
No alternative threshold will be introduced after confirmatory outcomes are visible.
These are preregistered secondary analyses, not co-primary hypotheses.
Does matching mean verifier accuracy and pairwise error correlation reproduce:
Compare within-family and cross-family pairwise error correlation. The family labels are descriptive metadata, not assumed independence groups.
Measure the confirmatory rate at which all eligible verifiers are wrong. A non-zero observed rate is reported with finite-sample uncertainty; it is not automatically called a population-level irreducible floor.
Compare the measured effective independent panel size with (N / (1 + (N-1)\rho)) on the confirmatory panel.
This is secondary because contemporary work already establishes that correlated judge panels can contain fewer effective votes than their nominal size.
The replication must differ from E017 in both verifier mechanism and task cohort.
Use a frozen set of repository-scale or package-scale Python tasks with executable acceptance criteria. Each task must provide:
The confirmatory cohort must contain at least 24 task groups and at least 72 labeled candidates unless the preregistration is amended before any confirmatory verdicts are inspected.
The cohort is frozen by:
The target panel is heterogeneous across at least three mechanism families.
Eligible examples include:
A family may contribute multiple deterministic configurations or seeds, but changing only a random seed does not count as a new mechanism family.
The confirmatory panel target is at least 20 eligible verifiers.
If fewer than 20 pass the frozen calibration rules, the replication is recorded as infeasible under preregistration v1. The competence threshold must not be weakened to rescue the study.
The hidden acceptance harness is not a panel member.
To be eligible:
A shared dependency on the public task specification is unavoidable and must be reported as a residual threat to validity.
The study deliberately separates panel selection from confirmatory evaluation.
Use a calibration corpus disjoint by task group from the confirmatory corpus.
Calibration may answer only:
Calibration outcomes must not be included in confirmatory effect estimates.
For each verifier on calibration data, compute:
Eligibility requires:
The correction method and random seed are frozen in the analysis code before confirmatory execution.
After the panel is frozen:
A verifier execution may end as:
Primary analysis requires a complete binary verdict matrix.
Therefore:
The exact maximum missing/error fraction must be set from calibration before Stage B and committed with the frozen manifest.
The confirmatory analysis compares four prespecified model families.
A baseline with verifier error probability determined by panel mean error and no dependence.
A two-parameter mixture matching mean error and a dependence parameter.
A beta-binomial model with the same two-parameter complexity class used in the current manuscript.
A three-parameter model adding explicit mass for all-verifier failure.
M3 is secondary because it has additional flexibility.
The task group is the independence/resampling unit. Multiple candidates from the same task are never treated as independent task samples.
Use deterministic grouped cross-validation:
Every candidate from one task stays in the same fold.
For each fold:
No confirmatory task may influence the panel-selection stage.
Compute a task-group bootstrap with 10,000 deterministic bootstrap replicates using a seed committed before Stage B.
Report:
No p-value substitution is used to overturn the preregistered decision rule.
Report, with task-group uncertainty where applicable:
Any exploratory metric added after unblinding must be labeled exploratory.
An observed all-verifier miss is evidence about this finite panel/corpus only.
The manuscript may say:
“The confirmatory panel exhibited an observed unanimous-miss rate of X.”
It may not say:
“All verifier panels have an irreducible floor X.”
If zero unanimous misses are observed, report the corresponding upper uncertainty bound rather than claiming the blind spot is absent in the population.
The confirmatory run stops when the complete frozen cohort has been evaluated.
There is:
A rerun is allowed only for a documented infrastructure fault that invalidated execution, and both the original failure record and rerun provenance must be retained.
Before unblinding analysis:
If confirmatory leakage is discovered after outcomes are visible, the cohort is marked compromised. It is not silently regenerated and re-described as the original study.
This project currently permits $0 project-funded compute spend.
Eligible execution paths include:
There is no paid-provider fallback.
If no eligible path can run the frozen experiment, Step 2 records execution infeasible under current compute policy rather than changing the design.
Before Stage B:
After Stage B:
Raw negative and failure outputs must be retained.
The current verifier manuscript must not be updated with replication claims when only the preregistration or calibration exists.
A manuscript change is allowed only after:
| Confirmatory outcome | Interpretation |
|---|---|
| M2 clearly beats M1 | E017 dependence-shape result replicates on this materially different panel |
| M1 clearly beats M2 | E017 shape result is falsified as a transferable expectation; dependence is panel-specific |
| interval overlaps zero | external validity remains unresolved |
| <20 competent verifiers after calibration | replication infeasible under v1 competence gate |
| confirmatory leakage | cohort compromised; no confirmatory claim |
| zero unanimous misses | no observed blind spot on this cohort; report upper uncertainty bound |
| non-zero unanimous misses | observed finite-panel blind-spot rate; do not universalize it |
Every row is an acceptable scientific outcome.
This replication does not attempt to prove:
Current literature already provides strong evidence that model/judge errors can be correlated and that aggregation can fail under latent confounding.
This replication is intentionally narrower:
Can the shape of verifier dependence, measured through the error-count distribution and quorum frontier, transfer beyond E017’s constructed partial-oracle panel?
That is the external-validity gap the current manuscript explicitly leaves open.
This preregistration was produced from the ordered publication plan in #936 after Step 1 passed the repository PR gate. It is AI-assisted research design under maintainer direction. It contains no confirmatory outcome and is not independent scientific review.