IDKMesh

Preregistration v1 — forecasting the value of the next coding agent

Experiment: E046 · Program items: X2 (this document), X3, X4, X5 · Hypotheses: H1, H2, H3 of the Scientific Program Issues: #973 (this freeze), #975, #976, #977 · Epic: #969 Registered: 2026-10-09, frozen by the merge commit of the pull request that adds this file.

1. Frozen artefacts

artefact value
analysis code experiments/agent_value_forecast.py
analysis code SHA-256 67aa3068b52cf67baa6c3862f03fee8ae93a6ad0568de3d26f7b5a1848f28f7d
upstream data SWE-bench/experiments at commit 40f164d5b8f1d249bf95a6df8b74b577fd8e519d (the same commit as E045)
seed 20261009

Any change to the analysis code after the merge is an amendment. The results record must list it, give its reason, and report the frozen version’s output as well.

2. What has and has not been seen

3. Instruments and unit of analysis

4. Estimands

5. Hypotheses, tests and per-split verdicts

H1 — shape (X3, #975). The beta-binomial (per-item difficulty, two parameters) forecasts the curve better than the two-parameter comparators, Kish and shared shock.

H2 — forecast (X4, #976). From a 10-agent pilot, at least one prespecified item-level forecaster — beta_binomial, rasch or incidence — forecasts C(M) within 0.03.

H3 — selection (X5, #977). Greedy complementarity selection beats top-k-by-accuracy at equal k.

6. Cross-split decision rules

These apply over the feasible held-out splits only (lite, test, multilingual, multimodal). verified is reported separately and never counts.

7. Program consequences

This section restates the Scientific Program’s kill criteria.

The program’s headline claim is “agent value is forecastable from measured dependence”. It loses its worker-side support if H1 is falsified and no H2 forecaster beats pilot_only. A falsified H3 removes the selection algorithm (A4) from the sizing advisor (T2, #982).

Every outcome — supported, falsified, unresolved or infeasible — is published in an E046 results record. The record updates the program’s §3–§5 and the issues listed above.

8. Execution procedure

for split in lite test multilingual multimodal; do
  python experiments/agent_value_forecast.py fetch \
    --cache .cache/e046/$split --split $split \
    --commit 40f164d5b8f1d249bf95a6df8b74b577fd8e519d
done
python experiments/agent_value_forecast.py analyze \
  .cache/e046/lite .cache/e046/test .cache/e046/multilingual .cache/e046/multimodal \
  --out experiments/results/E046-confirmatory.json

analysis_sha256 in the output must equal §1’s digest. If it does not, the run is an amendment.