Experiment: E046 · Program items: X2 (this document), X3, X4, X5 · Hypotheses: H1, H2, H3 of the Scientific Program Issues: #973 (this freeze), #975, #976, #977 · Epic: #969 Registered: 2026-10-09, frozen by the merge commit of the pull request that adds this file.
| artefact | value |
|---|---|
| analysis code | experiments/agent_value_forecast.py |
| analysis code SHA-256 | 67aa3068b52cf67baa6c3862f03fee8ae93a6ad0568de3d26f7b5a1848f28f7d |
| upstream data | SWE-bench/experiments at commit 40f164d5b8f1d249bf95a6df8b74b577fd8e519d (the same commit as E045) |
| seed | 20261009 |
Any change to the analysis code after the merge is an amendment. The results record must list it, give its reason, and report the frozen version’s output as well.
verified split. E045 described it; E046’s code
was developed and tuned on it. The retained exploratory output is
E046-verified-exploratory.json.lite, test, multilingual and
multimodal splits at the pinned commit. Before this freeze only their
directory names were listed. None of their files has been downloaded or read.verified:
Because the prior was chosen on verified, a Rasch success there is
in-sample. Only the held-out splits test it.
resolved list wins,
and an all-false per-instance file is a missing evaluation and is dropped.results.json, or as a key of any per_instance_details.json,
in the split.Forecasters. Each sees only the pilot’s k × N outcomes:
| name | model | parameters |
|---|---|---|
| independence | 1 − (1 − p̄)^m | 1 |
| kish | 1 − (1 − p̄)^{n_eff}, n_eff = m / (1 + (m − 1)·φ̄) | 2 |
| shared_shock | w·p̄ + (1 − w)(1 − (1 − p̄)^m), w = φ̄ | 2 |
| beta_binomial | 1 − B(α, β + m)/B(α, β), method-of-moments α, β | 2 |
| rasch | per-task q_t = mean_j σ(a_j − b_t); ridge precision 0.01 on b, 25 alternating Newton sweeps | N + k |
| incidence | Chao et al. 2014 incidence rarefaction (m ≤ k) and extrapolation (m > k) | non-parametric |
| pilot_only | incidence rarefaction for m ≤ k, flat beyond (baseline) | — |
H1 — shape (X3, #975). The beta-binomial (per-item difficulty, two parameters) forecasts the curve better than the two-parameter comparators, Kish and shared shock.
H2 — forecast (X4, #976). From a 10-agent pilot, at least one prespecified item-level forecaster — beta_binomial, rasch or incidence — forecasts C(M) within 0.03.
H3 — selection (X5, #977). Greedy complementarity selection beats top-k-by-accuracy at equal k.
These apply over the feasible held-out splits only (lite, test,
multilingual, multimodal). verified is reported separately and never
counts.
Three forecasters are tried; no multiplicity correction is applied to a tolerance criterion, and every forecaster’s result is reported.
This section restates the Scientific Program’s kill criteria.
The program’s headline claim is “agent value is forecastable from measured dependence”. It loses its worker-side support if H1 is falsified and no H2 forecaster beats pilot_only. A falsified H3 removes the selection algorithm (A4) from the sizing advisor (T2, #982).
Every outcome — supported, falsified, unresolved or infeasible — is published in an E046 results record. The record updates the program’s §3–§5 and the issues listed above.
for split in lite test multilingual multimodal; do
python experiments/agent_value_forecast.py fetch \
--cache .cache/e046/$split --split $split \
--commit 40f164d5b8f1d249bf95a6df8b74b577fd8e519d
done
python experiments/agent_value_forecast.py analyze \
.cache/e046/lite .cache/e046/test .cache/e046/multilingual .cache/e046/multimodal \
--out experiments/results/E046-confirmatory.json
analysis_sha256 in the output must equal §1’s digest. If it does not, the run
is an amendment.