Research Index
This directory contains executable research plans, experiment interpretations,
and calibration evidence. Results are evidence for bounded claims, not automatic
policy changes or merge decisions.
Program and Measurement
- First Research Program — staged path from
deterministic foundations to held-out real-task evidence.
- Top 20 Questions — prioritized open questions about
scaling, verification, governance, and community growth.
- Metric Uncertainty v0.1 — uncertainty and
reporting rules for repository metrics.
- Collaboration Observables v0.1 —
deterministic latency, concentration, recurrence, queue, CI, and debt
measurements plus the preregistered E023 community analysis.
- Preregistration v1: First-Review Latency and Contributor Recurrence —
the specification, the randomized design it would take to earn a causal
claim, and the threats to validity, registered against zero observations.
- Phase 0: Executable Research Foundation — schemas,
fixtures, and replay requirements beneath later experiments.
- Randomness Research Roadmap and
experiment status — sequence and current
state of the bio-inspired scheduling program.
Routing and Orchestration Experiments
Verification Research
- Coordination Criticality and Finite-Difference Response
— matched small-load probes compared with utilization and backlog baselines.
- Evaluator Plan Binding — binds verifier-owned
evaluation intent to exact task evidence.
- Executable Independent Verifier MVP — defines
the first executable authority-separated verifier.
- Verification Debt and Risk-Weighted Backpressure
— controller model for limiting unverified work.
- Verification Backpressure Temporal Benchmark
— multi-window benchmark for that controller.
- E022 Seven-Mode Verification Scaling Matrix
— matched comparison of every verification condition required by issue #14.
- E024 Matched-Budget Emergence
— equal-evaluation comparison of random, fixed-scalar, and Quality-Diversity search.
- E026 Imperfect Verifier Panel
— E024 rerun with E017/E020’s measured correlated panel and blind-spot floor.
- E027 Defect Propagation
— gives an accepted defect a cost, so verifier error can reach the outcome
metric, and sweeps the cost knob across its whole range.
- E028 Latent Defect Dimension
— removes E027’s confound by moving viability into a dimension the goals cannot
see; the archive’s survival holds in 18 of 20 cells and breaks only under the
stress panel at full defect cost.
- E029 First Real Model Attempts
— 60 sandboxed attempts by a pinned 0.5B open-weight producer on the frozen
benchmark: 0 accepted, 56 of 60 failing the diff protocol before any
repository content was consulted.
- E030 Supplied-Goal Membership
— removes E024’s supplied-oracle confound by switching the environment to a
parity-matched goal the arms do not hold; the archive keeps all but 1.6-4.4%
of its lead and stays 0/100 catastrophic, while the majority-vote swarm loses
its whole lead in every panel.
- E031 Learned Goal Filter
— the other half of E024’s caveat: gives the consensus swarm a particle filter
that learns the goal from ordinal evidence. Learning from generation 0 roughly
doubles its catastrophic seeds; learning from post-change evidence alone is the
best variant in all eight cells on both the tail and the mean. An evidence-free
rescue — perturbing each agent’s hypothesis once at initialisation — takes
38/100 catastrophic seeds to 0/100 while the new goal is one of the four
supplied, and to 71/100 when it is not.
- E032 Population Scaling
— answers issue 13’s success criterion, at a fixed budget when is another
agent worth adding, by running the sweep with the budget held and with it
free. The two disagree: the archive gains on every doubling when the budget is
allowed to grow with the population, and gains nothing at all (0.03 AUC across
a 16x change) when it is held. The scalar hill-climber is the opposite — it
gains near-linearly to N=256 and goes from 100/100 to 0/100 catastrophic — so
returns to population run inversely to how much diversity an arm already
retains. No arm shows the negative return hypothesis 1 predicts; the resolved
negative returns are on the other two axes, archive capacity past bins=8 and
budget spent on generations rather than agents.
- E033 Goal Distance
— turns E030’s single substitute goal into a ladder of rings, six goals each,
at a matched change size. The archive’s lead over the arms that hold no
hypothesis decays smoothly rather than off a cliff, and is fully gone by 0.35
from the supplied set — about one and a half times the set’s own spread. It
closes not because the archive gets worse (22.4 to 22.1) but because the
simple arms get better (18.9 to 21.9), and distant goals are measurably more
discriminating rather than less, so ‘nothing helps out there’ is falsified.
E030’s published point ranks second of seven at its own distance: retention
there is 78.3% on average, not the 95.6% one goal reports. Sweeping the same
axis without holding the change size returns ‘unresolved’ and would have
missed the decay entirely.
- E034 Goal Direction
- E035 Direction Across Shells
- E036 Adversarial Contributors
- E037 Ladder Under Panels
- E038 Symmetric Gate
- E039 Content Addressed Blind Spot
— holds E033’s distance still (0.30 from the supplied set, 0.392 of change)
and sweeps direction instead, 385 goals on one shell. Direction is worth more
than distance: the archive’s lead runs from -4.894 to +4.471, a spread of
9.365 against the 3.309 E033’s whole distance sweep moved, and 24.2% of
directions leave the archive behind the arms that hold no hypothesis. The
mechanism E033 proposed for this — the viability floor — is falsified by its
own preregistered test: the control trait ‘simplicity’ was predicted flat and
instead carries the lead from +2.257 to -0.371, and the two traits sim.viable
floors identically do not behave alike. The structural trait categories are
not a valid grouping either; the two descriptor traits move in opposite
directions and average to nothing. E033’s post-hoc ‘security’ observation
(-1.362) does not survive the control, which gives +0.279 [-1.095, +1.653].
- E040 Diversity Correlation Threshold
— sweeps the assumed worker-error correlation the R1 scaling runner had fixed
at 0.25, and connects the result to issue 13’s hypothesis 2. There is no
threshold: the equal-budget advantage is proportional to retained
independence, 1 - rho, in 17 of 18 curves at an uncentered R-squared of at
least 0.99 through the origin, so this grid returns ‘hypothesis 2 holds’ at
every correlation short of 1.0 and cannot falsify it. The issue 30 help/hurt
sweep corroborates that and bounds it: across its whole correlation ladder it
reports 0 hurts cells at quality penalty 0, and all 13 of its hurts cells
carry a non-zero penalty. Correlation sets the size of the advantage; worker
quality sets its sign. Splitting the two arms also splits the hypothesis:
randomizing verifier assignment raises the slope in 5 of 9 cells, by at most
0.0181 against worker-diversity slopes running to 0.5504, so the whole
measured effect is on the worker side and none of it belongs to the ‘with
independent verification’ half.
- E025 Learned Verifier Reliability
— calibration/held-out evidence for reliability and dependence-aware aggregation.
- IDKGraph P1 Independent Review Protocol
— frozen-cohort human review and attention measurement for issue #152.
Benchmark Calibration Evidence
Calibration candidates are not scored benchmark outcomes. Preserve exact
source revisions, digests, and lifecycle status when citing these records.