Reached the result it set out to test, at its own stated scope and no wider.
Read the status before the claim
Every artifact in this repository sits on an explicit evidence ladder, and most of these experiments deliberately stop short of the top rung. A benchmark contract plus a synthetic fixture means the experiment is runnable — never that a strategy has been shown superior.
Qualified, narrowed or falsified an earlier experiment — including several of this project’s own headline claims.
The mechanism did not work, or could not be measured at all. Retained in full.
Synthetic mechanism test or preregistration. Tests the mechanics; proves nothing about real collaboration.
The atlas
Filter by what the experiment actually concluded. Each card links to its record in the repository, which carries the module that produced it, the machine-readable artifact, and the tests that pin it.
Emergence from vague goals
Can exploratory agents move toward a coherent system when the objective is incomplete, given hard constraints, diversity preservation, evidence-driven selection and shared memory? Deliberately synthetic — mechanisms tested before any real engineering task.
Read the record ↗
Correlated verification failure
Replaces the perfect viability oracle with an imperfect panel. As verifier errors correlate, majority-vote error rises while panel disagreement falls — so the panel looks more confident exactly as it becomes less reliable.
Read the record ↗
Independence-aware aggregation
At the same nominal review cost and the same votes, when should a large correlated verifier cluster be discounted rather than counted vote-by-vote? The rule should help when a cluster shares failure modes — and must not be assumed to help when reviewers are genuinely independent.
Read the record ↗
ACO stigmergic task routing
Proposed, with an executable synthetic baseline: can evidence-backed pheromone routing allocate tasks better than simpler rules without causing herding, duplication or review concentration?
Read the record ↗
Verification phase diagram
Sweeps E012's single operating point into a phase diagram: given panel size, accuracy, correlation and quorum, how many statistically independent verifiers is a panel actually worth? Cost-asymmetric quorums help, and the N_eff heuristic is optimistic.
Read the record ↗
Live verifier correlation
Negative result, and an honest one — its own status line reads “the verifiers it deployed do not verify”. The 20 LLM verifiers averaged Youden J = +0.0487, and their majority vote scored 0.514 accuracy against 0.639 for a trivial “always reject”. Correlation could not be estimated from votes that carry almost no signal. Screen a panel for discrimination before trusting anything it votes on — and note what this leaves open: no AI review panel has been measured in this repository.
Read the record ↗
Item difficulty and quorum
The first observed measurement of verifier error correlation here, on 25 independently seeded partial test oracles — programs, graded on real defects — over a 72-candidate corpus: ρ = 0.5873 across 300 pairs, an effective panel size of 1.00 of 25 under majority vote, and the shared-shock model shown to be the wrong shape. Fixing the aggregation rule cut error 3.7×; growing the panel bought nothing.
Read the record ↗
Dependence model shape
Which E015 conclusions depend on the shape of the dependence model? One correction to E015 and one strengthening of it — the direction of the N_eff result survives, the magnitude does not.
Read the record ↗
Group independence under item difficulty
E013's crossover is robust to the model's shape — it sits at ρ = 0.25 under both — and disappears entirely when the declared independence groups are not actually independent. Declared structure is not measured structure.
Read the record ↗
Quorum frontier under the measured shape
Extends E017 §6 to the full frontier with the corpus base rate, and qualifies it twice: the recommended beta-binomial under-predicts the unanimity floor by 1.77×, and the shared-shock model reports that no quorum beats majority — getting the highest-leverage decision on that panel exactly backwards.
Read the record ↗
Coordination criticality
Matched small-load susceptibility probes give earlier overload warning at a measurable false-alarm cost — but no evidence that susceptibility dominates an ordinary utilisation threshold.
Read the record ↗
Seven-mode verification scaling matrix
A matched comparison of every verification condition issue #14 names: can candidate generation and independent verification scale together without collapsing into unchecked acceptance or an unbounded review queue?
Read the record ↗
First-review latency and contributor recurrence
Preregistered against zero observations: the specification, the randomised design it would take to earn a causal claim, and the threats to validity — written before any data, precisely so the analysis cannot be chosen after seeing it.
Read the record ↗
Matched-budget emergence
Every strategy issue #22 names gets exactly the same proposal and verification-attempt budget. The Quality-Diversity archive's surviving advantage is reliability: it never fails catastrophically, where the majority-vote swarm fails in 44 of 100 seeds. Its own stated limitation drove E030 and E031.
Read the record ↗
Learned verifier reliability
Learned weighting helps in stable regimes and fails materially under plausible shift. Explicitly not production-ready reputation — the failure mode is recorded rather than smoothed over.
Read the record ↗
Imperfect verifier panel
E024 rerun with E017/E020's measured correlated panel and blind-spot floor. Nothing moved: sweeping the panel to 45% wrong in both directions changed the archive's post-change AUC by 1.5% and left catastrophes at 0/100. A null result, kept.
Read the record ↗
Defect propagation
Gives an accepted defect an actual cost, so verifier error can finally reach the outcome metric, and sweeps that cost across its whole range. The constraint-guided archive holds at 0/100 in all twenty cells while unconstrained random search goes 0/100 to 94/100.
Read the record ↗
Latent defect dimension
Removes E027's confound by moving viability into a dimension the goals cannot see. The archive's survival holds in 18 of 20 cells and breaks only under the stress panel at full defect cost.
Read the record ↗
First real model attempts
The first negative result this repository was entitled to: 60 sandboxed attempts by a pinned 0.5B open-weight producer on the frozen benchmark produced 0 accepted candidates, and 56 of the 60 failures were failures of the unified-diff protocol before any repository content was consulted.
Read the record ↗
Supplied-goal membership
Removes E024's supplied-oracle confound by switching to a parity-matched goal the arms do not hold. The archive keeps all but 1.6–4.4% of its lead and stays 0/100 catastrophic; the majority-vote swarm loses its whole lead in every panel.
Read the record ↗
Learned goal filter
Gives the consensus swarm a particle filter that learns the goal from ordinal evidence. Learning from generation 0 roughly doubles its catastrophic seeds; learning from post-change evidence alone is the best variant in all eight cells. An evidence-free rescue takes 38/100 catastrophic seeds to 0/100 — and to 71/100 once the new goal is outside the supplied set.
Read the record ↗
Population scaling
Answers issue #13's question — at a fixed budget, when is another agent worth adding? — by running the sweep with the budget held and with it free. The two disagree: the archive gains on every doubling with a free budget and gains nothing at all when it is held.
Read the record ↗
Goal distance
Turns E030's single substitute goal into a ladder of rings at a matched change size. The archive's lead decays smoothly and is gone by 0.35 from the supplied set. E030's published point ranks second of seven at its own distance — retention there averages 78.3%, not the 95.6% one goal reported.
Read the record ↗
Goal direction
At a fixed distance, direction matters as much as distance: on a single shell the archive's lead runs from −4.894 to +4.471 and is negative for 24.2% of directions. The mechanism E033 proposed for that spread is falsified by its own preregistered test.
Read the record ↗
Direction across shells
Re-runs E034's ladder at further distances. The spread does not grow with distance as conjectured, only one ladder in five replicates, and the geometry admits very few valid shells at all — a resolved effect on one shell is not a general one.
Read the record ↗
Adversarial contributors
No adversary made the constraint-guided archive worse than another arm in any of 48 cells — but the preregistered reason for expecting failure was wrong. Two panels with identical accuracy, differing only in verifier correlation, behaved completely differently: one immune, one not. Correlation, not accuracy, is the attack surface.
Read the record ↗
The ladder under imperfect panels
E030–E035 all shared one unvaried setting: verification was perfect. Running the ladder on real panels produced a result E037 could not explain — weakening the panel made the archive's lead go up — with the two obvious explanations ruled out and the question recorded as open.
Read the record ↗
A symmetric gate is not a symmetric burden
Explains E037's anomaly. The panel errs at the same rate for every arm, but base viability decides the direction of a review error, so an arm-blind gate is not an arm-neutral gate: it charges the reference arm most and inflates every other arm's lead.
Read the record ↗
Content-addressed blind spot
The blind spot has no address, and giving it one baits the optimiser rather than the attacker. Panel errors here are content-independent, which makes 'can an attacker learn to pass the gate?' unanswerable in this arena — a limitation of the benchmark, stated rather than papered over.
Read the record ↗
Diversity–correlation threshold
Issue #13's hypothesis 2 has no threshold to find, and its verification half is worth nothing: the equal-budget scaling grid never charges for diversity, so correlation cannot flip its sign there.
Read the record ↗
Verifier strictness shock
The R1 lab's 'verifier error correlation' is not verifier correlation at all — it is a within-task strictness shock on a single verifier. It is flat in N and disappears entirely once a second verifier exists.
Read the record ↗
Worker dependence shape
E040's hedge points the wrong way: the beta-binomial puts less mass on total failure than the shared shock, not more. Slopes rise in 16 of 18 cells and proportionality falls from 17 of 18 to 8 of 18 — the conclusion was a property of the assumed shape.
Read the record ↗
The programmes behind the numbers
Individual experiments sit inside longer research tracks. These are the specifications and sweeps they answer to.
Research index
Every programme, experiment interpretation and calibration record, with what each one is and is not allowed to claim.
Open ↗
Swarm diversity vs replication ↗
When do diverse attempts help and when do they hurt? With a help/hurt sweep, a collective-scaling extension and a real-corpus readiness gate.
Randomised scheduling under churn ↗
Capability, staleness, failure and scale effects, with a scale/regime sweep and a capability-rarity sweep.
Evolutionary and stigmergic routing ↗
Evolve policies on training tasks and confirm on held-out work; and route from verified outcomes with evaporation and newcomer exploration.
Debt and backpressure ↗
The controller model for limiting unverified work, plus its multi-window temporal benchmark.
Metric uncertainty
The reporting rules for repository metrics — how much a number is allowed to say, and how uncertainty must travel with it.
Open ↗
Top 20 questions
Prioritised open questions about scaling, verification, governance and community growth. The falsifiable ones are the invitation.
Open ↗
Research questions
The project-level questions the whole programme is organised around, and the flagship matched-budget comparison family.
Open ↗
Reproduce any of it
Every experiment is deterministic given its seed, and each record names the module that produced it, the committed artifact, and the tests that pin the result.
# the simulations behind E024-E039
ls sim/
# the R1/R2 scaling and dependence labs behind E040-E042
ls randomness_lab/
# committed artifacts
ls experiments/results/ results/experiments/
Implemented infrastructure is evidence of a capability to run experiments — not evidence that the research hypotheses are true. A simulator can validate an algorithm’s implementation without proving that the algorithm improves real collaboration. E029 is the reminder: the first attempt at real model attempts on the frozen benchmark produced zero accepted candidates.