IDKMesh · evidence

Every experiment, with the result it actually got

Thirty-two numbered experiments, each linked to the record in the repository that produced it. The negative ones are here because they are here in the tree: a mechanism that fails under a frozen test is evidence, not something to hide — and several of the entries below exist only to falsify an earlier one.

32
numbered experiments, E011 to E042
8
reached a positive result at their stated scope
15
corrected, qualified or falsified an earlier experiment
4
negative results, kept in the tree

Read the status before the claim

Every artifact in this repository sits on an explicit evidence ladder, and most of these experiments deliberately stop short of the top rung. A benchmark contract plus a synthetic fixture means the experiment is runnable — never that a strategy has been shown superior.

positive

Reached the result it set out to test, at its own stated scope and no wider.

corrective

Qualified, narrowed or falsified an earlier experiment — including several of this project’s own headline claims.

negative

The mechanism did not work, or could not be measured at all. Retained in full.

prototype

Synthetic mechanism test or preregistration. Tests the mechanics; proves nothing about real collaboration.

The full evidence ladder, as a figure →

The atlas

Filter by what the experiment actually concluded. Each card links to its record in the repository, which carries the module that produced it, the machine-readable artifact, and the tests that pin it.

CSS only — this site runs no JavaScript.
E011 prototype

Emergence from vague goals

Can exploratory agents move toward a coherent system when the objective is incomplete, given hard constraints, diversity preservation, evidence-driven selection and shared memory? Deliberately synthetic — mechanisms tested before any real engineering task.

Read the record ↗

E012 prototype

Correlated verification failure

Replaces the perfect viability oracle with an imperfect panel. As verifier errors correlate, majority-vote error rises while panel disagreement falls — so the panel looks more confident exactly as it becomes less reliable.

Read the record ↗

E013 prototype

Independence-aware aggregation

At the same nominal review cost and the same votes, when should a large correlated verifier cluster be discounted rather than counted vote-by-vote? The rule should help when a cluster shares failure modes — and must not be assumed to help when reviewers are genuinely independent.

Read the record ↗

E014 prototype

ACO stigmergic task routing

Proposed, with an executable synthetic baseline: can evidence-backed pheromone routing allocate tasks better than simpler rules without causing herding, duplication or review concentration?

Read the record ↗

E015 positive

Verification phase diagram

Sweeps E012's single operating point into a phase diagram: given panel size, accuracy, correlation and quorum, how many statistically independent verifiers is a panel actually worth? Cost-asymmetric quorums help, and the N_eff heuristic is optimistic.

Read the record ↗

E016 negative

Live verifier correlation

Negative result, and an honest one — its own status line reads “the verifiers it deployed do not verify”. The 20 LLM verifiers averaged Youden J = +0.0487, and their majority vote scored 0.514 accuracy against 0.639 for a trivial “always reject”. Correlation could not be estimated from votes that carry almost no signal. Screen a panel for discrimination before trusting anything it votes on — and note what this leaves open: no AI review panel has been measured in this repository.

Read the record ↗

E017 positive

Item difficulty and quorum

The first observed measurement of verifier error correlation here, on 25 independently seeded partial test oracles — programs, graded on real defects — over a 72-candidate corpus: ρ = 0.5873 across 300 pairs, an effective panel size of 1.00 of 25 under majority vote, and the shared-shock model shown to be the wrong shape. Fixing the aggregation rule cut error 3.7×; growing the panel bought nothing.

Read the record ↗

E018 corrective

Dependence model shape

Which E015 conclusions depend on the shape of the dependence model? One correction to E015 and one strengthening of it — the direction of the N_eff result survives, the magnitude does not.

Read the record ↗

E019 corrective

Group independence under item difficulty

E013's crossover is robust to the model's shape — it sits at ρ = 0.25 under both — and disappears entirely when the declared independence groups are not actually independent. Declared structure is not measured structure.

Read the record ↗

E020 corrective

Quorum frontier under the measured shape

Extends E017 §6 to the full frontier with the corpus base rate, and qualifies it twice: the recommended beta-binomial under-predicts the unanimity floor by 1.77×, and the shared-shock model reports that no quorum beats majority — getting the highest-leverage decision on that panel exactly backwards.

Read the record ↗

E021 corrective

Coordination criticality

Matched small-load susceptibility probes give earlier overload warning at a measurable false-alarm cost — but no evidence that susceptibility dominates an ordinary utilisation threshold.

Read the record ↗

E022 positive

Seven-mode verification scaling matrix

A matched comparison of every verification condition issue #14 names: can candidate generation and independent verification scale together without collapsing into unchecked acceptance or an unbounded review queue?

Read the record ↗

E023 prototype

First-review latency and contributor recurrence

Preregistered against zero observations: the specification, the randomised design it would take to earn a causal claim, and the threats to validity — written before any data, precisely so the analysis cannot be chosen after seeing it.

Read the record ↗

E024 positive

Matched-budget emergence

Every strategy issue #22 names gets exactly the same proposal and verification-attempt budget. The Quality-Diversity archive's surviving advantage is reliability: it never fails catastrophically, where the majority-vote swarm fails in 44 of 100 seeds. Its own stated limitation drove E030 and E031.

Read the record ↗

E025 corrective

Learned verifier reliability

Learned weighting helps in stable regimes and fails materially under plausible shift. Explicitly not production-ready reputation — the failure mode is recorded rather than smoothed over.

Read the record ↗

E026 positive

Imperfect verifier panel

E024 rerun with E017/E020's measured correlated panel and blind-spot floor. Nothing moved: sweeping the panel to 45% wrong in both directions changed the archive's post-change AUC by 1.5% and left catastrophes at 0/100. A null result, kept.

Read the record ↗

E027 positive

Defect propagation

Gives an accepted defect an actual cost, so verifier error can finally reach the outcome metric, and sweeps that cost across its whole range. The constraint-guided archive holds at 0/100 in all twenty cells while unconstrained random search goes 0/100 to 94/100.

Read the record ↗

E028 corrective

Latent defect dimension

Removes E027's confound by moving viability into a dimension the goals cannot see. The archive's survival holds in 18 of 20 cells and breaks only under the stress panel at full defect cost.

Read the record ↗

E029 negative

First real model attempts

The first negative result this repository was entitled to: 60 sandboxed attempts by a pinned 0.5B open-weight producer on the frozen benchmark produced 0 accepted candidates, and 56 of the 60 failures were failures of the unified-diff protocol before any repository content was consulted.

Read the record ↗

E030 positive

Supplied-goal membership

Removes E024's supplied-oracle confound by switching to a parity-matched goal the arms do not hold. The archive keeps all but 1.6–4.4% of its lead and stays 0/100 catastrophic; the majority-vote swarm loses its whole lead in every panel.

Read the record ↗

E031 corrective

Learned goal filter

Gives the consensus swarm a particle filter that learns the goal from ordinal evidence. Learning from generation 0 roughly doubles its catastrophic seeds; learning from post-change evidence alone is the best variant in all eight cells. An evidence-free rescue takes 38/100 catastrophic seeds to 0/100 — and to 71/100 once the new goal is outside the supplied set.

Read the record ↗

E032 corrective

Population scaling

Answers issue #13's question — at a fixed budget, when is another agent worth adding? — by running the sweep with the budget held and with it free. The two disagree: the archive gains on every doubling with a free budget and gains nothing at all when it is held.

Read the record ↗

E033 corrective

Goal distance

Turns E030's single substitute goal into a ladder of rings at a matched change size. The archive's lead decays smoothly and is gone by 0.35 from the supplied set. E030's published point ranks second of seven at its own distance — retention there averages 78.3%, not the 95.6% one goal reported.

Read the record ↗

E034 corrective

Goal direction

At a fixed distance, direction matters as much as distance: on a single shell the archive's lead runs from −4.894 to +4.471 and is negative for 24.2% of directions. The mechanism E033 proposed for that spread is falsified by its own preregistered test.

Read the record ↗

E035 corrective

Direction across shells

Re-runs E034's ladder at further distances. The spread does not grow with distance as conjectured, only one ladder in five replicates, and the geometry admits very few valid shells at all — a resolved effect on one shell is not a general one.

Read the record ↗

E036 corrective

Adversarial contributors

No adversary made the constraint-guided archive worse than another arm in any of 48 cells — but the preregistered reason for expecting failure was wrong. Two panels with identical accuracy, differing only in verifier correlation, behaved completely differently: one immune, one not. Correlation, not accuracy, is the attack surface.

Read the record ↗

E037 corrective

The ladder under imperfect panels

E030–E035 all shared one unvaried setting: verification was perfect. Running the ladder on real panels produced a result E037 could not explain — weakening the panel made the archive's lead go up — with the two obvious explanations ruled out and the question recorded as open.

Read the record ↗

E038 positive

A symmetric gate is not a symmetric burden

Explains E037's anomaly. The panel errs at the same rate for every arm, but base viability decides the direction of a review error, so an arm-blind gate is not an arm-neutral gate: it charges the reference arm most and inflates every other arm's lead.

Read the record ↗

E039 corrective

Content-addressed blind spot

The blind spot has no address, and giving it one baits the optimiser rather than the attacker. Panel errors here are content-independent, which makes 'can an attacker learn to pass the gate?' unanswerable in this arena — a limitation of the benchmark, stated rather than papered over.

Read the record ↗

E040 negative

Diversity–correlation threshold

Issue #13's hypothesis 2 has no threshold to find, and its verification half is worth nothing: the equal-budget scaling grid never charges for diversity, so correlation cannot flip its sign there.

Read the record ↗

E041 negative

Verifier strictness shock

The R1 lab's 'verifier error correlation' is not verifier correlation at all — it is a within-task strictness shock on a single verifier. It is flat in N and disappears entirely once a second verifier exists.

Read the record ↗

E042 corrective

Worker dependence shape

E040's hedge points the wrong way: the beta-binomial puts less mass on total failure than the shared shock, not more. Slopes rise in 16 of 18 cells and proportionality falls from 17 of 18 to 8 of 18 — the conclusion was a property of the assumed shape.

Read the record ↗

The programmes behind the numbers

Individual experiments sit inside longer research tracks. These are the specifications and sweeps they answer to.

Reproduce any of it

Every experiment is deterministic given its seed, and each record names the module that produced it, the committed artifact, and the tests that pin the result.

# the simulations behind E024-E039
ls sim/

# the R1/R2 scaling and dependence labs behind E040-E042
ls randomness_lab/

# committed artifacts
ls experiments/results/ results/experiments/
One caveat worth stating plainly

Implemented infrastructure is evidence of a capability to run experiments — not evidence that the research hypotheses are true. A simulator can validate an algorithm’s implementation without proving that the algorithm improves real collaboration. E029 is the reminder: the first attempt at real model attempts on the frozen benchmark produced zero accepted candidates.