Status: synthetic reference experiment
Date: 2026-08-28
Related: issue #14, VERIFICATION_DEBT_AND_BACKPRESSURE.md, ADR-0007
When candidate generation becomes faster than independent verification, can verification-debt backpressure keep the verification queue in a bounded operating region without pretending that generated work is already trustworthy?
This experiment advances the existing one-window Risk-Weighted Verification Backpressure (RWVB) controller into a multi-window queue benchmark. It is a mechanism test, not evidence that RWVB is optimal on real software work.
The benchmark implements five conditions:
The implementation is in:
experiments/verification_backpressure_benchmark.py
Dedicated regression tests are in:
tests/test_verification_backpressure_benchmark.py
Each generated candidate receives deterministic seeded attributes:
The stream is generated from (seed, candidate_index). Therefore fixed policies at the same fan-out receive the exact same candidates in the exact same generation order. Each run records a SHA-256 digest of the generated stream.
Adaptive RWVB intentionally consumes only a deterministic prefix when it throttles generation. Comparing its raw candidate count to a fixed policy is therefore a comparison of closed-loop system behavior, not a claim that it solved the same number of generated candidates.
For every time window:
age pending candidates
-> generate fanout candidates
-> allocate bounded verifier-cost capacity
-> record verifier evidence
-> remove verified candidates
-> measure queue + verification debt
-> adapt next fanout only for rwvb-adaptive
-> next window
No simulated verifier decision grants merge, acceptance, or integration authority.
The experiment preserves raw control/evidence metrics rather than collapsing them into one reward:
risk × (1 + impact) for synthetic defects still waiting);This makes negative trade-offs visible. A scheduler can clear many cheap candidates while retaining more risk-weighted debt, for example.
The checked-in reference result uses:
20 seeds (0..19)
100 verification windows per run
verification-cost capacity = 8 per window
initial fanout = 2, 4, 8, 12
5 policies
Compact aggregate results are stored at:
experiments/results/E014-verification-backpressure-20-seed-summary.json
The full per-run JSON can be reproduced with:
python experiments/verification_backpressure_benchmark.py \
--benchmark \
--seeds 20 \
--steps 100 \
--fanouts 2,4,8,12 \
--capacity 8 \
--output /tmp/e014-full.json
Mean values across the 20 seeds:
| Initial fanout | Policy | Verified | Pending at end | Final debt | Pending defect exposure | Mean wait | Final fanout |
|---|---|---|---|---|---|---|---|
| 8 | FIFO | 503.25 | 296.75 | 366.36 | 96.47 | 18.58 | 8.0 |
| 8 | highest-risk-first | 539.65 | 260.35 | 143.66 | 21.08 | 1.49 | 8.0 |
| 8 | cheapest-first | 605.80 | 194.20 | 416.63 | 61.38 | 2.28 | 8.0 |
| 8 | RWVB fixed | 508.00 | 292.00 | 361.77 | 92.45 | 15.13 | 8.0 |
| 8 | RWVB adaptive | 490.30 | 15.90 | 7.99 | 1.06 | 2.49 | 4.8 |
| 12 | FIFO | 503.75 | 696.25 | 869.84 | 228.61 | 29.03 | 12.0 |
| 12 | highest-risk-first | 537.10 | 662.90 | 475.02 | 90.69 | 1.53 | 12.0 |
| 12 | cheapest-first | 733.70 | 466.30 | 906.75 | 151.68 | 1.18 | 12.0 |
| 12 | RWVB fixed | 506.00 | 694.00 | 866.38 | 226.32 | 25.04 | 12.0 |
| 12 | RWVB adaptive | 490.55 | 15.80 | 8.15 | 1.11 | 2.49 | 4.6 |
The important result is not that adaptive RWVB has the largest verified count. It does not. The result is that under the synthetic overload regime it refuses to keep generating work far above evidence capacity and contracts toward roughly 4–5 candidates/window, leaving a much smaller risk-weighted queue.
Cheapest-first is a useful counterexample: at fanout 12 it verifies substantially more candidates than the other fixed policies, but still ends with very high verification debt. This is why throughput alone is not a sufficient objective.
Highest-risk-first also performs strongly on pending synthetic defect exposure in these runs. That is evidence that the RWVB scheduler itself should continue to be compared against simpler baselines rather than assumed superior.
At initial fanout 2, all fixed policies clear all 200 generated candidates in the reference sweep. Adaptive RWVB detects spare verification capacity and expands generation, reaching about 501 generated / 485 verified candidates on average and a final fanout around 5.1.
This demonstrates the other side of the negative-feedback controller: it can increase generation when verification debt is low instead of permanently suppressing parallelism.
The seven-mode synthetic comparison requested by issue #14 is now implemented
in E022-verification-scaling-matrix.md.
It adds the no-verification, one-reviewer, fixed-quorum, independent-test, and
test-plus-adversarial-review conditions while retaining fixed and adaptive RWVB.
Its findings remain synthetic, so the real-evidence step below is unchanged.
The next #14 experiment should replace synthetic quantities with measured signals from the emerging local IDKMesh loop:
real WorkUnit attempts
-> ResultManifest
-> independent VerificationResult
-> measured verifier cost/time
-> measured queue/backlog
-> candidate risk class
-> human attention where available
Then replay fixed vs adaptive generation policies over the same frozen candidate/verification corpus. The controller should be allowed to fail the experiment. A negative result is useful evidence for changing or removing the policy.
A particularly important follow-up is to compare:
Generation is supply; independent verification is trust capacity. Backpressure may decide how much work to generate or verify, but it never decides that a candidate is true, accepted, or mergeable merely because the queue is stable.