Issue: #31
Depends on: randomness_lab.r2
This layer asks a different question from a single scheduling run:
How do local randomized schedulers behave as the worker population grows, and where does limited local information stop being enough?
The default CLI covers:
1
10
100
1,000
10,000
100,000 workers
This matches the scale ladder in #31.
The global least-loaded oracle performs a full required-capability-pool scan on every routing attempt. That is intentionally an O(n)-information reference. By default it runs through 10,000 workers and is marked skipped at 100,000 workers.
This is not a missing result: the skip is itself part of the scalability model. The 100,000-worker cell exists to exercise the O(1)-sample local policies without pretending that a full scan is an efficient decentralized option.
The threshold is configurable with --oracle-max-workers.
Three named presets are included.
freshThis approximates an information-rich, stable environment.
moderatestaleThese presets change several dimensions together. They are descriptive stress regimes, not causal one-factor experiments. If a result changes between fresh and stale, a later controlled sweep should isolate which factor caused the change.
To keep 100,000-worker simulations computationally bounded, arrivals do not scale linearly without limit.
The first rule is:
arrivals_per_tick = clamp(ceil(worker_count / arrival_divisor), 1, max_arrivals_per_tick)
Defaults:
arrival_divisor = 1,000
max_arrivals_per_tick = 100
This means the very large cells are scalability/protocol-overhead probes, not saturation tests of all available worker capacity.
A later capacity-saturation experiment should scale workload intensity separately and report the cost.
--seeds accepts multiple trace seeds. Each worker-count/regime/seed combination creates one exact R2 trace and runs the applicable policies on that same trace.
The output retains every raw seed-level policy metric and trace digest.
Aggregates report mean/min/max across seeds for:
With two or more seeds, every numeric aggregate also reports the sample standard
deviation and a two-sided 95% Student-t interval for the mean. A one-seed run
reports both fields as null rather than inventing uncertainty evidence. The
interval is clipped to the metric’s physical domain: zero for all reported
metrics and one for rates, utilization, and Jain fairness.
When the oracle is available, every local policy gets a seed-matched comparison.
The first diagnostic loses_badly flag is intentionally simple:
oracle completion advantage >= 0.10
OR
local p95 response >= 2 * oracle p95 response
This is a triage signal, not a scientific significance test. Raw metrics remain authoritative.
The comparison also reports:
local metadata probes / oracle metadata probes
This exposes the quality-versus-information trade-off directly.
A local policy may be somewhat worse in latency while using orders of magnitude less global state. Whether that is a good trade depends on the application and should not be collapsed into one undocumented scalar score.
With one worker, the scale runner adds every task-required capability to that one synthetic worker. This prevents the 1-node baseline from being dominated by an arbitrary impossible capability mismatch.
Larger populations retain the heterogeneous capability distribution from the underlying R2 generator.
python -m randomness_lab.r2_scale \
--worker-counts 1,10,100,1000,10000,100000 \
--seeds 42 \
--regimes fresh,moderate,stale \
--ticks 30 \
--oracle-max-workers 10000 \
--output results/r2-scale.json
For a stronger repeated experiment:
python -m randomness_lab.r2_scale \
--worker-counts 1,10,100,1000,10000,100000 \
--seeds 41,42,43,44,45 \
--regimes fresh,moderate,stale \
--ticks 50 \
--oracle-max-workers 10000 \
--output results/r2-scale-5-seeds.json
The 100,000-worker runs may be materially more expensive than the smaller cells. Keep the raw configuration with every published result.
The first five-seed full-ladder evidence and guarded interpretation are retained
in ../../results/experiments/r2/reference-scale-seeds41-45.md.
The factor-isolated capability-prevalence design is specified in
R2_CAPABILITY_RARITY_SWEEP.md.
The sweep is specifically designed to allow several outcomes:
All are useful findings.
Issue #84 Phase B/C is retained in
reference-factor-isolation-seeds41-45.md.
It varies availability lag, load lag, regional failure correlation, and offered
load separately across five seeds. It also reports directory operations,
modeled messages/bytes, state entries, locality mismatch, and a separate
host-specific CPU/peak-memory profile. The preceding capability-prevalence
benchmark supplies the other Phase B factor.
The combined follow-ups expose both useful and negative regimes. Capability-aware two-choice avoids blind rarity failures, while its latency advantage over simpler local policies narrows as offered load reaches and exceeds capacity.
Remaining extensions include real network packet measurements, distributed directory consistency, multi-resource matching, checkpointing, replicated tasks, malicious workers, and real-fleet CPU/memory measurements. Those require different systems or experiments and are not inferred from the synthetic factor sweep.
Do not claim that power-of-two is universally optimal because of its classic balls-into-bins result or because it wins one R2 cell.
IDKMesh needs the regime map: where the local rule works, where additional capability/state information is worth paying for, and where the assumptions behind a simple randomized scheduler fail.