Status: Experimental product-facing aggregation layer
Related: #4, #5, #16, PR #72, PR #78, schemas/result-manifest-v0.1.schema.json, schemas/verification-result-v0.1.schema.json
The Verified Swarm Runner needs one place where a human can inspect the evidence from a multi-attempt run without collapsing worker claims and verifier evidence into a vote.
The current canonical trust pipeline is:
WorkUnit
-> worker attempt
-> ResultManifest (worker self-report)
-> independent verifier
-> VerificationResult (evidence + recommendation)
-> human / governance integration decision
Run Evidence Report v0.1 is a read-only aggregation view over that pipeline. It does not replace any canonical protocol object above and does not recreate the closed candidate-level Evidence Report proposal from PR #42. The independent verifier protocol is VerificationResult v0.1.
The report answers a narrower product question:
What happened to every attempt in this run, what independent evidence exists for each attempt, where do verifiers disagree or fail, and is the saved run reproducible?
A generated report MUST preserve these boundaries:
worker success != verified correctness
verifier recommendation != integration decision
multiple recommendations != majority truth
report generation != candidate selection
replay equality != correctness
The generated report therefore has:
{
"human_decision": {
"status": "pending",
"selected_attempt_id": null,
"integration_authority": "external_human_or_governance"
},
"authority": {
"canonical_state_write": false,
"git_push": false,
"merge": false,
"automatic_candidate_selection": false
}
}
A later human/integration-decision record may refer to this report, but the generated report itself cannot be mutated into acceptance authority.
The first implementation consumes the deterministic run record emitted by:
experiments/two_attempt_orchestrator.py
That run record already retains:
Before rendering, run_evidence_report.py validates those relationships again and fails closed on binding drift.
In particular:
VerificationResult.result_manifest_digest == summarized ResultManifest digest
VerificationResult.work_unit_digest == run WorkUnit digest
If either relationship is inconsistent, no report is produced.
Every attempt is mapped to exactly one report state:
| Evidence state | Meaning |
|---|---|
supported |
independent verifier recommendation is accept_candidate |
rejected |
independent verifier recommendation is reject_candidate |
inconclusive |
verifier says escalate / insufficient_evidence, or no conclusive recommendation exists |
worker_error |
worker adapter failed before producing usable candidate evidence |
result_manifest_error |
candidate ResultManifest could not be collected/parsed |
verification_error |
candidate exists but verifier control path failed |
The report never silently drops an errored attempt. One failed worker must not erase the evidence produced by a surviving peer.
verification_disagreement=true when the run contains more than one distinct independent verifier recommendation.
Example:
attempt A -> worker says succeeded -> verifier says accept_candidate
attempt B -> worker says succeeded -> verifier says reject_candidate
The correct product behavior is:
preserve both histories
-> show disagreement
-> leave integration decision pending
not:
count votes -> choose winner
This becomes more important later when verifier failures are correlated. IDKMesh already has research showing that raw verifier count is not equivalent to independent evidence.
Replay is a provenance/reproducibility check, not a correctness proof.
For a saved deterministic run record R and its replayed record R':
ReplayMatch = SHA256(canonical(R)) == SHA256(canonical(R'))
The current fixture orchestration has no timestamps or nondeterministic worker execution, so complete-record equality is expected.
When real workers are added, replay semantics will need to distinguish:
Do not weaken v0.1 silently. Introduce an explicit later replay mode/version when real-worker evidence requires it.
Self-test:
python experiments/run_evidence_report.py self-test
Generate one deterministic run plus JSON and Markdown evidence views:
python experiments/run_evidence_report.py generate \
--config examples/orchestration/two-attempt-good-vs-bad.json \
--run-output results/orchestration/demo.run.json \
--report-json results/orchestration/demo.evidence.json \
--report-markdown results/orchestration/demo.evidence.md
Re-render an existing saved run:
python experiments/run_evidence_report.py report \
--run-record results/orchestration/demo.run.json \
--report-json results/orchestration/demo.evidence.json \
--report-markdown results/orchestration/demo.evidence.md
Check deterministic replay:
python experiments/run_evidence_report.py replay-check \
--config examples/orchestration/two-attempt-good-vs-bad.json \
--run-record results/orchestration/demo.run.json
All generated files are constrained to results/. The tool does not write canonical repository state.
See:
schemas/run-evidence-report-v0.1.schema.json
The schema is intentionally an aggregation schema. It references digests and summaries of canonical worker/verifier evidence; it is not a second verifier protocol.
The built-in self-test checks that:
pending;This closes one control-plane gap in #16 without pretending the real worker path is complete.
The remaining critical path is still:
canonical local node (#34)
-> controlled Docker acceptance (#37)
-> repository-candidate verification (#5 B1 / patch verifier work)
-> plug real node adapter into the already-landed orchestration kernel
-> use this evidence/replay surface over real attempts
-> add one trivial heterogeneous real adapter
VerificationResult v0.1;