Status: experimental, required before a successor to the burned Phase B2 first-five pilot.
The first frozen Phase B2 pilot exposed an ambiguity that must not be repaired by changing history.
EvaluatorPlan v0.2 uses:
"required_added_text": ["..."]
Deterministic patch verifier v0.1.1 interprets each value as an exact full added-line requirement. The burned pilot’s plans instead contained semantic fragments. Once a real task solution exposed that mismatch, changing v0.2/v0.1.1 to substring matching would have changed the meaning of already-frozen evidence.
Therefore v0.2 and verifier v0.1.1 remain unchanged. Substring semantics use a new schema and verifier version.
| EvaluatorPlan | Verifier adapter version | Semantic field | Meaning |
|---|---|---|---|
| v0.2 | deterministic-patch-verifier 0.1.1 |
required_added_text |
every configured value must equal one complete added line |
| v0.3 | deterministic-patch-verifier 0.2.0 |
required_added_substrings |
every configured value must occur within at least one individual added line |
The canonical experiments/evaluator_plan_runner.py routes both versions. It does not reinterpret a v0.2 plan as v0.3.
For each required_added_substrings[i] = s, verification succeeds for that semantic requirement iff there exists an added line L parsed from a structurally valid unified-diff hunk such that:
s is a contiguous, case-sensitive substring of L
More formally:
match(s, A) = 1 iff exists L in A such that s ⊑ L
where A is the set/list of validated-hunk added lines and ⊑ means contiguous substring containment.
All configured substrings are required:
semantic_pass = AND_s match(s, A)
The operation is deliberately simple:
experiments/substring_patch_verifier.py is a versioned semantic adapter over the existing hardened metadata-only patch-verifier core.
It independently parses the same candidate patch, maps each required substring to a concrete matching added line when one exists, and delegates structural diff validation, artifact hashing, log integrity, WorkUnit scope, and required-check construction to the unchanged v0.1.1 core.
The resulting VerificationResult adds an explicit added-substring-semantic-observation evidence item containing:
This makes the new semantic decision auditable without changing the historical verifier implementation.
A v0.3 VerificationResult must preserve:
deterministic-patch-verifier;0.2.0;provenance.verifier_config_digest after canonical runner binding;added_line_substring_all;0.1.1 as an implementation provenance detail.A plan/result version mismatch fails closed in the canonical runner.
The repository uses the same candidate patch for the version boundary test. The patch adds:
<!-- patch-evaluator expected -->
Both contrast plans require the fragment:
patch-evaluator expected
Expected behavior:
EvaluatorPlan v0.2 + verifier 0.1.1 -> reject
EvaluatorPlan v0.3 + verifier 0.2.0 -> support
The v0.2 rejection is intentional evidence that historical exact-line meaning has not changed. The v0.3 support proves that substring meaning is introduced only under the new version.
The self-test also requires both versions to report the same candidate-patch digest, proving the semantic comparison did not change candidate bytes.
A successor Phase B2 cohort may use v0.3 only after:
The burned first-five cohort must remain burned with its original definition digest. Task 001 is already known and cannot be presented as untouched held-out evidence in a successor cohort.
Neither v0.2 nor v0.3 executes candidate code. Neither verifier recommendation grants:
Verification remains decision-support evidence for a later integration/human-governance stage.
Related: issue #157, PR #153, PR #160, benchmarks/phase-b2-first-five/BURN_NOTICE.md, schemas/evaluator-plan-v0.2.schema.json, and schemas/evaluator-plan-v0.3.schema.json.