Date: 2026-10-07
Status: proposed convergence plan
Assessment source: PR #935
Scope: strengthen the parts of IDKMesh that are currently weakest as a professional engineering product and as a falsifiable research program, without broadening the architecture unnecessarily.
IDKMesh already has a strong verification/evidence foundation.
The next goal is not to add more concepts. It is to make the project’s strongest engineering claims true end to end, reproducible on external repositories, and easy for a newcomer to verify.
The central claim to strengthen is:
IDKMesh should increase verified useful software work per unit of human attention and compute, while preserving explicit authority, provenance, and reproducibility boundaries.
The work in this plan is complete only when a skeptical external engineer can reproduce the product path and inspect evidence for that claim without relying on project-authored interpretation.
The 2026-10-07 engineering assessment identified six gaps that materially limit the product today.
Current experiments establish important verifier and coordination findings, but they do not yet prove that the IDKMesh workflow beats a strong single-agent/simple baseline on real external software work.
The local runner correctly avoids claiming that raw subprocess execution is a hostile-code sandbox. A production sandbox gate remains open in #804.
Issue #921 identifies the readiness/claim/input-binding time-of-check/time-of-use gap. A professional executor must prevent stale work from being dispatched or canonically submitted.
The repository can explain and exercise many components, but a newcomer should not have to learn the whole architecture to execute:
real issue
-> bounded WorkUnit
-> admitted worker
-> candidate
-> independent verification
-> evidence
-> explicit human decision
Issue #713 already tracks the correct API program. The missing work is convergence and qualification, not another API design.
The package, website, and external installation evidence should reflect current supported behavior and make experimental versus production-ready capability obvious.
This plan deliberately does not require:
The plan is a convergence program.
The engineering-strengthening program is done when all of the following are true.
A new user can install a released IDKMesh package and complete one documented real-repository run that reaches:
bounded task
-> exact admitted execution binding
-> two materially different worker attempts
-> canonical candidate normalization
-> independent verification
-> non-selecting evidence report
-> human accept/reject/escalate decision
without either worker or verifier gaining merge authority.
For local hostile-code execution:
A preregistered real-task benchmark compares at least:
The comparison uses matched task snapshots and transparent budgets.
The primary result reports verified useful work per unit of human attention and compute, not agent count or task count.
At least one second repository outside the IDKMesh codebase completes a bounded pilot with:
The declared beta API scope has:
Do not attack all gaps independently.
Use this dependency order:
W0 capability truth + claim freeze
|
v
W1 one golden path
|
+--------------------+
| |
v v
W2 runtime trust W3 benchmark protocol
| |
+---------+----------+
|
v
W4 external pilot
|
+-------+-------+
| |
v v
W5 API beta W6 release/distribution
| |
+-------+-------+
|
v
W7 public evidence
Some work can execute in parallel, but no later claim should be made before its upstream evidence gate is satisfied.
Create one canonical map of what IDKMesh can prove today and prevent the website, README, package metadata, paper, and API docs from drifting apart.
Add a maintained capability matrix with one row per public engineering question.
Minimum columns:
implemented;experimental;planned;Example:
| Question | Status | User surface | Evidence | Limitation |
|---|---|---|---|---|
| Can I measure verifier independence? | implemented | idkmesh gate-audit |
E015-E017 + tests | requires ground-truthed verdict data |
| Can I safely run hostile local code? | planned/experimental | local agent runner | #804 | production sandbox not yet complete |
| Can I deploy a public multi-user API? | planned | API program | #713 | local profile is not public-service auth |
A newcomer can answer, in under one document:
Turn the current Product Spine, local loop, evidence, and decision concepts into one obvious reference workflow.
This is the product path every later experiment and release should reuse.
1. intake real bounded issue/request
2. derive/freeze WorkUnit
3. compute exact source + input binding
4. explain routing/admission
5. dispatch two isolated attempts
6. observe exact candidate revisions/artifacts
7. normalize to ResultManifest
8. run verifier-owned EvaluatorPlan
9. produce VerificationResult per candidate
10. produce non-selecting Run Evidence Report
11. inspect in Control Tower
12. record explicit human accept/reject/escalate
13. normal protected integration remains outside worker/verifier authority
The reference workflow should be executable with a small number of supported commands.
The user should not need to call internal Python modules manually.
The final command vocabulary may evolve, but it should feel like one product:
idkmesh work preview ...
idkmesh route explain ...
idkmesh run create/execute ...
idkmesh run status ...
idkmesh run evidence ...
idkmesh decision record ...
idkmesh control-tower ...
Do not create duplicate commands if an existing command can be extended safely.
Retain one small repository task that includes:
The golden path must visibly demonstrate:
A newcomer can complete the full path from one bounded request to an inspectable human decision artifact without reading historical conversation documents.
This work is mandatory before stronger claims about safe autonomous local coding.
Owner: #804.
subprocess.Popen can remain an internal process primitive but must never satisfy the hostile-code sandbox capability contract.
Tests must prove:
Owner: #921.
project scope
+ WorkUnit digest
+ source revision
+ input/dependency digest
+ readiness state
+ task claim
+ fencing epoch
+ resource/admission decision
= one execution grant
Revalidate the exact bound inputs:
Use explicit failures such as:
task_not_ready;task_already_integrated;binding_mismatch;inputs_changed;stale_epoch;not_claim_owner.Do not silently refresh a stale binding.
No supported local coding-agent path can execute without a conforming sandbox, and no execution/candidate can cross the canonical boundary after relevant inputs changed.
This is the most important research-strengthening work.
Answer:
For real repository tasks, when does IDKMesh produce more verified useful work per human attention and compute than simpler alternatives?
The experiment should be capable of disproving the IDKMesh advantage.
One strong coding agent receives the same bounded task and budget.
It may use ordinary tests/tools, but does not receive extra parallel candidate generation.
Multiple workers attempt the task.
Use a simple aggregation/selection strategy such as:
without IDKMesh’s correlation-aware/evidence-first strategy.
Use:
Use #5 as the bootstrap cohort owner.
Start with 5 replayable tasks, then expand only after the first cohort is valid.
The full evidence target should later include enough tasks for meaningful paired analysis; do not choose the final size merely to obtain statistical significance.
Task families should include:
Avoid tasks that can only be judged by subjective style.
For every task retain:
Do not change acceptance rules after seeing candidate outputs.
Use a transparent multi-objective record, with this leading quantity:
verified useful work
-------------------------------
human reviewer attention + compute/provider cost
Do not hide the numerator or denominator behind one opaque score.
Record the components separately.
At minimum:
Measure actual review/decision minutes where practical.
Do not count owner-controlled autonomous compute as human attention.
Record:
Per task/arm record:
Where multiple verifiers are used:
Use paired task comparisons where possible.
Report:
Do not discard failed tasks because they make the system look worse.
Preregister the expansion rule.
Example:
Do not stop early simply because Arm C leads.
The repository can make a bounded empirical statement such as:
On task families X/Y/Z under the declared budget and verifier configuration, the IDKMesh workflow changed verified useful work per reviewer-minute by N relative to the single-worker baseline, with these confidence intervals and failure modes.
A valid result may be positive, neutral, or negative.
Move the central result off the IDKMesh repository.
Owner: #599.
Use a fresh, non-safety-critical repository.
Required evidence:
Publish a retrospective with four sections:
Do not classify the pilot as successful merely because 10 WorkUnits ran.
Success means the product path was usable, trustworthy, recoverable, and inspectable.
A different repository reaches a real release using the same public IDKMesh interfaces.
Make the API support the proven product path, not expand into a second product.
Owner: #713 and its child issues.
The first beta does not need every future control-plane capability.
It does need stable support for:
“API v1 beta” refers to a tagged, reproducible contract with a qualification artifact, not merely endpoints on main.
Make the strongest current capabilities consumable without cloning the research repository.
Owner: #769.
Required:
Owner: #374.
The release should include one bounded canonical end-to-end example with:
Owner: #401.
Collect genuine external evidence across:
Record failures as evidence.
A user can install the released package from PyPI and complete the supported reference workflow without a source checkout, except where a real repository checkout is inherently part of the task.
The public story should be generated from or checked against current supported capability.
The front door should answer only:
Then link into research/architecture depth.
CI should execute or parse every copy-paste command shown on the main start page.
Avoid manually duplicated commands when generated snippets are practical.
Every major public claim should expose one of:
Show the release version or “verified against source revision” for command-heavy pages.
Publish only evidence-backed cases:
A visitor does not need GitHub issue archaeology to know what is supported.
Do not create duplicate implementation issues for these responsibilities.
| Need | Canonical owner |
|---|---|
| central multi-agent scientific benchmark | #1 |
| first 5–10 benchmark tasks | #5 |
| local Verified Swarm Runner | #16 |
| runner/package reproducibility gate | #374 |
| docs/paper/repo synchronization | #471 |
| second-project no-server pilot | #599 |
| API professionalization | #713 |
| first PyPI release | #769 |
| external install/platform evidence | #401 |
| production local sandbox | #804 |
| atomic executor admission/stale-input checks | #921 |
| engineering assessment | PR #935 |
New issues should be created only for a bounded missing slice that has no existing owner.
These should not distract from proving the current core claim.
Use hard gates to avoid architecture expansion without evidence.
Pass only when #804 + #921 requirements are satisfied for the declared local execution profile.
If not, continue to label local hostile-code execution experimental.
Pass only when the golden path is executable through supported interfaces.
If not, do not add provider breadth.
Pass only when the baseline benchmark is complete.
Possible decisions:
Pass only after #599 produces inspectable external-project evidence.
If not, treat self-hosting evidence as insufficient for general adoption claims.
Pass only after the declared API scope is schema-frozen, qualified, and tagged.
Pass only when released install, docs, benchmark, and external-pilot evidence agree on the same capability boundary.
Do not use raw agent activity as the top-level success metric.
Track these categories.
Use this hierarchy in README, website, release notes, and papers.
Specified but not implemented.
Code + deterministic tests exist.
Mechanism executed against real runtime/work, but within owner-controlled project conditions.
Compared against frozen baselines under preregistered metrics.
Independently or externally operated use confirms the behavior.
Security/reliability/deployment qualification exists for the declared production profile.
Do not use a higher-level wording for lower-level evidence.
Example:
The single most valuable result would be a table like this, backed by replayable evidence:
| Metric | Strong single agent | Naive multi-agent | IDKMesh |
|---|---|---|---|
| tasks attempted | N | N | N |
| verifier-accepted regression-free changes | … | … | … |
| post-integration defects | … | … | … |
| median reviewer minutes | … | … | … |
| compute/provider cost | … | … | … |
| wall time | … | … | … |
| nominal verifier count | … | … | … |
| effective verifier votes | … | … | … |
| control failures | … | … | … |
| stale/unsafe executions blocked | … | … | … |
The point is not to force the IDKMesh column to win every row.
The useful scientific result is to discover:
which task/risk/verification regimes justify the extra coordination and verification cost.
That answer would be stronger than a universal “multi-agent is better” claim.
Until the external benchmark and pilot graduate, use:
IDKMesh is a verification-first control layer for human and AI software work. It turns bounded tasks into untrusted candidates, binds them to exact provenance, measures independent verification evidence, and keeps final integration authority separate from workers and verifiers.
Avoid presenting the project as a generally proven large-scale swarm system.
After the benchmark/pilot, strengthen the wording only to the level supported by the data.
The next concrete sequence should be:
Do not add another major architectural layer until one of these gates demonstrates a concrete need for it.
The shortest path to making IDKMesh stronger is now:
fewer new concepts
+ stronger runtime boundaries
+ one simple product path
+ real baseline comparison
+ external repository evidence
+ qualified release
= professional credibility