Date: 2026-10-07
Status: publication strategy / research agenda
Repository baseline inspected: main@7ac46d39ba629fce7dafc4b25c69cbcee5c1f03e
Evidence rule: implemented mechanisms, synthetic experiments, observed real runs, and independently reviewed empirical evidence are different evidence classes and must never be merged rhetorically.
IDKMesh should not try to publish the claim that “swarms are better.”
The scientifically useful thesis is narrower and falsifiable:
Under fixed compute, verification, communication, and human-attention budgets, when does adding heterogeneous human/AI participation increase independently verified useful work, when does it merely duplicate correlated error, and when does coordination or verification cost make the collective worse?
That thesis creates a coherent research program around four quantities:
[ ext{useful generation} imes ext{composability} imes ext{verification quality} - ext{coordination cost} - ext{failure cost}. ]
The expression is a scaffold, not a proven law. Each paper should operationalize only the terms it actually measures.
Do not optimize raw agents, commits, votes, tokens, or accepted candidates. Report Verified Useful Work per Unit of Scarce Resource (VUWSR) as a vector:
A paper may define one primary endpoint, but it must keep the underlying cost/quality vector visible.
A question should become a manuscript only when it has all six:
This is the filter used below.
This is a targeted novelty check as of 2026-10-07, not a systematic literature review.
Several adjacent results now exist, so IDKMesh should position papers around what its evidence uniquely adds rather than around broad slogans.
| Related work | What is already established | Consequence for IDKMesh |
|---|---|---|
| Kim et al., Correlated Errors in Large Language Models, ICML 2025 | Large-scale evidence that LLM errors are correlated, including across strong models/providers. | “LLM errors are correlated” is not novel enough by itself. |
| Kohli, Nine Judges, Two Effective Votes, 2026 | A 9-judge LLM panel can contain only about two independent votes; the best single judge can match/outperform the panel. | “Reviewer count is not evidence count” now has direct prior art. IDKMesh must lead with dependence shape, executable ground truth, partial failures, blind spots, and quorum consequences. |
| Zhao et al., CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation, ICML 2026 | Correlated judge errors can arise from latent confounders; confounder-aware aggregation is an active method direction. | A new IDKMesh verifier paper must compare model/aggregation assumptions, not imply majority vote is the only alternative. |
| Iyer et al., Multiagent Quality-Diversity for Effective Adaptation, ECAI 2025 | Quality-diversity can improve adaptation in multi-agent learning. | “QD helps adaptation” is not enough. IDKMesh’s differentiator is goal geometry + imperfect/correlated verification + latent defects + adversarial gate pressure. |
| Krentsel et al., Reality Is the Final Verifier, 2026 | Agentic software engineering has unavoidable requirement/model gaps; verification is a resource-allocation problem. | IDKMesh’s verification-capacity work should provide measured/control evidence, not only repeat the verification-first argument. |
The canonical paper at paper/main.tex is still the closest IDKMesh asset to submission, but its strongest novelty is not the headline fact that correlated panels contain fewer effective votes.
The more defensible contribution stack is:
A stronger positioning/title candidate is therefore:
Dependence Shape Matters: Partial Failures, Blind Spots, and Quorum Limits in Executable Verification Panels
The current title may remain for continuity, but the abstract/introduction should make the six items above—not “many reviewers are not many independent votes”—the novelty claim.
| Rank | Research question | Null / falsifier | Current evidence | Missing gate | Publication state |
|---|---|---|---|---|---|
| 1 | RQ1 — What dependence structure actually governs verifier-panel failure? | Pairwise correlation plus a simple shared-shock model predicts the observed failure-count/quorum behavior adequately. | E017–E020; 25 real executable partial oracles; measured failures and blind spots. | Replication on at least one materially different verifier/task family would strengthen external validity. | Closest to submission. Reposition novelty. |
| 2 | RQ2 — Under fixed total budget, when is another coding agent worth adding? | After matching compute/verification budget, heterogeneous multi-agent systems provide no positive marginal verified value over the best single/replicated baseline. | R1/E032/E040 synthetic mechanism work. | Frozen real held-out coding corpus (#70), independent verification, real dependence measurement. | Highest-impact next empirical paper. |
| 3 | RQ3 — When does diversity beat replication, and what does diversity cost? | Any apparent gain disappears when candidate quality, budget, and correlation are controlled. | R1 help/hurt sweeps show correlation controls effect size while worker-quality penalty can flip the sign. | Same as RQ2; must measure real diversity/correlation instead of configuring it. | Combine with RQ2, do not salami-slice. |
| 4 | RQ4 — How should verification capacity control generation pressure? | Adaptive backpressure does not improve the defect/backlog/cost Pareto frontier over simpler fixed policies. | E021/E022/E043/E044; AVE ablations; deterministic queue/control mechanisms. | Replay/shadow evaluation on real workload traces and observed review/service times. | Strong systems/control paper after one real trace layer. |
| 5 | RQ5 — Does diversity-preserving memory improve adaptation after goal change under imperfect verification? | Once verifier imperfection, latent defects, goal distance/direction, and adversarial contributors are introduced, the archive no longer has robust adaptation advantage. | E024, E026–E039 contain a deep synthetic sequence with counterexamples. | Consolidated preregistered analysis and sharper comparison with current QD literature. | Strong mechanism/simulation paper. |
| 6 | RQ6 — What WorkUnit granularity minimizes total verified completion cost? | No interior optimum exists; finer or coarser decomposition monotonically dominates. | Protocol and decomposition infrastructure exists. | Controlled real benchmark that varies granularity on the same tasks. | Promising but under-measured. |
| 7 | RQ7 — When is adaptive routing worth its own state/exploration overhead? | Adaptive routing never beats simpler policies enough to repay route/state burden under matched load. | R2, ACO/stigmergic work, PHY-0/PHY-1; stationary Physarum negative control already useful. | Complete matched stress matrix: topology, churn, failures, stale estimates, nonstationarity. | Medium readiness. |
| 8 | RQ8 — Is review capacity a carrying-capacity limit for agentic repositories? | More generated activity continues to increase verified throughput even when review service is saturated. | ACE simulation, E043/E044, collaboration observables. | Longitudinal real repository data; avoid causal language without intervention/randomization. | Mechanism paper possible; causal community paper not ready. |
| 9 | RQ9 — Can hierarchical coordination preserve decisions with sublinear global communication? | Coarse-graining materially degrades scheduling/verification decisions before it saves enough coordination traffic. | Architecture/roadmap hypothesis. | Multi-scale simulator + controlled deployment evidence. | Future. |
| 10 | RQ10 — Which trust/provenance mechanism is sufficient across independent organizations? | Simpler signed/transparency-log designs satisfy the demonstrated threat model; stronger consensus/ledger machinery adds cost without needed assurance. | Architecture questions and provenance foundations. | Real multi-organization threat model/deployment. | Future; do not build paper around hypothetical scale. |
Preferred positioning:
Dependence Shape Matters: Partial Failures, Blind Spots, and Quorum Limits in Executable Verification Panels
Core RQs
Primary hypotheses
Why it is publishable
The paper has observed executable evidence, falsified modeling assumptions, an explicit negative result, and reproducible tooling. It also connects modern agent verification to the older N-version programming / input-difficulty literature in a technically meaningful way.
Main risk
External validity. One constructed 25-oracle panel over 72 items cannot justify universal claims about LLM judges, static analysers, code reviewers, or all verification systems.
Best strengthening experiment
Repeat the exact analysis protocol on one second panel with a genuinely different failure mechanism. Prefer a real code-analysis/test-generation/LLM-review setting with execution-decided labels.
Do not split RQ1 and quorum into separate thin papers. E017/E018/E020 are strongest as one coherent dependence-shape story.
Working title:
When Is Another Agent Worth Adding? Fixed-Budget Scaling, Diversity, and Correlated Failure in AI Software Engineering
This should combine population scaling and diversity-vs-replication.
At fixed candidate-generation and verification budgets, what is the marginal independently verified value of agent (N+1), and is that marginal value explained by worker quality, error independence, or coordination/verification overhead?
For each frozen real coding task:
Use population sizes that the budget can support, for example (Nin{1,2,4,8}). Do not increase total allowed tokens/compute simply because (N) increases in the fixed-budget analysis.
The strong “collective advantage” claim fails if:
Any of those outcomes is still publishable.
Working title:
Diversity as Memory Under Goal Drift: Quality-Diversity Search with Correlated Verification and Latent Defects
Do not claim generic novelty for “quality-diversity helps adaptation”; related multi-agent QD work already exists.
The IDKMesh-specific scientific contribution should be the interaction of:
The E024–E039 series is unusually valuable because it already contains results that break simple narratives. The paper should center those boundary conditions, not only the positive archive result.
Kill criterion: if the archive’s apparent advantage disappears under a reasonable verifier/latent-defect model or is not robust across goal geometry, publish that boundary rather than rescuing the headline.
Working title:
Generation Is Cheap, Verification Is Not: Backpressure for Agentic Software Engineering
Question
When candidate generation can scale faster than verification, which admission/verification policy maximizes verified throughput subject to an escaped-defect and backlog constraint?
Baselines
Primary outcomes
Required strengthening step
Replay or shadow these policies on real issue/PR/agent-attempt traces using observed arrival and review/service distributions. Simulation alone should not be sold as empirical repository behavior.
Working title:
When Adaptive Routing Is Worth Its Complexity: Falsification-First Scheduling for Heterogeneous Agent and Compute Pools
The scientific question is:
Which non-stationarity regimes create enough benefit for adaptive routing to repay its additional state, exploration, and route burden?
The stationary Physarum result is exactly the right negative control: if a simple policy is already adequate, sophistication should lose.
The paper needs a factorial stress matrix covering at least:
Report the region of superiority, not a universal winner.
Scores are 1 (weak) to 5 (strong). “Extra work” is reversed: 5 means little additional evidence is required.
| Paper | Novelty after 2026 literature | Current evidence | Reproducibility | External validity | Extra work | Recommendation |
|---|---|---|---|---|---|---|
| A — dependence shape / blind spots | 4 | 5 | 5 | 2 | 4 | Finish first; sharpen novelty. |
| B — fixed-budget agent scaling | 5 | 2 | 4 | 1 | 1 | Highest-impact next experiment. |
| C — QD under goal drift + imperfect verification | 4 | 4 | 5 | 2 | 3 | Strong simulation/mechanism paper. |
| D — verification backpressure | 4 | 4 | 5 | 2 | 3 | Add real trace replay, then write. |
| E — adaptive routing | 3 | 3 | 4 | 2 | 2 | Complete stress matrix first. |
This is a fit map, not a deadline recommendation.
| Contribution shape | Natural venue families |
|---|---|
| empirical AI software-engineering / agent evaluation | ICSE, FSE, ASE, MSR, TOSEM, TSE, EMSE |
| multi-agent coordination / collective intelligence | AAMAS; relevant NeurIPS/ICML/ICLR tracks or workshops depending maturity |
| verification/reliability/control of agentic software | ICSE/FSE/ASE, TSE/TOSEM, reliability/autonomic-computing venues |
| quality-diversity / adaptive multi-agent mechanisms | AAMAS, GECCO, ECAI, ALIFE; ML venues if the empirical contribution is strong enough |
| distributed scheduling / federated coordination | ICDCS, Middleware, CCGrid/HPDC-type venues when real systems evidence exists |
Choose the venue after fixing the scientific contribution and evidence class, not the reverse.
Do not publish these as established conclusions:
These are hypotheses, synthetic mechanisms, implemented contracts, or future evidence gates.
If IDKMesh can fund/obtain only one major new scientific dataset, it should be the real fixed-budget collective-coding experiment for Paper B.
Why this one:
Do not optimize for a huge task count before methodology is sound. A credible pilot should first demonstrate:
Use the pilot to estimate variance and then perform a real power/sample-size calculation for the confirmatory cohort. Do not invent a sample size from convention.
Every IDKMesh manuscript should ship with a table mapping each quantitative claim to:
The existing paper/CLAIM_EVIDENCE_MAP.md is the right pattern. New papers should use the same discipline.
This document is a publication-oriented projection, not a replacement for:
RESEARCH_QUESTIONS.md — broad long-term question inventory;docs/research/TOP_20_QUESTIONS.md — prioritized project questions;docs/research/FIRST_RESEARCH_PROGRAM.md — initial joint program;experiments/README.md — experiment/evidence index;docs/research/README.md — research navigation;paper/README.md and paper/CLAIM_EVIDENCE_MAP.md — manuscript stewardship.This document preserves and improves the outcome of the 2026-10-07 maintainer request to identify the scientific questions IDKMesh can answer and turn into papers.
The improvement pass checked current repository evidence and a targeted sample of current external literature before revising the ranking. It specifically corrected the risk of presenting “reviewer count is not evidence count” as if no closely related 2026 LLM-panel result existed.
This is owner-directed AI-assisted research planning, not independent scientific peer review and not a systematic literature review.
A publication roadmap should make it easier for contributors to answer three questions:
That makes reproducibility, replication, falsification, and negative evidence first-class contribution paths rather than secondary work.