Per-die CPU core-harvest maps from BMC telemetry

Recovering which CPU cores are disabled on each individual processor die, from the out-of-band sensor telemetry supercomputers already collect. Reproducible results for 1,962 POWER9 sockets of CINECA's Marconi100.

View the Project on GitHub MSKazemi/m100-silicon-binning

Limitations

This project publishes its open issues rather than burying them. The full register — nineteen issues, each with what is known, what would settle it, and where to look — is LIMITATIONS.md in the repository.

The largest one: a single system

Every number comes from one machine, Marconi100. Whether the lot structure or the clustering generalises to other POWER9 installations, or to other vendors, is untested. A replication on a second AC922 installation is the single highest-value addition anyone could make to this work, and it would also settle the defect-versus-policy question below.

If you have per-core BMC telemetry from any comparable fleet, we would like to hear from you — including if your results disagree. A disagreement would delimit exactly when the technique works, which is worth as much as a confirmation.

Defect-driven or policy-driven? Not identified

Slice 0 is harvested on 88.2% of one lot’s sockets but 39.3% of the other’s. Two hypotheses: harvesting follows killer defects, or slice 0 abuts a shared structure and is preferentially disabled for reasons of binning policy. Telemetry shows which positions are disabled, not why: a partial-good record, a field failure or a licensing override can each disable a core.

The fitted model cannot separate them, and we no longer claim it can. A correlated defect process and a binning policy that disables contiguous blocks both predict the clustering we measure. An earlier draft said the decomposition showed “both”; that overstated the identification and was withdrawn.

Parametric model-replicate checks add a further constraint: the model is misspecified. It reproduces the distinct-pattern count but not the pattern frequencies, and it misplaces adjacency across the die — over-predicting co-harvest at some slice pairs and under-predicting at others, with simulation-calibrated residuals up to |z| = 9.2 and 10 of 66 cells outside the simultaneous 95% band. Whatever drives the clustering is not position-independent: it sits mainly on slice pairs that share a POWER9 quad.

The power claim was withdrawn, and replaced by a bound

An earlier version reported a small correlation and concluded that harvest maps are “operationally invisible”. That does not survive two corrections: “no effect” was argued from a small r rather than an equivalence test against a stated margin, and the interval treated node-days within a rack as independent.

Re-run over the steady-state record with contemporaneous maps, one map summary per model and rack-clustered intervals, the three summaries shift p0−p1 socket power by +2.7, −1.5 and −1.0 W across their full observed ranges, 90% CIs within ±5.6 W. Two are equivalent at ±5 W; in every sensitivity design, including idle node-days and a window adjusted for SLURM socket placement, every interval stays within ±7.4 W. The summary is bounded and small, but not provably negligible. (An intermediate version reported +7.9 W; that came from fitting three collinear summaries jointly.)

We record this because it is the most transferable thing we learned: analysis choices moved our numbers more often than the data did, five times, and the first four ran in the direction we would have preferred.

Other open items


Back to home · FAQ · full register