Recovering which CPU cores are disabled on each individual processor die, from the out-of-band sensor telemetry supercomputers already collect. Reproducible results for 1,962 POWER9 sockets of CINECA's Marconi100.
This project publishes its open issues rather than burying them. The full register — nineteen
issues, each with what is known, what would settle it, and where to look — is
LIMITATIONS.md in
the repository.
Every number comes from one machine, Marconi100. Whether the lot structure or the clustering generalises to other POWER9 installations, or to other vendors, is untested. A replication on a second AC922 installation is the single highest-value addition anyone could make to this work, and it would also settle the defect-versus-policy question below.
If you have per-core BMC telemetry from any comparable fleet, we would like to hear from you — including if your results disagree. A disagreement would delimit exactly when the technique works, which is worth as much as a confirmation.
Slice 0 is harvested on 88.2% of one lot’s sockets but 39.3% of the other’s. Two hypotheses: harvesting follows killer defects, or slice 0 abuts a shared structure and is preferentially disabled for reasons of binning policy. Telemetry shows which positions are disabled, not why: a partial-good record, a field failure or a licensing override can each disable a core.
The fitted model cannot separate them, and we no longer claim it can. A correlated defect process and a binning policy that disables contiguous blocks both predict the clustering we measure. An earlier draft said the decomposition showed “both”; that overstated the identification and was withdrawn.
Parametric model-replicate checks add a further constraint: the model is misspecified. It reproduces the distinct-pattern count but not the pattern frequencies, and it misplaces adjacency across the die — over-predicting co-harvest at some slice pairs and under-predicting at others, with simulation-calibrated residuals up to |z| = 9.2 and 10 of 66 cells outside the simultaneous 95% band. Whatever drives the clustering is not position-independent: it sits mainly on slice pairs that share a POWER9 quad.
An earlier version reported a small correlation and concluded that harvest maps are “operationally invisible”. That does not survive two corrections: “no effect” was argued from a small r rather than an equivalence test against a stated margin, and the interval treated node-days within a rack as independent.
Re-run over the steady-state record with contemporaneous maps, one map summary per model and rack-clustered intervals, the three summaries shift p0−p1 socket power by +2.7, −1.5 and −1.0 W across their full observed ranges, 90% CIs within ±5.6 W. Two are equivalent at ±5 W; in every sensitivity design, including idle node-days and a window adjusted for SLURM socket placement, every interval stays within ±7.4 W. The summary is bounded and small, but not provably negligible. (An intermediate version reported +7.9 W; that came from fitting three collinear summaries jointly.)
We record this because it is the most transferable thing we learned: analysis choices moved our numbers more often than the data did, five times, and the first four ran in the direction we would have preferred.
Back to home · FAQ · full register