Task catalog¶
AOBench contains 88 tasks across 10 question categories (QCATs) and 5 operator roles. 8 of them are grounded in real Marconi100 ExaData rather than synthesised.
Splits: 67 dev (open) and 21 test (held out behind AOBENCH_UNLOCK_TEST=1).
Reproduce this page from your own checkout:
Coverage at a glance¶
| QCAT | Tasks | What it probes |
|---|---|---|
AIOPS | 7 | Anomaly detection and incident response |
ARCH | 6 | Architecture and capability questions |
DATA | 5 | Data movement, storage, and filesystem operations |
DOCS | 5 | Documentation lookup and policy grounding |
ENERGY | 15 | Power and energy reasoning |
FAC | 5 | Facility and cooling operations |
JOB | 14 | Job submission, scheduling, and queue reasoning |
MON | 16 | Monitoring and telemetry interpretation |
PERF | 7 | Performance analysis and bottleneck attribution |
SEC | 8 | Security posture and access questions |
| Role | Tasks |
|---|---|
facility_admin | 21 |
researcher | 11 |
scientific_user | 21 |
sysadmin | 24 |
system_designer | 11 |
AIOPS¶
Anomaly detection and incident response
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
AIOPS_DES_001 | system_designer | dev | env_16 | 3 | Anomaly detector threshold trade-off analysis |
AIOPS_FAC_001 | facility_admin | test | env_16 | 2 | Thermal anomaly false-positive rate trend |
AIOPS_FAC_002 | facility_admin | dev | env_18 | 2 | Proactive TIMEOUT-rate reduction plan |
AIOPS_RES_001 | researcher | dev | env_14 | 2 | Interpreting a predictive failure alert for a node used by my job |
AIOPS_SYS_001 | sysadmin | dev | env_11 | 2 | Predictive failure list for next 72 hours |
AIOPS_SYS_002 | sysadmin | test | env_14 | 3 | Root-cause attribution for AIOps P1 alert |
AIOPS_USR_001 | scientific_user | dev | env_04 | 1 | Walltime overrun forecast for own job |
ARCH¶
Architecture and capability questions
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
ARCH_DES_001 | system_designer | test | env_23 | 3 | Partition expansion prioritization |
ARCH_DES_002 | system_designer | dev | env_13 | 3 | Forecasting cluster saturation and producing a procurement timeline |
ARCH_FAC_001 | facility_admin | dev | env_23 | 2 | Rack row power headroom assessment |
ARCH_RES_001 | researcher | dev | env_23 | 2 | High-memory partition memory bandwidth topology |
ARCH_SYS_001 | sysadmin | dev | env_23 | 2 | Network topology between login nodes and storage |
ARCH_USR_001 | scientific_user | dev | env_23 | 1 | Available GPU partition inventory |
DATA¶
Data movement, storage, and filesystem operations
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
DATA_DES_001 | system_designer | test | env_21 | 3 | Storage capacity expansion timeline |
DATA_FAC_001 | facility_admin | dev | env_21 | 3 | Storage retention policy compliance review |
DATA_RES_001 | researcher | dev | env_21 | 2 | Job I/O throughput pattern analysis |
DATA_SYS_001 | sysadmin | dev | env_21 | 2 | Projects near quota limit identification |
DATA_USR_001 | scientific_user | dev | env_21 | 1 | User storage quota check |
DOCS¶
Documentation lookup and policy grounding
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
DOCS_DES_001 | system_designer | test | env_23 | 2 | Node acceptance test criteria retrieval |
DOCS_FAC_001 | facility_admin | dev | env_22 | 1 | ASHRAE thermal envelope limits documentation |
DOCS_RES_001 | researcher | dev | env_21 | 2 | Lustre I/O optimization documentation retrieval |
DOCS_SYS_001 | sysadmin | dev | env_22 | 1 | CRAC fault response procedure lookup |
DOCS_USR_001 | scientific_user | dev | env_21 | 1 | Scratch space data archival policy lookup |
ENERGY¶
Power and energy reasoning
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
ENERGY_DES_001 | system_designer | dev | env_22 | 3 | Specifying power infrastructure requirements for GPU cluster expansion |
ENERGY_FAC_001 | facility_admin | dev | env_03 | 1 | Highest power consumer |
ENERGY_FAC_002 | facility_admin | test | env_03 | 3 | Energy trend explanation |
ENERGY_FAC_003 | facility_admin | dev | env_04 | 2 | Rack energy comparison |
ENERGY_FAC_004 | facility_admin | dev | env_05 | 2 | Identify cooling failure impact |
ENERGY_FAC_005 | facility_admin | dev | env_04 | 3 | Rack B energy investigation and recommendation |
ENERGY_FAC_006 | facility_admin | dev | env_05 | 3 | Disputed power reading during cooling failure |
ENERGY_RES_001 | researcher | dev | env_03 | 2 | Partition energy efficiency analysis |
ENERGY_SYS_001 | sysadmin | dev | env_03 | 1 | Nodes exceeding power threshold |
ENERGY_SYS_002 | sysadmin | test | env_05 | 2 | Power impact of cooling failure |
ENERGY_USR_001 | scientific_user | dev | env_03 | 1 | Job power consumption estimate |
ENERGY_USR_002 | scientific_user | test | env_01 | 2 | Energy cost of a failed OOM job |
M100_ENERGY_FAC_001 | facility_admin | dev | env_m100_02 | 2 | M100 power overshoot quantification |
M100_ENERGY_FAC_002 | facility_admin | dev | env_m100_03 | 2 | M100 rack cooling fault diagnosis |
M100_ENERGY_SYS_001 | sysadmin | dev | env_m100_02 | 2 | M100 node power anomaly identification |
FAC¶
Facility and cooling operations
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
FAC_DES_001 | system_designer | test | env_22 | 3 | Cooling capacity assessment for GPU expansion |
FAC_FAC_001 | facility_admin | dev | env_22 | 1 | Active critical cooling alarms identification |
FAC_RES_001 | researcher | dev | env_22 | 2 | Rack power vs inlet temperature correlation |
FAC_SYS_001 | sysadmin | dev | env_22 | 2 | Hot rack compute node drain assessment |
FAC_USR_001 | scientific_user | dev | env_22 | 1 | Planned maintenance impact on user jobs |
JOB¶
Job submission, scheduling, and queue reasoning
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
JOB_DES_001 | system_designer | dev | env_01 | 3 | Scheduling policy impact on queue wait time |
JOB_FAC_001 | facility_admin | test | env_03 | 2 | Job workload vs power budget |
JOB_FAC_002 | facility_admin | dev | env_04 | 1 | Rack serving most active jobs |
JOB_RES_001 | researcher | dev | env_01 | 2 | Array job variance analysis |
JOB_SYS_001 | sysadmin | dev | env_02 | 2 | Cluster queue bottleneck |
JOB_SYS_002 | sysadmin | test | env_02 | 3 | QoS violation investigation |
JOB_SYS_003 | sysadmin | dev | env_01 | 1 | Failed jobs on a node |
JOB_USR_001 | scientific_user | dev | env_01 | 1 | Failed job diagnosis |
JOB_USR_002 | scientific_user | dev | env_02 | 2 | Long queue wait explanation |
JOB_USR_003 | scientific_user | dev | env_01 | 1 | OOM cause analysis |
JOB_USR_004 | scientific_user | test | env_01 | 3 | Memory request recommendation after OOM |
JOB_USR_005 | scientific_user | dev | env_02 | 2 | Estimate GPU job wait time |
M100_JOB_USR_001 | scientific_user | dev | env_m100_05 | 2 | M100 job failure telemetry correlation |
M100_JOB_USR_002 | scientific_user | dev | env_m100_06 | 2 | M100 real out-of-memory diagnosis |
MON¶
Monitoring and telemetry interpretation
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
M100_MON_SYS_001 | sysadmin | dev | env_m100_01 | 2 | M100 GPU thermal hotspot detection |
M100_MON_SYS_002 | sysadmin | dev | env_m100_04 | 2 | M100 unreachable node detection |
M100_MON_USR_001 | scientific_user | dev | env_m100_01 | 2 | M100 user job GPU health check |
MON_DES_001 | system_designer | dev | env_23 | 3 | Designing a monitoring strategy for a new GPU partition |
MON_FAC_001 | facility_admin | dev | env_03 | 1 | Active thermal alerts summary |
MON_FAC_002 | facility_admin | test | env_05 | 2 | Cooling incident timeline |
MON_RES_001 | researcher | test | env_01 | 3 | Telemetry correlation matrix analysis |
MON_RES_002 | researcher | dev | env_08 | 2 | Identifying a thermally-throttled node from job performance telemetry |
MON_SYS_001 | sysadmin | dev | env_02 | 2 | Node anomaly detection |
MON_SYS_002 | sysadmin | dev | env_02 | 1 | Queue pressure trend |
MON_SYS_003 | sysadmin | dev | env_03 | 2 | Rack thermal summary |
MON_SYS_004 | sysadmin | dev | env_05 | 2 | Rack thermal anomaly root cause |
MON_SYS_005 | sysadmin | test | env_01 | 3 | Nodes with high memory pressure |
MON_SYS_006 | sysadmin | dev | env_05 | 3 | Conflicting thermal alerts disambiguation |
MON_USR_001 | scientific_user | dev | env_01 | 1 | Node availability check |
MON_USR_002 | scientific_user | test | env_02 | 2 | Cluster load status for job submission |
PERF¶
Performance analysis and bottleneck attribution
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
PERF_DES_001 | system_designer | dev | env_17 | 3 | Next-generation node LINPACK scaling projection |
PERF_FAC_001 | facility_admin | dev | env_12 | 2 | Energy efficiency of GPU partition |
PERF_RES_001 | researcher | dev | env_03 | 2 | Profiling CPU efficiency for a completed ensemble run |
PERF_SYS_001 | sysadmin | dev | env_05 | 2 | Cluster-wide GPU utilisation snapshot |
PERF_SYS_002 | sysadmin | test | env_09 | 3 | Strong-scaling regression diagnosis |
PERF_USR_001 | scientific_user | dev | env_03 | 1 | Own-job parallel efficiency lookup |
PERF_USR_002 | scientific_user | test | env_07 | 2 | Own-job MPI communication fraction |
SEC¶
Security posture and access questions
| Task ID | Role | Split | Environment | Difficulty | Title |
|---|---|---|---|---|---|
SEC_DES_001 | system_designer | dev | env_06 | 3 | Designing RBAC policy for a new multi-tenant partition |
SEC_FAC_001 | facility_admin | dev | env_15 | 2 | 30-day privilege-escalation audit |
SEC_FAC_002 | facility_admin | test | env_20 | 3 | Partition-policy bypass root cause and remediation |
SEC_RES_001 | researcher | dev | env_02 | 1 | Understanding partition access restrictions for high-memory jobs |
SEC_SYS_001 | sysadmin | dev | env_06 | 2 | Unauthorised highmem submission audit |
SEC_SYS_002 | sysadmin | test | env_19 | 3 | Project membership change with conflict scan |
SEC_USR_001 | scientific_user | test | env_02 | 1 | Self-scope check for elevated telemetry |
SEC_USR_002 | scientific_user | dev | env_10 | 1 | Understanding why a job was held due to partition access policy |
Adding a task¶
The corpus is versioned JSON — see Adding a task for the schema, the fidelity gate a new task must pass, and the review checklist. New tasks are genuinely welcome; corpus breadth is the main thing limiting what AOBench can measure.