AOBench is an open-source benchmark for AI agents that operate High-Performance Computing systems: 88 tasks across 10 question categories and 5 operator roles, 29 deterministic environment snapshots — 6 of them rebuilt from real CINECA Marconi100 ExaData — role-based access control enforced so that a policy violation hard-fails the task, and 12 scorers over 7 weighted dimensions. It runs offline, with no cluster and no API key.
Not to be confused with aobench, the ambient-occlusion ray-tracing microbenchmark by Syoyo Fujita.