"""Benchmark suite harness (#13/#21/#22): run a task set through all 5 systems → report. Default backend is OFFLINE (hermetic, key-free) — it validates the whole pipeline but, by design, shows NO swarm quality advantage (every system gets the same reference solution). REAL G_E numbers require `--backend openai` with OPENAI_API_KEY set and a sufficiently large, frozen task set. Grading executes generated code in the fail-closed sandbox, so HEICODE_SANDBOX_ISOLATED must be set (only inside an isolated pod / ephemeral CI runner — see security-boundary §8.1). Usage: HEICODE_SANDBOX_ISOLATED=1 python scripts/run-benchmark-suite.py --taskset coding-set-1 \ [--backend offline|openai] [--archive runs/