Files
Agentswarm/benchmark/runners/__init__.py
T
Songhaoz666andClaude Opus 4.8 baa67350e6 benchmark Group C:基线运行器 + 统一 BenchmarkRunRecord + 报告/回放(Closes #21)
在去中心化重构之上落地 benchmark 对比管线:5 个系统(single/strong/chain/sub_agent/swarm)
跑同一任务集、同一执行后端,产出统一 BenchmarkRunRecord → 评估器算 G_E/G_E,c → 报告 + 回放。

- benchmark/runners/:backend(Offline 确定性 / OpenAI 真实)+ base + 5 个 runner。各 runner
  用 held-out fixture 测试在 Group B 沙箱里评分得 TestPassRate(权威,非自评)。
- benchmark/tasksets/:统一任务集 + 加载器(coding-set-1,1 个 fixture)。
- benchmark/reports/、benchmark/replay/:G_E/G_E,c/coverage/confidence + 归档。
- benchmark/baselines/comparison.py:BenchmarkRunRecord 的 CodeReview/UserAcceptance 改为
  Optional(掩码归一,未采集即 None,规则 #9)。
- scripts/run-benchmark-suite.py harness + scripts/test-benchmark-runners.py。

与去中心化重构对齐:swarm runner 拓扑已**重指向去中心化流程**(种子→自选→自主分解→竞争→
同伴交叉评审→收敛,calls=6/review=1),非旧 Master「分解→派发→单评审」。仍用同一离线后端
建模以保证公平对比(驱动活体编排器会换后端→记录不可比;活体全流程由 test-workflow-e2e 验证)。

沙箱适配:runner 评分走 fail-closed 沙箱(#24),故 test + CI 步骤设 HEICODE_SANDBOX_ISOLATED=1
(仅 CI/隔离 Pod)。

影响范围:agent_swarm(benchmark 层 + 测试 + docs + CI)。不碰 orchestrator 编排逻辑、
不改 Manager↔Swarm 契约、不影响 Client/计费/密钥/审计/发布链路。

诚实边界:
- **离线后端只验证管线**:所有系统拿同一参考解 → quality 相同 → G_E=0、swarm_valid=False,
  刻意不显示蜂群优势(反造假)。真实 G_E>0 需 --backend openai + 足量冻结任务集 + 多次运行。
- 故 Closes #21(运行器 + 统一记录已落地并产出合规非 NaN 记录);Refs #20(仅 1/5 场景)、
  Refs #22(评估器/报告/回放已建,但 Quality 仅 TestPassRate,CodeReview/UserAcceptance 缺)、
  Refs #13(验收 EPIC,需真实 run 证明 Swarm>baselines,未满足)。

Closes #21
Refs #20
Refs #22
Refs #13

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 17:44:34 +08:00

45 lines
1.7 KiB
Python

"""Benchmark runners (#21): run a shared task set through each system → BenchmarkRunRecord.
Same backend + same task set for all systems (fairness); only topology differs. `run_all` is the
convenience entry the harness uses; `strong_backend` lets the Strong baseline use a costlier/larger
model than the others (its definitional difference).
"""
from __future__ import annotations
from typing import Dict, List, Optional
from ..baselines import BenchmarkRunRecord
from ..tasksets import TaskSet
from .backend import ExecutionBackend, GenerationResult, OfflineBackend, OpenAIBackend
from .base import BaseRunner, Topology
from .single import SingleRunner
from .strong import StrongRunner
from .chain import ChainRunner
from .sub_agent import SubAgentRunner
from .swarm import SwarmRunner
RUNNERS = {
"single": SingleRunner,
"strong": StrongRunner,
"chain": ChainRunner,
"sub": SubAgentRunner,
"swarm": SwarmRunner,
}
def run_all(taskset: TaskSet, backend: ExecutionBackend, *,
strong_backend: Optional[ExecutionBackend] = None) -> Dict[str, BenchmarkRunRecord]:
"""Run every system on the task set; return {system: record}. swarm + 4 baselines."""
out: Dict[str, BenchmarkRunRecord] = {}
for system, runner_cls in RUNNERS.items():
be = strong_backend if (system == "strong" and strong_backend is not None) else backend
out[system] = runner_cls().run(taskset, be)
return out
__all__ = [
"BenchmarkRunRecord", "ExecutionBackend", "GenerationResult", "OfflineBackend", "OpenAIBackend",
"BaseRunner", "Topology", "SingleRunner", "StrongRunner", "ChainRunner", "SubAgentRunner",
"SwarmRunner", "RUNNERS", "run_all",
]