四块互相交织的 benchmark 覆盖增量,统一提交: 1) 通信遥测(#23):orchestrator 路由 peer 消息时按 correlation_id 计请求/应答到 SwarmRun.collaboration(内部状态,不进 Manager 事件流);collector 算 s_communication。 治理计数由 run.approvals 派生(合规/总数)→ s_governance。 2) Q_quality 掩码归一(v2.1 裁定):metrics.quality_score 改为对 present 输入加权归一, 非编码任务自动忽略 TestPassRate,全缺 → NaN(不伪造)。 3) 质量插桩 / Group B:新增 Pod 内代码测试沙箱(orchestrator/sandbox.py,环境清洗 + 超时强杀 + 资源限额 + 路径越界校验,门控 ENABLE_QUALITY_EVAL)与 held-out fixture (benchmark/fixtures/);run 完成时用留出测试评分得 TestPassRate → Q_quality → collector 合成 reward。安全边界见 docs/integration/security-boundary.md §8.1。 4) 决策引擎 / Group A(#10,Option A score-at-pull):新增 orchestrator/decision_engine.py —— 信息素 τ(Redis 持久、(role,agent) 键控、冷启动 0.5、ρ 蒸发、夹紧、学习常开)+ η 启发式评分 + ε-greedy 概率采样;每次 dispatch 产一条 DecisionTrace → SwarmRun.decisions;collector 算 tau/eta/p_decision。概率选择门控 ENABLE_ACO_DISPATCH (默认关,CI 用 ACO_SEED 固定)。 覆盖:单次 run 真实可算字段由 4 提升至最多 10/15(新增 communication/reward/tau/eta/ p_decision,外加 governance 有条件)。 测试:新增 test-sandbox / test-quality / test-decision-engine;扩充 collector/metrics 用例; CI 纳入全部 benchmark 套件 + flag-on 的 ACO e2e。本地 11 项 gate 全绿。 诚实边界(未越界声称): - Group A 为单边匹配(Option B 待 Group C);概率派发优于贪心未证;默认关闭。 - reward 的 CodeReview/UserAcceptance 未采集(掩码忽略);P_risk 为审批派生低估。 - s_gain/s_swarm/g_e/g_e_cost/benchmark 仍 NaN —— 需基线(#21/#13),本 PR 不动验收。 影响范围:Swarm(orchestrator + benchmark + docs + CI)。不改 Manager↔Swarm 事件契约 (遥测均为运行时内部状态);不影响 Client/计费/密钥/发布链路。新增 ENABLE_QUALITY_EVAL / ENABLE_ACO_DISPATCH 两个开关,默认关闭。 Closes #10 Closes #23 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
89 lines
2.8 KiB
YAML
89 lines
2.8 KiB
YAML
name: CI
|
|
|
|
on:
|
|
push:
|
|
branches: ["**"]
|
|
pull_request:
|
|
branches: ["**"]
|
|
|
|
jobs:
|
|
guardrails:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
|
|
# Defense in depth: fail if secrets or heavy/generated dirs were ever committed.
|
|
- name: Block secrets & node_modules
|
|
run: |
|
|
if git ls-files | grep -E '(^|/)\.env($|\.)|(^|/)secrets/|\.pem$|\.key$|\.p12$|\.pfx$|(^|/)id_rsa$|(^|/)id_ed25519$'; then
|
|
echo "::error::Secret-like files are tracked — remove them and rotate any exposed credential."; exit 1
|
|
fi
|
|
if git ls-files | grep -E '(^|/)node_modules/'; then
|
|
echo "::error::node_modules is tracked — it must be gitignored."; exit 1
|
|
fi
|
|
|
|
- name: Required standards files present
|
|
run: |
|
|
for f in CLAUDE.md PROJECT_STANDARD.md README.md; do
|
|
test -f "$f" || { echo "::error::Missing required file: $f"; exit 1; }
|
|
done
|
|
|
|
tests:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
|
|
- uses: actions/setup-python@v5
|
|
with:
|
|
python-version: "3.13"
|
|
|
|
- name: Install dependencies
|
|
run: |
|
|
python -m pip install --upgrade pip
|
|
pip install -r orchestrator/requirements.txt -r agent/requirements.txt
|
|
|
|
# Hermetic: in-memory store, planner forced offline by the tests — no model key needed.
|
|
- name: Manager contract test
|
|
env: { REDIS_FAKE: "1" }
|
|
run: python scripts/test-runtime-contract.py
|
|
|
|
- name: Workflow mechanism smoke test
|
|
env: { REDIS_FAKE: "1" }
|
|
run: python scripts/test-merge-smoke.py
|
|
|
|
- name: End-to-end workflow test
|
|
env: { REDIS_FAKE: "1" }
|
|
run: python scripts/test-workflow-e2e.py
|
|
|
|
- name: Manager event-contract test
|
|
env: { REDIS_FAKE: "1" }
|
|
run: python scripts/test-contract-events.py
|
|
|
|
- name: Benchmark metric formulas (v2.1)
|
|
env: { REDIS_FAKE: "1" }
|
|
run: python scripts/test-benchmark-metrics.py
|
|
|
|
- name: Benchmark collector
|
|
env: { REDIS_FAKE: "1" }
|
|
run: python scripts/test-benchmark-collector.py
|
|
|
|
- name: Baseline comparison
|
|
env: { REDIS_FAKE: "1" }
|
|
run: python scripts/test-baseline-comparison.py
|
|
|
|
- name: Code sandbox (in-pod test runner)
|
|
run: python scripts/test-sandbox.py
|
|
|
|
- name: Quality instrumentation (Group B)
|
|
env: { REDIS_FAKE: "1" }
|
|
run: python scripts/test-quality.py
|
|
|
|
- name: ACO decision engine (Group A)
|
|
env: { REDIS_FAKE: "1" }
|
|
run: python scripts/test-decision-engine.py
|
|
|
|
# Same e2e workflow, but through the probabilistic ACO dispatch path (seeded).
|
|
- name: End-to-end workflow test (ACO dispatch on)
|
|
env: { REDIS_FAKE: "1", ENABLE_ACO_DISPATCH: "1", ACO_SEED: "42" }
|
|
run: python scripts/test-workflow-e2e.py
|