Remove duplicate test docs

Delete redundant and pure-English test documentation so the docs set has one clear Chinese source of truth for standards, scenarios, model I/O, and industry mapping.

Constraint: The user asked to delete unnecessary duplicate docs and pure-English docs.

Rejected: Keeping parallel English and article-derived matrices | they duplicate the current Chinese standard reports and confuse the active S07 FastAPI evidence.

Confidence: high

Scope-risk: narrow

Directive: Keep future Agent test documentation in docs/ and prefer Chinese source-of-truth reports over parallel duplicate matrices.

Tested: .venv/bin/python -B -m unittest discover -s tests; .venv/bin/python -B -m py_compile swarm_minimal/*.py examples/*.py tests/*.py; git diff --check; duplicate-doc reference scan; docs secret pattern scan.

Not-tested: Live S07 matrix was not rerun because this change only deletes duplicate docs and updates indexes.

Co-authored-by: OmX <omx@oh-my-codex.dev>
This commit is contained in:
gongzhiyong
2026-05-16 15:49:24 +08:00
co-authored by OmX
parent dc2b166e60
commit f4292e5999
7 changed files with 2 additions and 984 deletions
+1 -2
View File
@@ -23,9 +23,8 @@
主要报告入口都放在 `docs/`,完整索引见 `docs/README.md`:
- `docs/AGENT_SWARM_QUALITY_STANDARD.zh-CN.md`:Agent / 蜂群 Agent 质量标准、测试标准来源、S07 外部 GitHub 代码任务分配和接手机制。
- `docs/ACADEMIC_STANDARD_TEST_REPORT.zh-CN.md`:测试场景、测试过程、测试结果、算法清单和马尔可夫过程判定。
- `docs/ACADEMIC_STANDARD_TEST_REPORT.zh-CN.md`:S01-S08 测试场景、测试过程、测试结果、算法清单和马尔可夫过程判定。
- `docs/MODEL_AGNET_IO_REPORT.zh-CN.md`:每个模型 / Agnet 的任务、输入、输出、评分和交接过程。
- `docs/STANDARD_TEST_MATRIX.md`:S01-S08 标准矩阵。
- `docs/INDUSTRY_STANDARD_AGNET_TEST_PLAN.md`:行业参考框架到本项目标准的映射。
目标是先把蜂群的四个共享资源跑通:
@@ -1,139 +0,0 @@
# 基于掘金《Agent 蜂群模式(Swarm)》的测试标准
依据文章:[Agent 蜂群模式(Swarm)](https://juejin.cn/post/7603575399255949352)
本文档把该文章中的蜂群定义、核心特征、执行流程、通信机制、系统组件和监控指标,转换成 `swarm-minimal` 的测试标准。后续判断“像不像蜂群”,以本文档为主;LangGraph Swarm 只作为 handoff 实现参考,不作为蜂群标准本身。
## 1. 文章核心定义
文章对 Agent 蜂群模式的核心定义可以归纳为:
```text
每个 Agent 仅依据局部环境信息和简单规则进行决策,
通过信息素、消息广播或状态共享等方式间接通信,
最终在整体层面涌现出复杂、高效的智能行为。
```
因此,测试不能只看“多个模型被调用”或“最终有一个结果”。必须证明:
1. Agent 有局部感知。
2. Agent 有自主决策。
3. Agent 通过共享环境间接协作。
4. 单个 Agent 失败不破坏整体结果。
5. 多个局部行为能产生整体收敛。
6. 系统能观测收敛过程和关键指标。
## 2. 文章核心特征到测试项
| 编号 | 文章特征 | 测试标准 | 当前状态 |
| --- | --- | --- | --- |
| J-SW-01 | 去中心化 | Task Orchestrator 只能注入初始任务,不能逐步命令每个 Agent 做什么;Agent 必须通过任务池和共享环境自主 claim/decide。 | 部分满足:当前 `SwarmCoordinator` 仍是中心循环,需继续弱化。 |
| J-SW-02 | 自组织 | Agent 根据任务池、信息素、共享状态自主形成执行顺序,而不是固定链路。 | 部分满足:已有 pheromone claim 测试,但 live 路径仍偏顺序。 |
| J-SW-03 | 涌现性 | 多个局部弱信号聚合后,群体结果强于任一单个局部判断。 | 已有确定性测试 B02/C02。 |
| J-SW-04 | 鲁棒性 | 单个 Agent 出错后,其他 Agent 继续完成,整体仍可收敛。 | 已有确定性测试 B01/C01。 |
| J-SW-05 | 可扩展性 | Agent 数量应能在 3-7 个范围内增减,任务池和环境结构不变。 | 待补规模测试。 |
| J-SW-06 | 隐式协作 | Agent 不通过中央指令直接串联,而通过信息素、共享状态、消息流协作。 | 部分满足:PostgreSQL/Redis/Blob live 闭环已有,隐式协作测试还需加强。 |
## 3. 文章执行流程到测试项
文章给出的执行循环是:
```text
① 初始化 -> ② 感知 -> ③ 决策 -> ④ 交互与更新 -> ⑤ 收敛判断
未收敛则继续迭代,已收敛则输出结果。
```
对应测试标准:
| 阶段 | 测试标准 | 当前状态 |
| --- | --- | --- |
| 初始化 | 创建 N 个 Agent、初始任务池、信息素地图、共享状态和终止条件。 | 已满足基础形态。 |
| 感知 | 每个 Agent 读取局部环境:可用任务、信息素、共享状态、上一轮摘要。 | 部分满足:代码有 shared_state 传入,但缺少显式 `perceive()` 抽象。 |
| 决策 | Agent 基于感知结果选择任务、候选答案或 handoff 对象。 | 部分满足:claim 基于 capability/pheromone;缺少策略对象。 |
| 交互更新 | Agent 完成动作后更新任务、信息素、共享状态、事件流。 | 已在 PG/Redis/Blob live 路径中实现基础闭环。 |
| 收敛判断 | 不能只是“取最高分”;应有阈值、margin、多轮或任务完成率判定。 | 已新增本地多轮共识收敛;尚未接入 Azure live 主路径。 |
## 4. 三种通信机制到测试项
文章列出的通信机制:
1. 信息素通信(Stigmergy)
2. 广播通信(Broadcast)
3. 交接通信(Handoff)
对应标准:
| 机制 | 合格标准 | 当前状态 |
| --- | --- | --- |
| 信息素通信 | Agent 修改环境分数,后续 Agent 感知并调整选择概率或 claim 顺序。 | 已有 B03/C03 确定性测试;Redis sorted set 已承担 live score。 |
| 广播通信 | Agent 的状态变化进入消息流,其他 Agent 可观察。 | Redis Stream 记录事件,但尚未做“其他 Agent 消费广播后改变决策”的测试。 |
| 交接通信 | Agent A 把上下文交给 Agent B,B 基于 payload 继续执行。 | 已有 B04 测试;但按文章它只是蜂群通信方式之一,不是唯一核心。 |
## 5. 系统组件到测试项
文章架构组件:
| 组件 | 文章职责 | 当前实现 / 缺口 |
| --- | --- | --- |
| Agent Pool | 管理所有 Agent,每个 Agent 有状态、行为规则和通信接口。 | 目前是 `Agent` 列表;缺少状态生命周期和利用率统计。 |
| Environment | 任务池、信息素地图、共享状态、约束条件。 | `InMemorySwarmStore` 和 `PostgresRedisBlobSwarmStore` 已有四资源模型。 |
| Message Bus / Shared Memory | 间接通信。 | Redis Stream + PG shared_state 已有基础;消费侧测试不足。 |
| Monitor / Logger | 状态监控、收敛指标、性能日志。 | 缺少统一 metrics 报告。 |
| Task Orchestrator | 只注入初始任务,不参与调度,以保持去中心化。 | 当前 `SwarmCoordinator` 仍参与执行循环,需进一步改造。 |
## 6. 文章监控指标到测试项
文章列出的关键监控指标:
| 指标 | 文章含义 | 测试标准 |
| --- | --- | --- |
| 收敛速度 | 达到目标质量所需迭代次数。 | 输出 `rounds_to_converge`,并验证小于最大轮数的 50% 或业务阈值。 |
| Agent 利用率 | 每个 Agent 忙碌时间占比。 | 输出每个 Agent claim/complete 次数,目标 > 60% 仅适合压测,不适合小样本。 |
| 任务完成率 | 成功完成任务比例。 | 小样本应为 100%;故障注入场景应证明整体仍收敛。 |
| 平均响应时间 | 从任务提交到结果产出的时间。 | live 路径需要记录每个 task started/completed 时间。 |
| 信息素分布 | 信息素浓度均匀程度,文章建议基尼系数 < 0.5。 | 待实现 pheromone Gini 计算和测试。 |
## 7. 当前测试结论,按文章标准重述
当前 `swarm-minimal` 已经证明:
1. 有四类共享资源:任务池、信息素/得分、共享状态、结果收敛。
2. 有真实 Azure-backed live 闭环:PostgreSQL、Redis、Blob、NewAPI。
3. 有多模型调用和任务拆分。
4. 有单 Agent 失败隔离的确定性测试。
5. 有局部弱信号聚合为群体更优结果的涌现测试。
6. 有信息素影响 claim 顺序的测试。
7. 有 handoff 上下文连续性测试。
8. 已新增本地多轮共识收敛,证明不是简单最高分选择。
但按文章标准,当前还不能说是完整蜂群系统,因为仍有缺口:
1. `SwarmCoordinator` 仍偏中心循环,不够“无单一控制节点”。
2. live 路径还没有真正的并发多 worker 自主 claim。
3. Agent 缺少显式 `perceive -> decide -> act` 生命周期。
4. Redis Stream 目前更多是日志/事件,不是已验证的广播决策输入。
5. 多轮共识收敛还没接入 Azure PG/Redis/Blob 主路径。
6. 缺少 Agent 利用率、响应时间、收敛速度、信息素 Gini 的统一 metrics。
7. 缺少 3-7 Agent 扩缩容下的稳定性测试。
## 8. 后续实现优先级
按文章标准,下一步不是再堆安全测试,而是改造蜂群运行时:
1. 把 Agent 拆成 `perceive()`、`decide()`、`act()` 三段。
2. 把 `SwarmCoordinator` 降级为只初始化任务和停止条件。
3. 引入 worker loop:每个 Agent 自主从任务池 claim。
4. 把 Redis Stream 的事件变成 Agent 可消费的广播输入。
5. 把多轮共识收敛接到 Azure live store。
6. 增加 metrics:收敛轮数、任务完成率、Agent 利用率、平均响应时间、信息素 Gini。
7. 增加 3、5、7 个 Agent 的扩缩容测试。
## 9. 严谨结论
```text
当前原型已经覆盖文章中蜂群模式的部分关键性质:
鲁棒性、涌现性、信息素间接协作、共享环境和 handoff。
但它还不是完整意义上的文章式蜂群系统:
去中心化、自组织 worker loop、广播驱动决策、完整监控指标和 live 多轮收敛仍需补齐。
```
+1 -6
View File
@@ -3,14 +3,9 @@
## Agent / 蜂群 Agent 测试
- `AGENT_SWARM_QUALITY_STANDARD.zh-CN.md`:AQS / SW-AQS v1 主标准,说明 Agent 和蜂群 Agent 的质量项、S07 任务分配、模型交接和当前结论。
- `ACADEMIC_STANDARD_TEST_REPORT.zh-CN.md`:测试场景、测试过程、测试结果、算法清单和马尔可夫过程判定。
- `ACADEMIC_STANDARD_TEST_REPORT.zh-CN.md`:S01-S08 测试场景、测试过程、测试结果、算法清单和马尔可夫过程判定。
- `MODEL_AGNET_IO_REPORT.zh-CN.md`:从 live run 导出的模型 / Agnet 任务、输入、输出、评分和交接过程。
- `STANDARD_TEST_MATRIX.md`:S01-S08 标准化测试矩阵。
- `INDUSTRY_STANDARD_AGNET_TEST_PLAN.md`:行业参考框架到本项目 Agent 质量标准的映射。
- `JUEJIN_SWARM_ARTICLE_TEST_STANDARD.zh-CN.md`:掘金蜂群文章对应的测试标准说明。
- `SWARM_BEHAVIOR_TEST_MATRIX.md`:蜂群行为测试矩阵。
- `SWARM_TEST_INPUT_OUTPUT_REPORT.md`:英文输入输出测试报告。
- `SWARM_TEST_INPUT_OUTPUT_REPORT.zh-CN.md`:中文输入输出测试报告。
## 交付计划
-38
View File
@@ -1,38 +0,0 @@
# Standard Test Matrix
This file defines the stricter acceptance shape for the standalone minimal swarm.
It separates deterministic checks from live Azure/NewAPI checks so a passing
result means more than "one long run completed".
## Layers
| Layer | Purpose | Command |
| --- | --- | --- |
| Static | Python syntax and importability | `./.venv/bin/python -B -m py_compile swarm_minimal/*.py examples/*.py tests/*.py` |
| Unit | Existing closed-loop and NewAPI mock behavior | `./.venv/bin/python -B -m unittest discover -s tests` |
| Deterministic scenarios | No-network Given/When/Then cases for continuity, policy, final scoring, and failure injection | `./.venv/bin/python -B -m unittest tests.test_standard_scenarios` |
| Academic Markov fit | No-network checks for Markov-style transition behavior and non-MDP limitations | `./.venv/bin/python -B -m unittest tests.test_markov_process_properties` |
| Live integration | Real NewAPI + Azure PostgreSQL/Redis/Blob continuous reasoning chain against external GitHub code | `./.venv/bin/python -u -B examples/run_continuous_reasoning_acceptance.py` |
| Report audit | Human-auditable model I/O report with scenario, input, output, handoff and secret-safety evidence | `./.venv/bin/python -B -m unittest tests.test_model_io_report_audit` |
## Required Scenarios
| ID | Scenario | Given | When | Then |
| --- | --- | --- | --- | --- |
| S01 | Syntax/import sanity | All prototype Python files | Compile every module | No syntax/import compile errors |
| S02 | Unit regression | Mock NewAPI and in-memory store | Run all unit tests | All tests pass |
| S03 | Chain continuity | Seven deterministic step outputs | Remove one previous-step link | Continuity checker fails the broken chain |
| S04 | Dependency boundary | Text mentions NATS/Cosmos | Use negative and positive dependency wording | Negative wording passes; required dependency wording fails |
| S05 | Final convergence scoring | STEP-06 and STEP-07 both produce rich outputs | Score both outputs | STEP-07 is strictly selected as final convergence |
| S06 | Failure injection | One deterministic worker raises | Coordinator runs to convergence | Failed task is marked failed, pheromone is negative, completed task still converges |
| S07 | External GitHub code reasoning | Real Azure/NewAPI env is loaded and target is `fastapi/fastapi` at a pinned commit | Run seven-step code reasoning against external GitHub files | PostgreSQL, Redis, Blob, stream events, model discovery, external file references, and final artifact all pass |
| S08 | Model I/O report audit | Generated model/Agnet I/O report | Audit scenario coverage, task input/output, handoff evidence, and obvious secret patterns | Report is human-auditable and contains no obvious secret values |
| M01 | Markov-style claim transition | Same current tasks and pheromone values but different prior paths | Claim next task | The same high-pheromone task is claimed |
| M02 | Markov-style score transition | Same current task, agent, output, and score but different prior paths | Complete task | Task status, output, score, and pheromone update match |
| M03 | Formal MDP limitation | Project metadata is evaluated | Assess Markov fit | Prototype is classified as Markov-style state machine, not a formal Markov process or MDP |
## Stop Rule
The standard acceptance result is PASS only if every scenario passes. Any FAIL
must include the scenario ID, the failed assertion, and enough evidence to
reproduce it without printing secrets.
-32
View File
@@ -1,32 +0,0 @@
# Swarm Behavior Test Matrix
This matrix tests swarm behavior itself, not security compliance. Its primary
basis is the Juejin article "Agent 蜂群模式(Swarm)": decentralized behavior,
self-organization, emergence, robustness, scalability, implicit collaboration,
and the initialize/perceive/decide/interact/converge loop. LangGraph Swarm is
used only as a handoff reference, not as the main swarm standard.
## Core Swarm Claims
| Claim | Why It Matters | Observable Evidence |
| --- | --- | --- |
| Fault isolation | A swarm should not lose the whole result when one Agnet fails. | A failed agent produces a failed observation while other agents still complete and convergence is written. |
| Stigmergy / pheromone coordination | Agents should coordinate indirectly through the shared environment, not only through a central planner. | Pheromone score changes claim order and later convergence preference. |
| Emergent consensus | Group-level output should be stronger than any single local signal in selected scenarios. | Multiple local weak signals aggregate into the accepted result even though no single weak signal is the strongest individual observation. |
| Handoff continuity | LangGraph-style swarm handoff must preserve control target and context across active agents. | `transfer_to_<agent>` naming, active-agent state changes, and payload continuity are checked step by step. |
## Scenarios
| ID | Scenario | Given | When | Then | Different From |
| --- | --- | --- | --- | --- | --- |
| B01 | Single Agnet failure isolation | Three competing route tasks; one Agnet always raises | The swarm coordinator runs to convergence | The failed task is recorded, two alternative tasks finish, and the accepted output comes from a healthy Agnet | Failure handling, not normal success |
| B02 | Emergent consensus from local evidence | Several local evidence tasks, where no single weak signal beats the strongest individual alternative | Agents update shared state with local candidate evidence | The accumulated candidate wins through shared-state aggregation | Emergence, not simple highest single answer |
| B03 | Pheromone-biased task selection | Multiple pending tasks with different pheromone values | Agents claim tasks through the shared task pool | The highest pheromone task is claimed first and positive feedback is recorded | Indirect coordination, not explicit handoff |
| B04 | LangGraph-style handoff continuity | Collector, analyst, and reporter agents with shared active-agent state | Each stage writes a `transfer_to_<agent>` handoff and payload | The next agent observes the previous handoff, continues context, and final convergence uses the reporter output | Control transfer, not parallel competition |
## Passing Bar
The behavior acceptance result is PASS only if all B01-B04 scenarios pass. A
passing run proves the current prototype has minimal swarm behavior under
deterministic conditions. It does not prove large-scale production performance,
Byzantine-agent resistance, or certified LangGraph package compatibility.
-363
View File
@@ -1,363 +0,0 @@
# Swarm Test Input / Output Report
This report shows the actual deterministic test content for the minimal swarm
prototype. It separates the swarm-behavior tests from the traditional-vs-swarm
comparison benchmark.
## Commands Executed
```bash
./.venv/bin/python -u -B examples/run_swarm_behavior_acceptance.py
./.venv/bin/python -u -B examples/run_swarm_vs_traditional_benchmark.py
```
Both commands returned `status: PASS`.
## Behavior Test Content
### B01: Single Agnet Failure Isolation
Purpose: prove one failed Agnet does not collapse the swarm result.
Input:
```json
{
"run_id": "behavior-fault-isolation",
"goal": "fault isolation with redundant routes",
"tasks": [
{"kind": "route", "input": "fragile-route"},
{"kind": "route", "input": "robust-route-a"},
{"kind": "route", "input": "robust-route-b"}
],
"agents": [
{"id": "crashing-agnet", "capability": "route", "behavior": "raise RuntimeError('single agnet crashed')"},
{"id": "backup-agnet-a", "capability": "route", "behavior": "return healthy result, score 0.91"},
{"id": "backup-agnet-b", "capability": "route", "behavior": "return alternative result, score 0.86"}
]
}
```
Expected output:
```json
{
"failed_tasks": 1,
"done_tasks": 2,
"run_status": "converged",
"accepted_output_contains": "healthy result",
"observations_include": ["route:failed", "route:done"]
}
```
Actual output:
```json
{
"passed": true,
"evidence": "Ran 1 test ... OK"
}
```
Concrete assertions are in
`tests/test_swarm_behavior_academic.py::test_single_agnet_failure_isolated_by_redundant_convergence`.
### B02: Emergent Consensus From Local Evidence
Purpose: prove group-level aggregation can beat the best single local signal.
Input:
```json
{
"run_id": "behavior-emergent-consensus",
"goal": "local evidence should create group consensus",
"tasks": [
{"kind": "evidence", "input": "alpha:0.31"},
{"kind": "evidence", "input": "beta:0.33"},
{"kind": "evidence", "input": "beta:0.34"},
{"kind": "evidence", "input": "gamma:0.45"}
],
"rule": "each agent adds its local value into shared_state['candidate:<name>:score']"
}
```
Traditional baseline output:
```json
{
"best_single_candidate": "gamma",
"best_single_score": 0.45
}
```
Expected swarm output:
```json
{
"accepted_candidate": "beta",
"aggregated_beta_score": 0.67,
"accepted_score_greater_than_best_single_weak_signal": true,
"run_status": "converged"
}
```
Actual output:
```json
{
"passed": true,
"evidence": "Ran 1 test ... OK"
}
```
Concrete assertions are in
`tests/test_swarm_behavior_academic.py::test_emergent_consensus_accumulates_local_evidence`.
### B03: Pheromone / Stigmergy Selection
Purpose: prove the shared pheromone environment changes task selection order.
Input:
```json
{
"run_id": "behavior-pheromone",
"goal": "pheromone should bias task selection",
"tasks": [
{"kind": "probe", "input": "low-signal", "initial_pheromone": 0.1, "completion_score": 0.31},
{"kind": "probe", "input": "high-signal", "initial_pheromone": 0.9, "completion_score": 0.82},
{"kind": "probe", "input": "medium-signal", "initial_pheromone": 0.4, "completion_score": 0.62}
],
"claim_rule": "claim_next sorts pending tasks by pheromone descending"
}
```
Expected output:
```json
{
"first_claimed_task": "high-signal",
"accepted_output_contains": "high-signal",
"final_pheromone_order": ["high-signal", "medium-signal", "low-signal"]
}
```
Actual output:
```json
{
"passed": true,
"evidence": "Ran 1 test ... OK"
}
```
Concrete assertions are in
`tests/test_swarm_behavior_academic.py::test_pheromone_biases_claim_order_and_records_positive_feedback`.
### B04: LangGraph-Style Handoff Continuity
Purpose: prove handoff preserves active agent and payload continuity.
Input:
```json
{
"run_id": "behavior-handoff",
"goal": "handoff continuity",
"initial_active_agent": "collector",
"tasks": [
{"kind": "collector", "input": "collect code facts"},
{"kind": "analyst", "input": "analyze code facts"},
{"kind": "reporter", "input": "write final report"}
],
"handoffs": [
{"from": "collector", "to": "analyst", "tool_name": "transfer_to_analyst"},
{"from": "analyst", "to": "reporter", "tool_name": "transfer_to_reporter"}
],
"payload": "facts: task_pool pheromone shared_state convergence; analysis: ready"
}
```
Expected output:
```json
{
"handoff_collector_to_analyst": "transfer_to_analyst",
"handoff_analyst_to_reporter": "transfer_to_reporter",
"final_output_contains": "reporter final result",
"completed_tasks": 3
}
```
Actual output:
```json
{
"passed": true,
"evidence": "Ran 1 test ... OK"
}
```
Concrete assertions are in
`tests/test_swarm_behavior_academic.py::test_langgraph_style_handoff_preserves_active_agent_and_payload`.
## Traditional vs Swarm Benchmark Content
### Baseline Definitions
Traditional Agnet:
```json
[
"single agent fails closed when its one route fails",
"best-of local answers without shared-state aggregation",
"FIFO task selection without pheromone feedback",
"stateless handoff without active-agent/payload continuity"
]
```
Swarm Agnet:
```json
[
"redundant agents share task pool and converge despite a failed individual",
"local observations accumulate through shared_state",
"pheromone scores bias claim order and final selection",
"handoff records active agent, transfer target, and payload continuity"
]
```
### C01: Fault Isolation Comparison
Input:
```json
{
"traditional": {
"tasks": [{"kind": "route", "input": "fragile-route"}],
"agents": [{"id": "single-agnet", "behavior": "raise RuntimeError('single agnet crashed')"}]
},
"swarm": {
"tasks": ["fragile-route", "robust-route-a", "robust-route-b"],
"agents": ["crashing-agnet", "backup-agnet-a", "backup-agnet-b"]
}
}
```
Actual output:
```json
{
"traditional": 0.0,
"swarm": 1.0,
"traditional_result": "failed before convergence",
"swarm_completed_tasks": 2,
"swarm_failed_tasks": 1
}
```
### C02: Emergent Consensus Comparison
Input:
```json
{
"local_values": {
"alpha": [0.31],
"beta": [0.33, 0.34],
"gamma": [0.45]
}
}
```
Actual output:
```json
{
"traditional_best_single": {"candidate": "gamma", "score": 0.45},
"swarm_aggregated_best": {"candidate": "beta", "score": 0.67},
"relative_gain_percent": 48.9
}
```
### C03: Pheromone Efficiency Comparison
Input:
```json
{
"traditional_order": ["low-signal", "medium-signal", "high-signal"],
"swarm_pheromone_order": ["high-signal", "medium-signal", "low-signal"],
"quality": {
"low-signal": 0.31,
"medium-signal": 0.62,
"high-signal": 0.82
}
}
```
Actual output:
```json
{
"traditional_steps_to_best": 3,
"swarm_steps_to_best": 1,
"best_route_efficiency_traditional": 0.3333,
"best_route_efficiency_swarm": 1.0,
"relative_gain_percent": 200.0,
"first_claim_quality_gain_percent": 164.5,
"steps_to_best_reduction_percent": 66.7
}
```
### C04: Handoff Context Retention Comparison
Input:
```json
{
"required_context": [
"task_pool",
"pheromone",
"shared_state",
"convergence",
"analysis: ready"
],
"traditional_payload": "final report",
"swarm_payload": "facts: task_pool pheromone shared_state convergence; analysis: ready"
}
```
Actual output:
```json
{
"traditional_retained_ratio": 0.0,
"swarm_retained_ratio": 1.0
}
```
## Full Benchmark Output
```json
{
"standard": "swarm-vs-traditional-deterministic-benchmark-v1",
"status": "PASS",
"overall_normalized_score": {
"traditional": 0.1958,
"swarm": 0.9175,
"relative_gain_percent": 368.5,
"ratio": 4.69,
"note": "This aggregate is a deterministic academic benchmark over four selected swarm properties, not a universal production claim."
}
}
```
## Verification Boundary
This report proves deterministic behavior for the four selected swarm properties:
fault isolation, emergence, pheromone coordination, and handoff continuity. It
does not prove all production-scale properties, large-cluster performance,
malicious-agent resistance, or certified compatibility with the external
LangGraph package.
@@ -1,404 +0,0 @@
# 蜂群测试输入 / 输出报告
这份报告展示 `swarm-minimal` 最小蜂群原型的确定性测试内容。报告分成两部分:
1. 蜂群行为本体测试:验证它是否真的具备蜂群特征。
2. 传统 Agnet vs 蜂群 Agnet 对比测试:量化蜂群机制相对传统方式强多少。
## 已执行命令
```bash
./.venv/bin/python -u -B examples/run_swarm_behavior_acceptance.py
./.venv/bin/python -u -B examples/run_swarm_vs_traditional_benchmark.py
```
两个命令都返回:
```json
{
"status": "PASS"
}
```
## 一、蜂群行为测试内容
### B01:单个 Agnet 故障隔离
测试目的:证明单个 Agnet 出错不会导致整个蜂群失败。
输入:
```json
{
"run_id": "behavior-fault-isolation",
"目标": "使用冗余路线验证故障隔离",
"任务": [
{"类型": "route", "输入": "fragile-route"},
{"类型": "route", "输入": "robust-route-a"},
{"类型": "route", "输入": "robust-route-b"}
],
"Agnet": [
{"id": "crashing-agnet", "能力": "route", "行为": "抛出 RuntimeError('single agnet crashed')"},
{"id": "backup-agnet-a", "能力": "route", "行为": "返回 healthy result,分数 0.91"},
{"id": "backup-agnet-b", "能力": "route", "行为": "返回 alternative result,分数 0.86"}
]
}
```
预期输出:
```json
{
"失败任务数": 1,
"完成任务数": 2,
"运行状态": "converged",
"最终采纳结果包含": "healthy result",
"观测记录包含": ["route:failed", "route:done"]
}
```
实际输出:
```json
{
"passed": true,
"evidence": "Ran 1 test ... OK"
}
```
对应测试代码:
```text
tests/test_swarm_behavior_academic.py::test_single_agnet_failure_isolated_by_redundant_convergence
```
结论:传统单点失败会中断;蜂群中一个 Agnet 失败后,其他 Agnet 仍然完成任务并收敛。
### B02:群体涌现
测试目的:证明多个局部弱信号通过共享状态聚合后,可以产生强于单个局部判断的整体结果。
输入:
```json
{
"run_id": "behavior-emergent-consensus",
"目标": "局部证据形成群体共识",
"任务": [
{"类型": "evidence", "输入": "alpha:0.31"},
{"类型": "evidence", "输入": "beta:0.33"},
{"类型": "evidence", "输入": "beta:0.34"},
{"类型": "evidence", "输入": "gamma:0.45"}
],
"规则": "每个 Agnet 把局部值累加到 shared_state['candidate:<name>:score']"
}
```
传统 Agnet 输出:
```json
{
"最佳单点候选": "gamma",
"最佳单点分数": 0.45
}
```
蜂群预期输出:
```json
{
"最终采纳候选": "beta",
"beta 聚合分数": 0.67,
"采纳分数高于最佳单点弱信号": true,
"运行状态": "converged"
}
```
实际输出:
```json
{
"passed": true,
"evidence": "Ran 1 test ... OK"
}
```
对应测试代码:
```text
tests/test_swarm_behavior_academic.py::test_emergent_consensus_accumulates_local_evidence
```
结论:单看局部值,传统 Agnet 会选择 `gamma=0.45`;蜂群通过共享状态把两个 `beta` 局部信号聚合成 `0.67`,最终选择 `beta`。这就是最小形式的群体涌现。
### B03:信息素 / 间接协作
测试目的:证明任务选择不是由中央调度器硬编码,而是由共享环境中的信息素影响。
输入:
```json
{
"run_id": "behavior-pheromone",
"目标": "信息素影响任务选择",
"任务": [
{"类型": "probe", "输入": "low-signal", "初始信息素": 0.1, "完成分数": 0.31},
{"类型": "probe", "输入": "high-signal", "初始信息素": 0.9, "完成分数": 0.82},
{"类型": "probe", "输入": "medium-signal", "初始信息素": 0.4, "完成分数": 0.62}
],
"claim 规则": "claim_next 按信息素从高到低选择 pending 任务"
}
```
预期输出:
```json
{
"第一个被 claim 的任务": "high-signal",
"最终采纳结果包含": "high-signal",
"最终信息素顺序": ["high-signal", "medium-signal", "low-signal"]
}
```
实际输出:
```json
{
"passed": true,
"evidence": "Ran 1 test ... OK"
}
```
对应测试代码:
```text
tests/test_swarm_behavior_academic.py::test_pheromone_biases_claim_order_and_records_positive_feedback
```
结论:信息素高的 `high-signal` 先被选择,并在完成后形成正反馈。这个场景验证的是蜂群的间接协作机制。
### B04:LangGraph 风格 Handoff 连续性
测试目的:证明 agent 交接时保留 active agent 状态、目标 agent 名称和上下文 payload。
输入:
```json
{
"run_id": "behavior-handoff",
"目标": "handoff continuity",
"初始 active_agent": "collector",
"任务": [
{"类型": "collector", "输入": "collect code facts"},
{"类型": "analyst", "输入": "analyze code facts"},
{"类型": "reporter", "输入": "write final report"}
],
"handoff": [
{"from": "collector", "to": "analyst", "tool_name": "transfer_to_analyst"},
{"from": "analyst", "to": "reporter", "tool_name": "transfer_to_reporter"}
],
"payload": "facts: task_pool pheromone shared_state convergence; analysis: ready"
}
```
预期输出:
```json
{
"collector 到 analyst 的 handoff": "transfer_to_analyst",
"analyst 到 reporter 的 handoff": "transfer_to_reporter",
"最终输出包含": "reporter final result",
"完成任务数": 3
}
```
实际输出:
```json
{
"passed": true,
"evidence": "Ran 1 test ... OK"
}
```
对应测试代码:
```text
tests/test_swarm_behavior_academic.py::test_langgraph_style_handoff_preserves_active_agent_and_payload
```
结论:这个测试对应 LangGraph Swarm 的核心语义:通过 `transfer_to_<agent>` 将控制权交给指定 agent,并保留上下文。
## 二、传统 Agnet vs 蜂群 Agnet 对比测试
### 基线定义
传统 Agnet:
```json
[
"单个 agent 路线失败后整体失败",
"只选择最佳单点答案,不做共享状态聚合",
"按 FIFO 选择任务,不使用信息素反馈",
"无状态 handoff,不保留 active_agent 和 payload 连续性"
]
```
蜂群 Agnet:
```json
[
"多个 agent 共享任务池,单个个体失败后仍可收敛",
"局部观察通过 shared_state 聚合",
"信息素分数影响任务 claim 顺序和最终选择",
"handoff 记录 active_agent、目标 agent 和上下文 payload"
]
```
### C01:故障隔离对比
输入:
```json
{
"传统 Agnet": {
"任务": [{"类型": "route", "输入": "fragile-route"}],
"Agnet": [{"id": "single-agnet", "行为": "抛出 RuntimeError('single agnet crashed')"}]
},
"蜂群 Agnet": {
"任务": ["fragile-route", "robust-route-a", "robust-route-b"],
"Agnet": ["crashing-agnet", "backup-agnet-a", "backup-agnet-b"]
}
}
```
实际输出:
```json
{
"传统 Agnet": 0.0,
"蜂群 Agnet": 1.0,
"传统结果": "收敛前失败",
"蜂群完成任务数": 2,
"蜂群失败任务数": 1
}
```
### C02:群体涌现对比
输入:
```json
{
"局部值": {
"alpha": [0.31],
"beta": [0.33, 0.34],
"gamma": [0.45]
}
}
```
实际输出:
```json
{
"传统最佳单点": {"候选": "gamma", "分数": 0.45},
"蜂群聚合最佳": {"候选": "beta", "分数": 0.67},
"提升比例": "48.9%"
}
```
### C03:信息素效率对比
输入:
```json
{
"传统顺序": ["low-signal", "medium-signal", "high-signal"],
"蜂群信息素顺序": ["high-signal", "medium-signal", "low-signal"],
"质量分": {
"low-signal": 0.31,
"medium-signal": 0.62,
"high-signal": 0.82
}
}
```
实际输出:
```json
{
"传统找到最优所需步数": 3,
"蜂群找到最优所需步数": 1,
"传统最优路线效率": 0.3333,
"蜂群最优路线效率": 1.0,
"相对提升": "200.0%",
"首次 claim 质量提升": "164.5%",
"找到最优步数减少": "66.7%"
}
```
### C04:Handoff 上下文保留对比
输入:
```json
{
"必须保留的上下文": [
"task_pool",
"pheromone",
"shared_state",
"convergence",
"analysis: ready"
],
"传统 payload": "final report",
"蜂群 payload": "facts: task_pool pheromone shared_state convergence; analysis: ready"
}
```
实际输出:
```json
{
"传统上下文保留率": 0.0,
"蜂群上下文保留率": 1.0
}
```
## 三、完整对比结果
```json
{
"standard": "swarm-vs-traditional-deterministic-benchmark-v1",
"status": "PASS",
"overall_normalized_score": {
"传统 Agnet": 0.1958,
"蜂群 Agnet": 0.9175,
"相对提升": "368.5%",
"倍率": "4.69x",
"说明": "这是四个蜂群核心性质下的确定性学术 benchmark,不是所有生产任务的通用结论。"
}
}
```
## 四、验证边界
本报告证明的是这四个蜂群核心性质:
1. 单个 Agnet 故障隔离。
2. 群体涌现。
3. 信息素间接协作。
4. Handoff 状态连续性。
它没有证明:
1. 大规模生产集群性能。
2. 恶意 / 拜占庭 Agnet 对抗能力。
3. 所有业务任务都能提升 4.69 倍。
4. 官方 LangGraph 包兼容性认证。
所以更严谨的结论是:
```text
在当前最小蜂群原型的四个核心蜂群机制场景下,蜂群 Agnet 相比传统 Agnet 的综合归一化得分约为 4.69x。
```