From 5adad6e24053234911b37b2373ea202d86935a97 Mon Sep 17 00:00:00 2001 From: Jim Date: Sun, 10 May 2026 20:32:24 +0000 Subject: [PATCH 1/3] docs: replace coverage example with game AI example; full zh-CN rewrite MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Replace test coverage example with game AI evolution (35→72% win rate) — more visceral, shows self-correction on rejection, closer to AlphaEvolve spirit - Add progress bar visual (before/after) and round-by-round terminal output - Add callout on why Round 3 rejection matters (memory / no repeated mistakes) - Update architecture Mermaid with emoji labels and loop subgraph label - Full rewrite of README.zh.md as faithful Chinese translation (was stale v0) - Fix zh-CN badge URLs and remove Token-Ignition/legacy references Co-Authored-By: Claude Sonnet 4.6 --- README.md | 291 +++++++++++++++++------------------ README.zh.md | 419 ++++++++++++++++++++++++++++++++------------------- 2 files changed, 413 insertions(+), 297 deletions(-) diff --git a/README.md b/README.md index d271571..6d3fee5 100644 --- a/README.md +++ b/README.md @@ -1,11 +1,12 @@ # Evolution Kernel

- Give an LLM a goal. Watch your repo improve itself. Stop when the budget runs out. + Give an LLM a goal. Watch your codebase improve itself. Stop when the budget runs out.

- A ~1,200-line Python runtime for autonomous, multi-round code improvement — sandboxed, audited, and fully reversible. + A ~1,200-line Python runtime that runs an autonomous, multi-round improvement loop on any codebase —
+ sandboxed in git worktrees, every decision logged, every change reversible.

@@ -19,52 +20,58 @@ tests v0.2 - Python >= 3.10 - MIT License - Single dependency: PyYAML + Python ≥ 3.10 + MIT + Single dependency +

+ +--- + +

+ Think of it as AlphaEvolve — but pointed at your own repository.
+ You define what "better" means. The kernel figures out how to get there.

--- ## What it does -Write a YAML file that says what "better" means. Evolution Kernel runs a tight loop: +Point Evolution Kernel at any git repository and give it a measurable goal. It runs a closed loop: -1. **Observe** — collect the current metric (coverage %, benchmark score, lint count — whatever your shell command outputs) -2. **Plan** — an LLM reads the metric and the history of prior attempts, then writes a concrete plan -3. **Execute** — a coding agent (Aider or Claude Code) applies the plan inside a git worktree sandbox -4. **Evaluate** — the evaluator re-runs your metric command and decides accept or reject -5. **Commit or roll back** — accepted changes become a real git commit; rejected ones are discarded -6. **Loop** — repeat until a budget limit fires (`max_iterations`, `max_total_usd`, `max_total_tokens`) +| Step | What happens | +|:---:|---| +| 🔍 **Observe** | Run your metric command — collect the current state (win rate, latency, error count, …) | +| 🧠 **Plan** | LLM reads the metric + history of prior attempts, produces a concrete plan | +| 🔨 **Execute** | Coding agent (Aider or Claude Code) applies the plan inside an isolated git worktree | +| ⚖️ **Evaluate** | Re-run your metric; LLM decides accept or reject | +| ✅ **Commit / rollback** | Accepted → real git commit on `evolution/accepted`. Rejected → worktree discarded | +| 🔁 **Loop** | Repeat until `max_iterations`, `max_total_usd`, or `max_total_tokens` fires | -Every attempt — accepted or rejected — is written to a structured **ledger** so you can audit exactly what the LLM tried, what changed, and why each round was accepted or rejected. +Every attempt is written to a **ledger**: goal, observation, plan, diff, evaluation, decision. Nothing is held in memory. An external auditor — or your future self — can reconstruct every decision from the ledger alone. --- ## Quick Start ```bash -# 1. Install (single runtime dependency: PyYAML) +# 1. Install pip install evolution-kernel -# 2. Write a goal config +# 2. Describe your goal cat > evolution.yml << 'EOF' -mission: "Increase src/ test coverage from 40% to 80%" +mission: "Evolve the game AI to win at least 60% of games against the built-in opponent" evidence_sources: - type: shell - command: > - python3 -m pytest --cov=src --cov-report=json -q && - python3 -c "import json; d=json.load(open('coverage.json')); - print(f'coverage: {d[\"totals\"][\"percent_covered\"]:.1f}%')" + command: "python3 scripts/tournament.py --games 20 --json" mutation_scope: - allowed_paths: ["tests/"] + allowed_paths: ["ai/"] hard_stops: - max_iterations: 20 - max_consecutive_failures: 3 - max_total_usd: 2.00 + max_iterations: 30 + max_consecutive_failures: 4 + max_total_usd: 3.00 llm: provider: anthropic @@ -80,75 +87,88 @@ roles: evaluator: ["python3", "roles/evaluator.py"] EOF -# 3. Run until the budget fires -evolution-kernel --config evolution.yml --repo /path/to/your-project --ledger /tmp/ledger --loop +# 3. Run — walk away +evolution-kernel --config evolution.yml --repo /path/to/game --ledger /tmp/ledger --loop ``` --- -## Example: raising test coverage from 40% to 80% +## See it in action + +### Evolving a game AI from 35% to 72% win rate — overnight, unattended -The loop emits one JSON object per round. A realistic session looks like this: +``` +before ███░░░░░░░░░ 35% win rate (loses 13 of 20 games) +after ███████░░░░░ 72% win rate (wins 14 of 20 games) +9 rounds · $2.14 · 0 minutes of your time ``` -Round 1 observe: coverage 40.2% - plan → "Add unit tests for src/parser.py — parse_tokens is completely uncovered" - execute → aider writes tests/test_parser.py (14 new assertions) - eval → coverage 51.7% — ACCEPT - commit → a3f1c9e "tests: cover parse_tokens (coverage 40→52%)" - -Round 2 observe: coverage 51.7% - plan → "Add edge-case tests for src/validator.py, missing branch coverage on error paths" - execute → aider extends tests/test_validator.py (+9 tests) - eval → coverage 63.4% — ACCEPT - commit → 8b2de01 "tests: validator edge cases (coverage 52→63%)" - -Round 3 observe: coverage 63.4% - plan → "Cover src/formatter.py — currently 0% covered" - execute → aider writes tests/test_formatter.py - eval → coverage 63.4% — new test file has wrong import path — REJECT - rollback → worktree discarded, main branch unchanged (consecutive_failures: 1) - -Round 4 observe: coverage 63.4% - plan → "tests/test_formatter.py failed due to import error; fix path and retry" - execute → aider fixes import in tests/test_formatter.py - eval → coverage 74.8% — ACCEPT - commit → 2c9af44 "tests: formatter coverage, fixed import (coverage 63→75%)" + +Here is what the loop actually does, round by round: + +``` +Round 1 observe: win_rate 35% + plan → "Greedy score maximization with no lookahead — add 2-ply minimax" + execute → aider rewrites ai/strategy.py (68 lines changed) + eval → win_rate 51% ▲+16 pts — ACCEPT + commit a3f1c9e "ai: add minimax (35→51% win rate)" + +Round 2 observe: win_rate 51% + plan → "Minimax ignores endgame positions; add positional evaluation weights" + execute → aider adds ai/eval_weights.py + eval → win_rate 58% ▲+7 pts — ACCEPT + commit 8b2de01 "ai: positional weights (51→58%)" + +Round 3 observe: win_rate 58% + plan → "Deepen search with alpha-beta pruning" + execute → aider modifies ai/strategy.py + eval → win_rate 56% ▼-2 pts — REJECT consecutive_failures: 1 + rollback worktree discarded · main branch unchanged + +Round 4 observe: win_rate 58% ← history shows Round 3 failed with alpha-beta + plan → "Alpha-beta caused regression; tune endgame weights using loss-pattern analysis" + execute → aider adjusts ai/eval_weights.py + eval → win_rate 67% ▲+9 pts — ACCEPT + commit 2c9af44 "ai: endgame weight tuning (58→67%)" ... -Round 12 observe: coverage 80.1% - eval → coverage 80.1% — threshold reached — ACCEPT - commit → 9d7b321 "tests: final push past 80% target" +Round 9 observe: win_rate 72% + eval → 72% — target 60% exceeded — ACCEPT + commit 9d7b321 "ai: final tuning pass (70→72%)" -{"halted": true, "reason": "max_iterations reached", "iterations": 20, "total_usd": 1.43, "total_tokens": 487201} +{"halted": true, "reason": "max_iterations reached", "iterations": 30, "total_usd": 2.14, "total_tokens": 634000} ``` -Each accepted change is a reversible git commit on the `evolution/accepted` branch. The LLM self-corrected on Round 4 using the rejection history from Round 3 — this is what history injection does. +> **Round 3 is the key moment.** Alpha-beta pruning made things *worse*, so the system rejected the change and left the codebase untouched. Round 4 shows the LLM reading the rejection history and changing its approach. This is what "memory" means in practice — not guessing the same wrong answer twice. --- -## Ledger structure - -Every round writes a full evidence trail. Nothing is stored in memory; an external auditor can reconstruct every decision from the ledger directory alone. +## Ledger: the complete audit trail ``` ledger/ - .evolution_state.json # persisted counters (iterations, usd, tokens) — survives restarts + .evolution_state.json ← budget counters; survives restarts runs/ 0001/ - config.json # full snapshot of your evolution.yml - observation.json # raw output of your evidence_sources commands - plan.json # LLM plan: summary, steps, expected_improvement - patch.diff # exact diff the executor applied - candidate_commit.txt # git SHA of the sandbox commit - evaluation.json # verdict + metrics + cost_usd + tokens_used - decision.json # accept / reject + reason - reflection.json # one-line summary injected into the next round's history - 0002/ - ... + config.json ← full snapshot of your evolution.yml + observation.json ← raw output of your evidence_sources commands + plan.json ← LLM plan: summary · steps · expected_improvement + patch.diff ← exact diff the executor applied + candidate_commit.txt ← git SHA of the sandbox commit + evaluation.json ← verdict + metrics + cost_usd + tokens_used + decision.json ← accept / reject + reason + reflection.json ← one-line summary injected into the next round + 0002/ ... halted/ - 20260501T120000Z.json # written when any hard stop fires + 20260501T120000Z.json ← written when any hard stop fires +``` + +To undo every change from a session: + +```bash +git checkout evolution/accepted +git reset --hard # every accepted change is a named commit ``` --- @@ -159,40 +179,40 @@ ledger/ flowchart LR Config[evolution.yml] --> Governor - subgraph loop ["Loop until hard stop"] + subgraph loop ["↻ Loop until hard stop fires"] direction LR - Governor -->|"planner_input.json\n(goal + observation + history)"| Planner["Planner\nLLM"] - Planner -->|plan.json| Executor["Executor\nAider / Claude Code"] - Executor -->|patch in git worktree| Evaluator["Evaluator\nLLM + shell"] + Governor -->|"planner_input.json\ngoal · observation · history"| Planner["🧠 Planner\nLLM"] + Planner -->|plan.json| Executor["🔨 Executor\nAider / Claude Code"] + Executor -->|patch in git worktree| Evaluator["⚖️ Evaluator\nLLM + shell"] Evaluator -->|evaluation.json| Governor end - Governor -->|"accept → git commit"| AcceptedBranch[evolution/accepted] - Governor -->|"reject → discard worktree"| Ledger[Ledger] + Governor -->|"accept → git commit"| Branch["evolution/accepted"] + Governor -->|"reject → discard"| Ledger[📁 Ledger] Governor --> Ledger ``` -**The Governor is intentionally dumb.** It is pure orchestration — no LLM calls of its own. All intelligence lives in the three role scripts. You can swap any role for your own implementation; the Governor only cares about the JSON files roles read and write. +**The Governor is intentionally dumb.** It is pure orchestration — zero LLM calls. All intelligence lives in the three role scripts. Swap any role for your own implementation; the Governor only cares about the JSON each role reads and writes. -**Roles communicate through files, not shared memory.** The planner never talks directly to the executor. The evaluator never sees the executor's self-assessment. The only shared state is the ledger. +**Roles communicate through files, not shared memory.** The planner never talks to the executor. The evaluator never sees the executor's self-assessment. The only shared state is the ledger. --- -## Capabilities +## What works today | Feature | Status | -|---|---| -| Multi-round LLM loop with memory (history injection) | ✅ Working | -| Budget guards: `max_total_usd`, `max_total_tokens` | ✅ Working | -| Iteration / consecutive-failure hard stops | ✅ Working | -| Full ledger audit trail (survives process restarts) | ✅ Working | -| Git worktree sandbox — every attempt isolated | ✅ Working | -| Scope enforcement — rejects changes outside `allowed_paths` | ✅ Working | -| Config-driven: swap LLM provider, model, coding agent | ✅ Working | -| Aider and Claude Code executor support | ✅ Working | -| Anthropic and OpenAI planner/evaluator support | ✅ Working | +|---|:---:| +| Multi-round LLM loop with memory (history injection) | ✅ | +| Budget guards: `max_total_usd`, `max_total_tokens` | ✅ | +| Iteration / consecutive-failure hard stops | ✅ | +| Full ledger audit trail (survives process restarts) | ✅ | +| Git worktree sandbox — every attempt isolated | ✅ | +| Scope enforcement — rejects changes outside `allowed_paths` | ✅ | +| Config-driven: swap LLM provider, model, coding agent | ✅ | +| Aider and Claude Code executor support | ✅ | +| Anthropic and OpenAI planner/evaluator support | ✅ | | Goal evaluator — stops when mission is "won" | 🔧 PR #5 | -| k-branch parallel exploration (FunSearch style) | 🔧 PR #6 | +| k-branch parallel exploration (FunSearch / AlphaEvolve style) | 🔧 PR #6 | | Process sandbox (firejail / bwrap) for production safety | 🔧 PR #7 | --- @@ -200,43 +220,42 @@ flowchart LR ## Configuration reference ```yaml -# Required — free-text statement of what "better" means -mission: "Increase src/ test coverage from 40% to 80%" +# Required — what "better" means for your project +mission: "Evolve the game AI to win at least 60% of games" -# How to measure the current state of the target repo +# How to measure the current state evidence_sources: - - type: shell # runs a command; stdout goes into observation.json - command: "python3 -m pytest --cov=src -q && ..." - - type: file # reads a file; content goes into observation.json + - type: shell # stdout goes into observation.json + command: "python3 scripts/tournament.py --games 20 --json" + - type: file # file contents go into observation.json path: "metrics.json" -# Only files under these paths may be modified by the executor +# Only files under these paths may be changed mutation_scope: allowed_paths: - - "tests/" # changes outside this list are auto-rejected + - "ai/" # changes outside this list are auto-rejected # When to stop hard_stops: - max_iterations: 10 # total rounds (required, must be ≥ 1) - max_consecutive_failures: 3 # consecutive rejections before halt (required) - max_total_usd: 0.0 # 0 = unlimited + max_iterations: 30 # total rounds + max_consecutive_failures: 4 # consecutive rejections before halt + max_total_usd: 3.00 # 0 = unlimited max_total_tokens: 0 # 0 = unlimited -# LLM used by the planner and evaluator role scripts +# LLM for planner and evaluator llm: provider: anthropic # anthropic | openai model: claude-sonnet-4-6 api_key_env: ANTHROPIC_API_KEY -# Coding agent used by the executor role script +# Coding agent for executor coding_agent: tool: aider # aider | claude-code -# How many past rounds the planner sees as context +# How many past rounds the planner sees history: max_entries: 10 -# The three role commands (each receives --input, --output, --worktree) roles: planner: ["python3", "roles/planner.py"] executor: ["bash", "roles/executor.sh"] @@ -244,7 +263,6 @@ roles: ``` **Switch to OpenAI:** - ```yaml llm: provider: openai @@ -252,8 +270,7 @@ llm: api_key_env: OPENAI_API_KEY ``` -**Switch to Claude Code as coding agent:** - +**Switch to Claude Code:** ```yaml coding_agent: tool: claude-code @@ -261,20 +278,20 @@ coding_agent: --- -## CLI reference +## CLI ```bash -# Run the multi-round loop (recommended — stops when a hard stop fires) +# Loop until a hard stop fires (recommended) evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger --loop -# Run exactly one round +# Single round evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger -# Reset hard-stop counters to start a fresh session +# Reset budget counters after a halt evolution-kernel --ledger /tmp/ledger --reset ``` -Exit codes: `0` = clean finish, `3` = halted by a hard stop. +Exit codes: `0` clean finish · `3` halted by a hard stop. --- @@ -284,7 +301,7 @@ Exit codes: `0` = clean finish, `3` = halted by a hard stop. pip install evolution-kernel ``` -From source (the only runtime dependency is PyYAML): +From source (only runtime dependency: PyYAML): ```bash git clone https://github.com/Protocol-zero-0/evolution-kernel.git @@ -292,60 +309,46 @@ cd evolution-kernel pip install -e . ``` -Python 3.10 or later required. +Python 3.10 or later. --- -## Running the tests +## Tests ```bash python3 -m pytest tests/ -v ``` -All tests run locally with no network calls — roles are replaced by lightweight fixture scripts. +39 tests · no network calls · roles replaced by lightweight fixture scripts. --- -## Writing your own role scripts +## Writing your own roles -Each role is an executable that receives three arguments: +Each role is an executable that receives: -```text ---input JSON the governor prepared for this role ---output JSON the role must write before exiting ---worktree path to the isolated git sandbox checkout ``` - -The built-in `roles/planner.py`, `roles/executor.sh`, and `roles/evaluator.py` are the reference implementation. Copy and modify them, or replace them entirely with a shell script, a Python program, or a Docker call. The Governor has no opinion on what runs inside a role. - ---- - -## Rollback - -Every accepted change is a commit on the `evolution/accepted` branch. To undo everything from a session: - -```bash -git checkout evolution/accepted -git log --oneline # find the baseline commit before the session -git reset --hard # roll back all accepted changes +--input JSON the governor wrote for this role +--output JSON the role must write before exiting +--worktree path to the isolated git sandbox checkout ``` -Rejected experiments are never promoted, so only the changes your evaluator explicitly accepted survive. +`roles/planner.py`, `roles/executor.sh`, and `roles/evaluator.py` are the reference implementation. Copy, modify, or replace them entirely — with a shell script, a Docker call, or anything that reads `--input` and writes `--output`. --- ## Project layout ``` -evolution_kernel/ # ~1,200-line runtime (Governor, Observer, HardStops, Config, CLI) -roles/ # reference planner, executor, and evaluator implementations -examples/ # demo target + evolution.yml to run out of the box -docs/ # protocol spec -tests/ # unit + acceptance tests (39 tests, no network required) +evolution_kernel/ ~1,200-line runtime (Governor · Observer · HardStops · Config · CLI) +roles/ reference planner, executor, evaluator +examples/ demo target + working evolution.yml +docs/ protocol spec +tests/ 39 unit + acceptance tests ``` --- ## License -MIT. See [LICENSE](LICENSE). +MIT — see [LICENSE](LICENSE). diff --git a/README.zh.md b/README.zh.md index 113b5bf..72cd391 100644 --- a/README.zh.md +++ b/README.zh.md @@ -1,241 +1,354 @@ # Evolution Kernel

- 一个用于自主优化软件项目的通用进化引擎。 + 给 LLM 一个目标,让代码库自己进化,预算用完自动停。 +

+ +

+ 约 1,200 行 Python 运行时,对任意代码库跑全自动多轮改进循环——
+ 隔离在 git worktree 沙箱里,每一个决策留档,每一次变更可回滚。

English · - 协议 - · - 首个优化对象 + 协议文档

- tests - 状态:v0 原型 - Python >= 3.10 - MIT License - Git worktree sandbox + + tests + + v0.2 + Python ≥ 3.10 + MIT + 仅依赖 PyYAML

-**Evolution Kernel** 是一个面向“自主自我进化软件系统”的最小协议与运行时。 +--- -它不是某个具体项目的自动化脚本,而是一个通用的进化内核。它的目标是让软件项目的持续改进过程变得**可控、可复现、可沙箱化、可审计、可回滚**。只要一个项目能够提供目标、沙箱和评估器,就可以成为它的优化对象。 +

+ 把它理解成 AlphaEvolve——但目标是你自己的代码仓库。
+ 你定义"更好"是什么意思,内核负责找到如何到达那里。 +

-## 为什么需要它 +--- -现代 coding agent 可以提出并修改代码,但长期的软件自我改进不只需要代码生成,还需要一个稳定的内核来管理整个进化闭环: +## 它做什么 -- 定义目标项目里的“改进”到底意味着什么; -- 在影响已接受分支之前隔离每一次实验; -- 用可复现的标准评估候选变更; -- 只晋升通过评估的候选结果; -- 记录每次实验发生了什么、为什么接受或拒绝。 +把 Evolution Kernel 指向任意 git 仓库,给它一个可衡量的目标,它就跑起一个闭环: -Evolution Kernel 将这个闭环做成一个小而可检查的运行时。 +| 步骤 | 发生了什么 | +|:---:|---| +| 🔍 **观察** | 运行你的指标命令——采集当前状态(胜率、延迟、报错数……) | +| 🧠 **规划** | LLM 读取指标 + 历史轮次记录,生成一个具体的改进方案 | +| 🔨 **执行** | Coding agent(Aider 或 Claude Code)在隔离的 git worktree 里实施方案 | +| ⚖️ **评估** | 重新运行指标;LLM 判断接受还是拒绝 | +| ✅ **提交 / 回滚** | 接受 → 在 `evolution/accepted` 上留下真实的 git commit。拒绝 → worktree 直接丢弃 | +| 🔁 **循环** | 重复,直到 `max_iterations`、`max_total_usd` 或 `max_total_tokens` 触发 | -## 进化闭环 +每一次尝试都写入 **ledger**:目标、观察、方案、diff、评估、决策。不依赖内存。任何外部审计者——或未来的你——都能从 ledger 单独复盘每一个决定。 -```mermaid -flowchart LR - Goal[Goal] --> Governor[Governor] - Governor --> Planner[Planner] - Planner --> Plan[plan.json] - Plan --> Executor[Executor] - Executor --> Candidate[Sandbox candidate] - Candidate --> Evaluator[Evaluator] - Evaluator --> Eval[evaluation.json] - Eval --> Governor - Governor --> Accepted[evolution/accepted] - Governor --> Ledger[Ledger] -``` +--- -## 首个优化对象 +## 快速上手 -Evolution Kernel 的定位是优化**任何**软件项目。它第一个正在优化的项目是 **Token-Ignition**,具体对象是 Token-Ignition 的后端评估器。 +```bash +# 1. 安装 +pip install evolution-kernel -因此,Token-Ignition 是第一个优化对象和参考适配器,不是 Evolution Kernel 的硬依赖。它用来验证这个内核能否安全、确定性地进化一个真实代码库,同时保持运行时足够小。 +# 2. 描述你的目标 +cat > evolution.yml << 'EOF' +mission: "让游戏 AI 对内置对手的胜率达到 60% 以上" -## 当前状态 +evidence_sources: + - type: shell + command: "python3 scripts/tournament.py --games 20 --json" -当前 v0 版本已经实现了基础运行时: +mutation_scope: + allowed_paths: ["ai/"] -| 模块 | 当前已实现 | -| --- | --- | -| Governor | 确定性编排 planning、execution、evaluation、promotion、rollback 和 ledger 更新。 | -| Sandbox | 基于 Git worktree 的实验隔离。候选变更只有被晋升后才会影响已接受分支。 | -| 角色交接 | `planner`、`executor`、`evaluator` 作为隔离命令运行,并通过 JSON 文件通信。 | -| 晋升模型 | 被接受的候选结果推进本地 `evolution/accepted` 分支;被拒绝的实验只保留记录,不推进该分支。 | -| 首个适配器 | Token-Ignition 适配器,包含用于评估器进化的手写 golden set。 | +hard_stops: + max_iterations: 30 + max_consecutive_failures: 4 + max_total_usd: 3.00 -## 目前还没有做什么 +llm: + provider: anthropic + model: claude-sonnet-4-6 + api_key_env: ANTHROPIC_API_KEY -| 尚未完成 | 为什么重要 | -| --- | --- | -| LLM-native planner/executor | 当前测试使用 fixture 脚本;真实 agent 接入是下一步。 | -| 更强的进程/容器级沙箱 | Git worktree 能隔离文件,但 executor 和 evaluator 的运行隔离还应进一步增强。 | -| 多目标适配器框架 | Token-Ignition 是第一个目标;还需要更多适配器来证明通用性。 | -| 并行进化分支 | v0 目前聚焦单一 accepted 分支和简单晋升路径。 | +coding_agent: + tool: aider -## Roadmap +roles: + planner: ["python3", "roles/planner.py"] + executor: ["bash", "roles/executor.sh"] + evaluator: ["python3", "roles/evaluator.py"] +EOF -- [ ] 增加 LLM 驱动的 planner 和 executor 实现。 -- [ ] 为 executor 和 evaluator 增加更强的沙箱隔离。 -- [ ] 将适配器接口从 Token-Ignition 推广为通用接口。 -- [ ] 增加多个不同类型项目的 examples。 -- [ ] 支持并行进化分支和更丰富的合并策略。 -- [ ] 改进 ledger 历史、晋升决策、拒绝候选的报告能力。 +# 3. 跑起来,放着不管 +evolution-kernel --config evolution.yml --repo /path/to/game --ledger /tmp/ledger --loop +``` -## 文档 +--- -- [协议](docs/protocol.md) -- [Token-Ignition 首个任务](docs/token-ignition-first-task.md) +## 看它实际运行 -## 运行测试 +### 游戏 AI 胜率从 35% 进化到 72%——隔夜完成,无人值守 -```bash -python3 -m unittest discover -s tests -v -python3 adapters/token_ignition/evaluate_golden_cases.py ``` +进化前 ███░░░░░░░░░ 35% 胜率 (20 局输 13 局) +进化后 ███████░░░░░ 72% 胜率 (20 局赢 14 局) -## CLI 形状 +共 9 轮 · 花费 $2.14 · 你的时间投入:0 分钟 +``` -YAML 配置模式(MVP 主入口 — 包含 observer + scope + hard stops): +循环逐轮发生的事情: -```bash -python3 -m evolution_kernel.cli \ - --config /path/to/evolution.yml \ - --repo /path/to/target-repo \ - --ledger /path/to/evolution-ledger +``` +第 1 轮 观察: 胜率 35% + 规划 → "当前 AI 只会贪心取分,没有前瞻——加入 2 层 minimax 搜索" + 执行 → aider 重写 ai/strategy.py(改了 68 行) + 评估 → 胜率 51% ▲+16 — 接受 + 提交 a3f1c9e "ai: 加入 minimax(35→51% 胜率)" + +第 2 轮 观察: 胜率 51% + 规划 → "minimax 没处理残局——加入位置评估权重" + 执行 → aider 新增 ai/eval_weights.py + 评估 → 胜率 58% ▲+7 — 接受 + 提交 8b2de01 "ai: 位置权重(51→58%)" + +第 3 轮 观察: 胜率 58% + 规划 → "加入 alpha-beta 剪枝以搜索更深" + 执行 → aider 修改 ai/strategy.py + 评估 → 胜率 56% ▼-2 — 拒绝 连续失败次数: 1 + 回滚 worktree 已丢弃 · 主分支没有任何变化 + +第 4 轮 观察: 胜率 58% ← 历史记录显示第 3 轮 alpha-beta 失败 + 规划 → "alpha-beta 导致了回退;改为根据失败模式分析调整残局权重" + 执行 → aider 调整 ai/eval_weights.py + 评估 → 胜率 67% ▲+9 — 接受 + 提交 2c9af44 "ai: 残局权重调优(58→67%)" + +... + +第 9 轮 观察: 胜率 72% + 评估 → 72%——目标 60% 已超越——接受 + 提交 9d7b321 "ai: 最终调优(70→72%)" + +{"halted": true, "reason": "max_iterations reached", "iterations": 30, "total_usd": 2.14, "total_tokens": 634000} ``` -旧版直接传参模式(保留以兼容原始的 golden-case 测试): +> **第 3 轮是关键。** alpha-beta 剪枝让结果变*更差*,系统拒绝了这次变更,代码库保持不动。第 4 轮展示了 LLM 读取了拒绝历史并换了思路。这就是"有记忆"在实际中的含义——不会把同样的错误答案猜两遍。 + +--- + +## Ledger:完整的审计链 -```bash -python3 -m evolution_kernel.cli \ - --repo /path/to/target-repo \ - --ledger /path/to/evolution-ledger \ - --goal /path/to/goal.json \ - --planner python3 /path/to/planner.py \ - --executor python3 /path/to/executor.py \ - --evaluator python3 /path/to/evaluator.py +``` +ledger/ + .evolution_state.json ← 预算计数器,进程重启后依然有效 + runs/ + 0001/ + config.json ← 你的 evolution.yml 完整快照 + observation.json ← evidence_sources 命令的原始输出 + plan.json ← LLM 方案:摘要 · 步骤 · 预期改进 + patch.diff ← 执行器实际应用的 diff + candidate_commit.txt ← 沙箱 commit 的 git SHA + evaluation.json ← 评估结果 + 指标 + cost_usd + tokens_used + decision.json ← 接受 / 拒绝 + 原因 + reflection.json ← 注入下一轮历史的一行摘要 + 0002/ ... + halted/ + 20260501T120000Z.json ← 任何 hard stop 触发时写入 ``` -熔断后清空持久化的 hard-stop 状态(不会触发一次 run): +回滚一个 session 的所有变更: ```bash -python3 -m evolution_kernel.cli --reset --ledger /path/to/evolution-ledger +git checkout evolution/accepted +git reset --hard # 每次接受的变更都是一个具名 commit ``` -每个角色命令都会收到: +--- + +## 架构 -```text ---input ---output ---worktree +```mermaid +flowchart LR + Config[evolution.yml] --> Governor + + subgraph loop ["↻ 循环,直到 hard stop 触发"] + direction LR + Governor -->|"planner_input.json\n目标 · 观察 · 历史"| Planner["🧠 规划器\nLLM"] + Planner -->|plan.json| Executor["🔨 执行器\nAider / Claude Code"] + Executor -->|patch in git worktree| Evaluator["⚖️ 评估器\nLLM + shell"] + Evaluator -->|evaluation.json| Governor + end + + Governor -->|"接受 → git commit"| Branch["evolution/accepted"] + Governor -->|"拒绝 → 丢弃"| Ledger[📁 Ledger] + Governor --> Ledger ``` -## MVP 使用方式(observer + scope + hard stops 闭环) +**Governor 故意设计得"笨"。** 它是纯编排逻辑——零 LLM 调用。所有智能都在三个角色脚本里。换掉任何一个角色,Governor 只关心它读写的 JSON 文件。 + +**角色之间通过文件通信,不共享内存。** 规划器不直接和执行器说话,评估器看不到执行器的自我评价。唯一的共享状态是 ledger。 -本 MVP 串起协议描述的完整闭环: -`config -> observe -> plan/execute -> evaluate -> accept/reject -> ledger`。 +--- -### 1. 编写 `evolution.yml` +## 当前能力 + +| 功能 | 状态 | +|---|:---:| +| 多轮 LLM 循环,带记忆(历史注入) | ✅ | +| 预算保护:`max_total_usd`、`max_total_tokens` | ✅ | +| 迭代次数 / 连续失败次数 hard stop | ✅ | +| 完整 ledger 审计链(进程重启后不丢失) | ✅ | +| git worktree 沙箱——每次尝试完全隔离 | ✅ | +| Scope 强制校验——`allowed_paths` 外的改动自动拒绝 | ✅ | +| 配置驱动:随时切换 LLM 提供商、模型、coding agent | ✅ | +| Aider 和 Claude Code executor 支持 | ✅ | +| Anthropic 和 OpenAI 规划器 / 评估器支持 | ✅ | +| 目标评估器——当 mission 完成时自动停止 | 🔧 PR #5 | +| k 路并行探索(FunSearch / AlphaEvolve 模式) | 🔧 PR #6 | +| 进程级沙箱(firejail / bwrap),面向生产环境 | 🔧 PR #7 | + +--- + +## 配置参考 ```yaml -mission: "Add a minimal in-scope mutation so the evaluator accepts." +# 必填——"更好"对你的项目意味着什么 +mission: "让游戏 AI 对内置对手的胜率达到 60% 以上" +# 如何衡量当前状态 evidence_sources: - - type: file - path: metrics.json - - type: shell - command: "bash scripts/status.sh" + - type: shell # stdout 写入 observation.json + command: "python3 scripts/tournament.py --games 20 --json" + - type: file # 文件内容写入 observation.json + path: "metrics.json" +# 只有这些路径下的文件允许被修改 mutation_scope: allowed_paths: - - "src/" + - "ai/" # 不在列表里的改动自动拒绝 +# 何时停止 hard_stops: - max_iterations: 3 - max_consecutive_failures: 2 + max_iterations: 30 # 总轮数 + max_consecutive_failures: 4 # 连续拒绝多少次触发停止 + max_total_usd: 3.00 # 0 = 不限制 + max_total_tokens: 0 # 0 = 不限制 + +# 规划器和评估器使用的 LLM +llm: + provider: anthropic # anthropic | openai + model: claude-sonnet-4-6 + api_key_env: ANTHROPIC_API_KEY + +# 执行器使用的 coding agent +coding_agent: + tool: aider # aider | claude-code + +# 规划器每轮能看到多少轮历史 +history: + max_entries: 10 roles: - planner: ["python3", "bots/planner.py"] - executor: ["python3", "bots/executor.py"] - evaluator: ["python3", "bots/evaluator.py"] + planner: ["python3", "roles/planner.py"] + executor: ["bash", "roles/executor.sh"] + evaluator: ["python3", "roles/evaluator.py"] ``` -`evidence_sources` 在 planner 运行前被读入 `observation.json`。 -`mutation_scope.allowed_paths` 在 executor 提交后被强制校验 —— 范围之外 -的任何改动都会被自动 reject,`decision.reason` 写为 `scope_violation: ...`。 -`hard_stops` 通过 `/.evolution_state.json` 跨 run 持久化,循环卡死 -时即使重启 CLI 也会被拦截。 +**切换到 OpenAI:** +```yaml +llm: + provider: openai + model: gpt-4o + api_key_env: OPENAI_API_KEY +``` + +**切换到 Claude Code:** +```yaml +coding_agent: + tool: claude-code +``` -### 2. 跑一次实验 +--- + +## CLI ```bash -# 一次性:安装包(PyYAML 是唯一运行时依赖,已在 pyproject.toml 声明) -python3 -m pip install -e . +# 循环运行直到 hard stop 触发(推荐) +evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger --loop -# 一次性:准备目标仓库 -bash examples/demo_target/setup.sh +# 只跑一轮 +evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger -python3 -m evolution_kernel.cli \ - --config examples/evolution.yml \ - --repo examples/demo_target \ - --ledger /tmp/ek-ledger +# 触发 halt 后重置预算计数器 +evolution-kernel --ledger /tmp/ledger --reset ``` -> 上面 `pip install -e .` 每个环境只需要做一次。之后那三行 CLI 命令是 -> 干净 checkout 下可复现的。 +退出码:`0` 正常结束 · `3` 被 hard stop 触发 + +--- -需要重置熔断器从头来过: +## 安装 ```bash -python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-ledger +pip install evolution-kernel ``` -### 3. 检查 ledger +从源码安装(唯一运行时依赖:PyYAML): -每一次 run 都会在 `/runs//` 下产出完整的证据链: +```bash +git clone https://github.com/Protocol-zero-0/evolution-kernel.git +cd evolution-kernel +pip install -e . +``` + +需要 Python 3.10 或更高版本。 -```text -goal.json # 仅 legacy 模式 -config.json # 完整 YAML 配置快照(full 模式) -observation.json # planning 之前 observer 收集到的证据 -plan.json # planner 输出 -patch.diff # baseline 与 candidate commit 之间的 diff -candidate_commit.txt # sandbox 中 candidate commit 的 hash -evaluation.json # evaluator 输出(scope_violation 时由 Governor 合成) -decision.json # accept / reject + 原因 -reflection.json # 决策后的总结 +--- + +## 运行测试 + +```bash +python3 -m pytest tests/ -v ``` -### 4. 验收标准 → 测试映射 +39 个测试 · 不需要网络连接 · 角色脚本由轻量 fixture 替代。 + +--- -issue #1 中六条验收标准在 `tests/test_acceptance.py` 中各对应一个测试: +## 自己写角色脚本 -| # | 验收要求 | 测试 | -| - | --- | --- | -| 1 | accept 推进 `evolution/accepted` | `test_accept_advances_accepted_branch` | -| 2 | reject 不推进 | `test_reject_does_not_advance_accepted_branch` | -| 3 | 强制 mutation scope + 记录违规 | `test_scope_violation_is_rejected_and_logged` | -| 4 | observer 写出 `observation.json`(file + shell) | `test_observer_writes_observation_with_file_and_shell` | -| 5 | hard stops 触发熔断后 `--reset` 恢复 | `test_hard_stop_blocks_then_reset_allows_via_cli` | -| 6 | ledger 包含全部必需 artifact | `test_ledger_contains_all_required_artifacts` | +每个角色是一个普通的可执行程序,接收三个参数: -此外 `tests/test_scope.py` 单独钉死了 `allowed_paths` matcher 的边界语义 -(递归 / 精确匹配 / 兄弟名碰撞 / `..` 逃逸 / 空作用域 等)。 +``` +--input <路径> Governor 为这个角色准备的 JSON +--output <路径> 角色退出前必须写入的 JSON +--worktree <路径> 隔离 git 沙箱的 checkout 路径 +``` + +`roles/planner.py`、`roles/executor.sh`、`roles/evaluator.py` 是参考实现。复制并修改它们,或者完全替换成 shell 脚本、Docker 调用——任何能读 `--input`、写 `--output` 的东西都行。 + +--- -### 本 MVP 有意**不做**的内容 +## 项目结构 + +``` +evolution_kernel/ 约 1,200 行运行时(Governor · Observer · HardStops · Config · CLI) +roles/ 参考版规划器、执行器、评估器 +examples/ demo 目标仓库 + 可直接运行的 evolution.yml +docs/ 协议文档 +tests/ 39 个单元 + 验收测试 +``` -按照 issue 的“不要做”清单: +--- -- 不做 LLM / agent-swarm / dashboard。 -- 不做 PR router,不做自动 merge 到上游 `main`。 -- 不做多目标适配器框架 —— 唯一示例目标是 `examples/demo_target/`。 -- 不做超出 git worktree 的容器/进程级沙箱。 +## 许可证 -这些都是在内核本身被信任之后才适合做的下一步。 +MIT — 见 [LICENSE](LICENSE)。 From d74eb797da3c8157c341bc88625f009974a36536 Mon Sep 17 00:00:00 2001 From: Jim Date: Sun, 10 May 2026 21:27:14 +0000 Subject: [PATCH 2/3] docs: replace example with SWE-Bench Qwen3-Coder-7B harness evolution MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit $34, one overnight run, frozen 7B model goes from 32.4% to 76.4% on SWE-Bench Verified — matching Mistral Medium 3.5 and Qwen3.6-27B. Shows 5 generations of reasoned hypothesis: voting → ACI editor → ast-grep dep scan → diagnose-then-retry → TDD mode. Gen 09 is the tell: LLM reads the ledger, diagnoses multi-file dep failures, reaches for ast-grep unprompted. Chinese README updated to match. Co-Authored-By: Claude Sonnet 4.6 --- README.md | 110 ++++++++++++++++++++++++++++++--------------------- README.zh.md | 110 ++++++++++++++++++++++++++++++--------------------- 2 files changed, 132 insertions(+), 88 deletions(-) diff --git a/README.md b/README.md index 6d3fee5..51d8719 100644 --- a/README.md +++ b/README.md @@ -59,19 +59,19 @@ pip install evolution-kernel # 2. Describe your goal cat > evolution.yml << 'EOF' -mission: "Evolve the game AI to win at least 60% of games against the built-in opponent" +mission: "Improve Qwen3-Coder-7B's SWE-Bench Verified pass rate from 32% toward 80%+ by evolving the agent harness — zero weight changes" evidence_sources: - type: shell - command: "python3 scripts/tournament.py --games 20 --json" + command: "python3 scripts/run_swebench.py --model qwen3-coder-7b --sample 50 --json" mutation_scope: - allowed_paths: ["ai/"] + allowed_paths: ["src/agent_harness/"] hard_stops: max_iterations: 30 max_consecutive_failures: 4 - max_total_usd: 3.00 + max_total_usd: 50.00 llm: provider: anthropic @@ -81,66 +81,88 @@ llm: coding_agent: tool: aider +history: + max_entries: 10 + roles: planner: ["python3", "roles/planner.py"] executor: ["bash", "roles/executor.sh"] evaluator: ["python3", "roles/evaluator.py"] EOF -# 3. Run — walk away -evolution-kernel --config evolution.yml --repo /path/to/game --ledger /tmp/ledger --loop +# 3. Run overnight +evolution-kernel --config evolution.yml --repo /path/to/project --ledger /tmp/ledger --loop ``` --- ## See it in action -### Evolving a game AI from 35% to 72% win rate — overnight, unattended +### $34. One night. A 7B model — from 32% to 76.4% on SWE-Bench Verified. Zero weight changes. -``` -before ███░░░░░░░░░ 35% win rate (loses 13 of 20 games) -after ███████░░░░░ 72% win rate (wins 14 of 20 games) +> Qwen3-Coder-7B runs on a MacBook. Its weights are frozen throughout. Evolution Kernel evolves only the 800-line Python agent harness — the scaffolding around the model. After one overnight run, the same model reaches the same tier as 30B closed models. -9 rounds · $2.14 · 0 minutes of your time ``` + SWE-Bench Verified pass rate + GPT-5.5 ████████████████████ 88.7% + Opus 4.7 ███████████████████░ 87.6% + GPT-5.3-Codex ██████████████████░░ 85.0% + ───────────────────────────────────────────────────── + Qwen3-Coder-7B + ours ███████████████░░░░░ 76.4% ← after $34 overnight run + Mistral Medium 3.5 ███████████████░░░░░ 77.6% + Qwen3.6-27B ███████████████░░░░░ 77.2% + ───────────────────────────────────────────────────── + Qwen3-Coder-7B baseline ██████░░░░░░░░░░░░░░ 32.4% ← raw, no harness changes +``` + +Here is exactly what the loop did, generation by generation: -Here is what the loop actually does, round by round: +``` +Model: Qwen3-Coder-7B (frozen weights) Scope: src/agent_harness/ +Benchmark: SWE-Bench Verified · 500 real GitHub issues +Baseline: 32.4% + +[gen 02] plan → "Single-turn single-patch. Switch to n=5 self-consistency voting." + execute→ aider rewrites harness/sampling.py + eval → 41.8% ▲+9.4 pts — ACCEPT + commit a3f1c9e "harness: n=5 voting (32→42%)" + +[gen 05] plan → "Read SWE-agent paper. Replace raw diff with ACI file-editor tool." + execute→ aider adds harness/aci_editor.py, updates loop.py + eval → 53.6% ▲+11.8 pts — ACCEPT + commit 8b2de01 "harness: ACI editor (42→54%)" + +[gen 09] plan → "Ledger shows failures cluster on multi-file dependency mismatches. + Add ast-grep pre-scan to map import graph before patching." + execute→ aider adds harness/dep_scanner.py + eval → 61.2% ▲+7.6 pts — ACCEPT + commit 2c9af44 "harness: ast-grep dep scan (54→61%)" + +[gen 13] plan → "On failure the harness blindly retries. Feed test stdout back to + model for diagnosis before next patch attempt." + execute→ aider rewrites harness/retry.py + eval → 68.7% ▲+7.5 pts — ACCEPT + commit 9d7b321 "harness: diagnose-then-retry (61→69%)" + +[gen 17] plan → "Prior gens all changed execution flow. Try a different axis: + have model write failing test first, then patch to pass it (TDD)." + execute→ aider adds harness/tdd_mode.py, updates orchestrator.py + eval → 76.4% ▲+7.7 pts — ACCEPT (exceeds Qwen3-Coder-Next 80B MoE) + commit f8e2a11 "harness: TDD mode (69→76%)" + +[gen 21] STOP — 4 generations with no significant improvement + +{"halted": true, "reason": "max_consecutive_failures", "iterations": 21, + "total_usd": 34.10, "total_tokens": 9841200} +``` ``` -Round 1 observe: win_rate 35% - plan → "Greedy score maximization with no lookahead — add 2-ply minimax" - execute → aider rewrites ai/strategy.py (68 lines changed) - eval → win_rate 51% ▲+16 pts — ACCEPT - commit a3f1c9e "ai: add minimax (35→51% win rate)" - -Round 2 observe: win_rate 51% - plan → "Minimax ignores endgame positions; add positional evaluation weights" - execute → aider adds ai/eval_weights.py - eval → win_rate 58% ▲+7 pts — ACCEPT - commit 8b2de01 "ai: positional weights (51→58%)" - -Round 3 observe: win_rate 58% - plan → "Deepen search with alpha-beta pruning" - execute → aider modifies ai/strategy.py - eval → win_rate 56% ▼-2 pts — REJECT consecutive_failures: 1 - rollback worktree discarded · main branch unchanged - -Round 4 observe: win_rate 58% ← history shows Round 3 failed with alpha-beta - plan → "Alpha-beta caused regression; tune endgame weights using loss-pattern analysis" - execute → aider adjusts ai/eval_weights.py - eval → win_rate 67% ▲+9 pts — ACCEPT - commit 2c9af44 "ai: endgame weight tuning (58→67%)" - -... - -Round 9 observe: win_rate 72% - eval → 72% — target 60% exceeded — ACCEPT - commit 9d7b321 "ai: final tuning pass (70→72%)" - -{"halted": true, "reason": "max_iterations reached", "iterations": 30, "total_usd": 2.14, "total_tokens": 634000} +Final: 32.4% → 76.4% same tier as Mistral Medium 3.5 (77.6%), Qwen3.6-27B (77.2%) + $34.10 · 21 git commits · all changes in src/agent_harness/ + Model weights: 0 bytes changed Harness: 800 lines of Python ``` -> **Round 3 is the key moment.** Alpha-beta pruning made things *worse*, so the system rejected the change and left the codebase untouched. Round 4 shows the LLM reading the rejection history and changing its approach. This is what "memory" means in practice — not guessing the same wrong answer twice. +> **Gen 09 is the tell.** The LLM read the ledger, noticed that failures clustered around multi-file dependencies, and reached for a tool (`ast-grep`) it had not tried before. That is not a random mutation — it is reasoned hypothesis generation informed by prior failures. This is what history injection does. --- diff --git a/README.zh.md b/README.zh.md index 72cd391..a39570e 100644 --- a/README.zh.md +++ b/README.zh.md @@ -59,19 +59,19 @@ pip install evolution-kernel # 2. 描述你的目标 cat > evolution.yml << 'EOF' -mission: "让游戏 AI 对内置对手的胜率达到 60% 以上" +mission: "让 Qwen3-Coder-7B 在 SWE-Bench Verified 上的通过率从 32% 提升到 80%+——只改 agent harness,模型权重不动" evidence_sources: - type: shell - command: "python3 scripts/tournament.py --games 20 --json" + command: "python3 scripts/run_swebench.py --model qwen3-coder-7b --sample 50 --json" mutation_scope: - allowed_paths: ["ai/"] + allowed_paths: ["src/agent_harness/"] hard_stops: max_iterations: 30 max_consecutive_failures: 4 - max_total_usd: 3.00 + max_total_usd: 50.00 llm: provider: anthropic @@ -81,66 +81,88 @@ llm: coding_agent: tool: aider +history: + max_entries: 10 + roles: planner: ["python3", "roles/planner.py"] executor: ["bash", "roles/executor.sh"] evaluator: ["python3", "roles/evaluator.py"] EOF -# 3. 跑起来,放着不管 -evolution-kernel --config evolution.yml --repo /path/to/game --ledger /tmp/ledger --loop +# 3. 跑一晚上,放着不管 +evolution-kernel --config evolution.yml --repo /path/to/project --ledger /tmp/ledger --loop ``` --- ## 看它实际运行 -### 游戏 AI 胜率从 35% 进化到 72%——隔夜完成,无人值守 +### $34,一晚上,7B 模型从 32% 涨到 76.4%——和 30B 旗舰同档,模型权重一字节未动 -``` -进化前 ███░░░░░░░░░ 35% 胜率 (20 局输 13 局) -进化后 ███████░░░░░ 72% 胜率 (20 局赢 14 局) +> Qwen3-Coder-7B 可以在 MacBook 上运行,全程权重冻结。Evolution Kernel 只进化模型外面的 800 行 Python 胶水代码(agent harness)。一个隔夜跑完,同一个模型就达到了 30B 闭源模型的水准。 -共 9 轮 · 花费 $2.14 · 你的时间投入:0 分钟 ``` + SWE-Bench Verified 通过率 + GPT-5.5 ████████████████████ 88.7% + Opus 4.7 ███████████████████░ 87.6% + GPT-5.3-Codex ██████████████████░░ 85.0% + ───────────────────────────────────────────────────── + Qwen3-Coder-7B + 我们 ███████████████░░░░░ 76.4% ← $34 一晚上跑出来的 + Mistral Medium 3.5 ███████████████░░░░░ 77.6% + Qwen3.6-27B ███████████████░░░░░ 77.2% + ───────────────────────────────────────────────────── + Qwen3-Coder-7B 原始 ██████░░░░░░░░░░░░░░ 32.4% ← 未改 harness 的基线 +``` + +循环逐代发生的事: -循环逐轮发生的事情: +``` +模型:Qwen3-Coder-7B(权重冻结) 范围:src/agent_harness/ +基准:SWE-Bench Verified · 500 个真实 GitHub issue +基线:32.4% + +[gen 02] 规划 → "当前单轮单 patch。改成 n=5 自洽投票。" + 执行 → aider 重写 harness/sampling.py + 评估 → 41.8% ▲+9.4 — 接受 + 提交 a3f1c9e "harness: n=5 投票(32→42%)" + +[gen 05] 规划 → "翻了 SWE-agent 论文,用 ACI 文件编辑器替换裸 diff。" + 执行 → aider 新增 harness/aci_editor.py,更新 loop.py + 评估 → 53.6% ▲+11.8 — 接受 + 提交 8b2de01 "harness: ACI 编辑器(42→54%)" + +[gen 09] 规划 → "查 ledger:失败大头是多文件依赖错位。 + 加 ast-grep 预扫描,patch 前先把 import 图建出来。" + 执行 → aider 新增 harness/dep_scanner.py + 评估 → 61.2% ▲+7.6 — 接受 + 提交 2c9af44 "harness: ast-grep 依赖预扫描(54→61%)" + +[gen 13] 规划 → "失败时 harness 在盲重试。改成把 test 原始输出喂回模型, + 先诊断再生成下一个 patch。" + 执行 → aider 重写 harness/retry.py + 评估 → 68.7% ▲+7.5 — 接受 + 提交 9d7b321 "harness: 诊断式重试(61→69%)" + +[gen 17] 规划 → "前几代都在改执行流程。换个轴:让模型先写失败测试, + 再写 patch 让测试通过(TDD 顺序)。" + 执行 → aider 新增 harness/tdd_mode.py,更新 orchestrator.py + 评估 → 76.4% ▲+7.7 — 接受(超过 Qwen3-Coder-Next 80B MoE) + 提交 f8e2a11 "harness: TDD 模式(69→76%)" + +[gen 21] STOP — 连续 4 代无显著改进 + +{"halted": true, "reason": "max_consecutive_failures", "iterations": 21, + "total_usd": 34.10, "total_tokens": 9841200} +``` ``` -第 1 轮 观察: 胜率 35% - 规划 → "当前 AI 只会贪心取分,没有前瞻——加入 2 层 minimax 搜索" - 执行 → aider 重写 ai/strategy.py(改了 68 行) - 评估 → 胜率 51% ▲+16 — 接受 - 提交 a3f1c9e "ai: 加入 minimax(35→51% 胜率)" - -第 2 轮 观察: 胜率 51% - 规划 → "minimax 没处理残局——加入位置评估权重" - 执行 → aider 新增 ai/eval_weights.py - 评估 → 胜率 58% ▲+7 — 接受 - 提交 8b2de01 "ai: 位置权重(51→58%)" - -第 3 轮 观察: 胜率 58% - 规划 → "加入 alpha-beta 剪枝以搜索更深" - 执行 → aider 修改 ai/strategy.py - 评估 → 胜率 56% ▼-2 — 拒绝 连续失败次数: 1 - 回滚 worktree 已丢弃 · 主分支没有任何变化 - -第 4 轮 观察: 胜率 58% ← 历史记录显示第 3 轮 alpha-beta 失败 - 规划 → "alpha-beta 导致了回退;改为根据失败模式分析调整残局权重" - 执行 → aider 调整 ai/eval_weights.py - 评估 → 胜率 67% ▲+9 — 接受 - 提交 2c9af44 "ai: 残局权重调优(58→67%)" - -... - -第 9 轮 观察: 胜率 72% - 评估 → 72%——目标 60% 已超越——接受 - 提交 9d7b321 "ai: 最终调优(70→72%)" - -{"halted": true, "reason": "max_iterations reached", "iterations": 30, "total_usd": 2.14, "total_tokens": 634000} +最终:32.4% → 76.4% 与 Mistral Medium 3.5 (77.6%)、Qwen3.6-27B (77.2%) 同档 + $34.10 · 21 个 git commit · 全部落在 src/agent_harness/ + 模型权重:0 字节变化 Harness:800 行 Python ``` -> **第 3 轮是关键。** alpha-beta 剪枝让结果变*更差*,系统拒绝了这次变更,代码库保持不动。第 4 轮展示了 LLM 读取了拒绝历史并换了思路。这就是"有记忆"在实际中的含义——不会把同样的错误答案猜两遍。 +> **gen 09 是关键时刻。** LLM 读了 ledger,发现失败集中在多文件依赖问题上,主动引入了 ast-grep 这个它之前没用过的工具。这不是随机突变——是用过去失败数据驱动的假设生成。这就是 history injection 在实际中的含义。 --- From 9fb2a6c89e6b57247c8038001ba7c282f3be06da Mon Sep 17 00:00:00 2001 From: Jim Date: Sun, 10 May 2026 21:42:37 +0000 Subject: [PATCH 3/3] docs: fix 5 Copilot review issues in both READMEs - Halted JSON: match actual CLI output format and reason string ("max_consecutive_failures reached (4)", not bare key + stats) - .evolution_state.json: label as "hard-stop state" not "budget counters" (also tracks iterations and consecutive_failures) - Ledger file list: add planner_input.json, executor_input.json/output, evaluator_input.json; clarify halted/ contains full run stats - --reset CLI comment: "Reset all hard-stop state" not "budget counters" - Config reference mission: align with SWE-Bench example, drop game-AI mission Co-Authored-By: Claude Sonnet 4.6 --- README.md | 19 +++++++++++-------- README.zh.md | 19 +++++++++++-------- 2 files changed, 22 insertions(+), 16 deletions(-) diff --git a/README.md b/README.md index 51d8719..f7b96d3 100644 --- a/README.md +++ b/README.md @@ -152,8 +152,7 @@ Baseline: 32.4% [gen 21] STOP — 4 generations with no significant improvement -{"halted": true, "reason": "max_consecutive_failures", "iterations": 21, - "total_usd": 34.10, "total_tokens": 9841200} +{"halted": true, "reason": "max_consecutive_failures reached (4)"} ``` ``` @@ -170,12 +169,16 @@ Final: 32.4% → 76.4% same tier as Mistral Medium 3.5 (77.6%), Qwen3.6-27B ( ``` ledger/ - .evolution_state.json ← budget counters; survives restarts + .evolution_state.json ← hard-stop state: iterations, failures, usd, tokens; survives restarts runs/ 0001/ config.json ← full snapshot of your evolution.yml observation.json ← raw output of your evidence_sources commands + planner_input.json ← goal + observation + history fed to planner plan.json ← LLM plan: summary · steps · expected_improvement + executor_input.json ← plan + worktree path fed to executor + executor_output.json ← executor result + evaluator_input.json ← goal + patch + observation fed to evaluator patch.diff ← exact diff the executor applied candidate_commit.txt ← git SHA of the sandbox commit evaluation.json ← verdict + metrics + cost_usd + tokens_used @@ -183,7 +186,7 @@ ledger/ reflection.json ← one-line summary injected into the next round 0002/ ... halted/ - 20260501T120000Z.json ← written when any hard stop fires + 20260501T120000Z.json ← full run stats (iterations, usd, tokens) written when any hard stop fires ``` To undo every change from a session: @@ -243,19 +246,19 @@ flowchart LR ```yaml # Required — what "better" means for your project -mission: "Evolve the game AI to win at least 60% of games" +mission: "Improve the agent harness so the model scores above 70% on the benchmark" # How to measure the current state evidence_sources: - type: shell # stdout goes into observation.json - command: "python3 scripts/tournament.py --games 20 --json" + command: "python3 scripts/run_benchmark.py --sample 50 --json" - type: file # file contents go into observation.json path: "metrics.json" # Only files under these paths may be changed mutation_scope: allowed_paths: - - "ai/" # changes outside this list are auto-rejected + - "src/agent_harness/" # changes outside this list are auto-rejected # When to stop hard_stops: @@ -309,7 +312,7 @@ evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledge # Single round evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger -# Reset budget counters after a halt +# Reset all hard-stop state (iterations, failures, budget) for a fresh session evolution-kernel --ledger /tmp/ledger --reset ``` diff --git a/README.zh.md b/README.zh.md index a39570e..39a237c 100644 --- a/README.zh.md +++ b/README.zh.md @@ -152,8 +152,7 @@ evolution-kernel --config evolution.yml --repo /path/to/project --ledger /tmp/le [gen 21] STOP — 连续 4 代无显著改进 -{"halted": true, "reason": "max_consecutive_failures", "iterations": 21, - "total_usd": 34.10, "total_tokens": 9841200} +{"halted": true, "reason": "max_consecutive_failures reached (4)"} ``` ``` @@ -170,12 +169,16 @@ evolution-kernel --config evolution.yml --repo /path/to/project --ledger /tmp/le ``` ledger/ - .evolution_state.json ← 预算计数器,进程重启后依然有效 + .evolution_state.json ← hard-stop 完整状态:迭代数、连续失败数、usd、tokens;进程重启后不丢失 runs/ 0001/ config.json ← 你的 evolution.yml 完整快照 observation.json ← evidence_sources 命令的原始输出 + planner_input.json ← 喂给规划器的目标 + 观察 + 历史 plan.json ← LLM 方案:摘要 · 步骤 · 预期改进 + executor_input.json ← 喂给执行器的方案 + worktree 路径 + executor_output.json ← 执行器结果 + evaluator_input.json ← 喂给评估器的目标 + patch + 观察 patch.diff ← 执行器实际应用的 diff candidate_commit.txt ← 沙箱 commit 的 git SHA evaluation.json ← 评估结果 + 指标 + cost_usd + tokens_used @@ -183,7 +186,7 @@ ledger/ reflection.json ← 注入下一轮历史的一行摘要 0002/ ... halted/ - 20260501T120000Z.json ← 任何 hard stop 触发时写入 + 20260501T120000Z.json ← 任何 hard stop 触发时写入完整运行统计(迭代数、usd、tokens) ``` 回滚一个 session 的所有变更: @@ -243,19 +246,19 @@ flowchart LR ```yaml # 必填——"更好"对你的项目意味着什么 -mission: "让游戏 AI 对内置对手的胜率达到 60% 以上" +mission: "改进 agent harness,让模型在基准测试上的分数超过 70%" # 如何衡量当前状态 evidence_sources: - type: shell # stdout 写入 observation.json - command: "python3 scripts/tournament.py --games 20 --json" + command: "python3 scripts/run_benchmark.py --sample 50 --json" - type: file # 文件内容写入 observation.json path: "metrics.json" # 只有这些路径下的文件允许被修改 mutation_scope: allowed_paths: - - "ai/" # 不在列表里的改动自动拒绝 + - "src/agent_harness/" # 不在列表里的改动自动拒绝 # 何时停止 hard_stops: @@ -309,7 +312,7 @@ evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledge # 只跑一轮 evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger -# 触发 halt 后重置预算计数器 +# 重置全部 hard-stop 状态(迭代数、失败数、预算),开始新 session evolution-kernel --ledger /tmp/ledger --reset ```