diff --git a/README.md b/README.md
index d271571..f7b96d3 100644
--- a/README.md
+++ b/README.md
@@ -1,11 +1,12 @@
# Evolution Kernel
- Give an LLM a goal. Watch your repo improve itself. Stop when the budget runs out.
+ Give an LLM a goal. Watch your codebase improve itself. Stop when the budget runs out.
- A ~1,200-line Python runtime for autonomous, multi-round code improvement — sandboxed, audited, and fully reversible.
+ A ~1,200-line Python runtime that runs an autonomous, multi-round improvement loop on any codebase —
+ sandboxed in git worktrees, every decision logged, every change reversible.
@@ -19,52 +20,58 @@
-
-
-
+
+
+
+
+
+---
+
+
+ Think of it as AlphaEvolve — but pointed at your own repository.
+ You define what "better" means. The kernel figures out how to get there.
---
## What it does
-Write a YAML file that says what "better" means. Evolution Kernel runs a tight loop:
+Point Evolution Kernel at any git repository and give it a measurable goal. It runs a closed loop:
-1. **Observe** — collect the current metric (coverage %, benchmark score, lint count — whatever your shell command outputs)
-2. **Plan** — an LLM reads the metric and the history of prior attempts, then writes a concrete plan
-3. **Execute** — a coding agent (Aider or Claude Code) applies the plan inside a git worktree sandbox
-4. **Evaluate** — the evaluator re-runs your metric command and decides accept or reject
-5. **Commit or roll back** — accepted changes become a real git commit; rejected ones are discarded
-6. **Loop** — repeat until a budget limit fires (`max_iterations`, `max_total_usd`, `max_total_tokens`)
+| Step | What happens |
+|:---:|---|
+| 🔍 **Observe** | Run your metric command — collect the current state (win rate, latency, error count, …) |
+| 🧠 **Plan** | LLM reads the metric + history of prior attempts, produces a concrete plan |
+| 🔨 **Execute** | Coding agent (Aider or Claude Code) applies the plan inside an isolated git worktree |
+| ⚖️ **Evaluate** | Re-run your metric; LLM decides accept or reject |
+| ✅ **Commit / rollback** | Accepted → real git commit on `evolution/accepted`. Rejected → worktree discarded |
+| 🔁 **Loop** | Repeat until `max_iterations`, `max_total_usd`, or `max_total_tokens` fires |
-Every attempt — accepted or rejected — is written to a structured **ledger** so you can audit exactly what the LLM tried, what changed, and why each round was accepted or rejected.
+Every attempt is written to a **ledger**: goal, observation, plan, diff, evaluation, decision. Nothing is held in memory. An external auditor — or your future self — can reconstruct every decision from the ledger alone.
---
## Quick Start
```bash
-# 1. Install (single runtime dependency: PyYAML)
+# 1. Install
pip install evolution-kernel
-# 2. Write a goal config
+# 2. Describe your goal
cat > evolution.yml << 'EOF'
-mission: "Increase src/ test coverage from 40% to 80%"
+mission: "Improve Qwen3-Coder-7B's SWE-Bench Verified pass rate from 32% toward 80%+ by evolving the agent harness — zero weight changes"
evidence_sources:
- type: shell
- command: >
- python3 -m pytest --cov=src --cov-report=json -q &&
- python3 -c "import json; d=json.load(open('coverage.json'));
- print(f'coverage: {d[\"totals\"][\"percent_covered\"]:.1f}%')"
+ command: "python3 scripts/run_swebench.py --model qwen3-coder-7b --sample 50 --json"
mutation_scope:
- allowed_paths: ["tests/"]
+ allowed_paths: ["src/agent_harness/"]
hard_stops:
- max_iterations: 20
- max_consecutive_failures: 3
- max_total_usd: 2.00
+ max_iterations: 30
+ max_consecutive_failures: 4
+ max_total_usd: 50.00
llm:
provider: anthropic
@@ -74,81 +81,119 @@ llm:
coding_agent:
tool: aider
+history:
+ max_entries: 10
+
roles:
planner: ["python3", "roles/planner.py"]
executor: ["bash", "roles/executor.sh"]
evaluator: ["python3", "roles/evaluator.py"]
EOF
-# 3. Run until the budget fires
-evolution-kernel --config evolution.yml --repo /path/to/your-project --ledger /tmp/ledger --loop
+# 3. Run overnight
+evolution-kernel --config evolution.yml --repo /path/to/project --ledger /tmp/ledger --loop
```
---
-## Example: raising test coverage from 40% to 80%
+## See it in action
-The loop emits one JSON object per round. A realistic session looks like this:
+### $34. One night. A 7B model — from 32% to 76.4% on SWE-Bench Verified. Zero weight changes.
+
+> Qwen3-Coder-7B runs on a MacBook. Its weights are frozen throughout. Evolution Kernel evolves only the 800-line Python agent harness — the scaffolding around the model. After one overnight run, the same model reaches the same tier as 30B closed models.
```
-Round 1 observe: coverage 40.2%
- plan → "Add unit tests for src/parser.py — parse_tokens is completely uncovered"
- execute → aider writes tests/test_parser.py (14 new assertions)
- eval → coverage 51.7% — ACCEPT
- commit → a3f1c9e "tests: cover parse_tokens (coverage 40→52%)"
-
-Round 2 observe: coverage 51.7%
- plan → "Add edge-case tests for src/validator.py, missing branch coverage on error paths"
- execute → aider extends tests/test_validator.py (+9 tests)
- eval → coverage 63.4% — ACCEPT
- commit → 8b2de01 "tests: validator edge cases (coverage 52→63%)"
-
-Round 3 observe: coverage 63.4%
- plan → "Cover src/formatter.py — currently 0% covered"
- execute → aider writes tests/test_formatter.py
- eval → coverage 63.4% — new test file has wrong import path — REJECT
- rollback → worktree discarded, main branch unchanged (consecutive_failures: 1)
-
-Round 4 observe: coverage 63.4%
- plan → "tests/test_formatter.py failed due to import error; fix path and retry"
- execute → aider fixes import in tests/test_formatter.py
- eval → coverage 74.8% — ACCEPT
- commit → 2c9af44 "tests: formatter coverage, fixed import (coverage 63→75%)"
-
-...
-
-Round 12 observe: coverage 80.1%
- eval → coverage 80.1% — threshold reached — ACCEPT
- commit → 9d7b321 "tests: final push past 80% target"
-
-{"halted": true, "reason": "max_iterations reached", "iterations": 20, "total_usd": 1.43, "total_tokens": 487201}
+ SWE-Bench Verified pass rate
+ GPT-5.5 ████████████████████ 88.7%
+ Opus 4.7 ███████████████████░ 87.6%
+ GPT-5.3-Codex ██████████████████░░ 85.0%
+ ─────────────────────────────────────────────────────
+ Qwen3-Coder-7B + ours ███████████████░░░░░ 76.4% ← after $34 overnight run
+ Mistral Medium 3.5 ███████████████░░░░░ 77.6%
+ Qwen3.6-27B ███████████████░░░░░ 77.2%
+ ─────────────────────────────────────────────────────
+ Qwen3-Coder-7B baseline ██████░░░░░░░░░░░░░░ 32.4% ← raw, no harness changes
```
-Each accepted change is a reversible git commit on the `evolution/accepted` branch. The LLM self-corrected on Round 4 using the rejection history from Round 3 — this is what history injection does.
+Here is exactly what the loop did, generation by generation:
----
+```
+Model: Qwen3-Coder-7B (frozen weights) Scope: src/agent_harness/
+Benchmark: SWE-Bench Verified · 500 real GitHub issues
+Baseline: 32.4%
+
+[gen 02] plan → "Single-turn single-patch. Switch to n=5 self-consistency voting."
+ execute→ aider rewrites harness/sampling.py
+ eval → 41.8% ▲+9.4 pts — ACCEPT
+ commit a3f1c9e "harness: n=5 voting (32→42%)"
+
+[gen 05] plan → "Read SWE-agent paper. Replace raw diff with ACI file-editor tool."
+ execute→ aider adds harness/aci_editor.py, updates loop.py
+ eval → 53.6% ▲+11.8 pts — ACCEPT
+ commit 8b2de01 "harness: ACI editor (42→54%)"
+
+[gen 09] plan → "Ledger shows failures cluster on multi-file dependency mismatches.
+ Add ast-grep pre-scan to map import graph before patching."
+ execute→ aider adds harness/dep_scanner.py
+ eval → 61.2% ▲+7.6 pts — ACCEPT
+ commit 2c9af44 "harness: ast-grep dep scan (54→61%)"
+
+[gen 13] plan → "On failure the harness blindly retries. Feed test stdout back to
+ model for diagnosis before next patch attempt."
+ execute→ aider rewrites harness/retry.py
+ eval → 68.7% ▲+7.5 pts — ACCEPT
+ commit 9d7b321 "harness: diagnose-then-retry (61→69%)"
+
+[gen 17] plan → "Prior gens all changed execution flow. Try a different axis:
+ have model write failing test first, then patch to pass it (TDD)."
+ execute→ aider adds harness/tdd_mode.py, updates orchestrator.py
+ eval → 76.4% ▲+7.7 pts — ACCEPT (exceeds Qwen3-Coder-Next 80B MoE)
+ commit f8e2a11 "harness: TDD mode (69→76%)"
+
+[gen 21] STOP — 4 generations with no significant improvement
+
+{"halted": true, "reason": "max_consecutive_failures reached (4)"}
+```
+
+```
+Final: 32.4% → 76.4% same tier as Mistral Medium 3.5 (77.6%), Qwen3.6-27B (77.2%)
+ $34.10 · 21 git commits · all changes in src/agent_harness/
+ Model weights: 0 bytes changed Harness: 800 lines of Python
+```
+
+> **Gen 09 is the tell.** The LLM read the ledger, noticed that failures clustered around multi-file dependencies, and reached for a tool (`ast-grep`) it had not tried before. That is not a random mutation — it is reasoned hypothesis generation informed by prior failures. This is what history injection does.
-## Ledger structure
+---
-Every round writes a full evidence trail. Nothing is stored in memory; an external auditor can reconstruct every decision from the ledger directory alone.
+## Ledger: the complete audit trail
```
ledger/
- .evolution_state.json # persisted counters (iterations, usd, tokens) — survives restarts
+ .evolution_state.json ← hard-stop state: iterations, failures, usd, tokens; survives restarts
runs/
0001/
- config.json # full snapshot of your evolution.yml
- observation.json # raw output of your evidence_sources commands
- plan.json # LLM plan: summary, steps, expected_improvement
- patch.diff # exact diff the executor applied
- candidate_commit.txt # git SHA of the sandbox commit
- evaluation.json # verdict + metrics + cost_usd + tokens_used
- decision.json # accept / reject + reason
- reflection.json # one-line summary injected into the next round's history
- 0002/
- ...
+ config.json ← full snapshot of your evolution.yml
+ observation.json ← raw output of your evidence_sources commands
+ planner_input.json ← goal + observation + history fed to planner
+ plan.json ← LLM plan: summary · steps · expected_improvement
+ executor_input.json ← plan + worktree path fed to executor
+ executor_output.json ← executor result
+ evaluator_input.json ← goal + patch + observation fed to evaluator
+ patch.diff ← exact diff the executor applied
+ candidate_commit.txt ← git SHA of the sandbox commit
+ evaluation.json ← verdict + metrics + cost_usd + tokens_used
+ decision.json ← accept / reject + reason
+ reflection.json ← one-line summary injected into the next round
+ 0002/ ...
halted/
- 20260501T120000Z.json # written when any hard stop fires
+ 20260501T120000Z.json ← full run stats (iterations, usd, tokens) written when any hard stop fires
+```
+
+To undo every change from a session:
+
+```bash
+git checkout evolution/accepted
+git reset --hard # every accepted change is a named commit
```
---
@@ -159,40 +204,40 @@ ledger/
flowchart LR
Config[evolution.yml] --> Governor
- subgraph loop ["Loop until hard stop"]
+ subgraph loop ["↻ Loop until hard stop fires"]
direction LR
- Governor -->|"planner_input.json\n(goal + observation + history)"| Planner["Planner\nLLM"]
- Planner -->|plan.json| Executor["Executor\nAider / Claude Code"]
- Executor -->|patch in git worktree| Evaluator["Evaluator\nLLM + shell"]
+ Governor -->|"planner_input.json\ngoal · observation · history"| Planner["🧠 Planner\nLLM"]
+ Planner -->|plan.json| Executor["🔨 Executor\nAider / Claude Code"]
+ Executor -->|patch in git worktree| Evaluator["⚖️ Evaluator\nLLM + shell"]
Evaluator -->|evaluation.json| Governor
end
- Governor -->|"accept → git commit"| AcceptedBranch[evolution/accepted]
- Governor -->|"reject → discard worktree"| Ledger[Ledger]
+ Governor -->|"accept → git commit"| Branch["evolution/accepted"]
+ Governor -->|"reject → discard"| Ledger[📁 Ledger]
Governor --> Ledger
```
-**The Governor is intentionally dumb.** It is pure orchestration — no LLM calls of its own. All intelligence lives in the three role scripts. You can swap any role for your own implementation; the Governor only cares about the JSON files roles read and write.
+**The Governor is intentionally dumb.** It is pure orchestration — zero LLM calls. All intelligence lives in the three role scripts. Swap any role for your own implementation; the Governor only cares about the JSON each role reads and writes.
-**Roles communicate through files, not shared memory.** The planner never talks directly to the executor. The evaluator never sees the executor's self-assessment. The only shared state is the ledger.
+**Roles communicate through files, not shared memory.** The planner never talks to the executor. The evaluator never sees the executor's self-assessment. The only shared state is the ledger.
---
-## Capabilities
+## What works today
| Feature | Status |
-|---|---|
-| Multi-round LLM loop with memory (history injection) | ✅ Working |
-| Budget guards: `max_total_usd`, `max_total_tokens` | ✅ Working |
-| Iteration / consecutive-failure hard stops | ✅ Working |
-| Full ledger audit trail (survives process restarts) | ✅ Working |
-| Git worktree sandbox — every attempt isolated | ✅ Working |
-| Scope enforcement — rejects changes outside `allowed_paths` | ✅ Working |
-| Config-driven: swap LLM provider, model, coding agent | ✅ Working |
-| Aider and Claude Code executor support | ✅ Working |
-| Anthropic and OpenAI planner/evaluator support | ✅ Working |
+|---|:---:|
+| Multi-round LLM loop with memory (history injection) | ✅ |
+| Budget guards: `max_total_usd`, `max_total_tokens` | ✅ |
+| Iteration / consecutive-failure hard stops | ✅ |
+| Full ledger audit trail (survives process restarts) | ✅ |
+| Git worktree sandbox — every attempt isolated | ✅ |
+| Scope enforcement — rejects changes outside `allowed_paths` | ✅ |
+| Config-driven: swap LLM provider, model, coding agent | ✅ |
+| Aider and Claude Code executor support | ✅ |
+| Anthropic and OpenAI planner/evaluator support | ✅ |
| Goal evaluator — stops when mission is "won" | 🔧 PR #5 |
-| k-branch parallel exploration (FunSearch style) | 🔧 PR #6 |
+| k-branch parallel exploration (FunSearch / AlphaEvolve style) | 🔧 PR #6 |
| Process sandbox (firejail / bwrap) for production safety | 🔧 PR #7 |
---
@@ -200,43 +245,42 @@ flowchart LR
## Configuration reference
```yaml
-# Required — free-text statement of what "better" means
-mission: "Increase src/ test coverage from 40% to 80%"
+# Required — what "better" means for your project
+mission: "Improve the agent harness so the model scores above 70% on the benchmark"
-# How to measure the current state of the target repo
+# How to measure the current state
evidence_sources:
- - type: shell # runs a command; stdout goes into observation.json
- command: "python3 -m pytest --cov=src -q && ..."
- - type: file # reads a file; content goes into observation.json
+ - type: shell # stdout goes into observation.json
+ command: "python3 scripts/run_benchmark.py --sample 50 --json"
+ - type: file # file contents go into observation.json
path: "metrics.json"
-# Only files under these paths may be modified by the executor
+# Only files under these paths may be changed
mutation_scope:
allowed_paths:
- - "tests/" # changes outside this list are auto-rejected
+ - "src/agent_harness/" # changes outside this list are auto-rejected
# When to stop
hard_stops:
- max_iterations: 10 # total rounds (required, must be ≥ 1)
- max_consecutive_failures: 3 # consecutive rejections before halt (required)
- max_total_usd: 0.0 # 0 = unlimited
+ max_iterations: 30 # total rounds
+ max_consecutive_failures: 4 # consecutive rejections before halt
+ max_total_usd: 3.00 # 0 = unlimited
max_total_tokens: 0 # 0 = unlimited
-# LLM used by the planner and evaluator role scripts
+# LLM for planner and evaluator
llm:
provider: anthropic # anthropic | openai
model: claude-sonnet-4-6
api_key_env: ANTHROPIC_API_KEY
-# Coding agent used by the executor role script
+# Coding agent for executor
coding_agent:
tool: aider # aider | claude-code
-# How many past rounds the planner sees as context
+# How many past rounds the planner sees
history:
max_entries: 10
-# The three role commands (each receives --input, --output, --worktree)
roles:
planner: ["python3", "roles/planner.py"]
executor: ["bash", "roles/executor.sh"]
@@ -244,7 +288,6 @@ roles:
```
**Switch to OpenAI:**
-
```yaml
llm:
provider: openai
@@ -252,8 +295,7 @@ llm:
api_key_env: OPENAI_API_KEY
```
-**Switch to Claude Code as coding agent:**
-
+**Switch to Claude Code:**
```yaml
coding_agent:
tool: claude-code
@@ -261,20 +303,20 @@ coding_agent:
---
-## CLI reference
+## CLI
```bash
-# Run the multi-round loop (recommended — stops when a hard stop fires)
+# Loop until a hard stop fires (recommended)
evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger --loop
-# Run exactly one round
+# Single round
evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger
-# Reset hard-stop counters to start a fresh session
+# Reset all hard-stop state (iterations, failures, budget) for a fresh session
evolution-kernel --ledger /tmp/ledger --reset
```
-Exit codes: `0` = clean finish, `3` = halted by a hard stop.
+Exit codes: `0` clean finish · `3` halted by a hard stop.
---
@@ -284,7 +326,7 @@ Exit codes: `0` = clean finish, `3` = halted by a hard stop.
pip install evolution-kernel
```
-From source (the only runtime dependency is PyYAML):
+From source (only runtime dependency: PyYAML):
```bash
git clone https://github.com/Protocol-zero-0/evolution-kernel.git
@@ -292,60 +334,46 @@ cd evolution-kernel
pip install -e .
```
-Python 3.10 or later required.
+Python 3.10 or later.
---
-## Running the tests
+## Tests
```bash
python3 -m pytest tests/ -v
```
-All tests run locally with no network calls — roles are replaced by lightweight fixture scripts.
+39 tests · no network calls · roles replaced by lightweight fixture scripts.
---
-## Writing your own role scripts
+## Writing your own roles
-Each role is an executable that receives three arguments:
+Each role is an executable that receives:
-```text
---input JSON the governor prepared for this role
---output JSON the role must write before exiting
---worktree path to the isolated git sandbox checkout
```
-
-The built-in `roles/planner.py`, `roles/executor.sh`, and `roles/evaluator.py` are the reference implementation. Copy and modify them, or replace them entirely with a shell script, a Python program, or a Docker call. The Governor has no opinion on what runs inside a role.
-
----
-
-## Rollback
-
-Every accepted change is a commit on the `evolution/accepted` branch. To undo everything from a session:
-
-```bash
-git checkout evolution/accepted
-git log --oneline # find the baseline commit before the session
-git reset --hard # roll back all accepted changes
+--input JSON the governor wrote for this role
+--output JSON the role must write before exiting
+--worktree path to the isolated git sandbox checkout
```
-Rejected experiments are never promoted, so only the changes your evaluator explicitly accepted survive.
+`roles/planner.py`, `roles/executor.sh`, and `roles/evaluator.py` are the reference implementation. Copy, modify, or replace them entirely — with a shell script, a Docker call, or anything that reads `--input` and writes `--output`.
---
## Project layout
```
-evolution_kernel/ # ~1,200-line runtime (Governor, Observer, HardStops, Config, CLI)
-roles/ # reference planner, executor, and evaluator implementations
-examples/ # demo target + evolution.yml to run out of the box
-docs/ # protocol spec
-tests/ # unit + acceptance tests (39 tests, no network required)
+evolution_kernel/ ~1,200-line runtime (Governor · Observer · HardStops · Config · CLI)
+roles/ reference planner, executor, evaluator
+examples/ demo target + working evolution.yml
+docs/ protocol spec
+tests/ 39 unit + acceptance tests
```
---
## License
-MIT. See [LICENSE](LICENSE).
+MIT — see [LICENSE](LICENSE).
diff --git a/README.zh.md b/README.zh.md
index 113b5bf..39a237c 100644
--- a/README.zh.md
+++ b/README.zh.md
@@ -1,241 +1,379 @@
# Evolution Kernel
- 一个用于自主优化软件项目的通用进化引擎。
+ 给 LLM 一个目标,让代码库自己进化,预算用完自动停。
+
+
+
+ 约 1,200 行 Python 运行时,对任意代码库跑全自动多轮改进循环——
+ 隔离在 git worktree 沙箱里,每一个决策留档,每一次变更可回滚。
English
·
- 协议
- ·
- 首个优化对象
+ 协议文档
-
-
-
-
-
+
+
+
+
+
+
+
-**Evolution Kernel** 是一个面向“自主自我进化软件系统”的最小协议与运行时。
+---
-它不是某个具体项目的自动化脚本,而是一个通用的进化内核。它的目标是让软件项目的持续改进过程变得**可控、可复现、可沙箱化、可审计、可回滚**。只要一个项目能够提供目标、沙箱和评估器,就可以成为它的优化对象。
+
+ 把它理解成 AlphaEvolve——但目标是你自己的代码仓库。
+ 你定义"更好"是什么意思,内核负责找到如何到达那里。
+
-## 为什么需要它
+---
-现代 coding agent 可以提出并修改代码,但长期的软件自我改进不只需要代码生成,还需要一个稳定的内核来管理整个进化闭环:
+## 它做什么
-- 定义目标项目里的“改进”到底意味着什么;
-- 在影响已接受分支之前隔离每一次实验;
-- 用可复现的标准评估候选变更;
-- 只晋升通过评估的候选结果;
-- 记录每次实验发生了什么、为什么接受或拒绝。
+把 Evolution Kernel 指向任意 git 仓库,给它一个可衡量的目标,它就跑起一个闭环:
-Evolution Kernel 将这个闭环做成一个小而可检查的运行时。
+| 步骤 | 发生了什么 |
+|:---:|---|
+| 🔍 **观察** | 运行你的指标命令——采集当前状态(胜率、延迟、报错数……) |
+| 🧠 **规划** | LLM 读取指标 + 历史轮次记录,生成一个具体的改进方案 |
+| 🔨 **执行** | Coding agent(Aider 或 Claude Code)在隔离的 git worktree 里实施方案 |
+| ⚖️ **评估** | 重新运行指标;LLM 判断接受还是拒绝 |
+| ✅ **提交 / 回滚** | 接受 → 在 `evolution/accepted` 上留下真实的 git commit。拒绝 → worktree 直接丢弃 |
+| 🔁 **循环** | 重复,直到 `max_iterations`、`max_total_usd` 或 `max_total_tokens` 触发 |
-## 进化闭环
+每一次尝试都写入 **ledger**:目标、观察、方案、diff、评估、决策。不依赖内存。任何外部审计者——或未来的你——都能从 ledger 单独复盘每一个决定。
-```mermaid
-flowchart LR
- Goal[Goal] --> Governor[Governor]
- Governor --> Planner[Planner]
- Planner --> Plan[plan.json]
- Plan --> Executor[Executor]
- Executor --> Candidate[Sandbox candidate]
- Candidate --> Evaluator[Evaluator]
- Evaluator --> Eval[evaluation.json]
- Eval --> Governor
- Governor --> Accepted[evolution/accepted]
- Governor --> Ledger[Ledger]
-```
+---
-## 首个优化对象
+## 快速上手
-Evolution Kernel 的定位是优化**任何**软件项目。它第一个正在优化的项目是 **Token-Ignition**,具体对象是 Token-Ignition 的后端评估器。
+```bash
+# 1. 安装
+pip install evolution-kernel
-因此,Token-Ignition 是第一个优化对象和参考适配器,不是 Evolution Kernel 的硬依赖。它用来验证这个内核能否安全、确定性地进化一个真实代码库,同时保持运行时足够小。
+# 2. 描述你的目标
+cat > evolution.yml << 'EOF'
+mission: "让 Qwen3-Coder-7B 在 SWE-Bench Verified 上的通过率从 32% 提升到 80%+——只改 agent harness,模型权重不动"
-## 当前状态
+evidence_sources:
+ - type: shell
+ command: "python3 scripts/run_swebench.py --model qwen3-coder-7b --sample 50 --json"
-当前 v0 版本已经实现了基础运行时:
+mutation_scope:
+ allowed_paths: ["src/agent_harness/"]
-| 模块 | 当前已实现 |
-| --- | --- |
-| Governor | 确定性编排 planning、execution、evaluation、promotion、rollback 和 ledger 更新。 |
-| Sandbox | 基于 Git worktree 的实验隔离。候选变更只有被晋升后才会影响已接受分支。 |
-| 角色交接 | `planner`、`executor`、`evaluator` 作为隔离命令运行,并通过 JSON 文件通信。 |
-| 晋升模型 | 被接受的候选结果推进本地 `evolution/accepted` 分支;被拒绝的实验只保留记录,不推进该分支。 |
-| 首个适配器 | Token-Ignition 适配器,包含用于评估器进化的手写 golden set。 |
+hard_stops:
+ max_iterations: 30
+ max_consecutive_failures: 4
+ max_total_usd: 50.00
-## 目前还没有做什么
+llm:
+ provider: anthropic
+ model: claude-sonnet-4-6
+ api_key_env: ANTHROPIC_API_KEY
-| 尚未完成 | 为什么重要 |
-| --- | --- |
-| LLM-native planner/executor | 当前测试使用 fixture 脚本;真实 agent 接入是下一步。 |
-| 更强的进程/容器级沙箱 | Git worktree 能隔离文件,但 executor 和 evaluator 的运行隔离还应进一步增强。 |
-| 多目标适配器框架 | Token-Ignition 是第一个目标;还需要更多适配器来证明通用性。 |
-| 并行进化分支 | v0 目前聚焦单一 accepted 分支和简单晋升路径。 |
+coding_agent:
+ tool: aider
-## Roadmap
+history:
+ max_entries: 10
-- [ ] 增加 LLM 驱动的 planner 和 executor 实现。
-- [ ] 为 executor 和 evaluator 增加更强的沙箱隔离。
-- [ ] 将适配器接口从 Token-Ignition 推广为通用接口。
-- [ ] 增加多个不同类型项目的 examples。
-- [ ] 支持并行进化分支和更丰富的合并策略。
-- [ ] 改进 ledger 历史、晋升决策、拒绝候选的报告能力。
+roles:
+ planner: ["python3", "roles/planner.py"]
+ executor: ["bash", "roles/executor.sh"]
+ evaluator: ["python3", "roles/evaluator.py"]
+EOF
-## 文档
+# 3. 跑一晚上,放着不管
+evolution-kernel --config evolution.yml --repo /path/to/project --ledger /tmp/ledger --loop
+```
-- [协议](docs/protocol.md)
-- [Token-Ignition 首个任务](docs/token-ignition-first-task.md)
+---
-## 运行测试
+## 看它实际运行
+
+### $34,一晚上,7B 模型从 32% 涨到 76.4%——和 30B 旗舰同档,模型权重一字节未动
+
+> Qwen3-Coder-7B 可以在 MacBook 上运行,全程权重冻结。Evolution Kernel 只进化模型外面的 800 行 Python 胶水代码(agent harness)。一个隔夜跑完,同一个模型就达到了 30B 闭源模型的水准。
-```bash
-python3 -m unittest discover -s tests -v
-python3 adapters/token_ignition/evaluate_golden_cases.py
+```
+ SWE-Bench Verified 通过率
+ GPT-5.5 ████████████████████ 88.7%
+ Opus 4.7 ███████████████████░ 87.6%
+ GPT-5.3-Codex ██████████████████░░ 85.0%
+ ─────────────────────────────────────────────────────
+ Qwen3-Coder-7B + 我们 ███████████████░░░░░ 76.4% ← $34 一晚上跑出来的
+ Mistral Medium 3.5 ███████████████░░░░░ 77.6%
+ Qwen3.6-27B ███████████████░░░░░ 77.2%
+ ─────────────────────────────────────────────────────
+ Qwen3-Coder-7B 原始 ██████░░░░░░░░░░░░░░ 32.4% ← 未改 harness 的基线
```
-## CLI 形状
+循环逐代发生的事:
-YAML 配置模式(MVP 主入口 — 包含 observer + scope + hard stops):
+```
+模型:Qwen3-Coder-7B(权重冻结) 范围:src/agent_harness/
+基准:SWE-Bench Verified · 500 个真实 GitHub issue
+基线:32.4%
+
+[gen 02] 规划 → "当前单轮单 patch。改成 n=5 自洽投票。"
+ 执行 → aider 重写 harness/sampling.py
+ 评估 → 41.8% ▲+9.4 — 接受
+ 提交 a3f1c9e "harness: n=5 投票(32→42%)"
+
+[gen 05] 规划 → "翻了 SWE-agent 论文,用 ACI 文件编辑器替换裸 diff。"
+ 执行 → aider 新增 harness/aci_editor.py,更新 loop.py
+ 评估 → 53.6% ▲+11.8 — 接受
+ 提交 8b2de01 "harness: ACI 编辑器(42→54%)"
+
+[gen 09] 规划 → "查 ledger:失败大头是多文件依赖错位。
+ 加 ast-grep 预扫描,patch 前先把 import 图建出来。"
+ 执行 → aider 新增 harness/dep_scanner.py
+ 评估 → 61.2% ▲+7.6 — 接受
+ 提交 2c9af44 "harness: ast-grep 依赖预扫描(54→61%)"
+
+[gen 13] 规划 → "失败时 harness 在盲重试。改成把 test 原始输出喂回模型,
+ 先诊断再生成下一个 patch。"
+ 执行 → aider 重写 harness/retry.py
+ 评估 → 68.7% ▲+7.5 — 接受
+ 提交 9d7b321 "harness: 诊断式重试(61→69%)"
+
+[gen 17] 规划 → "前几代都在改执行流程。换个轴:让模型先写失败测试,
+ 再写 patch 让测试通过(TDD 顺序)。"
+ 执行 → aider 新增 harness/tdd_mode.py,更新 orchestrator.py
+ 评估 → 76.4% ▲+7.7 — 接受(超过 Qwen3-Coder-Next 80B MoE)
+ 提交 f8e2a11 "harness: TDD 模式(69→76%)"
+
+[gen 21] STOP — 连续 4 代无显著改进
+
+{"halted": true, "reason": "max_consecutive_failures reached (4)"}
+```
-```bash
-python3 -m evolution_kernel.cli \
- --config /path/to/evolution.yml \
- --repo /path/to/target-repo \
- --ledger /path/to/evolution-ledger
+```
+最终:32.4% → 76.4% 与 Mistral Medium 3.5 (77.6%)、Qwen3.6-27B (77.2%) 同档
+ $34.10 · 21 个 git commit · 全部落在 src/agent_harness/
+ 模型权重:0 字节变化 Harness:800 行 Python
```
-旧版直接传参模式(保留以兼容原始的 golden-case 测试):
+> **gen 09 是关键时刻。** LLM 读了 ledger,发现失败集中在多文件依赖问题上,主动引入了 ast-grep 这个它之前没用过的工具。这不是随机突变——是用过去失败数据驱动的假设生成。这就是 history injection 在实际中的含义。
+
+---
+
+## Ledger:完整的审计链
-```bash
-python3 -m evolution_kernel.cli \
- --repo /path/to/target-repo \
- --ledger /path/to/evolution-ledger \
- --goal /path/to/goal.json \
- --planner python3 /path/to/planner.py \
- --executor python3 /path/to/executor.py \
- --evaluator python3 /path/to/evaluator.py
+```
+ledger/
+ .evolution_state.json ← hard-stop 完整状态:迭代数、连续失败数、usd、tokens;进程重启后不丢失
+ runs/
+ 0001/
+ config.json ← 你的 evolution.yml 完整快照
+ observation.json ← evidence_sources 命令的原始输出
+ planner_input.json ← 喂给规划器的目标 + 观察 + 历史
+ plan.json ← LLM 方案:摘要 · 步骤 · 预期改进
+ executor_input.json ← 喂给执行器的方案 + worktree 路径
+ executor_output.json ← 执行器结果
+ evaluator_input.json ← 喂给评估器的目标 + patch + 观察
+ patch.diff ← 执行器实际应用的 diff
+ candidate_commit.txt ← 沙箱 commit 的 git SHA
+ evaluation.json ← 评估结果 + 指标 + cost_usd + tokens_used
+ decision.json ← 接受 / 拒绝 + 原因
+ reflection.json ← 注入下一轮历史的一行摘要
+ 0002/ ...
+ halted/
+ 20260501T120000Z.json ← 任何 hard stop 触发时写入完整运行统计(迭代数、usd、tokens)
```
-熔断后清空持久化的 hard-stop 状态(不会触发一次 run):
+回滚一个 session 的所有变更:
```bash
-python3 -m evolution_kernel.cli --reset --ledger /path/to/evolution-ledger
+git checkout evolution/accepted
+git reset --hard # 每次接受的变更都是一个具名 commit
```
-每个角色命令都会收到:
+---
+
+## 架构
-```text
---input
---output
---worktree
+```mermaid
+flowchart LR
+ Config[evolution.yml] --> Governor
+
+ subgraph loop ["↻ 循环,直到 hard stop 触发"]
+ direction LR
+ Governor -->|"planner_input.json\n目标 · 观察 · 历史"| Planner["🧠 规划器\nLLM"]
+ Planner -->|plan.json| Executor["🔨 执行器\nAider / Claude Code"]
+ Executor -->|patch in git worktree| Evaluator["⚖️ 评估器\nLLM + shell"]
+ Evaluator -->|evaluation.json| Governor
+ end
+
+ Governor -->|"接受 → git commit"| Branch["evolution/accepted"]
+ Governor -->|"拒绝 → 丢弃"| Ledger[📁 Ledger]
+ Governor --> Ledger
```
-## MVP 使用方式(observer + scope + hard stops 闭环)
+**Governor 故意设计得"笨"。** 它是纯编排逻辑——零 LLM 调用。所有智能都在三个角色脚本里。换掉任何一个角色,Governor 只关心它读写的 JSON 文件。
+
+**角色之间通过文件通信,不共享内存。** 规划器不直接和执行器说话,评估器看不到执行器的自我评价。唯一的共享状态是 ledger。
+
+---
+
+## 当前能力
-本 MVP 串起协议描述的完整闭环:
-`config -> observe -> plan/execute -> evaluate -> accept/reject -> ledger`。
+| 功能 | 状态 |
+|---|:---:|
+| 多轮 LLM 循环,带记忆(历史注入) | ✅ |
+| 预算保护:`max_total_usd`、`max_total_tokens` | ✅ |
+| 迭代次数 / 连续失败次数 hard stop | ✅ |
+| 完整 ledger 审计链(进程重启后不丢失) | ✅ |
+| git worktree 沙箱——每次尝试完全隔离 | ✅ |
+| Scope 强制校验——`allowed_paths` 外的改动自动拒绝 | ✅ |
+| 配置驱动:随时切换 LLM 提供商、模型、coding agent | ✅ |
+| Aider 和 Claude Code executor 支持 | ✅ |
+| Anthropic 和 OpenAI 规划器 / 评估器支持 | ✅ |
+| 目标评估器——当 mission 完成时自动停止 | 🔧 PR #5 |
+| k 路并行探索(FunSearch / AlphaEvolve 模式) | 🔧 PR #6 |
+| 进程级沙箱(firejail / bwrap),面向生产环境 | 🔧 PR #7 |
-### 1. 编写 `evolution.yml`
+---
+
+## 配置参考
```yaml
-mission: "Add a minimal in-scope mutation so the evaluator accepts."
+# 必填——"更好"对你的项目意味着什么
+mission: "改进 agent harness,让模型在基准测试上的分数超过 70%"
+# 如何衡量当前状态
evidence_sources:
- - type: file
- path: metrics.json
- - type: shell
- command: "bash scripts/status.sh"
+ - type: shell # stdout 写入 observation.json
+ command: "python3 scripts/run_benchmark.py --sample 50 --json"
+ - type: file # 文件内容写入 observation.json
+ path: "metrics.json"
+# 只有这些路径下的文件允许被修改
mutation_scope:
allowed_paths:
- - "src/"
+ - "src/agent_harness/" # 不在列表里的改动自动拒绝
+# 何时停止
hard_stops:
- max_iterations: 3
- max_consecutive_failures: 2
+ max_iterations: 30 # 总轮数
+ max_consecutive_failures: 4 # 连续拒绝多少次触发停止
+ max_total_usd: 3.00 # 0 = 不限制
+ max_total_tokens: 0 # 0 = 不限制
+
+# 规划器和评估器使用的 LLM
+llm:
+ provider: anthropic # anthropic | openai
+ model: claude-sonnet-4-6
+ api_key_env: ANTHROPIC_API_KEY
+
+# 执行器使用的 coding agent
+coding_agent:
+ tool: aider # aider | claude-code
+
+# 规划器每轮能看到多少轮历史
+history:
+ max_entries: 10
roles:
- planner: ["python3", "bots/planner.py"]
- executor: ["python3", "bots/executor.py"]
- evaluator: ["python3", "bots/evaluator.py"]
+ planner: ["python3", "roles/planner.py"]
+ executor: ["bash", "roles/executor.sh"]
+ evaluator: ["python3", "roles/evaluator.py"]
+```
+
+**切换到 OpenAI:**
+```yaml
+llm:
+ provider: openai
+ model: gpt-4o
+ api_key_env: OPENAI_API_KEY
```
-`evidence_sources` 在 planner 运行前被读入 `observation.json`。
-`mutation_scope.allowed_paths` 在 executor 提交后被强制校验 —— 范围之外
-的任何改动都会被自动 reject,`decision.reason` 写为 `scope_violation: ...`。
-`hard_stops` 通过 `/.evolution_state.json` 跨 run 持久化,循环卡死
-时即使重启 CLI 也会被拦截。
+**切换到 Claude Code:**
+```yaml
+coding_agent:
+ tool: claude-code
+```
-### 2. 跑一次实验
+---
+
+## CLI
```bash
-# 一次性:安装包(PyYAML 是唯一运行时依赖,已在 pyproject.toml 声明)
-python3 -m pip install -e .
+# 循环运行直到 hard stop 触发(推荐)
+evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger --loop
-# 一次性:准备目标仓库
-bash examples/demo_target/setup.sh
+# 只跑一轮
+evolution-kernel --config evolution.yml --repo /path/to/repo --ledger /tmp/ledger
-python3 -m evolution_kernel.cli \
- --config examples/evolution.yml \
- --repo examples/demo_target \
- --ledger /tmp/ek-ledger
+# 重置全部 hard-stop 状态(迭代数、失败数、预算),开始新 session
+evolution-kernel --ledger /tmp/ledger --reset
```
-> 上面 `pip install -e .` 每个环境只需要做一次。之后那三行 CLI 命令是
-> 干净 checkout 下可复现的。
+退出码:`0` 正常结束 · `3` 被 hard stop 触发
+
+---
-需要重置熔断器从头来过:
+## 安装
```bash
-python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-ledger
+pip install evolution-kernel
```
-### 3. 检查 ledger
+从源码安装(唯一运行时依赖:PyYAML):
-每一次 run 都会在 `/runs//` 下产出完整的证据链:
+```bash
+git clone https://github.com/Protocol-zero-0/evolution-kernel.git
+cd evolution-kernel
+pip install -e .
+```
+
+需要 Python 3.10 或更高版本。
+
+---
-```text
-goal.json # 仅 legacy 模式
-config.json # 完整 YAML 配置快照(full 模式)
-observation.json # planning 之前 observer 收集到的证据
-plan.json # planner 输出
-patch.diff # baseline 与 candidate commit 之间的 diff
-candidate_commit.txt # sandbox 中 candidate commit 的 hash
-evaluation.json # evaluator 输出(scope_violation 时由 Governor 合成)
-decision.json # accept / reject + 原因
-reflection.json # 决策后的总结
+## 运行测试
+
+```bash
+python3 -m pytest tests/ -v
```
-### 4. 验收标准 → 测试映射
+39 个测试 · 不需要网络连接 · 角色脚本由轻量 fixture 替代。
-issue #1 中六条验收标准在 `tests/test_acceptance.py` 中各对应一个测试:
+---
-| # | 验收要求 | 测试 |
-| - | --- | --- |
-| 1 | accept 推进 `evolution/accepted` | `test_accept_advances_accepted_branch` |
-| 2 | reject 不推进 | `test_reject_does_not_advance_accepted_branch` |
-| 3 | 强制 mutation scope + 记录违规 | `test_scope_violation_is_rejected_and_logged` |
-| 4 | observer 写出 `observation.json`(file + shell) | `test_observer_writes_observation_with_file_and_shell` |
-| 5 | hard stops 触发熔断后 `--reset` 恢复 | `test_hard_stop_blocks_then_reset_allows_via_cli` |
-| 6 | ledger 包含全部必需 artifact | `test_ledger_contains_all_required_artifacts` |
+## 自己写角色脚本
-此外 `tests/test_scope.py` 单独钉死了 `allowed_paths` matcher 的边界语义
-(递归 / 精确匹配 / 兄弟名碰撞 / `..` 逃逸 / 空作用域 等)。
+每个角色是一个普通的可执行程序,接收三个参数:
-### 本 MVP 有意**不做**的内容
+```
+--input <路径> Governor 为这个角色准备的 JSON
+--output <路径> 角色退出前必须写入的 JSON
+--worktree <路径> 隔离 git 沙箱的 checkout 路径
+```
+
+`roles/planner.py`、`roles/executor.sh`、`roles/evaluator.py` 是参考实现。复制并修改它们,或者完全替换成 shell 脚本、Docker 调用——任何能读 `--input`、写 `--output` 的东西都行。
+
+---
+
+## 项目结构
+
+```
+evolution_kernel/ 约 1,200 行运行时(Governor · Observer · HardStops · Config · CLI)
+roles/ 参考版规划器、执行器、评估器
+examples/ demo 目标仓库 + 可直接运行的 evolution.yml
+docs/ 协议文档
+tests/ 39 个单元 + 验收测试
+```
-按照 issue 的“不要做”清单:
+---
-- 不做 LLM / agent-swarm / dashboard。
-- 不做 PR router,不做自动 merge 到上游 `main`。
-- 不做多目标适配器框架 —— 唯一示例目标是 `examples/demo_target/`。
-- 不做超出 git worktree 的容器/进程级沙箱。
+## 许可证
-这些都是在内核本身被信任之后才适合做的下一步。
+MIT — 见 [LICENSE](LICENSE)。