diff --git a/README.md b/README.md index f3e65e67..8aaa4a1e 100644 --- a/README.md +++ b/README.md @@ -12,7 +12,7 @@ [![Python](https://img.shields.io/badge/python-3.10+-blue)](https://python.org) [![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE) [![Tests](https://github.com/he-yufeng/CoreCoder/actions/workflows/ci.yml/badge.svg)](https://github.com/he-yufeng/CoreCoder/actions) -[![engine](https://img.shields.io/badge/engine-1308_LoC-blue)](article/00-index_EN.md) +[![engine](https://img.shields.io/badge/engine-1324_LoC-blue)](article/00-index_EN.md) [![essays](https://img.shields.io/badge/source--reading-8_bilingual-orange)](article/00-index_EN.md) @@ -25,7 +25,7 @@ | | CoreCoder | Claude Code | aider | nanoGPT | |---|---|---|---|---| -| Lines of code | ~1,308 engine / 2,594 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) | +| Lines of code | ~1,324 engine / 2,621 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) | | Time to read it all | one afternoon | can't (closed) | a few days of slogging | one afternoon | | Breakpoint, change, rerun? | yes, every line | no | yes, but there's a lot | yes | | What it's for | understand one, then fork your own | production coding assistant | terminal pair-programming | minimal GPT for teaching | @@ -36,9 +36,9 @@ The nanoGPT column is there as a reference point: minimal, readable, but it teac I've always felt coding agents get talked about as if they were arcane. Strip a tool like Claude Code or Cursor all the way down and the core is a `while` loop wrapped around a large model, plus seven or eight tools that let it actually do things. The hard part was never the loop; it's everything the loop has to cope with once it meets the real world. CoreCoder is the minimal version that writes that core out honestly. -The engine (loop, model interface, context, tools, sessions) is 1,308 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 24 files: 2,594 physical lines, 2,089 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below. +The engine (loop, model interface, context, tools, sessions) is 1,324 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 24 files: 2,621 physical lines, 2,113 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below. -And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 171 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work. +And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 178 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work. The code came out of a public teardown: open analyses have already exposed a lot of the load-bearing architecture inside production agents like Claude Code. I took the most essential layer and rewrote it honestly, in as little code as I could. So reading CoreCoder is roughly like reading a runnable, annotated take on how that kind of agent works, except it's only a minimal reimplementation, sitting right there on your machine for you to take apart and change. diff --git a/README_CN.md b/README_CN.md index 74c14667..c355a2d6 100644 --- a/README_CN.md +++ b/README_CN.md @@ -12,7 +12,7 @@ [![Python](https://img.shields.io/badge/python-3.10+-blue)](https://python.org) [![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE) [![Tests](https://github.com/he-yufeng/CoreCoder/actions/workflows/ci.yml/badge.svg)](https://github.com/he-yufeng/CoreCoder/actions) -[![engine](https://img.shields.io/badge/engine-1308_LoC-blue)](article/) +[![engine](https://img.shields.io/badge/engine-1324_LoC-blue)](article/) [![源码导读](https://img.shields.io/badge/源码导读-8篇双语-orange)](article/) @@ -25,7 +25,7 @@ | | CoreCoder | Claude Code | aider | nanoGPT | |---|---|---|---|---| -| 代码量 | 引擎约 1308 行 / 整包 2594 行 | 几十万行(闭源) | 数万行 Python | 约 600 行(两个文件) | +| 代码量 | 引擎约 1324 行 / 整包 2621 行 | 几十万行(闭源) | 数万行 Python | 约 600 行(两个文件) | | 读完要多久 | 一个下午 | 读不了(闭源) | 得啃几天 | 一个下午 | | 能不能下断点改了再跑 | 能,每一行 | 不能 | 能,但量大 | 能 | | 定位 | 读懂并 fork 出你自己的 agent | 生产级编程助手 | 终端结对编程 | 教学用最小 GPT | @@ -36,9 +36,9 @@ nanoGPT 那一列是拿来对照的:它最小、可读,但教的是训一个 我一直觉得 coding agent 被讲得太玄了。把 Claude Code、Cursor 这类工具扒到底,核心是一个 while 循环套着一个大模型,外加七八个让它能真正动手的工具。难的从来不是这个循环,而是循环跑进真实世界以后要兜的那些底。CoreCoder 就是把这个核心老老实实写出来的最小版本。 -引擎部分(循环、模型接口、上下文、工具、会话)去掉空行和注释是 1308 行。连最外层的 CLI、配置、打包一起算,整个包 24 个文件、物理 2594 行、净 2089 行,每个文件都短到能一口气读完。自 1161 行快照之后的增长都花在了看得见的功能上:plan mode、hooks、checkpoints,下文各有交代。 +引擎部分(循环、模型接口、上下文、工具、会话)去掉空行和注释是 1324 行。连最外层的 CLI、配置、打包一起算,整个包 24 个文件、物理 2621 行、净 2113 行,每个文件都短到能一口气读完。自 1161 行快照之后的增长都花在了看得见的功能上:plan mode、hooks、checkpoints,下文各有交代。 -它真能跑:读写文件、执行 shell、派子 agent、分三层压上下文,还能随时把这趟烧掉的 token 和美元数报给你。任何要动你磁盘、要跑命令的调用,都会先停下来等你点头,171 个测试是绿的。但能跑不是为了劝你拿去日用,而是为了让这份「注释」不撒谎:一个解释 agent 怎么运作的范例,自己得真能运作。 +它真能跑:读写文件、执行 shell、派子 agent、分三层压上下文,还能随时把这趟烧掉的 token 和美元数报给你。任何要动你磁盘、要跑命令的调用,都会先停下来等你点头,178 个测试是绿的。但能跑不是为了劝你拿去日用,而是为了让这份「注释」不撒谎:一个解释 agent 怎么运作的范例,自己得真能运作。 代码来自一次公开拆解。公开的源码分析里,Claude Code 这类生产级 agent 暴露出不少关键架构,我挑出最核心的一层,用尽量少的代码诚实地复写了一遍。所以读 CoreCoder,约等于读一份基于公开源码分析的「可运行注释版」:讲的是这类 agent 的核心思路,而它本身只是最小复写,就摆在你机器上,随你拆、随你改。 diff --git a/corecoder/hooks.py b/corecoder/hooks.py index 6b75e591..a1dde40e 100644 --- a/corecoder/hooks.py +++ b/corecoder/hooks.py @@ -13,9 +13,12 @@ import json import logging +import re import subprocess from pathlib import Path +from .tools.bash import _find_bash + log = logging.getLogger(__name__) HOOKS_FILE = Path.home() / ".corecoder" / "hooks.json" @@ -71,9 +74,15 @@ def _fire(hook: dict, payload: dict): matcher = hook.get("matcher", "") if matcher not in ("", "*", payload["tool_name"]): return None + bash_exe = _find_bash() + cmd_str = hook["command"] + if bash_exe and re.match(r"^[a-zA-Z]:\\", cmd_str): + parts = cmd_str.split(None, 1) + cmd_str = f'"{parts[0]}" {parts[1]}' if len(parts) > 1 else f'"{cmd_str}"' + cmd: list[str] | str = [bash_exe, "-c", cmd_str] if bash_exe else cmd_str try: proc = subprocess.run( - hook["command"], shell=True, check=False, input=json.dumps(payload), + cmd, shell=not bash_exe, check=False, input=json.dumps(payload), capture_output=True, text=True, encoding="utf-8", errors="replace", timeout=TIMEOUT, ) except (subprocess.TimeoutExpired, OSError) as e: diff --git a/corecoder/tools/bash.py b/corecoder/tools/bash.py index c1bc8fed..d6e970aa 100644 --- a/corecoder/tools/bash.py +++ b/corecoder/tools/bash.py @@ -7,14 +7,30 @@ - Working directory tracking (cd awareness) """ +import functools import os import re +import shutil import subprocess import threading from typing import ClassVar from .base import Tool + +@functools.lru_cache(maxsize=1) +def _find_bash() -> str | None: + """Locate Git Bash on Windows so POSIX commands and shell hooks run.""" + if os.name != "nt": + return None + for p in (r"C:\Program Files\Git\bin\bash.exe", r"C:\Program Files\Git\usr\bin\bash.exe"): + if os.path.isfile(p): + return p + git = shutil.which("git") + if git and os.path.isfile(b := os.path.join(os.path.dirname(os.path.dirname(git)), "bin", "bash.exe")): + return b + return b if (b := shutil.which("bash")) and "system32" not in b.lower() else None + # Track cwd across commands (Claude Code does this too). Thread-local, so that # when the agent executes tools in parallel two bash calls never race on one # shared global: each worker thread carries its own cwd. See article 05. @@ -82,10 +98,12 @@ def execute(self, command: str, timeout: int = 120) -> str: # use this thread's own tracked working directory cwd = get_tracked_cwd() or os.getcwd() + bash_exe = _find_bash() + cmd: list[str] | str = [bash_exe, "-c", command] if bash_exe else command try: proc = subprocess.run( - command, - shell=True, + cmd, + shell=not bash_exe, check=False, capture_output=True, text=True, diff --git a/tests/test_core.py b/tests/test_core.py index 1db41d8d..496c6fe1 100644 --- a/tests/test_core.py +++ b/tests/test_core.py @@ -1,5 +1,6 @@ """Tests for core modules: config, context, session, imports.""" +import os import re from pathlib import Path from typing import ClassVar @@ -244,6 +245,10 @@ def norm_dir(s: str) -> str: # pwd prints the shell's own form: git-bash gives /c/Users/... where # Python gives C:\Users\...; compare both in one canonical shape s = s.strip().replace("\\", "/").rstrip("/").lower() + if os.name == "nt" and s.startswith("/tmp"): + import tempfile + + s = tempfile.gettempdir().replace("\\", "/").rstrip("/").lower() + s[4:] if len(s) >= 3 and s[0] == "/" and s[2] == "/" and s[1].isalpha(): s = s[1] + ":/" + s[3:] return s @@ -267,7 +272,7 @@ def __init__(self, i, cmd): # a cd in the batch lands on the session afterwards; siblings in the # same batch still start from the pre-batch cwd (parallel, not serial) - results = agent._exec_tools_parallel([_TC(3, f"cd {target}"), _TC(4, "pwd")]) + results = agent._exec_tools_parallel([_TC(3, f"cd {target.as_posix()}"), _TC(4, "pwd")]) assert get_tracked_cwd() == str(target) assert norm_dir(str(tmp_path)) in norm_dir(results[1]) finally: @@ -714,7 +719,6 @@ def test_retry_exhausts_and_raises(self): ] from openai import APIConnectionError - with mock.patch("corecoder.llm.time.sleep"): - with pytest.raises(APIConnectionError): - llm.chat(messages=[{"role": "user", "content": "hi"}]) + with mock.patch("corecoder.llm.time.sleep"), pytest.raises(APIConnectionError): + llm.chat(messages=[{"role": "user", "content": "hi"}]) assert create.call_count == 3 diff --git a/tests/test_hooks.py b/tests/test_hooks.py index 6c517779..fc2420c7 100644 --- a/tests/test_hooks.py +++ b/tests/test_hooks.py @@ -34,7 +34,7 @@ def test_pre_hook_blocks_with_reason_and_the_tool_never_runs(tmp_path): asked = [] agent = _agent( tmp_path, - Hooks(pre=[{"matcher": "*", "command": f"sh {blocker}"}], post=[]), + Hooks(pre=[{"matcher": "*", "command": f"sh {blocker.as_posix()}"}], post=[]), permission=Permission(ask=lambda n, a: asked.append(n) or "once"), ) @@ -55,7 +55,7 @@ def test_pre_hook_passing_lets_the_call_through(tmp_path): def test_post_hook_observes_the_finished_call(tmp_path): marker = tmp_path / "seen.jsonl" - agent = _agent(tmp_path, Hooks(pre=[], post=[{"matcher": "*", "command": f"cat >> {marker}"}])) + agent = _agent(tmp_path, Hooks(pre=[], post=[{"matcher": "*", "command": f"cat >> {marker.as_posix()}"}])) assert agent.chat("go") == "done" seen = marker.read_text() @@ -117,8 +117,8 @@ def test_hooks_gate_each_call_of_a_parallel_batch(tmp_path): llm=ScriptedLLM([LLMResponse(tool_calls=calls), LLMResponse(content="done")]), tools=[WriteFileTool()], hooks=Hooks( - pre=[{"matcher": "*", "command": f"cat >> {pre_log}"}], - post=[{"matcher": "*", "command": f"cat >> {post_log}"}], + pre=[{"matcher": "*", "command": f"cat >> {pre_log.as_posix()}"}], + post=[{"matcher": "*", "command": f"cat >> {post_log.as_posix()}"}], ), )