Nicholas Clooney

AgentOS:越用越聪明的 agent 工作环境

所属系列

声明:和我最近在做的许多东西一样,这篇博客也是人类主导,agent 执笔。方向、想法、品味,都来自我这个有血有肉的人。我也反复校对和编辑过。但大部分苦力活由 agent 完成。现在你知道了,祝阅读愉快。😀

昨天和我一起工作的那个 agent,已经不存在了。💀

没那么戏剧化,模型还在。但那个实例、当时的工作上下文,以及它对“为什么重命名这个文件”或“为什么这个抽象看起来不对”形成了一半的理解,全都消失了。明天的 agent 从零开始,下周的也是。

把真正的工作交给 AI agent 后,我意识到,瓶颈会从模型能力,转向项目解释自身的能力。如果仓库不能告诉新来的 agent 我们在做什么、为什么做,每次会话都得从同样的考古工作开始。我之前写过的机械操作负担,也就是那些悄悄累积、消磨行动势头的成本,在这里以一种新形式出现:重建上下文。

所以,在 ProjectSpire 中,我开始把周边文档当作工程环境本身的一部分。这个项目是一个由 AI 辅助开发的多语言 monorepo,用于探索 Slay the Spire 2 的数据、工具、模组和配套应用实验。文档不只是留待以后看的笔记,也不只是更新记录,而是一套真正的系统,目标是让项目随着开发推进,越来越容易被继续开发。

这就是本文想表达的观点:值得关注的变化,不是更快地生成代码,而是构建能积累上下文的环境。即使 agent 已经不记得上周二,合作本身也能随着时间不断改善。

三个层次,再向上一层

我以前写过创造性工作的三层模型:思考、技艺、机械操作。思考是判断和意图,技艺是品味和标准,机械操作则是样板工作的负担,是从想法到结果之间的一次次按键。

同样的三个层次,也适用于 AI agent 周围的工程环境,只是向上移了一层:

  • 思考保存在计划中。我们要做什么,为什么采用这种形态,而不是另一种?
  • 技艺保存在协作日志中。我希望怎么做?我纠正过什么,哪些不该再重复纠正?
  • 机械操作保存在技能和工作流文件中。确切命令是什么,顺序是什么,可能怎样失败?

多数“AI 工作流”文章会把它们混在一起,通常放进一条巨大提示,或一份试图包办一切的 CLAUDE.md 中。项目还小时,这样能用;大到提示承受不住自身重量时,就需要拆开这些层次,因为它们确实承担不同的职责。

实践背后的原则

我开始把这组想法称为 AgentOS。它不是工具或框架,而是围绕 agent 有意构建的一层结构。等这些模式成熟,我计划把它发展成独立仓库。即使在现在这个阶段,底层原则也已经相当清楚:

上下文应该积累,而不是蒸发。 每次会话结束时没有留下重要内容,未来的 agent 和未来的你,就得部分重做一次。目标是让项目越做越省力,而不是越来越难。

各层保持分离。 思考、技艺、机械操作的有效期和受众不同。计划、技能和 Captain Log 不是同一种文档,混在一起只会让它们都变差。

文档维护是核心工作。 与现实脱节的文档比没有文档更糟,因为它会主动误导。系统需要的不只是产出文档的机制,还要能发现并修正过时内容。

系统应该在使用中改进。 这一点最难落地,却正是目标所在。每篇 Captain Log 都应让下一次会话稍微更合拍;每个固化的工作流都应消除一类重复劳动;每份计划都应减少一次考古。

这些原则大多已经落实,但仍有演进空间。多数机制已经就位,另一些仍在设计。一个想法处于积极发展中,就是这样的状态。下面看看实际做法。

AgentOS

操作系统并不亲自做工作,它创造让工作可靠发生的条件。AgentOS 把这个思路用在 AI 协作中。在 ProjectSpire 里,它是这样的。

把 agent 指令作为项目记忆

系统根基是仓库级 agent 指令文件。在 ProjectSpire 中,AGENTS.md 是指向 CLAUDE.md 的符号链接,CLAUDE.md 则是所有参与项目的 agent 的操作手册。

CLAUDE.md
view raw
		
  1. ## Plans
  2. Implementation-ready plans live in `Documentation/Plans/`.
  3. Use this folder for agreed plans that should be executable by another engineer or agent without rediscovering the intent, target commit, naming convention, verification steps, or important assumptions.
  4. ## Captain Logs
  5. At the end of each meaningful user-agent conversation, add or update the appropriate Captain Log in `Documentation/Captain Logs/`.
  6. Follow the detailed local workflow in `Documentation/Captain Logs/AGENTS.md`. Claude-compatible local instructions are available through the symlink `Documentation/Captain Logs/CLAUDE.md`.
  7. ## Builds
  8. When compiling with `xcodebuild`, pipe through `xcbeautify` to keep output compact and readable:
  9. ```
  10. xcodebuild ... | xcbeautify
  11. ```

这个文件刻意不尝试包含所有可能的指令。它是路由入口,指向稳定的位置:实施计划、Captain Logs、发布标签流程、快照流程、构建约定、时间线写作流程。

这已经成为我看重的一种模式。好的 CLAUDE.md 不是一条巨大的提示,而是进入项目记忆系统的入口。

计划:记录思考的地方

Documentation/Plans/ 是把达成一致的意图变成可执行内容的地方。

计划不是模糊的 TODO,也不是我独自写出来的东西。它是与 agent 进行规划会话的产物。我引导方向、提出问题、质疑方案形态,agent 负责写下来。人类主导,agent 执笔。这个区别很重要,因为最终计划反映的是真正的架构思考,而不只是随手敲出来的方便方案。我提供判断,agent 负责凝练。

结果是另一位工程师或 agent 可以直接接手,不必重新摸索目标、假设、命名约定或验证路径。

当我把 Neow's Cafe 从模拟卡牌推进到真实卡牌目录时,计划记录了核心决定:先用朴素的静态目录,不做 REST API。

		
  1. ## Summary
  2. Integrate `Lab/data/v0.103.2/cards`, `Lab/resources/images/packed`, and card localization data into `Apps/Apple/Neow's Cafe` through a generated, static, versioned catalog serving view.
  3. Use one compact card index for app-side search and filters, lazy-load portrait images, and avoid a REST API until there is a concrete need for server-side querying or sync.
  4. ## Catalog layout
  5. Generate a static catalog serving view under a versioned root:
  6. ```text
  7. Lab/catalog/v0.103.2/
  8. manifest.json
  9. cards.index.json
  10. cards -> ../../data/v0.103.2/cards
  11. images/card_portraits -> ../../../resources/images/packed/card_portraits
  12. ```
  13. `Lab/data/v0.103.2/cards` and `Lab/resources/images/packed/card_portraits` remain the source of truth. The catalog generator should create real app-facing files for `manifest.json` and `cards.index.json`, then use symlinks for existing raw card JSON and portrait assets to avoid duplicating data during local development.

这样的上下文,从聊天历史里重建很昂贵,从文件里读取却几乎不费事。计划告诉 agent 目录在哪里、什么是权威来源、哪些应生成、哪些应保留为符号链接,以及为什么暂时不是 API。每次会话都会遇到的同样问题,只需问一次。

计划也记录验证方式和假设。这很重要,否则 agent 生成的工作容易滑向“在我机器上能编译”的境地。

		
  1. ## Verification
  2. Add unit coverage for:
  3. - manifest decoding and version/checksum comparison
  4. - card index decoding for integer and X-cost cards
  5. - existing search, pool, type, and rarity filters against generated catalog card fixtures
  6. - missing portrait behavior using known cards without portrait assets
  7. - catalog symlink resolution for card JSON and portrait assets
  8. Run the Neow's Cafe test target.
  9. Build the app with compact output:
  10. ```sh
  11. xcodebuild ... | xcbeautify
  12. ```
  13. Manually verify the Cards tab loads all cards, filters still work, search still works, and portraits load lazily without blocking the grid.
  14. Verify the static server can fetch a sample card JSON and portrait through the catalog URL layout.
  15. ## Assumptions
  16. - `v0.103.2` is the first catalog version to expose to the app.
  17. - The card metadata payload is small enough to download as one index.
  18. - Images are the only payload that needs lazy loading.
  19. - English resolved card text from the existing card JSON is sufficient for v1.
  20. - A static catalog is preferred over REST until server-side search, sync, remote hosting policy, or multi-version browsing requires it.
  21. - `Lab/catalog/` is generated, rebuildable, and tracked.
  22. - Symlinks are for local development and repo hygiene; packaged/offline catalogs should contain real files.

我还在思考一个未解决的问题:计划和 devlog 之间是什么关系。计划记录工作前的意图,devlog 记录落地后真正发生的事。理论上两者清晰分离,实践中边界却会模糊。计划在执行前不断演进,devlog 最后记录了本可以写进计划的决定,有时也不清楚究竟需要哪种文档。我仍在积极摸索这两者最合适的形式。这确实是构建这套系统时的成长烦恼,我不打算假装没有。

**更新:**我想,我已经找到了计划和 devlog 的分界。

计划来自一次有意识的规划会话,适用于复杂到需要在动手前想清楚方案形态的修改。

但不是每个改动都需要这样。加入自定义排版系统或浅色与深色主题,并不涉及复杂决策,但值得简要记录怎么做、为什么做。这就是 devlog 的用途。

Captain Logs:记录技艺的地方

我认为,这是整套系统中探索最少的部分。如果重新开始,我会先做它。

Captain Log 不是逐字记录,也不是开发日志。它记录的是合作的方式:我提出什么要求,agent 如何回应,我在哪里纠正或引导它,以及未来的 agent 应该延续什么。

		
  1. ## 2026-05-06 - Meaningful Interaction Threshold
  2. **Context:** A purely appreciative follow-up had been logged after the daily blog-summary workflow was accepted.
  3. **User Direction:** The user clarified that Captain Logs are only for meaningful interactions and that purely positive acknowledgements such as "nice" do not need log entries.
  4. **Agent Response:** The agent removed the acknowledgement-only entry and updated the root Captain Logs instruction to skip purely appreciative or acknowledgement-only replies unless they include a decision, correction, new constraint, or requested change.
  5. **User Feedback:** The user explicitly wanted the feedback itself recorded in the original Captain Logs conversation file.
  6. **Outcome:** Captain Logs now have a clearer threshold: record meaningful collaboration, steering, decisions, corrections, constraints, requested changes, and outcomes; do not record empty praise or acknowledgement-only replies.
  7. **Carry Forward:** Future agents should avoid treating every user message as log-worthy. Log only interactions that preserve useful collaboration context.

大多数 AI 工作流文章都关注提示、上下文窗口或工具使用。几乎没人把引导过程本身当作一项重要产物来记录。但真正有价值的信息,大多就在这里。

如果我告诉 agent“这太正式了,写得更像一条时间线动态”,那不是实现细节,而是偏好、品味、工作流纠正。如果只留在聊天里,下一个 agent 还会犯同样的错。如果写进 Captain Log,系统就会好一点;关键是,改善发生在技艺层面,而不仅是技术层面。

我认为,这一步最重要。代码质量是一种记忆,协作质量是另一种。把它们混为一谈,最终就会得到一条 3,000 行的提示,产出的语气却仍然不对。😜

至少对我自己的使用而言,Captain Logs 的方向不止是记录纠正。我想把积累的日志当作原料,提炼出更精简的东西:一份“用户指南”,让未来 agent 了解我通常如何思考、会反对什么、什么样的输出在我看来算好。最终还能基于历史主动跟进:“根据之前的 Captain Logs,你是否希望我对刚做的内容进行 X、Y 或 Z 调整?”引导历史会成为一个可用的用户模型,而不只是会话记录。这仍是实验,但方向让我觉得对。

Devlogs:技术历史记录者的日志

相比之下,Devlogs 单独负责技术历史,本地指令也明确说明了这种区分:

		
  1. # Devlog Documentation Instructions
  2. Devlogs in this folder are high-level historical records for humans and future agents.
  3. When adding a devlog:
  4. - Use the filename pattern `NNNN - Short Title.md`.
  5. - Start the title with the log number and topic, for example `# 0002 - Neow's Cafe Card Catalog Integration`.
  6. - Start with the date in this format `Date: YYYY-MM-DD`.
  7. - Record why the work happened, what changed at a system level, and which commits/files are the durable source of detail.
  8. - Prefer commit hashes, plan names, and important paths over long code snippets.
  9. - Include verification commands and notable outcomes.
  10. - Capture decisions, assumptions, and follow-up work that would be hard to infer from diffs alone.
  11. - Keep implementation minutiae out unless it explains a decision or a future maintenance risk.

两种记录,回答两个问题:

  • Devlog:技术上改了什么,为什么改,如何验证?
  • Captain Log:人类与 agent 的合作如何演进,未来 agent 应该记住用户方向中的哪些内容?

我喜欢这种区分,因为它避免每份文档最终都变成同一种文档。项目决策、未解决问题、计划、实施历史和协作笔记,各有职责,不应该全采用同一种形式。

技能与工作流:吸收机械操作的地方

第三层最偏操作,也最容易在缺失时被发现。

在 Lab/.claude/skills/decompile-sts2/SKILL.md 中,我有一个用于反编译本地 Slay the Spire 2 DLL 的小型自定义技能。它描述适用场景、确切命令、选项、前提条件和失败方式。

		
  1. ---
  2. name: decompile-sts2
  3. description: Decompile the local Slay the Spire 2 DLL into C# source. Use when the user wants to decompile the game, regenerate decompiled source, or update to a new game version.
  4. ---
  5. # Decompile STS2
  6. Run the decompile script:
  7. ```bash
  8. scripts/decompile-sts2.sh
  9. ```
  10. Requires `ilspycmd` on PATH (`dotnet tool install --global ilspycmd`). The script auto-detects the game version from `release_info.json` and outputs to `decompiled/<version>/`.
  11. ## Options
  12. - `--dll PATH` — override the DLL path (default: Steam macOS ARM64 install)
  13. - `--out DIR` — override output directory
  14. - `--clean` — wipe the output directory before decompiling
  15. - `STS2_DLL_PATH` / `STS2_DECOMPILED_DIR` — env var equivalents
  16. ## Known failure modes
  17. - **`ilspycmd` not found**: run `dotnet tool install --global ilspycmd` and ensure `~/.dotnet/tools` is on PATH
  18. - **DLL not found**: pass `--dll PATH` or set `STS2_DLL_PATH` if Steam library is in a non-default location
  19. - **Output directory not empty**: pass `--clean` to wipe and re-decompile
  20. - **Version shows `unknown`**: `release_info.json` not found — verify the DLL path points to the real Steam install, not a copy

这不是思考,也不是技艺,而是纯粹的机械性知识:把确切顺序固化下来,让我不必再解释。快照标签有文档,发布标签有文档,时间线摘要流程也有文档和示例:快照标签、发布标签与页面、今日工作时间线摘要。

连写下工作进展也成了可重复流程:检查今天的提交,检查改动过的文档,区分已提交和未提交的工作,再写一段以叙述为主、自然融入少量来源引用的更新。

我定下的规则是:发现自己第二次做同一套操作,就把它固化下来。可以是脚本、技能、指令文件、计划模板或日志格式。目的不是为了文档而写文档,而是减少重复的认知劳动,让下次会话能把精力花在真正需要思考的地方。

让人不太安心的部分

对于仍在变化的部分,我想坦诚说明。

真正的风险不是文档是否有用,它们确实有用。agent 能有效浏览计划:用 rg 或 grep 搜关键词,从标题判断相关性,只阅读需要的部分。查找问题基本解决了,过时问题还没完全解决,至少还没有作为系统性机制内置在 AgentOS 里。但这方面也不是毫无希望。

如果先按一种方式构建,后来大幅重构,却没更新对应计划,下一个 agent 就会把过时上下文当作事实。第一次发生只是小问题,发现后修文档,继续往前走。但如果它在足够多的文档中悄悄累积,上下文层就会开始帮倒忙。

我正在推进的缓解方法,是把文档维护视为核心工作,而不是事后补救:在 agent 指令中要求重大变更后检查相关文档,最终形成 agent 主动标记过时文档的流程。系统应能发现自身偏移。理想情况下,这会成为 AgentOS 内置原则之一,而不是依赖人类记得去做。

Captain Logs 也有自己的同类问题:不是每次引导都值得保存。有些纠正可能只是我在某个周二心情不好。(🤣 Claude,我给你的感觉是这样吗?不过它确实是在周二写的,哈哈。)分清哪些信号值得记录,哪些噪声只会污染日志,是我还在培养的判断力。

所以,我不会把它包装成已经解决一切的方法论。它是一个正在实践的假设:结构化上下文的成本,低于每次会话重新建立上下文的成本;而且按层次维护这种结构,比放进一个不加区分的大团块更值得。

这也让我想,有没有办法具体检验这个假设?也许可以比较有无 AgentOS 时,实现同样结果所需的 token 成本。但说实话,我懒得做。哈。

会自行积累价值的工程系统

这一切背后的模式,是我最感兴趣的地方:项目被设计成会积累,而不只是记录。

agent 做一项工作,我引导它,引导变成 Captain Log。如果工作成为可重复流程,就变成 agent 工作流;如果达成共识但尚未完成,就变成计划;如果落地,就变成 devlog;如果暴露 bug,就变成问题笔记;如果涉及重复命令,就变成技能。

随着时间推移,仓库保存的不只是代码,还有代码周围的思考与协作环境。思考层、技艺层和机械操作层各有自己的位置,彼此边界也保持清晰。

这改变了我对 AI agent 的期待。不只是更快的代码,而是经过有意设计的上下文环境,让它们能有效工作,并且环境设计本身也随着项目使用不断改进。

结束了吗?

这些并没有改变人类的职责。agent 仍需要有人提供判断、把握品味、决定什么值得保存。AgentOS 不是要自动化这些,而是要让它们留下来。

AgentOS 本身也没完成。它大概永远不会彻底完成,这正是一个会自行积累价值的系统的意义。模式会随着项目使用继续演进,最终我计划为它建立独立仓库。

如果你也在做类似的东西,或者对它会在哪里失效有明确看法,我很想听听。整个探索过程的协作日志,还在继续写。