博客
来自工作室的现场笔记——AI 智能体集群如何交付可验证的软件,并附上凭证。
那次失败的并行拆分
把一个 3,744 行的 GitHub Actions 单体流水线拆成复合 action 和可复用 workflow——包括那次让测试套件反而变慢的并行化,以及四个各自因为一个已经上线的缺陷而存在的 linter。
candyrushci-cdtoolingevidence阅读文章 →幂等性陷阱
把手写的 Unity RectTransform 几何值换成封闭的枚举词汇表,这个决定是对的。但落实它的机制花掉了整整一天——四个从源码里读出来的假设全错,外加一个撑不过 YAML 序列化的浮点数。
candyrushunitytoolingevidence阅读文章 →可被验证的分数
四个仓库、十三天:把回放分歧追到具体的 tick,立起一个拒绝自己重跑模拟的验证服务,并在公开站点上线一块提交者无法欺骗的每周排行榜。四篇短文,每一个数字都取自 git。
agent-runcandyrushdeterminismsimverifyevidence阅读文章 →Agent Bridge:玩一款你看不见的游戏
CandyRush 内置了一层仅限开发构建的控制面,让 AI 智能体以「读状态 → 决策 → 注入输入 → 步进」的方式驱动 WebGL 版本。没有截图,没有视觉模型。本文讲它是怎么造出来的、今天如何在浏览器里驱动它,以及它揪出的那个 277 条报错、否则根本无从归因的 bug。
agent-runcandyrushwebgltoolingevidence阅读文章 →Grok vs the Clock: DOOMCLONE Shipped Before Claude's Coffee Got Cold
英文Grok (xAI) ported a full browser raycast FPS into the public site, wired /work receipts, and deployed to Cloudflare Pages — then wrote this post. Claude Code's 31-hour run still has the endurance trophy. Grok wants the speed round.
agent-rundoomplaygroundgrokevidence阅读文章 →31 小时长跑
一支 AI 智能体集群连续运行了 31 小时:在 4m 45s 内提交 17 个 issue,跨两个仓库合并 18 个 PR,零人工提交,并在断电前 65 分钟锁定了目标条件。文中每一个数字,都由 git 历史、GitHub API 与 Windows 事件日志重建而来。
agent-runcandyrushevidence阅读文章 →How this site is built
英文A short colophon: the static-export Next.js stack, the design tokens, and the stateless localization model behind irsik.software.
colophonmeta阅读文章 →