← 总览
文章用户精选· 09-04 · 07:09

基础模型如何在 Bash 上达到超人水平

How Foundational Models Became Superhuman in Bash

打开原文约 16 分钟读

How Foundational Models Became Superhuman in Bash

If you have used coding agents, you may have noticed that they started to use Bash for tasks that would have traditionally used dedicated tools such as read, edit, and search. Well, models became superhuman in bash and might consider dedicated tools not good enough.

Every few months, I rebuild an agent harness from scratch to learn what changed and what I can delete. This time I removed every model-facing tool except for a bash and a media viewer.

The outcome matched what I had already seen in daily use. Task completion stayed in the same range, while the agent moved through complex work and verifications with fewer tool boundaries.

Frontier models have become superhuman in Bash. I mean that in a narrow sense: they synthesize disposable command-line programs, in seconds, that most developers would need much longer to assemble and verify.

What “superhuman in Bash” means

You probably know the individual pieces git, rg, jq, Python, temporary files, process substitution, and test runners. Writing a one-off 40-line workflow that combines them under time pressure is different. Most of us work through the problem interactively, checking syntax and state between steps.

Coding agents assemble these workflows from patterns learned across POSIX utilities, programming languages, build systems, and public source code. They still make mistakes. The difference is how much correct orchestration they can attempt at once.

Bash is the routing layer here. The heavy work may happen in Python, Git, SQLite, a compiler, or a project-specific CLI. One shell interface lets the agent compose all of them without forcing the runtime designer to predict every useful operation.

The following are simplified examples from my personal agent traces.

Example 1: In-place multi-file edits

You work on a project and you want to rename a variable across an implementation, its callers, and the tests.

The script first counts every old snippet. If a file has the wrong number of matches, it exits before writing anything. Only then does it apply the four-file rename in one pass, so you never sit in a half-updated tree. After the write it formats, lints, type-checks, runs the focused tests, and prints a compact diff. Failed checks stay in the working tree, so the next turn can inspect and fix them. That verify-before-done cycle is the inner loop.

Example 2: Reproducing and bisecting a flaky regression

You have a test that fails only sometimes. You want the first commit that introduced the flake, without moving your current checkout.

The extra worktree keeps your local files in place. Each commit is classified with five seeds, and it only counts as good if fewer than three runs fail, so one unlucky flake does not poison the bisect. When it lands, the script prints the first bad commit, a truncated diff of the suspected change, and the tail of the log. You get a diagnosis, not every test run in the prompt.

Example 3: Correlating compressed production logs

You have compressed production logs that are too large to load into the model. You want the failing endpoints, their error classes, and P95 latency.

Python reads the rotated files in batches and never dumps them into the prompt. SQLite does the join, the P95 ranking, and the error-class aggregation. What comes back is five JSON records: endpoint, failure count, average latency, P95, and error classes. Context offloading keeps the intermediate data in the environment, and puts only the summary in the model.

Why atomic tools were right

A few months ago, I argued that coding agents should start with robust atomic tools and avoid shell commands such as cat, sed, and echo. A careless command could flood the context window, return an opaque error, or corrupt an edit through bad quoting. Text output also cannot carry visual information into a vision model.

Those limits still apply to an uninstrumented shell. Foundational models are now much better at composing small Python patchers, Git commands, quoted heredocs, and focused tests. The Harness also absorbs the protections that made atomic tools useful:

  • Output control: truncate large results and tell the agent how to request a narrower slice.
  • Diagnostics: return the exit status, duration, timeout state, and process information.
  • Isolation and policy: gate paths, network access, and destructive actions outside the model.
  • Asynchronous processes: let the agent start, inspect, and stop long-running commands without blocking a turn.

The one exception multimodal input:

Text output cannot make a screenshot/image visible. A shell command can render a page, capture a chart, or save a video frame, but the pixels still need to enter the model through a multimodal channel.

What this means for harness engineering

I compared a shell-centered setup with a configuration that exposed separate tools for file reading, writing, editing, and search. Both ran against the same coding task set under the same conditions. The shell-centered setup achieved on par or better performance.

That is the Bitter Lesson applied to harness design: general methods that scale with computation beat hand-built shortcuts. Bash is that general computation layer.

A good agent should have more tools when they provide a better interface to a capability. If the model can do something better with a tool, it should use the tool. Browser control is a good example.

A browser tool can navigate, click, type, wait for the page to settle, and return a screenshot in the same call. Doing this through Bash means invoking a CLI, locating the captured artifact, and passing it through view\_media in a second step. Service integrations can benefit in a similar way, or stay on the shell with something like mcp-cli so the schemas never sit in the prompt.

Here is what to try:

  • Delete the micro-tools: Try to use Bash to handle file reading, search, multi-file edits, diffs, and verification.
  • Use subagents as execution firewalls: Delegate messy exploration and debugging, then return a clean result to the parent context.
  • Keep the system instructions minimal: Give the agent more room for repository instructions, domain knowledge, and the task itself.

Our goal is a smaller interface with a larger action space.

基础模型如何在 Bash 上达到超人水平

How Foundational Models Became Superhuman in Bash

如果你用过编码智能体,可能已经注意到:它们开始用 Bash 来处理过去由 read、edit 和 search 这类专用工具负责的任务。没错——模型在 Bash 上已经达到超人水平,说不定还嫌专用工具不够好用。

每隔几个月,我都会从零重写一遍智能体框架(harness),看看有什么变化、哪些东西可以删掉。这一次,我把面向模型的工具全部删光,只留下一个 bash 和一个 media 查看器。

结果和我日常使用中的观察一致:任务完成率维持在原来的区间,而智能体在推进复杂工作和验证时,需要跨越的工具边界反而更少了。

前沿模型已经在 Bash 上达到超人水平。我是在狭义上这么说的:它们能在几秒钟内合成一次性的命令行程序,而大多数开发者要组装并验证同样的东西,得花长得多的时间。

“在 Bash 上达到超人水平”意味着什么

git、rg、jq、Python、临时文件、进程替换、测试运行器——这些零件你大概都认识。但在时间压力下写一个把它们组合起来的一次性 40 行工作流,完全是另一回事。我们大多数人是一边交互一边推进问题的,每走一步都要核对语法和状态。

编码智能体把这些工作流组装起来,凭的是从 POSIX 工具、编程语言、构建系统和公开源代码里学到的模式。它们仍会犯错。差别在于:一次能尝试多大规模的正确编排。

在这里,Bash 是路由层。重活可能发生在 Python、Git、SQLite、编译器或某个项目专属的 CLI 里。一个 shell 接口让智能体把它们统统编排起来,也让运行时设计者不必预先想到每一种有用的操作。

以下是一些简化后的示例,来自我个人的智能体运行记录。

示例 1:就地多文件编辑

你正在一个项目上工作,想把某个变量在实现代码、所有调用方和测试里统一改名。

脚本会先统计每处旧片段出现的次数。只要某个文件的匹配数不对,它就会在写入任何东西之前退出。确认无误后,它才一次性完成四个文件的重命名,这样你永远不会坐在一个改了一半的工作区里。写完之后,它会格式化、lint、类型检查、跑聚焦的测试,并打印一份紧凑的 diff。失败的检查会留在工作区,让下一轮可以查看并修复它们。这个“完成前先验证”的循环,就是内循环。

示例 2:复现并二分定位一个偶发失败(flaky)的回归

你有一个时不时才失败的测试。你想找到引入这个偶发失败(flake)的第一个提交,又不想动当前检出的代码。

额外加出来的工作树(git worktree)让你的本地文件原封不动。每个提交都用五个随机种子来分类,只有失败次数少于三次才算“好”,这样一次不走运的偶发失败(flake)不会污染二分定位。尘埃落定时,脚本会打印第一个坏提交、疑似变更的截断 diff,以及日志的末尾。你得到的是一份诊断,而不是把每一次测试运行都塞进提示词里。

示例 3:关联压缩的生产日志

你手头有压缩过的生产日志,大到没法直接装进模型。你想要的是失败的端点、它们的错误类别,以及 P95 延迟。

Python 分批读取轮转的日志文件,从不把它们整个倒进提示词。SQLite 负责连接(join)、P95 排名和错误类别聚合。返回的是五条 JSON 记录:端点、失败次数、平均延迟、P95 和错误类别。上下文卸载把中间数据留在环境里,只把摘要放进模型。

为什么原子工具曾经是对的

几个月前,我还主张编码智能体应该从健壮的原子工具起步,避免使用 cat、sed 和 echo 之类的 shell 命令。一条漫不经心的命令可能灌爆上下文窗口、返回一个不知所云的错误,或者因为引号用错而毁掉一次编辑。文本输出也无法把视觉信息带进视觉模型。

这些限制对一个未加防护的裸 shell 依然成立。但如今的基础模型已经非常擅长组合小型 Python 补丁脚本、Git 命令、加好引号的 heredoc 和聚焦测试。框架(Harness)本身也吸收了那些让原子工具好用的保护措施:

  • 输出控制: 截断过大的结果,并告诉智能体如何请求更窄的切片。
  • 诊断信息: 返回退出状态、耗时、超时状态和进程信息。
  • 隔离与策略: 在模型之外把守路径、网络访问和破坏性操作。
  • 异步进程: 让智能体启动、查看、停止长时间运行的命令,而不阻塞一个回合。

唯一的例外是多模态输入:

文本输出没法让截图或图像变得可见。shell 命令可以渲染页面、截取图表或保存视频帧,但像素终究得通过多模态通道进入模型。

这对框架工程意味着什么

我对比了以 shell 为中心的配置,和一套为文件读取、写入、编辑、搜索分别暴露独立工具的配置。两者在相同条件下跑同一组编码任务。以 shell 为中心的配置取得了持平甚至更好的表现。

这就是《苦涩的教训》在框架设计上的应用:能随算力扩展的通用方法,胜过手工搭建的捷径。Bash 就是那层通用计算。

一个好的智能体应该在工具能提供更好的能力接口时拥有更多工具。如果模型借助某个工具能把事情做得更好,它就应该用这个工具。浏览器控制就是个好例子。

浏览器工具可以在同一次调用里导航、点击、输入、等待页面稳定下来,然后返回截图。换成用 Bash 做,就得调用一个 CLI、找到捕获的产物,再在第二步通过 view\_media 传进去。服务集成也能以类似的方式受益,或者干脆留在 shell 上,用 mcp-cli 之类的方案,让 schema 永远不进提示词。

下面是一些值得尝试的做法:

  • 删掉微工具: 试着用 Bash 来处理文件读取、搜索、多文件编辑、diff 和验证。
  • 把子智能体当作执行防火墙: 把杂乱的探索和调试委托出去,再把干净的结果带回父上下文。
  • 保持系统指令精简: 给智能体留出更多空间,容纳仓库说明、领域知识和任务本身。

我们的目标是:更小的接口,更大的动作空间。