← 总览
文章用户精选· 09-18 · 07:50

如果编程已被解决,接下来怎么办?:衡量代码的草率度 | Earendil

If coding is solved, what now?: Measuring the sloppiness of code | Earendil

打开原文约 20 分钟读

If coding is solved, what now?: Measuring the sloppiness of code | Earendil

If coding is solved, what now?: Measuring the sloppiness of code | Earendil

Exploring how to measure code sloppiness, why correct code can still erode a codebase, and why human intuition and taste still matter.

Date:Thu, 10 Sep 2026

From:Sebastian <sebastian@earendil.com\>

To:You

Subject:If coding is solved, what now?: Measuring the sloppiness of code

LLMs have become almost perfect at generating code, but that isn’t the end of the story. Just because the code is formally correct doesn’t mean that it is not introducing unnecessary abstractions, creating duplicates, or just making bad decisions overall. This is not a groundbreaking observation, most people who have vibe-coded a project, have realized that each additional feature can sometimes lead to an explosion of lines of code (LOC).

This results in a loss of human agency, because in projects that are adding millions of LOC per month, it is hard for humans to keep up.1 Some people might say that that is not an issue at all, because they trust their agents to deal with it. I have bad news for you, agents can't really deal with the slop either.

Coming from a physics background, I always had an experimental/quantitative approach to solving problems. When I started at Earendil, with the task of figuring out how to measure code sloppiness, my natural instinct was to first take a deep dive into the literature and then check what other companies were doing.

To be frank, with the exception of a few insightful research papers, I was disappointed at how “vibes based” the industry seems at the moment. In my research and on X, I was constantly bombarded with messages such as “End-to-end coding agents”, “AI that doesn't just suggest code—it ships it” or “Human-level evaluation without human-level cost”. Which like all good tales, have a grain of truth in them.

LLMs are able to write almost perfectly correct code. This is because of the scalability and the verifiability of code. It is pretty straightforward to let LLMs generate code and then let that code be checked by hidden tests, which results in a clear reward signal. In stark contrast to that, checking the ‘sloppiness’ of this code often requires human intuition and taste, and is an extremely difficult task in general. I think the best way to illustrate why that is, is by going through possible ways of measuring slop.

AI as a judge: This is probably the most common way of evaluating code quality in the industry and from my observations it rarely works. The most naive way of doing it, namely asking the models how good the code is on a scale from 1-10, is basically equivalent to a random number generator. The more sophisticated approach, namely trying to give the judge model two solutions A and B, and then letting it decide which solution it prefers, has the downside of the model changing its preference, when you rename the solutions. I am being a bit facetious here and the effect isn’t as pronounced with larger models, but the main point still stands. Asking LLMs to judge the code they write is not a substitute for a proper evaluation. Even though there are some interesting approaches with rubrics or the LLMs writing tests, they are still a far shot from actually getting rid of the slop.

Human judges the AI: If we ignore the fact that there is huge diversity in the quality of software-engineers, this would be the best solution to assure that the code stays human readable. With the downside being that this is not scalable for training AI or having large benchmarks with multiple model providers and harnesses.2

The simplest method: In my research and tests simply taking the change in the number of LOCs has been a surprisingly effective metric for sloppiness, with the ironic caveat that if we started optimizing for it, it would cease to be a meaningful measure.

The next two measures were introduced to me by the paper SlopCodeBench, and seemed promising because they were able to separate legacy code bases from LLM-slop quite well.

Verbosity: Tries to measure the amount of duplicated and unnecessary verbose lines.3

Verbosity=|AST-Grep flagged lines∪clone lines|LOC

Erosion: Tries to measure how much of a codebase's mass is concentrated in a few large and complex functions.

mass(f)=CC(f)√SLOC(f)

Here f is the function, SLOC is the source lines of code and CC(f) is the cyclomatic complexity of the function.

Erosion=∑f: CC(f)>10mass(f)∑fmass(f)

The erosion is then simply the fraction between the mass of the functions with a cyclomatic complexity larger than 10, by the mass of all functions.

If we look at the average verbosity and erosion of the code generated during the SlopCodeBench evaluation and compare that to a set of established repos there is a stark difference between them. On average the verbosity in the repos is 0.15 ± 0.06 and in the agents code is 0.33 ± 0.10. For erosion the repos achieve 0.31 ± 0.17 and the agents 0.68 ± 0.20. The agent's code is on average roughly twice as verbose and eroded as human code. I then investigated some vibe coded projects of my own and a lot of them had a verbosity of up to 0.4 and erosion as high as 0.75, so these results probably weren't just an artifact of the evaluation.4

To come back to the point of why agents can’t (really) deal with the slop themselves, we need to look at the evaluation of SlopCodeBench. In contrast to other coding benchmarks, which give the agent a complete list of instructions at the start and then have a set of hidden tests the program needs to pass, they do the opposite. They create multiple rounds of instruction and test iterations, where in between checkpoints the context of the models is erased. Thereby mimicking much more closely an iterative process, like how coding agents are actually used by humans. The result of that is that bad coding decisions accumulate over time and for the strict solve rate, where all tests have to be passed at all checkpoints, even state of the art models achieve 0% pass rate.5 Which should be a warning sign to everyone who happily adds tens of thousands or even hundreds of thousands of LOC a day. Obviously there are the usual caveats about too strict of tests or one or the other slightly ambiguous problem statement, but the general trend holds.

In exploring these metrics I hope you now have a clearer picture of why it is challenging to evaluate code sloppiness and why human intuition and taste are still either implicitly or explicitly baked into the evaluation.

There are some promising other directions I want to explore, such as coupledness of functions, code churn, cohesion and so on. If you are working on evals and would like to talk, I would be happy to do that: sebastian@earendil.com

  • * *
  1. This reminds me of the saying: “Measuring programming progress by lines of code is like measuring aircraft building progress by weight.” ↩
  2. I wouldn't want to force anyone to review millions of LOC just to get an ever changing ranking of model providers. ↩
  3. The rules for this are a set of handcrafted heuristics implemented via AST-Grep, which again goes to show the human aspect of it all. ↩
  4. Though a notoriously vibey, open project didn’t score too high on either metric, probably coming from really high coupledness of functions and/or just the sheer mass of unrelated functions reducing the averages. ↩
  5. Not tested on Fable 5.1 or Astra yet, but on GPT 5.6 sol xhigh etc. ↩

If coding is solved, what now?: Measuring the sloppiness of code | Earendil

如果编程已被解决,接下来怎么办?:衡量代码的草率度 | Earendil

探讨如何衡量代码的草率程度,为什么正确的代码仍然会侵蚀代码库,以及为什么人类的直觉和品味依然重要。

日期:2026 年 9 月 10 日(周四)

发件人:Sebastian <sebastian@earendil.com>

收件人:你

主题:如果编程已被解决,接下来怎么办?:衡量代码的草率度

大语言模型生成代码的能力已经几乎完美,但这并不是故事的终点。代码在形式上正确,并不意味着它没有引入不必要的抽象、制造重复代码,或者总体上做出糟糕的决策。这并不是什么惊天动地的观察——大多数曾用氛围编程做过项目的人都意识到,每新增一个功能,有时都会导致代码行数(LOC)爆炸式增长。

这会导致人类主导权的丧失,因为在每月新增数百万行代码的项目里,人类很难跟上节奏。1 有些人可能会说这根本不是问题,因为他们信任自己的智能体能处理好这些。但我有一个坏消息要告诉你们:智能体其实也对付不了这些潦草代码。

我出身物理学背景,解决问题时一向采用实验性/定量化的方法。当我加入 Earendil、接到弄清楚如何衡量代码草率度这个任务时,我的第一反应自然是先深入研究文献,然后再看看其他公司在做什么。

坦率地说,除了少数几篇有见地的研究论文之外,我对业界目前如此“凭感觉行事”的现状感到失望。在我的研究和 X 平台上,我不断被这样的口号轰炸:“端到端编程智能体”、“不仅建议代码——还能直接交付代码的 AI”,或者“人类水平的评测,无需人类水平的成本”。就像所有精彩的故事一样,这些话里都有一丝真实的成分。

大语言模型已经能够写出几乎完美正确的代码。这得益于代码的可扩展性和可验证性。让大语言模型生成代码,再用隐藏测试来检验这些代码,从而得到清晰的奖励信号,这是相当简单直接的做法。与此形成鲜明对比的是,检验这些代码的‘草率度’往往需要人类的直觉和品味,而且总体而言是一项极其困难的任务。我认为,要说明为什么会这样,最好的办法就是逐一梳理衡量潦草代码的几种可能方法。

AI 当裁判:这可能是业界评估代码质量最常见的方式,据我观察,它很少奏效。最天真的做法是让模型按 1-10 分给代码打分,这基本上等同于一个随机数生成器。更复杂的做法是给裁判模型两个方案 A 和 B,让它决定自己更偏好哪一个,其缺点在于,当你重命名这些方案时,模型会改变自己的偏好。我在这里说得略带戏谑,而且这一效应在更大的模型上并不那么明显,但核心观点仍然成立。让大语言模型评判它们自己写的代码,并不能替代真正的评测。尽管有一些有趣的方法,比如使用评分细则或让大语言模型编写测试,但它们离真正清除潦草代码还差得很远。

由人来评判 AI:如果我们忽略软件工程师水平存在巨大差异这一事实,那么这将是确保代码保持人类可读的最佳方案。其缺点在于,这种方式无法扩展到训练 AI,也无法支撑拥有多个模型供应商和评测脚手架的大型基准测试。2

最简单的方法:在我的研究和测试中,仅仅取代码行数(LOC)的变化量,就是一个出人意料有效的草率度指标,只是有一个带有讽刺意味的前提:如果我们开始针对这一指标进行优化,它就会不再是有效的衡量标准。

接下来的两个指标来自 SlopCodeBench 这篇论文,它们看起来很有前景,因为能够相当好地区分遗留代码库与大语言模型生成的潦草代码。

冗长度:试图衡量重复且不必要的冗长代码行的数量。3

Verbosity=|AST-Grep flagged lines∪clone lines|LOC

侵蚀度:试图衡量代码库的质量有多大比例集中在少数几个庞大而复杂的函数里。

mass(f)=CC(f)√SLOC(f)

其中 f 是函数,SLOC 是源代码行数,CC(f) 是该函数的圈复杂度。

Erosion=∑f: CC(f)>10mass(f)∑fmass(f)

侵蚀度其实只是圈复杂度大于 10 的函数的质量之和,除以所有函数的质量之和所得的比值。

如果我们看看 SlopCodeBench 评测期间生成的代码的平均冗长度和侵蚀度,并将它们与一组成熟代码库进行比较,就会发现两者之间存在巨大差异。代码库的平均冗长度为 0.15 ± 0.06,而智能体生成的代码为 0.33 ± 0.10。就侵蚀度而言,代码库为 0.31 ± 0.17,智能体为 0.68 ± 0.20。智能体生成的代码,其平均冗长度和侵蚀度大约是人类代码的两倍。随后,我调查了自己的一些氛围编程项目,其中很多项目的冗长度高达 0.4,侵蚀度高达 0.75,所以这些结果很可能不只是评测造成的人为现象。4

回到智能体为什么(真的)无法自己对付潦草代码这一问题上,我们需要看看 SlopCodeBench 的评测方式。与其他编程基准不同,那些基准一开始就给智能体一份完整的指令清单,随后准备一组程序必须通过的隐藏测试;他们则反其道而行之。他们设置多轮指令和测试的迭代,在各个检查点之间,模型的上下文会被清空,从而更贴近地模拟一个迭代过程,就像人类实际使用编程智能体的方式一样。这样做的结果是,糟糕的编码决策会随时间不断累积;对于严格求解率——要求在所有检查点都通过全部测试——即使是最先进的模型,通过率也只有 0%。5 对于每个每天开开心心新增数万甚至数十万行代码的人来说,这应该是一个警示信号。当然,通常那些保留意见依然存在,例如测试过于严格,或者某个问题描述略有些歧义,但总体趋势是成立的。

通过探讨这些指标,我希望你现在能更清楚地理解,为什么评估代码草率度如此具有挑战性,以及为什么人类的直觉和品味仍然以或明或暗的方式植根于评测之中。

我还想探索其他一些有前景的方向,比如函数间的耦合度、代码改动量、内聚性等等。如果你正在从事评测工作,并且愿意聊聊,我非常乐意交流:sebastian@earendil.com


  1. 这让我想起一句名言:“用代码行数来衡量编程进度,就像用重量来衡量飞机制造的进度。” ↩
  2. 我可不想强迫任何人仅仅为了一份不断变化的模型供应商排行,就去审查数百万行代码。 ↩
  3. 这背后的规则是一套手工编写的启发式方法,通过 AST-Grep 来实现,这再次体现了其中人的因素。 ↩
  4. 不过,一个出了名氛围感十足的开源项目在这两个指标上得分都不算高,这很可能是因为其函数耦合度极高,和/或仅仅因为大量无关函数拉低了平均值。 ↩
  5. 尚未在 Fable 5.1 或 Astra 上测试,但已在 GPT 5.6 sol xhigh 等模型上测试过。 ↩