The Unmeasured Fleet: Performance as a Practice
There is a number nobody publishes: the fraction of global cloud spend that exists to accommodate code no one has ever measured. Every capacity planning spreadsheet, every autoscaling policy, every instance-size selection encodes an assumption that the observed resource consumption of a system reflects its intrinsic requirements, when in reality that consumption also carries the accumulated weight of allocation churn, retained caches nobody evicts, defensive copies, and a thousand small decisions made by developers who had no feedback loop that would ever surface them.
We have institutionalized this. What began as individual, invisible inefficiencies has hardened into infrastructure—codified in provisioning scripts, extrapolated in capacity plans, normalized in budgets. The debt reaches beyond the technical into the architectural, the economic, and increasingly the environmental, and like all debt carried forward unexamined, it compounds.
Emerging technologies, including but not limited to AI, result in changed fundamentals. The tools to measure our systems have existed for years in every major ecosystem, mostly free, but they remained the province of a small priesthood because the knowledge was fragmented and the learning path illegible. Emerging AI technologies now collapse both the cost of learning these tools and the cost of acting on what they reveal, which means the familiar justifications for leaving systems unmeasured—that the tools are arcane, that the expertise is rare, that remediation is expensive—are losing their premises almost simultaneously.
What remains is the one thing no technology supplies: the will to make performance a practice woven into how we build rather than a quality we aspire to when trouble arrives. We can continue sizing our infrastructure around what we have never examined, or we can begin measuring what we build. This paper argues that the first path, long defensible on economic grounds, can no longer be defended at all.
Slop Is an Equilibrium
The conventional framing of inefficient code is moral: someone was careless, someone did not know better, someone cut a corner. This framing survives because it flatters everyone not currently being blamed, but it fails to explain what we actually observe. Inefficient code is written by careful developers, reviewed by thoughtful reviewers, and shipped by disciplined teams—constantly, everywhere, in every language. The query loop hidden inside an iteration, the string concatenation in the hot path, the collection materialized when it should have streamed: these pass review because review examines correctness and clarity, and nothing in the standard development loop ever measures cost.
Systems settle into the states their feedback loops permit. A codebase with no performance tests, no allocation budgets, and no traces anyone reads will drift toward inefficiency for the same reason an untended garden drifts toward weeds—not because anyone planted them, but because nothing removes them. Inefficiency, in other words, is simply the natural resting state of any system in which nothing downstream enforces efficiency, and it will accumulate under those conditions regardless of who is producing the code or how much they care.
This reframing matters for the current anxiety about AI-generated code. The worry is usually framed as a production problem, with models emitting plausible-but-wasteful code faster than humans can vet it. That much is true, but it misdiagnoses the disease, because human-generated waste has flowed faster than review could catch it for decades—review was never the mechanism capable of catching it. The only mechanism that reliably catches inefficiency is measurement, and measurement is indifferent to authorship. A harness with enforced budgets fails the build identically whether the regression came from a junior developer, a seasoned architect on a deadline, or a model. In this light, the arrival of generative tools changes little about the underlying problem: the enforcement layer that would have caught waste was missing all along, and building it addresses waste from every producer at once.
When we do not measure:
-
We size infrastructure around waste and record the result as a requirement, converting inefficiency into standing cost.
-
We add debt—technical, architectural, and economic—whose interest accrues monthly in cloud invoices and whose principal nobody has appraised.
-
We deny ourselves the learning that measurement affords. Without seeing what our code actually does, our intuitions about it calcify into folklore.
-
We surrender entire classes of architectural options—as we will see with serverless—not because our workloads disqualify us, but because their unexamined inefficiency does.
The Priesthood Problem
Performance instruments have existed in most ecosystems that matter today for quite a while, and they are mostly free. The barrier that kept most developers from using them was that the knowledge required to find, run, and interpret them was never made legible to people outside a small circle of specialists.
Consider what already sits at our fingertips. The Java world is arguably the most mature: JDK Flight Recorder ships inside the runtime itself—continuous, production-safe recording of allocations, garbage collection behavior, lock contention, and method sampling—analyzed in Mission Control, complemented by async-profiler’s flame graphs and by JMH, the gold standard for statistically honest micro-benchmarking anywhere. A Java team that has never opened a flight recording is ignoring an instrument panel built into its own engine.
The .NET ecosystem tells a parallel story: the runtime emits rich diagnostic events consumed by a coherent family of command-line tools—counters to classify a problem, traces to localize it, heap dumps to inspect it—working identically on a laptop and inside a production container, with PerfView offering the deepest free memory analysis in any ecosystem and BenchmarkDotNet doing for .NET what JMH does for Java.
Go deserves special mention because it shows what happens when measurement is made native to a platform from the beginning. Profiling in Go never became a specialist discipline; pprof is built into the standard library, a profile is an HTTP endpoint away in most services, and flame graphs are a routine artifact of code review in a way they simply are not elsewhere. Because Go’s designers made measurement an ordinary part of the developer’s daily experience rather than an expert practice, a performance culture followed naturally—which suggests that culture in the other ecosystems has been limited by legibility rather than by any shortage of tooling.
Which brings us to JavaScript and Python, where the common perception that such tools do not exist is instructive precisely because it is wrong. Node.js inherits V8’s excellent sampling profiler; browser devtools can attach to a server process and produce CPU profiles and heap snapshots as good as anything in the JVM world; wrappers like clinic.js reduce the workflow to a single command. Python ships a profiler in its standard library and has produced a generation of modern instruments—py-spy, which samples a running production process without code changes; Scalene, whose line-level memory attribution arguably leads every ecosystem; memray for allocation tracing. The tools exist, yet few in those communities know they exist, and fewer still can interpret their output. When a capability is present but illegible to the people who need it, it might as well be absent, and the widespread belief that these ecosystems lack profilers is exactly that phenomenon in its purest form.
Why did this happen? Partly fragmentation—the knowledge lives in aging blog posts, conference talks, and the heads of the two people at any large company who “do performance.” Partly the learning path itself: the honest core of profiling competence is small. One tool to classify (is the problem compute-bound, memory-bound, or waiting?), one tool to localize, one harness to validate—perhaps a day of genuine learning. But nothing made that day legible, and so the material was rationed to teams fortunate enough to contain an apprenticeship. Everyone else provisioned bigger instances.
When Inefficiency Becomes Infrastructure
Unmeasured inefficiency does not remain a property of the code; over time it becomes a property of the infrastructure that hosts the code, recorded and perpetuated in the systems we use to plan capacity.
The mechanism is mundane. Sizing decisions are made empirically: watch the service under load, observe memory climb to six gigabytes with the processor stalling during collection pauses, provision eight with headroom. Nobody asks whether the six gigabytes are intrinsic, or whether four of them are churn and retention that a week of remediation would eliminate. The observed number enters the infrastructure repository as the requirement, future capacity planning extrapolates from it, and autoscaling policies scale outward around per-instance waste rather than confronting it—rendering the inefficiency invisible as well as tolerated. The slop has compounded from the code layer to the architecture layer, where it accrues interest monthly.
For a long while this was a defensible trade. Engineering hours cost more than compute; throwing hardware at the problem was the rational move, and the cloud made the throw frictionless when overprovisioning became a dropdown rather than a purchase order. But the rationality of that trade always depended on remediation being expensive, and it is precisely that premise which is now dissolving.
The most interesting consequences are threshold effects rather than linear ones, and serverless is the clearest case. What keeps workloads out of scale-to-zero deployment models is a specific triad: cold-start latency, memory footprint above the comfortable pricing tier, and execution time per invocation. All three are precisely what unoptimized code inflates. Allocation churn drives working set; working set drives the memory tier; the memory tier drives cost per unit of execution. Startup cost is dominated by loading and just-in-time compilation—which is why ahead-of-time compilation and trimming matter so much at this boundary, turning a service that starts in seconds at hundreds of megabytes into one that starts in tens of milliseconds at a few dozen.
A very large population of services sits in always-on containers, idle ninety-five percent of the time, not because their duty cycles demand residency but because an unoptimized prototype made serverless feel unusable years ago, and the verdict stuck. Optimization at that boundary does far more than shave the bill: it moves the workload across an architectural frontier into a categorically cheaper operational model, with no idle spend, no fleet to patch, and no capacity to plan. The full cost of unexamined inefficiency, then, includes not only the invoices we pay but the architectures we never consider because our own waste has quietly placed them out of reach.
Aggregate the effect across the industry and it becomes an energy story. Sustainability conversations in computing fixate on datacenter efficiency—power usage, cooling, siting—while the software layer above the datacenter is plausibly the larger untapped multiple, because it is not asymptoting toward a physical limit; it has simply gone unexamined. Much of what we could honestly call green computing consists of the unglamorous work of profiling the software we already run and removing the waste we find there.
New Ways to Work
Two things, both recent, both compounding, and both flowing from the affordances of emerging AI technologies, can help us to change our performance situations.
First, the learning path has collapsed. The barrier to profiling competence was never intelligence but legibility—knowing which tool, which command, what the output means, at the moment of need. That is precisely the shape of assistance that conversational AI provides: not “go learn the tracing subsystem,” but here is the one command for your situation, here is what this column means, here is the anomaly in your output. This is tutoring at the moment of need, which is how profiling was always actually learned—apprenticeship with someone who had done it—except the apprenticeship is now available to everyone, in every ecosystem, at once. What protected the priesthood all these years was not the difficulty of the material but the scarcity of teachers, and that scarcity has ended.
Second, remediation itself has become largely delegable. The modern profiling chain in every major ecosystem is command-line driven with parseable output, which makes it agent-legible: an AI agent can install the tools, drive a representative workload, collect a trace, read the hot stacks, form a hypothesis, apply a change, and re-measure—a closed loop with objective feedback, which is exactly where agents do their best work. The common performance pigs carry recognizable signatures and well-known remediations: iteration-scoped allocations that belong in pooled buffers, hidden copies at abstraction boundaries, closures capturing what they should not, materialized sequences that should stream, synchronous waits lurking in asynchronous paths.
Crucially, measurement collapses the trust problem that would otherwise make delegation risky. We need not believe an agent’s reasoning when we can verify its claimed fix against a controlled before-and-after measurement produced by our own harness, and verifying a measurement is far easier than writing optimal code from scratch—an asymmetry that favors the human reviewer at exactly the point where human attention is scarcest. The same capability that generates plausible-but-wasteful code thus inverts into the verification asset that catches it, and because the harness evaluates the code rather than its author, it catches machine-generated and human-generated waste with equal indifference.
We should be honest that this amounts to more than doing our old performance work faster. It changes the nature of the work itself: performance shifts from occasional archaeology, digging when something is slow, into a continuous practice woven into how we build.
The Practice: Installing the Ratchet
None of this self-executes. Tools do not want things; harnesses do not build themselves; budgets do not set themselves. The infrastructure that turns measurement into enforcement has three parts, all now inexpensive, none automatic.
A workload corpus, anchored in reality. We need a curated set of real scenarios at real scale, treated as a first-class asset alongside the code. Not unit tests, not synthetic benchmarks—actual representative work, repeatable and deterministic so regressions are attributable, runnable at one-times, ten-times, and one-hundred-times volume so the shape of the scaling curve is visible rather than a single point. Building this corpus is the one step that genuinely requires our domain knowledge—knowing what representative means for a given system—and it is the step no tool or model can supply for us. It is also, not coincidentally, the prerequisite that makes everything downstream, including agent-driven remediation, trustworthy.
Budgets as gates. Allocation and latency budgets per scenario, enforced in continuous integration, failing the build when an innocent refactor doubles the churn. This is the ratchet that converts performance from an occasional excavation into ordinary regression testing: once a scenario meets its budget, the build system holds that ground permanently, and the codebase can only move in the direction of the standard we have set.
A design posture, not edge-case heroics. Memory efficiency as standard course means thinking in lifetimes: laziness as deliberate policy rather than accident—eager for the always-needed, cached for the expensive-and-sometimes-needed, ephemeral for the cheap; scratch space pooled rather than freshly allocated per iteration; data layouts chosen for how the data is actually traversed. The point of a stated posture is not that every function must be optimal, but that a codebase with an articulated policy drifts far more slowly than one where each component accretes whatever pattern its author reached for.
The Work of Wanting It
Everything described here has been technically possible for years; what the arrival of AI assistance changes is that the cost of doing it has fallen so far that the traditional justifications for not doing it no longer hold. The cost of the infrastructure has fallen to nearly zero. What has not fallen to zero, and never will, is the decision to install it: someone with authority over a codebase deciding that performance is a gated property rather than an aspiration, that a meaningful regression blocks a merge, that the corpus is maintained as the product evolves. That decision is an act of organizational will, and organizational will is the one input no model can generate on our behalf.
This is, in the end, the optimistic reading. For decades the fleet grew around code nobody measured because measuring was genuinely hard and remediating was genuinely expensive, and those who might have willed it otherwise could plausibly say the price was too high. That defense no longer holds, and with the emerging technologies we no longer face a forced trade between velocity and discipline—we can have both. What we choose between now is carrying the unmeasured past forward or building infrastructure sized around what our systems actually require.
Five years from now, the organizations still provisioning around slop will not be the ones that lacked tools, or teachers, or agents to do the work; they will be the ones that never made the decision to install the ratchet. The organizations that thrive will be those that made measurement a norm rather than an exception—that treat every trace as an invitation to learn, every budget as a commitment to relevance, and every workload examined as debt retired. They will spend less, certainly, but the deeper reward is that they will build differently, learn faster, and free themselves to imagine architectures the unmeasured cannot reach.