Software teams once argued over tabs versus spaces. Now they debate something far more expensive. How much mess can a codebase absorb before changes become painful? One developer’s recent analysis offers a blunt answer. Track how much code grows when requirements stay the same.
Quantifying the Cost of Carelessness
That developer, writing on the personal site Earendil, calls it measuring code sloppiness. The core idea lands hard. When a feature request triggers a 300-line patch instead of a 30-line one, something has gone wrong. The author tested this on real projects. He found that sloppiness shows up first in unnecessary complexity, duplicated logic, and functions that do too many things at once.
But the post goes beyond anecdote. It proposes concrete signals. Lines of code added per feature. Cognitive complexity scores. The ratio of new functions to modified ones. These numbers, tracked over time, paint a picture of decay that traditional bug counts miss. Short. Direct. And surprisingly predictive.
Engineers have chased quality metrics for decades. Cyclomatic complexity. Test coverage. SonarQube debt estimates. Most teams ignore them after the first demo. The numbers feel abstract. They rarely tie back to payroll or delivery dates.
Yet something shifted in 2025. AI coding tools began writing more than half the code in many organizations. And the old metrics started flashing red. Not because the models failed tests. Because the code they produced bloated repositories and resisted change.
A September 2026 study from researchers Yuecai Zhu, Nikolaos Tsantalis, and Peter C. Rigby examined LLM and agent-driven development. Titled “AI-Generated Smells: An Analysis of Code and Architecture in LLM and Agent-Driven Development”, it delivered a clear warning. As models grow more capable, they generate larger, more coupled code. Code volume becomes a near-perfect predictor of structural degradation. Neither correctness nor clever prompts stop the slide.
The authors named it the Volume-Quality Inverse Law. Bigger repositories from agents show more architectural smells. Long methods. Poor encapsulation. Temporal fields that mix state badly. These patterns differ from typical human mistakes. They carry a machine signature.
But the problem isn’t limited to research labs. Real teams see it daily. A GitClear analysis of hundreds of millions of lines, cited in The Hans India on September 10, 2026, found duplicate code blocks exploded eightfold since AI assistants went mainstream. Active refactoring collapsed. Production incidents per pull request rose even as output surged.
And. The code often passes tests. It compiles. It ships. Then six months later, someone tries to add one field and the whole module cracks open. That’s the slop hangover.
Developers coined terms for it. Vibe coding. Workslop. SlopCode. All describe the same phenomenon. Plausible but shallow implementations that lack deep understanding of the domain. A September 11, 2026, SD Times article by Eli Lopian explains why coverage numbers mislead here. Tests may execute the lines but fail to assert meaningful behavior. Mutation testing, which injects faults and checks if tests catch them, reveals the gap. Many suites with 85% coverage detect fewer than 60% of mutations.
So teams need better signals. The Earendil post suggests starting simple. Measure change in lines of code for stable features. Watch for functions whose complexity grows faster than their inputs. Track duplication across commits. These capture what static analyzers often miss in AI output: the gradual accumulation of unnecessary abstraction.
Recent benchmarks echo this. SlopCodeBench, detailed in an August 2026 research summary, evaluates iterative code quality beyond pass rates. It tracks structural erosion and verbosity across agent trajectories. Agents consistently produce more verbose code that erodes faster than human-written equivalents. Human repositories stay relatively flat. Agent checkpoints show 43% median verbosity growth and erosion in 79% of cases.
The pattern repeats across studies. A Jellyfish report from late August 2026 lists core metrics for AI code quality. AI code rework rate. Cyclomatic complexity trend. Change failure rate paired with defect density. Security vulnerability density. Mutation score. Time in review. Each targets a failure mode unique to generated code. Verbose boilerplate. Shallow tests. Hidden dependencies.
BlueOptima’s September 8, 2026, analysis “AI Coding Tool Cost Is Getting Harder to Forecast. Here’s What To Measure.” ties it to dollars. Token costs matter less than the downstream maintenance burden. Poor quality code increases review time, raises operational risk, and slows both humans and future agents. Leaders must track maintainability, complexity, rework, and repair patterns as part of the cost model.
But. Not every organization measures the same way. A Catio guide published in May 2026 recommends 12 metrics across code, architecture, and team layers. It translates them into dollar estimates of technical debt. The goal is a dashboard that executives actually read. One that lets CTOs prioritize debt repayment against new features.
The Drexel University multivocal review from June 2026, “Faster Code, Deeper Debt?”, catalogs the new debts AI introduces. Fast-integration debt. Prompt debt. Provenance debt. Ethical and data concerns layered on top of classic code and design smells. No standardized LLM-specific metrics exist yet. That gap leaves teams flying partially blind.
Tools are emerging to fill it. The npm package code-gauge, updated in early September 2026, ranks files by refactoring priority using cognitive complexity, nesting, duplication, and other signals. It deliberately limits output so agents don’t wander into unrelated improvements. Designed explicitly for AI-agent workflows.
Still, metrics alone won’t solve the problem. Brian Houck’s September 2, 2026, newsletter for GetDX describes a quality paradox. AI can improve some maintainability scores in aggregate while hiding deeper issues in shared team understanding and intent. Much of what makes code good lives outside the file. In the conversations, the diagrams, the decisions not yet codified.
Margaret-Anne Storey’s triple debt model separates technical debt in code from cognitive debt in the team and intent debt in documentation. As AI writes more, that separation grows. Fifty-three percent of code was AI-authored in Q2 2026, up sharply from prior quarters. Developers translate less of their knowledge directly into the artifacts.
This creates new risks. Error-masking code that catches exceptions without addressing root causes rose 47% in AI-assisted commits, according to GitClear data. Duplicate implementations multiply. Reuse drops. Legacy refactoring nearly disappears.
The industry response splits in two directions. Some push for stricter gates in benchmarks. VulcanBench-SWE v4, discussed on X in early September 2026, now weights code quality at 20% of the score. A multi-model judge evaluates for slop. Others focus on human oversight. Reviews that examine not just correctness but structure, test depth, and future change cost.
The Earendil analysis offers a practical starting point. Calculate a sloppiness score per change. Compare it against historical baselines for that codebase. Flag outliers for deeper review. Over time, the trend reveals whether velocity comes at the expense of sustainability.
Because speed without discipline produces volume. Volume without quality produces debt. And that debt compounds faster when machines generate it at scale. Teams that measure sloppiness now may avoid the painful refactoring marathons later.
Mutation testing in 2026, covered by AstaQC on September 11, shows one path forward. It exposes tests that execute code without truly verifying it. Combine that with complexity trends and rework rates, and a clearer picture emerges. One that goes beyond pass-fail to actual maintainability.
The data keeps arriving. LeadDev’s July 2026 summary of GitClear and GitKraken research reports maintainability plummeting in the AI era. Duplication up 81%. Reuse down 70%. These aren’t theoretical concerns. They show up in incident reports and engineer burnout.
So what now? Start with the basics the Earendil post recommends. Track size changes against functional stability. Measure complexity growth. Look for duplication spikes after AI-heavy sprints. Layer on mutation scores and architectural smell detectors tuned for machine-generated patterns.
The goal isn’t perfect code. It’s code that stays understandable six months from now. Because the next feature always arrives. And the team that can add it without rewriting half the system wins.
The Slop Score: Why AI Is Flooding Codebases With Maintainable Nightmares first appeared on Web and IT News.
