DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
World desk5 min

Code Got Cheap. Quality Didn’t: Why “AI Makes Software Worthless” Gets the Cost Structure Wrong

AI can reduce the effort of producing code, but that does not make finished software worthless—or prove that its total delivery cost has fallen. Evidence varies by task, developer experience, project maturity, and the work required to review and maintain the result.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

No. AI can make producing a first draft of code cheaper, but a draft is not finished, dependable software. Requirements still have to be met; code must work with the existing system, be reviewed and tested, and remain safe and maintainable. Evidence on AI-assisted coding is mixed across settings, and it does not establish that software—or its economic value—has become worthless.

Cheap code is not the same as cheap software

“Software” can mean a generated code snippet, a working feature, or a system people can rely on and change over time. Those are not interchangeable outputs. A model may reduce the effort needed to produce a draft without reducing the effort required to determine whether that draft is correct, secure, compatible with its surroundings, and economical to maintain.

That distinction matters because generating code is only one part of delivery. Teams also need to clarify requirements, integrate changes, run tests, inspect security and other non-functional properties, review the result, and handle defects or future changes. If a draft needs extensive correction, the time saved at the start may be partly or wholly offset later. The available studies do not establish one universal breakdown of software’s lifecycle costs, so a fixed percentage for “coding” versus “everything else” would be misleading.

What the evidence says about productivity

The studies below examine different populations, tasks, and outcomes. Their percentages should not be averaged or treated as competing estimates of one universal AI productivity effect.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Study and setting Reported result What it does—and does not—show
Xu, Medappa, Tunç, Vroegindeweij, and Fransoo, 2025; analysis of open-source projects after GitHub Copilot adoption Core developers reviewed 6.5% more code after adoption, while their original-code productivity fell 19%. The authors report productivity increases concentrated among less-experienced peripheral contributors, alongside more review and maintenance burden for core developers. This is evidence from the studied OSS projects, not a universal estimate for company teams or all tasks. The university portal describes the output as a peer-reviewed conference contribution and records a submitted status dated July 16, 2025.
Becker, Rush, Barnes, and Rein / METR, 2025; randomized trial with 16 experienced developers completing 246 tasks in mature projects they already knew For the early-2025 AI tools tested, task completion took 19% longer when AI tools were allowed. Participants had expected the tools to reduce their time. This small, specialized trial measured task completion time in familiar, mature projects. The authors say experimental artifacts cannot be entirely ruled out. It does not predict results for novices, greenfield projects, later tools, or every kind of software work.

The contrast is informative: an AI-assisted workflow may help one group or task while shifting work to another role or slowing a different kind of task. “Productivity” therefore needs an explicit unit—such as a task completed to requirements, rather than lines produced—and a clear account of whose work is counted.

Quality has more than one dimension

Working output is not automatically good output. Correctness asks whether the program does what the requirements demand. Security asks whether it introduces exploitable weaknesses. Complexity and maintainability ask whether people can understand, verify, and safely change it later. These properties can trade off against speed, and none can be inferred from how quickly code appeared.

A 2024 peer-reviewed evaluation by Liu, Tang, Luo, Zhou, and Zhang tested ChatGPT-generated code across defined algorithm and software-weakness scenarios. It assessed correctness, complexity, and security, finding relevant vulnerabilities in some tested scenarios and variation attributable to nondeterminism. The authors also reported limits in direct repair ability in their multi-round fixing setup. In the study’s vulnerability scenarios, more than 89% of vulnerabilities were successfully addressed through that multi-round fixing process; this is a result of that evaluation, not a general production-code security rate.

The same study reported a 48.14 percentage-point accepted-rate advantage on problems from before 2021 compared with problems after 2021. That is a difference within its ChatGPT coding benchmark, not a 48.14% general improvement in coding performance. Benchmark results depend on the tasks and evaluation design, and do not establish the quality rate of current models in deployed systems.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why organizational conditions change the result

DORA’s 2025 report summarizes AI as an amplifier of an organization’s existing strengths and weaknesses, and says the greatest returns come from strategic attention to the underlying organizational system rather than tools alone. This is DORA’s report-level conclusion, not a guarantee of identical effects at every company or a standalone causal estimate.

In practical terms, an organization that can define work clearly, test changes reliably, review code promptly, and learn from failures is better positioned to tell useful assistance from plausible but flawed output. Weak practices can make defects, rework, and review queues harder to detect or absorb. A tool’s contribution therefore depends not just on what it generates, but on the delivery system that checks and supports it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to tell whether AI is lowering your software costs

Evaluate a representative workflow end to end, with a comparison that includes the work displaced or added—not just the time to first draft. Use tasks drawn from the codebase and team that will actually use the tool. Define completion before the comparison begins, and count the people who review or repair the output as well as the person prompting the model.

  • Time to a completed task: measure elapsed or labor time through review, tests, integration, and necessary rework, not only initial generation.
  • Correctness: check requirements and relevant tests, including whether fixes introduce regressions.
  • Security and other required properties: assess the risks relevant to the feature rather than assuming a passing functional test is enough.
  • Review and repair burden: record who handles additional review, corrections, and maintenance work, so apparent savings are not simply shifted to another role.
  • Maintainability in context: examine whether the change fits the project’s architecture and can be understood and modified by its maintainers.
  • Conditions of use: separate results by task type, developer experience, project maturity, and delivery practices; one aggregate result can conceal meaningful differences.

A useful comparison is the full cost and outcome of completing equivalent work with and without AI assistance. This is an analytical decision rule, not a validated universal formula. A faster draft that fails requirements or creates expensive review work is not a productivity gain; a slower first pass may still be worthwhile if the delivered result is better or easier to maintain. Measure those outcomes rather than assuming either result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What remains unresolved

The cited evidence does not settle the long-run effect of AI on software prices, vendor margins, labor demand, or the economy-wide value of software. Nor does it rank today’s coding tools head to head. The studies cover different periods and settings, including early-2025 tools in the METR trial and benchmark-specific ChatGPT scenarios in the 2024 quality evaluation. Broad predictions about software becoming free, developers becoming unnecessary, or generated code always being unsafe go beyond what these findings support.

The narrower conclusion is more useful: AI can lower the effort of producing some code, but that alone cannot establish lower total delivery cost or lower software value. Those depend on whether teams can turn generated output into correct, secure, reviewable, maintainable results—and on the additional work required to do so.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.