Technical Debt Measurement Tools Compared: What Each One Measures, and Whether It Predicts Bugs

Raghav Chamadiya10 min read

technical debt measurement tools · technical debt measurement · code climate alternatives · codacy alternatives · automated codebase health scores · sonarqube technical debt ratio

On this page

Technical debt measurement tools fall into two groups. SonarQube, Qlty (formerly Code Climate Quality), Codacy and NDepend count rule violations, estimate fix time and turn the ratio into a letter grade. CodeScene and repowise score files 1 to 10 and add git history. Only the second group publishes evidence that its score points at files that later get bug fixes.

ToolWhat it countsScore you seeReads git history?Published evidence the score points at bugs
SonarQube ServerRule violations; debt = minutes to fix maintainability issuesDebt ratio = debt / (30 min x lines of code); rating A to E (A is 5% or less)No, apart from its "new code" scopeNot from the vendor. Two 2019 to 2020 academic studies (Lenarduzzi et al.) on Java rules found weak links to faults
CodeScene25+ code factors, plus change frequency for hotspotsCode Health 1 to 10 per file; hotspot healthYesCode Red paper (2022): 39 proprietary codebases, 15x more defects in low-quality code
Qlty (formerly Code Climate Quality)Duplication and structure smells, cognitive and cyclomatic complexity, coverageDebt ratio against a COCOMO estimate; grade A to F (A under 5%)NoNone cited on the pages we checked
CodacyIssues, cyclomatic complexity, duplication, coverageGrade A to F, a weighted average of those fourNoNone cited on the pages we checked
NDepend (.NET)Rule issues, each with debt and yearly "interest" in man-timeDebt ratio and SQALE rating A to E; breaking point per issueCompares against a baselineNone cited on the pages we checked
repowise25 deterministic defect markers, several read from git history (churn, author count, change entropy)Health 1 to 10 per file, in three pillars: defect, maintainability, performanceYesOpen study: 21 repos, 9 languages, ROC AUC 0.74 [95% CI 0.68 to 0.79]; ties line count on AUC

Scroll the table sideways to see every column.

Two ways to put a number on technical debt

Most tools on this list descend from one idea, the SQALE method: give every rule violation a time-to-fix, add them up, and divide by an estimate of how long the code took to write. The result is a debt ratio, with a letter grade on top. It is a cost model that answers "how much cleanup is in here, relative to the size of the code?"

The other approach starts from a different question: which files will hurt us? It still looks at the code, but it also reads version control, because a complicated file nobody has touched in three years costs you nothing today, while a complicated file that changes every week costs you on every change. CodeScene built its product around that idea, which it calls hotspots. repowise uses the same family of signals.

Both approaches are useful, as long as you don't read one as if it were the other.

SonarQube: debt ratio and the maintainability rating

SonarQube's debt model is the one most people mean by "technical debt score". Its docs define technical debt as the sum of remediation costs of maintainability issues, in minutes, taken from the effort assigned to each rule. The technical debt ratio is that debt divided by the cost to develop one line of code times the number of lines, where one line defaults to 30 minutes. The default rating scale runs A (5% or less) through E (50% or more), and you can change the grid.

Its strengths are a large rule set, quality gates in CI, and a clear split between overall code and new code, so a team can hold the line on new changes without fixing ten years of history first.

Sonar does not present the ratio as a bug forecast, but independent researchers have tested that link anyway. A 2020 SANER paper on 21 Java projects found that among 202 SonarQube Java rules, only 25 had relatively low fault-proneness, and that violations SonarQube labels as bugs were generally not fault-prone. A second study by the same group, on 33 Apache projects, found no difference in fault-proneness between classes with and without technical debt items. Both used older rule sets, so treat them as a reason to check, and don't read them as a verdict on today's product.

CodeScene: code health and hotspots

CodeScene scores code from 1 to 10, where 10 is very maintainable, based on 25+ factors scanned from the source. It then weights that by where development actually happens: hotspot health is the code health of the files that change most. That combination is the core of its product.

It also has the strongest published evidence of the commercial tools here. The Code Red paper by Adam Tornhill and Markus Borg (TechDebt 2022) analysed 39 proprietary production codebases and 30,737 files, and found that low-quality code contains 15 times more defects than high-quality code, that resolving issues in it takes on average 124% more development time, and that maximum cycle times are 9 times longer. The data is proprietary, so you cannot rerun it, but it is peer-reviewed.

Qlty, formerly Code Climate Quality

If you are looking for Code Climate alternatives, start here: Code Climate Quality is now Qlty. In November 2024 Code Climate announced that its Quality product had been spun out into a new company, Qlty Software, with Code Climate itself focusing on its engineering-intelligence product, Velocity. Qlty's migration guide says the Code Climate API was disabled on July 18, 2025, and that paid Quality accounts get a migration timeline from Qlty.

Qlty measures maintainability through code smells in two categories, duplication and structure, each with an estimated fix effort in minutes. It uses cognitive complexity by default and reports cyclomatic complexity too. Technical debt is the sum of that effort; the debt ratio divides it by a Basic COCOMO estimate of the effort to write the code from scratch. Project grades run A (under 5%) to F (50% or more), and file grades use a logarithmic function of total debt. It is the same family as SonarQube, with a CLI (qlty smells) and, per its migration guide, broader language support than the old Quality product.

Codacy

Codacy reports four metrics: issues in ten categories (code style, error prone, code complexity, performance, compatibility, unused code, security, documentation, best practice and comprehensibility), cyclomatic complexity, duplication and coverage. Its grade is a weighted average of those four, from A to F, with thresholds you configure, such as when a file counts as complex or duplicated. Its strength is putting issues, complexity, duplication and coverage behind one grade and one PR check.

NDepend

NDepend is for .NET, and it has the most detailed cost model on this list. Every issue carries a technical debt (man-time to fix) and an annual interest (man-time lost per year if you leave it). Dividing one by the other gives a breaking point: if it is under a year, fixing the issue pays for itself within twelve months. The debt ratio and SQALE rating use the same A to E bands as SonarQube.

repowise

repowise scores every file from 1 to 10 using 25 deterministic markers for the headline defect score, with no model in the loop, so the same code always gets the same score. Some markers read the code's structure, such as nesting and function size; others read git history, such as churn, author count and change entropy. The score is split into three pillars, defect, maintainability and performance, so a performance finding never moves the defect number. All of it is open source under AGPL-3.0.

I build repowise, so read the evidence below with that in mind, and note that it includes the parts that don't flatter us.

Does a technical debt score predict bugs?

Two tools in this comparison publish evidence. For CodeScene it is the Code Red result above, where low-quality code had 15 times more defects across 39 proprietary codebases.

repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code. We ran an open, reproducible study, written up in Does code health predict bugs?. Across 21 open-source repositories, 9 languages and 2,770 labelled files, ranking files by health score reached a cross-project mean ROC AUC of 0.74 [95% CI 0.68 to 0.79] at predicting which files got a bug fix in the following six months. ROC AUC is the chance that a random buggy file scores worse than a random clean one: 0.5 is a coin flip, 1.0 is perfect. Scores were taken at a past commit, so no later fix could leak into them.

The caveats are as important as the headline:

  • On raw AUC the score ties a plain line-count baseline, 0.737 against 0.742, so it is about as good as ranking files by size alone; its extra value is in naming why a file is risky. Big files carry more bugs, and AUC rewards that. The score pulls ahead only when you account for reading effort (an effort-aware measure called Popt), where it beats line count by +0.134.
  • It is near-blind on small files. Within the smallest size bands most of the signal disappears, because a tidy 40-line file gives the markers nothing to fire on.
  • On pure triage ordering, "re-inspect what broke before" still wins on Popt, 0.609 against 0.524.
  • Per-repo results ranged from 0.55 to 0.86, so no single repo's number generalises.

A simpler version of this check runs on the repos we index. Across 899 public repos with enough fix history, the 20 lowest-scored files had a bug fix in the previous 180 days at a median of 3.56 times the base rate; pooled, 63% of flagged files had one against a 19% base rate. That is a look-back, not a forecast, so I would say "3.3 to 3.6 times better than picking files at random" and no more.

Which technical debt tool should you use?

Pick by the question you need answered.

  • "How much cleanup is there, and is new code adding to it?" Use SonarQube, Qlty or Codacy. Debt ratios and quality gates are built for this, and the new-code view is very practical.
  • "Which fixes pay for themselves first?" Use NDepend's interest and breaking-point model, if you are on .NET.
  • "Where will the next bug or slowdown come from?" Use a tool that reads history: CodeScene or repowise. Choose CodeScene for a mature commercial product with published proprietary evidence; choose repowise if you want the scoring open and the study reproducible on your own repo.

Whatever you pick, turn the score into a short ranked list. How to prioritize technical debt covers ranking by impact per effort, and the best code health tools in 2026 covers more tools in this space. For a deeper look at the rule-based tools, see SonarQube alternatives.

How we measured

Every description of a third-party tool comes from that vendor's own docs or announcement, read on 2026-10-06; the academic results come from the papers' published abstracts. We did not run SonarQube, Qlty, Codacy or NDepend on a shared corpus for this post, so the table compares how each tool defines its score and says nothing about accuracy. The repowise study figures come from our published 21-repo study (measured August 2026). The 899-repo figures come from our production index on 2026-10-06 and are a 180-day look-back on each repo's own history.

FAQ

What is the best tool to measure technical debt?

The answer depends on what you mean by debt. For a cost estimate of cleanup with CI gates, SonarQube and Qlty are the standard choices, and NDepend has the most detailed cost model for .NET. For finding the files most likely to cause bugs, tools that combine code analysis with git history, such as CodeScene and repowise, have published evidence behind them.

What happened to Code Climate Quality?

Code Climate spun its Quality product out into a new company, Qlty Software, announced in November 2024, and Code Climate now focuses on its Velocity engineering-intelligence product. The Code Climate API was disabled on July 18, 2025, and paid Quality accounts receive a migration timeline to Qlty.

What is the SonarQube technical debt ratio?

It is the estimated time to fix all maintainability issues divided by the estimated time to write the code, where each line defaults to 30 minutes. The maintainability rating converts that ratio to a letter: A for 5% or less, then B, C and D, and E for 50% or more. You can change both the cost per line and the rating grid.

Are automated codebase health scores reliable?

They are consistent, since most are deterministic. Debt ratios reliably measure rule violations against size. Health scores that read git history have shown a measurable link to future bug fixes, such as ROC AUC 0.74 across 21 repos in our study, but even those tie line count on raw AUC and miss small buggy files. Use them to rank where to look, and don't treat the score as a verdict on any single file.

What are good Codacy alternatives for technical debt?

For the same kind of grade and PR checks, SonarQube and Qlty are the closest matches. If you want the score to reflect how often code changes and who owns it, look at CodeScene or repowise. For .NET-only teams, NDepend adds a debt-and-interest cost model.