Compare Code Health Across Multiple Repos: A Dashboard Guide

Swati Ahuja9 min read

compare code health across multiple repos · multi repo code analysis · tracking code health metrics · code health dashboard · portfolio code quality · combine repo data

On this page

We have all built the dashboard with one health number per repo, and it hides the files that matter. To compare code health across many repos, score files inside each repo, then compare each repo's worst files and its trend. Group repos by language first, and track code-level signals such as hotspots and ownership, which DORA metrics don't cover.

ToolMulti-repo viewSignals it comparesUses git historySelf-hostPrice note (6 Oct 2026)
CodeSceneInter-project dashboard across "4 factors"Code health, knowledge distribution, team-code alignment, deliveryYesYes (on-prem)Per active author; see its pricing page
SonarQube ServerPortfolios (Enterprise edition and up)Quality gate status, reliability, security, maintainability ratingsNoYesPortfolios need Enterprise
CodacyOrganization overviewGrade, issues, complex files, duplication, coverage (last 100 updated repos)NoSelf-hosted optionPer-seat plans
Cortex (or another service catalog)Scorecards across servicesWhatever you feed it, e.g. SonarQube metrics, coverage, ownershipVia integrationsNoEnterprise pricing
repowiseTeams portfolio dashboard; workspaces in the open-source CLIHealth score per file, hotspots, co-change, ownership and bus factor, dead codeYesYes (AGPL-3.0)Teams $16 to $20 per seat a month (yearly or monthly), min 3 seats, 25 repos

Scroll the table sideways to see every column.

I work on repowise, so weigh its row accordingly. The method in the rest of this post works whichever tool produces the numbers.

We build repowise, so weigh our entry accordingly. The measurements we publish about it, with their methods, are on the benchmarks page.

Why one score per repo misleads

The first dashboard most of us build has one row per repository and one number per row. The engineer in me finds a table like that very calming, and it hides almost everything useful, because an average flattens the few files you actually care about.

A repository with 2,000 files and an average health of 8 can still contain the twenty files that cause most of its incidents. In our own validation study, bug fixes were concentrated: on a typical repo, 16 of the 20 lowest-health files had received a bug fix in the prior six months, 3.3 times the 24% base rate across all files. In the same study, ranking files by health score predicted which files would get bug fixes over the next six months with a mean ROC AUC of 0.74 across 21 repos (AUC is a 0-to-1 measure of how well a ranking separates buggy files from clean ones, and 0.5 is a coin flip). That is about the same as ranking files by size alone (0.737 vs 0.742); the score's extra value is in naming why a file is risky. The lowest-health files had 2.18 times the defect density of the average file, which is a finding about files that a per-repo average can't show.

So the unit of comparison should be the file, and the repo-level view should summarise how its files are distributed.

What our own index shows about comparing repos

Two things get assumed a lot: that big repos score worse, and that scores are comparable across languages. We have a few thousand public repos indexed on repowise.dev, so I took the popular, non-trivial ones (public, 500+ stars, at least 100 files, a health score on the latest snapshot, and the same repo indexed under two owners counted once), which left 162 repos.

repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code.

The spread across those 162 repos is wide. On the 1-to-10 scale, the 10th percentile is 5.75, the median 7.02 and the 90th percentile 8.59.

Size barely moves the median:

Repo size (files)ReposMedian health10th pct90th pct
100 to 499617.035.938.58
500 to 1,999556.915.688.56
2,000 and up467.066.058.58

Scroll the table sideways to see every column.

Language moves it much more:

Primary languageReposMedian health
C#88.94
Java117.44
TypeScript367.43
Python446.95
Go196.80
JavaScript126.52
Rust105.77

Scroll the table sideways to see every column.

Some of these groups are small (8 C# repos, 10 Rust repos), so one or two repos can move the median. I'd read the language gap as a property of the markers. A health score built from long functions, deep nesting and large files reads languages differently, because idioms that are normal in one language trip a marker in another. For a dashboard, that means part of the gap between a Rust service and a C# service comes from the language, and ranking them on one score blames the team for it.

The practical rule is to compare repos within a language group, and to compare each repo with its own past more than with its neighbours.

Signals for a multi-repo dashboard

DORA metrics (deployment frequency, lead time, change failure rate, time to restore) describe how a team ships. They are useful, and tools like LinearB, Swarmia and Cortex do them well, but they don't say where in the code the risk is. For that you need code-level signals, and these are the ones I'd put on the dashboard, with the question each one answers:

  • Hotspots are files that change often and are complex. Count them per repo and list the top five, which answers "where will the next bug probably be".
  • Hotspot health is the average health of the files that change most. A repo of healthy files with three terrible hotspots is riskier than its overall average suggests.
  • Ownership concentration is how much of the recent change to each hotspot came from one person. Bus factor (how many people would have to leave before nobody knows the code) is the repo-level version, and it answers "if one person is away, who can review this".
  • Hidden coupling is pairs of files that usually change in the same commit without importing each other. These are the changes that break in review because nobody knew to look at the other file.
  • Dead code is unused exports and unreachable files. It is rarely urgent, but a steady rise is a cheap early sign that cleanup has stopped.
  • Trend is each of the above compared with 30 and 90 days ago. A repo moving from 7.0 to 7.4 is heading the right way; a flat 8 may be fine, so judge each repo against its own trend.

I'd leave off lines of code per developer, commit counts per person and any per-engineer score. They push people to game the number, and none of them predict defects.

How to build it

Option 1: one tool that already does it

CodeScene's inter-project dashboard places projects across its four factors so outliers stand out, and because it uses git history, hotspots and knowledge distribution are built in. SonarQube's Portfolios (Enterprise edition and above) aggregate quality gate status and ratings across projects, which suits teams that already work through quality gates, though the view is static and has no history-based signals. Codacy's organization overview compares grade, issues, complex files, duplication and coverage across the repos in one Git provider organization, up to the 100 most recently updated. If you want the git-history signals from this post across private repos without running analyzers yourself, repowise's Teams plan has a portfolio health dashboard across up to 25 repos, with hotspots, owners and trends per repo, refreshed on every push. The open-source CLI also supports workspaces, where several repos are indexed together and an agent can ask questions across all of them.

Option 2: a service catalog fed by your scanners

If you already run Cortex, Backstage or a similar catalog, the cleanest route is to keep code analysis in the scanner and push the numbers in. Cortex's SonarQube integration, for example, puts code smells, bugs, coverage and vulnerabilities on each service and lets you write scorecard rules on them. A catalog is good at answering "which services are missing X", and on code risk it is only as good as the scanner feeding it, which usually doesn't read git history.

Option 3: build it yourself

For a handful of repos, a weekly CI job can run your analyzer of choice on each repo, write a JSON summary (hotspot count, hotspot health, top owners, dead export count) and append it to a table that a spreadsheet or a small dashboard reads. That costs an afternoon and gives you exactly the view you want. It stops being cheap at around 20 repos, when keeping the analyzers configured and the history consistent becomes somebody's job, usually without a ticket.

A review routine that uses the dashboard

A dashboard nobody opens is worse than none, because it gives the impression that someone is watching. The routine I'd suggest is short:

  1. Once a month, sort repos by change in hotspot health over 90 days, within each language group.
  2. For the three biggest drops, open the top hotspots and read the recent commits. Usually it is one feature that grew quickly, which is fine if someone decides it is fine.
  3. For any hotspot where one person did most of the recent work, pair someone else on the next change.
  4. Write down the decision, even when it is "leave it". Next month's drop is easier to read when you know last month's was deliberate.

For the wider picture of what a manager can see from the code, see the engineering manager's guide to codebase visibility. For what goes into a health score, see what is code health, and for whether it predicts bugs, the validation study.

How we measured

The repo figures come from repowise.dev's production database on 6 October 2026: the latest ready snapshot of every public repo with at least 500 GitHub stars and at least 100 files whose snapshot has a health score, with two duplicate entries removed, 162 repos in all. The score is repowise's average file health (1 to 10). Languages are GitHub's primary language for the repo, so a mixed-language repo counts once. The numbers can't tell you whether any repo is well run, because a lower score often reflects a language or a domain (parsers, numerical code) where long functions are normal. Repos we have indexed are not a random sample of open source. The tool comparison was checked against each vendor's docs on the same date.

It ends up as more columns than the one-number dashboard, and every one of them leads to a file someone can open.

FAQ

How do I compare code health across multiple repositories?

Score files inside each repo, then summarise each repo by its hotspots (frequently changed, complex files), the health of those hotspots, ownership concentration and trend. Compare repos within the same language, and compare each repo with its own history before comparing it with others.

What is the best tool for multi-repo code analysis?

It depends on what you already run. CodeScene and repowise use git history, so they show hotspots and ownership across repos. SonarQube Enterprise Portfolios fit teams that manage by quality gates. Codacy's organization overview fits teams already using Codacy. A service catalog like Cortex fits teams that want code metrics next to operational ones.

Which code health metrics should I track over time?

Track hotspot count, average health of hotspot files, ownership concentration on hotspots, hidden coupling (files that change together without importing each other) and dead code. Watch how each one changes over 30 and 90 days, because the trend says more than the absolute value.

How do I combine repository, CI and issue data in one view?

Use a service catalog (Cortex, Backstage) or a small warehouse table that each source writes to on a schedule: code signals from your analyzer, build and test results from CI, incident and bug counts from the issue tracker. Join on repository or service name. Keep the code signals at file level underneath, so a bad number can be traced to files.

Is DORA enough to track code health?

DORA alone isn't enough. It measures delivery (how often you deploy, how long changes take, how often they fail and how fast you recover), and it does not say which files carry the risk. Pair it with code-level signals from git history and static analysis.

Why do repos in different languages get different health scores?

Many health markers, such as function length, nesting depth and file size, count things that are normal in one language and unusual in another. In our index of 162 popular repos the median score ranged from 5.77 for Rust to 8.94 for C#, while repo size barely moved it, so compare within a language.