Code health

Find the files most likely to break. Before they do.

Every file gets a 1 to 10 score checked against real bug fixes, the specific refactor that would help, and a map of where the risk sits.

Deterministic. The same source and history always score the same.

See the score on a real repo
0.74
bug-prediction AUC
n = 2,826 · 21 repositories, 9 languages, 2,826 files · 2026-08
2.3x
more defects found
n = 2,770 · p = 0.003 · 2,770 files scored by both tools at the same leakage-free commit · 2026-08
3.3x
self-check lift
16 of the 20 lowest-health files had a bug fix in the prior six months in the published self-check.
25
deterministic markers
Structural, historical, ownership, test, and defect-history signals.

In plain words

One number, 1 to 10, for how likely a file is to produce the next bug.

Every file gets a score from deterministic markers of its structure, its history and who owns it. The weights were learned from real bug fixes, so a low score has predicted defects before. File size is controlled for, so a big file is not marked down for being big.

How to read it

What's a good score?

8.5 and up · Excellent
Easy to change safely. Leave it alone.
7 to 8.5 · Good
Healthy. A seven is not a warning.
5.5 to 7 · Fair
Watch the trend. Refactor it if it is also a hotspot.
4 to 5.5 · Needs work
Changes here are the ones that come back as bug fixes. Plan the refactor.
Under 4 · At risk
Where the repository's bug fixes concentrate. Start here, with the named fix.

Sound familiar?

Every tool says everything is wrong.

  • 01

    “The linter flags five thousand things.”

  • 02

    “Big files always look worst, whether or not they break.”

  • 03

    “I cannot tell risky from merely ugly.”

What you get

A score that means something, and the next thing to do about it.

  1. Defect risk

    A 1 to 10 score checked against real bug fixes.

    Weights are learned from a defect corpus, so a low score has predicted real bugs before.

  2. Refactoring plans

    The specific fix, and what moves with it.

    Extract this, split that, break this cycle: named, ranked by impact for effort.

  3. Performance risk

    N+1 queries and I/O in loops, traced across files.

    Few findings, high precision, kept separate from the defect score.

  4. Trends

    Know when a file starts declining.

    Alerts when health drops, and coverage joined in to find untested hotspots.

On a real repository

From one score to the exact file behind it.

The code health page keeps the repository score, its three dimensions, the health map and the ranked findings in one reading flow, so a reviewer can go from the headline number to one file without losing the context that made the number meaningful.

Capture: Repowise indexed at 3,687 files. The visible score and finding counts belong to that snapshot.

Repowise Code Health showing defect risk, maintainability, performance risk, hotspot health, repository scope, and a file map.
Code Health for the Repowise repository, with scope and interpretation adjacent to the score.

How it works

Measure, calibrate, act.

Three signals stay separate until the evidence earns a conclusion, then the score points at a file and a fix.
  1. 01 / Measure

    Extract deterministic markers

    Complexity, nesting, cohesion, clones, churn, ownership, prior defects and coverage.
  2. 02 / Calibrate

    Weigh them against real defects

    Offline learning against bug-fix history. Only the constants ship.
  3. 03 / Act

    Go from score to worklist

    Health map, ranked findings, and the refactor to do next.

Where it shows up

What changes about your day.

  • The repository score, its three dimensions, the health map and the fix-next queue in one reading flow. Trends and alerts when a file starts declining.

    Dashboard
  • One call returns a file's score, its markers and the refactor to do next. Copy a plan for the agent and it knows what moves with it.

    Coding agents
  • Every pull request shows the health delta of the files it touched, marker by marker, and whether the AI-written part moved it.

    Pull requests

Honest limitations

One number is not enough, so there are four.

The average hides the file you edit every day, so the page keeps hotspot health, the worst file and maintainability beside it. The score ranks files by the odds of a later fix; it is not a verdict on any one file. Performance findings are deliberately few and do not replace profiling.

Common questions

Everything people ask before they try it.

What does the code-health score actually measure?

Every file gets a single 1 to 10 score computed from 25 deterministic markers: McCabe complexity, deep nesting, brain methods, class cohesion (LCOM4), god classes, Rabin-Karp clone detection, change entropy, co-change scatter, ownership dispersion, prior-defect history, test-quality smells, and more. Lower scores mean the file is more likely to harbor defects.

How do you know the score actually predicts bugs?

It is defect-validated. On the same 2,770 files across 9 languages with real defect labels, ranking by repowise health surfaces 2.3x the defects of a leading commercial tool under the same review budget (recall 0.173 vs 0.074, effort-aware Popt 0.607 vs 0.462, ROC AUC 0.731 vs 0.705). Across 21 open-source repos the mean cross-project ROC AUC is 0.74 with a 95% confidence interval of 0.68 to 0.79, ranging from 0.55 to 0.86 across individual repos.

Does it use an LLM?

No. Scoring is fully deterministic: 25 markers with weights calibrated offline against a defect corpus. Only the learned constants ship, so the same code always produces the same score, in under 30 seconds on a 3,000-file repo. No cloud, no API calls, no drift.

How is this different from CodeScene?

Both score code health, but repowise is open source so every heuristic is inspectable and reproducible on your own repo, and the score ships inside a broader platform: an architecture-aware wiki, git intelligence, architectural decisions, agent provenance, and ten MCP tools for AI agents. repowise does not offer AI auto-refactoring.

Will it just flag big files?

No. The discrimination survives controlling for file size (partial Spearman rho of -0.16) and significantly out-discriminates both recent churn (+0.10 AUC) and prior-defect history (+0.12 AUC), with DeLong p below 1e-9.

Can it use my test coverage?

Yes. repowise ingests LCOV and Cobertura coverage to compute untested-hotspot risk (the intersection of low coverage and high hotspot score), alerts when a file's health starts declining, and ranks refactoring targets by impact for effort.

Can I prove these numbers on my own codebase?

Yes, that is the point. Every heuristic is open source under AGPL-3.0, and the validation runs on your own repo. On a typical project, 16 of the 20 lowest-health files had a bug fix in the last 6 months, 3.3x the 24% baseline.

Does repowise check performance?

Yes, as a static health pillar, not as an APM or profiler. repowise statically detects performance-risk shapes such as N+1 access and IO-in-loop patterns and scores them as a co-equal third pillar alongside defect risk and maintainability. It is high precision and low recall by design: it raises few findings, but the ones it raises are real. There is no runtime, no agent, and no tracing; it reads your source, like every other signal.

Does repowise refactor my code for me?

No, and that is deliberate. repowise ranks refactoring targets by impact for effort and alerts you when a file's health starts declining, but it hands you a human-readable, deterministic worklist; it does not open PRs or rewrite your code. The suggestions are template-based and inspectable, with no LLM editing your repo.

Last reviewed: October 2026

See whether the score finds your risky files.

Index a public repository free and check the lowest-scoring files against the ones that broke last quarter.