Code health
Find the files most likely to break. Before they do.
Every file gets a 1 to 10 score checked against real bug fixes, the specific refactor that would help, and a map of where the risk sits.
Deterministic. The same source and history always score the same.
See the score on a real repo- 0.74
- bug-prediction AUC
- n = 2,826 · 21 repositories, 9 languages, 2,826 files · 2026-08
- 2.3x
- more defects found
- n = 2,770 · p = 0.003 · 2,770 files scored by both tools at the same leakage-free commit · 2026-08
- 3.3x
- self-check lift
- 16 of the 20 lowest-health files had a bug fix in the prior six months in the published self-check.
- 25
- deterministic markers
- Structural, historical, ownership, test, and defect-history signals.
In plain words
One number, 1 to 10, for how likely a file is to produce the next bug.
Every file gets a score from deterministic markers of its structure, its history and who owns it. The weights were learned from real bug fixes, so a low score has predicted defects before. File size is controlled for, so a big file is not marked down for being big.
defect risk
Good · up 0.6
health map
fix next · payments/processor.py
Extract the 3 nested branches. 9 bug fixes here in 6 months.
How to read it
What's a good score?
- 8.5 and up · Excellent
- Easy to change safely. Leave it alone.
- 7 to 8.5 · Good
- Healthy. A seven is not a warning.
- 5.5 to 7 · Fair
- Watch the trend. Refactor it if it is also a hotspot.
- 4 to 5.5 · Needs work
- Changes here are the ones that come back as bug fixes. Plan the refactor.
- Under 4 · At risk
- Where the repository's bug fixes concentrate. Start here, with the named fix.
Sound familiar?
Every tool says everything is wrong.
- 01
“The linter flags five thousand things.”
- 02
“Big files always look worst, whether or not they break.”
- 03
“I cannot tell risky from merely ugly.”
What you get
A score that means something, and the next thing to do about it.
- Defect risk
A 1 to 10 score checked against real bug fixes.
Weights are learned from a defect corpus, so a low score has predicted real bugs before.
- Refactoring plans
The specific fix, and what moves with it.
Extract this, split that, break this cycle: named, ranked by impact for effort.
- Performance risk
N+1 queries and I/O in loops, traced across files.
Few findings, high precision, kept separate from the defect score.
- Trends
Know when a file starts declining.
Alerts when health drops, and coverage joined in to find untested hotspots.
On a real repository
From one score to the exact file behind it.
The code health page keeps the repository score, its three dimensions, the health map and the ranked findings in one reading flow, so a reviewer can go from the headline number to one file without losing the context that made the number meaningful.
Capture: Repowise indexed at 3,687 files. The visible score and finding counts belong to that snapshot.

How it works
Measure, calibrate, act.
- 01 / Measure
Extract deterministic markers
Complexity, nesting, cohesion, clones, churn, ownership, prior defects and coverage. - 02 / Calibrate
Weigh them against real defects
Offline learning against bug-fix history. Only the constants ship. - 03 / Act
Go from score to worklist
Health map, ranked findings, and the refactor to do next.
Where it shows up
What changes about your day.
The repository score, its three dimensions, the health map and the fix-next queue in one reading flow. Trends and alerts when a file starts declining.
DashboardOne call returns a file's score, its markers and the refactor to do next. Copy a plan for the agent and it knows what moves with it.
Coding agentsEvery pull request shows the health delta of the files it touched, marker by marker, and whether the AI-written part moved it.
Pull requests
Honest limitations
One number is not enough, so there are four.
The average hides the file you edit every day, so the page keeps hotspot health, the worst file and maintainability beside it. The score ranks files by the odds of a later fix; it is not a verdict on any one file. Performance findings are deliberately few and do not replace profiling.
Common questions
Everything people ask before they try it.
What does the code-health score actually measure?
Every file gets a single 1 to 10 score computed from 25 deterministic markers: McCabe complexity, deep nesting, brain methods, class cohesion (LCOM4), god classes, Rabin-Karp clone detection, change entropy, co-change scatter, ownership dispersion, prior-defect history, test-quality smells, and more. Lower scores mean the file is more likely to harbor defects.
How do you know the score actually predicts bugs?
It is defect-validated. On the same 2,770 files across 9 languages with real defect labels, ranking by repowise health surfaces 2.3x the defects of a leading commercial tool under the same review budget (recall 0.173 vs 0.074, effort-aware Popt 0.607 vs 0.462, ROC AUC 0.731 vs 0.705). Across 21 open-source repos the mean cross-project ROC AUC is 0.74 with a 95% confidence interval of 0.68 to 0.79, ranging from 0.55 to 0.86 across individual repos.
Does it use an LLM?
No. Scoring is fully deterministic: 25 markers with weights calibrated offline against a defect corpus. Only the learned constants ship, so the same code always produces the same score, in under 30 seconds on a 3,000-file repo. No cloud, no API calls, no drift.
How is this different from CodeScene?
Both score code health, but repowise is open source so every heuristic is inspectable and reproducible on your own repo, and the score ships inside a broader platform: an architecture-aware wiki, git intelligence, architectural decisions, agent provenance, and ten MCP tools for AI agents. repowise does not offer AI auto-refactoring.
Will it just flag big files?
No. The discrimination survives controlling for file size (partial Spearman rho of -0.16) and significantly out-discriminates both recent churn (+0.10 AUC) and prior-defect history (+0.12 AUC), with DeLong p below 1e-9.
Can it use my test coverage?
Yes. repowise ingests LCOV and Cobertura coverage to compute untested-hotspot risk (the intersection of low coverage and high hotspot score), alerts when a file's health starts declining, and ranks refactoring targets by impact for effort.
Can I prove these numbers on my own codebase?
Yes, that is the point. Every heuristic is open source under AGPL-3.0, and the validation runs on your own repo. On a typical project, 16 of the 20 lowest-health files had a bug fix in the last 6 months, 3.3x the 24% baseline.
Does repowise check performance?
Yes, as a static health pillar, not as an APM or profiler. repowise statically detects performance-risk shapes such as N+1 access and IO-in-loop patterns and scores them as a co-equal third pillar alongside defect risk and maintainability. It is high precision and low recall by design: it raises few findings, but the ones it raises are real. There is no runtime, no agent, and no tracing; it reads your source, like every other signal.
Does repowise refactor my code for me?
No, and that is deliberate. repowise ranks refactoring targets by impact for effort and alerts you when a file's health starts declining, but it hands you a human-readable, deterministic worklist; it does not open PRs or rewrite your code. The suggestions are template-based and inspectable, with no LLM editing your repo.
Last reviewed: October 2026
See whether the score finds your risky files.
Index a public repository free and check the lowest-scoring files against the ones that broke last quarter.