Every commit scored for change-risk against this repo's own history, so 'elevated' means elevated here rather than on some global curve.
Needs review
579 commits sit in this repo's top risk tercile, which is 34% of the 1,726scored. The cut is drawn against this codebase's own history rather than a global curve, so a quiet repo still fills its top band, and here it starts at 4.8 out of 10. What pushes a commit up is size and spread together: a large change confined to one area scores below a smaller one scattered across a dozen files.
Commit categories over time, read off the subject line. Fixes carry the accent because that is the series this chart exists to show.
Consistently other-driven across its history.
Ranked by change-risk, highest first. Priority is a tercile of this repo's own distribution, so a quiet repo still fills its top band.
| # | Commit | Author | When | Lines | Risk | Top driver |
|---|---|---|---|---|---|---|
| 1 | 64edc956Introduce Datalab (#614) | Elías Snorrason | 3y ago | +9.8K -20 | 100%Elevated | large diff (many lines added) |
| 2 | 31868642The glorious init commit. | Curtis Northcutt | 8y ago | +4.3K -0 | 100%Elevated | large diff (many lines added) |
| 3 | 460bb862replace w updated file 2 | Jonas Mueller | 2y ago | +3.2K -0 | 100%Elevated | large diff (many lines added) |
| 4 | e82a6072Added fix to test data inspection/cleaning, changed wording, cleaned up notebook and added more on hyperparameter optimization section. This section still needs to be improved. | Matt Turk | 2y ago | +1.8K -1.0K | 100%Elevated | large diff (many lines added) |
| 5 | 64488565Can ignore commented out code and also some code I pasted in from a previous version following the model eval on clean training + test data. Fixed section on using Datalab on training data to clean the data | Matt Turk | 2y ago | +2.4K -535 | 100%Elevated | large diff (many lines added) |
| 6 | b5c44b71Added WIP new CLOS train test split tutorial notebook | Matt Turk | 2y ago | +3.1K -0 | 100%Elevated | large diff (many lines added) |
| 7 | 5903c1b9Draft for new image tutorial | Sanjana Garg | 2y ago | +1.9K -0 | 100%Elevated | large diff (many lines added) |
| 8 | c8def3ecAddressed comments on PR | Sanjana Garg | 2y ago | +1.2K -214 | 99%Elevated | large diff (many lines added) |
| 9 | 9f302200Regression label quality scores (#572) | Mayank Kumar | 3y ago | +2.1K -0 | 99%Elevated | large diff (many lines added) |
| 10 | 35606fc4label error detection in semantic segmentation datasets (#677) | Vedang Lad | 3y ago | +1.8K -0 | 99%Elevated | large diff (many lines added) |
| 11 | da65a972identifying label errors in Object Detection data (#676) | Ulyana | 3y ago | +2.8K -2 | 99%Elevated | large diff (many lines added) |
| 12 | 9dfa0010Implementing get_ood_scores function (#338) | Ulyana | 3y ago | +1.4K -396 | 99%Elevated | large diff (many lines added) |
| 13 | 960c2b4aCL functionality for multiannotator data (#333) | Hui Wen | 3y ago | +2.1K -1 | 99%Elevated | large diff (many lines added) |
| 14 | 8f9f3f53Major API change. Introducing Cleanlab 2.0 (#128) | Curtis G. Northcutt | 4y ago | +2.3K -2.4K | 99%Elevated | large diff (many lines added) |
| 15 | a0f24addUse existing Datalab to audit additional new data (#1049) | ChG | 2y ago | +1.3K -5 | 99%Elevated | large diff (many lines added) |
| 16 | 885fc46eChanged filename | Sanjana Garg | 2y ago | +862 -1.4K | 99%Elevated | large diff (many lines added) |
| 17 | 3419ae5eAdded explanations to image tutorial | Sanjana Garg | 2y ago | +804 -736 | 99%Elevated | large diff (many lines added) |
| 18 | 3a09fb8eAdded fashion mnist image tutorial | Sanjana Garg | 2y ago | +1.1K -0 | 99%Elevated | large diff (many lines added) |
| 19 | f6146942Multiannotator Active Learning Support (#538) | Hui Wen | 3y ago | +1.2K -113 | 99%Elevated | large diff (many lines added) |
| 20 | e060f551supporting multilabel via one-vs-rest reductions (#483) | Aditya Thyagarajan | 3y ago | +856 -199 | 99%Elevated | large diff (many lines added) |
| 21 | 1bad2f82Adding functionality for cleanlab to find label errors in token classification datasets (#347) | Eric Wang | 3y ago | +1.3K -0 | 99%Elevated | large diff (many lines added) |
| 22 | 5b6d297bAdd audio tutorial to doc site (#165) | Wei Jing | 4y ago | +1.2K -1 | 99%Elevated | large diff (many lines added) |
| 23 | 11dff1faStandardize code style to Black (#107) | Anish Athalye | 4y ago | +1.4K -1.1K | 99%Elevated | large diff (many lines added) |
| 24 | d3fd6280Add functionality for finding spurious correlations for image(#1140) | Rahul Aditya | 2y ago | +772 -0 | 98%Elevated | large diff (many lines added) |
| 25 | a7aab16eAdd notebook with miscellaneous Datalab workflows (#1125) | Elías Snorrason | 2y ago | +1.1K -0 | 98%Elevated | large diff (many lines added) |
| 26 | 83d4209cUpdated train and test datasets used, fixed bug with not dropping rows from training data that are exact duplicat with test set, updated seed usage to be proper, and fixed unit tests accordingly | Matt Turk | 2y ago | +729 -963 | 98%Elevated | large diff (many lines added) |
| 27 | f954aa86Fixed datasets and added sections on checking for near duplicates/non iid issues and filtered training data based on exact duplicates between training and test sets | Matt Turk | 2y ago | +667 -1.5K | 98%Elevated | large diff (many lines added) |
| 28 | a6d13193Introduce regression support to Datalab (#796) | OrdoAbChao | 2y ago | +787 -231 | 98%Elevated | large diff (many lines added) |
| 29 | f8c1866cmove methods to multilabel_classification module (#657) | Aditya Thyagarajan | 3y ago | +1.1K -335 | 98%Elevated | large diff (many lines added) |
| 30 | d397cdebAdd in-depth tutorial [WIP] (#208) | Wei Jing | 4y ago | +970 -8 | 98%Elevated | large diff (many lines added) |
| 31 | 60be9f72Estimates and fully characterizes label noise | Curtis Northcutt | 7y ago | +772 -0 | 98%Elevated | large diff (many lines added) |
| 32 | 5bf6d689A PyTorch MNIST CNN wrapped in a partial sklearn template. Works with all algorithms. | Curtis Northcutt | 8y ago | +604 -0 | 98%Elevated | large diff (many lines added) |
| 33 | 69295dd6Updated tutorial hidden test thresholds, updated a few code blocks that were outdated with newest version of cleanlab package, and some wording in markdown | Matt Turk | 2y ago | +555 -486 | 98%Elevated | large diff (many lines added) |
| 34 | 25b7aabaImprove KNN Graph Construction for Handling Exact Duplicates and Numerical Precision (#1119) | Elías Snorrason | 2y ago | +872 -83 | 98%Elevated | large diff (many lines added) |
| 35 | 55409591Add existing issue managers to more tasks (#979) | Elías Snorrason | 2y ago | +896 -258 | 98%Elevated | large diff (many lines added) |
| 36 | dee32ad9Underperforming Group Issue Type (#838) | Ganesh Tata | 2y ago | +879 -9 | 98%Elevated | large diff (many lines added) |
| 37 | 686cbf63Method to estimate label issues with limited memory via mini-batches (#615) | Jonas Mueller | 3y ago | +821 -23 | 98%Elevated | large diff (many lines added) |
| 38 | c32335c7Tutorial for multi-label classification (#517) | Aditya Thyagarajan | 3y ago | +574 -0 | 98%Elevated | large diff (many lines added) |
| 39 | b12d76b6Multilabel code restructuring with aggregation/scorer functions (#509) | Aditya Thyagarajan | 3y ago | +696 -332 | 98%Elevated | large diff (many lines added) |
| 40 | c50836cfAdd label quality scoring functions and user API to choose the method (#131) | Johnson Kuan | 4y ago | +651 -94 | 98%Elevated | large diff (many lines added) |
| 41 | 6169fdcdAdd new documentation site | Wei Jing Lok | 4y ago | +720 -225 | 98%Elevated | large diff (many lines added) |
| 42 | 86115263GNU GPL License | Curtis G. Northcutt | 5y ago | +674 -27 | 98%Elevated | large diff (many lines added) |
| 43 | 0f558a69The RankPruning() class for learning with noisy labels. | Curtis Northcutt | 7y ago | +708 -246 | 98%Elevated | large diff (many lines added) |
| 44 | 2a68fd72Improve knn graph handling and outlier detection in issue managers (#1155) | Elías Snorrason | 2y ago | +520 -249 | 97%Elevated | large diff (many lines added) |
| 45 | d9f589eeAdd table of issue type info and relevant column name descriptions (#1100) | Elías Snorrason | 2y ago | +512 -5 | 97%Elevated | large diff (many lines added) |
| 46 | 71ba4b32Add a knn module (#1117) | Elías Snorrason | 2y ago | +667 -179 | 97%Elevated | large diff (many lines added) |
| 47 | e22ffd81Re-added tabular datalab tutorial | Matt Turk | 2y ago | +532 -0 | 97%Elevated | large diff (many lines added) |
| 48 | 51de7776Multilabel Issue Manager for Classification (#929) | Ganesh Tata | 2y ago | +627 -91 | 97%Elevated | large diff (many lines added) |
| 49 | b93fdebfAdd compatibility for tensorflow and pytorch Dataset objects (#311) | Jonas Mueller | 4y ago | +637 -80 | 97%Elevated | large diff (many lines added) |
| 50 | b8e85284Added outlier detection tutorial into docs (#310) | Ulyana | 4y ago | +629 -0 | 97%Elevated | large diff (many lines added) |
Two views of the same model: where the cuts fall, and what commit shape lands you above them.
Every scored commit, binned on the raw 0 to 10 score rather than the percentile. Percentile ranks are uniform by construction, so that axis has no shape to draw. The dashed lines are the tercile cuts behind each row's priority pill.
The 200 most recent commits, on their own recency sample rather than the feed above: that defaults to risk-sorted, so reusing it would plot only the top tercile and call it the spread. Big and scattered is what the model penalises. Click a dot to open it.
Repowise tracks change history across 155 files in 2233admin/cleanlab. No file in the repository changed in the last 90 days. Every commit is scored for change risk from its size, spread and the history of the files it touches.