repowiserepowise
Sign in
2233admin/cleanlab
OverviewDocsArchitectureKnowledge GraphFilesCode HealthRefactoring

People & History

CommitsContributorsDecisions
ChatPro
Stats
repowiserepowise
ExplorePricingDocs
Sign inIndex repoIndex your repo free
repowise2233admin/cleanlab

Commits

Every commit scored for change-risk against this repo's own history, so 'elevated' means elevated here rather than on some global curve.

Needs review

579of 1,726 scored

579 commits sit in this repo's top risk tercile, which is 34% of the 1,726scored. The cut is drawn against this codebase's own history rather than a global curve, so a quiet repo still fills its top band, and here it starts at 4.8 out of 10. What pushes a commit up is size and spread together: a large change confined to one area scores below a smaller one scattered across a dozen files.

Fix commits1267%Commits whose subject reads as a bug fix rather than new work.Change diffusion0.39bitsShannon entropy of a commit's churn across its files. Zero is a single file, and every extra bit is a doubling of how widely the change spread.Review threshold4.8out of 10Score a commit has to clear to land in this repo's top tercile.

How the work changed shape

Commit categories over time, read off the subject line. Fixes carry the accent because that is the series this chart exists to show.

Consistently other-driven across its history.

Other100%

Review-priority queue

Ranked by change-risk, highest first. Priority is a tercile of this repo's own distribution, so a quiet repo still fills its top band.

Commit review-priority queue
#CommitAuthorWhenLinesRiskTop driver
1
64edc956Introduce Datalab (#614)
Elías Snorrason
3y ago+9.8K -20
100%Elevated
large diff (many lines added)
2
31868642The glorious init commit.
Curtis Northcutt
8y ago+4.3K -0
100%Elevated
large diff (many lines added)
3
460bb862replace w updated file 2
Jonas Mueller
2y ago+3.2K -0
100%Elevated
large diff (many lines added)
4
e82a6072Added fix to test data inspection/cleaning, changed wording, cleaned up notebook and added more on hyperparameter optimization section. This section still needs to be improved.
Matt Turk
2y ago+1.8K -1.0K
100%Elevated
large diff (many lines added)
5
64488565Can ignore commented out code and also some code I pasted in from a previous version following the model eval on clean training + test data. Fixed section on using Datalab on training data to clean the data
Matt Turk
2y ago+2.4K -535
100%Elevated
large diff (many lines added)
6
b5c44b71Added WIP new CLOS train test split tutorial notebook
Matt Turk
2y ago+3.1K -0
100%Elevated
large diff (many lines added)
7
5903c1b9Draft for new image tutorial
Sanjana Garg
2y ago+1.9K -0
100%Elevated
large diff (many lines added)
8
c8def3ecAddressed comments on PR
Sanjana Garg
2y ago+1.2K -214
99%Elevated
large diff (many lines added)
9
9f302200Regression label quality scores (#572)
Mayank Kumar
3y ago+2.1K -0
99%Elevated
large diff (many lines added)
10
35606fc4label error detection in semantic segmentation datasets (#677)
Vedang Lad
3y ago+1.8K -0
99%Elevated
large diff (many lines added)
11
da65a972identifying label errors in Object Detection data (#676)
Ulyana
3y ago+2.8K -2
99%Elevated
large diff (many lines added)
12
9dfa0010Implementing get_ood_scores function (#338)
Ulyana
3y ago+1.4K -396
99%Elevated
large diff (many lines added)
13
960c2b4aCL functionality for multiannotator data (#333)
Hui Wen
3y ago+2.1K -1
99%Elevated
large diff (many lines added)
14
8f9f3f53Major API change. Introducing Cleanlab 2.0 (#128)
Curtis G. Northcutt
4y ago+2.3K -2.4K
99%Elevated
large diff (many lines added)
15
a0f24addUse existing Datalab to audit additional new data (#1049)
ChG
2y ago+1.3K -5
99%Elevated
large diff (many lines added)
16
885fc46eChanged filename
Sanjana Garg
2y ago+862 -1.4K
99%Elevated
large diff (many lines added)
17
3419ae5eAdded explanations to image tutorial
Sanjana Garg
2y ago+804 -736
99%Elevated
large diff (many lines added)
18
3a09fb8eAdded fashion mnist image tutorial
Sanjana Garg
2y ago+1.1K -0
99%Elevated
large diff (many lines added)
19
f6146942Multiannotator Active Learning Support (#538)
Hui Wen
3y ago+1.2K -113
99%Elevated
large diff (many lines added)
20
e060f551supporting multilabel via one-vs-rest reductions (#483)
Aditya Thyagarajan
3y ago+856 -199
99%Elevated
large diff (many lines added)
21
1bad2f82Adding functionality for cleanlab to find label errors in token classification datasets (#347)
Eric Wang
3y ago+1.3K -0
99%Elevated
large diff (many lines added)
22
5b6d297bAdd audio tutorial to doc site (#165)
Wei Jing
4y ago+1.2K -1
99%Elevated
large diff (many lines added)
23
11dff1faStandardize code style to Black (#107)
Anish Athalye
4y ago+1.4K -1.1K
99%Elevated
large diff (many lines added)
24
d3fd6280Add functionality for finding spurious correlations for image(#1140)
Rahul Aditya
2y ago+772 -0
98%Elevated
large diff (many lines added)
25
a7aab16eAdd notebook with miscellaneous Datalab workflows (#1125)
Elías Snorrason
2y ago+1.1K -0
98%Elevated
large diff (many lines added)
26
83d4209cUpdated train and test datasets used, fixed bug with not dropping rows from training data that are exact duplicat with test set, updated seed usage to be proper, and fixed unit tests accordingly
Matt Turk
2y ago+729 -963
98%Elevated
large diff (many lines added)
27
f954aa86Fixed datasets and added sections on checking for near duplicates/non iid issues and filtered training data based on exact duplicates between training and test sets
Matt Turk
2y ago+667 -1.5K
98%Elevated
large diff (many lines added)
28
a6d13193Introduce regression support to Datalab (#796)
OrdoAbChao
2y ago+787 -231
98%Elevated
large diff (many lines added)
29
f8c1866cmove methods to multilabel_classification module (#657)
Aditya Thyagarajan
3y ago+1.1K -335
98%Elevated
large diff (many lines added)
30
d397cdebAdd in-depth tutorial [WIP] (#208)
Wei Jing
4y ago+970 -8
98%Elevated
large diff (many lines added)
31
60be9f72Estimates and fully characterizes label noise
Curtis Northcutt
7y ago+772 -0
98%Elevated
large diff (many lines added)
32
5bf6d689A PyTorch MNIST CNN wrapped in a partial sklearn template. Works with all algorithms.
Curtis Northcutt
8y ago+604 -0
98%Elevated
large diff (many lines added)
33
69295dd6Updated tutorial hidden test thresholds, updated a few code blocks that were outdated with newest version of cleanlab package, and some wording in markdown
Matt Turk
2y ago+555 -486
98%Elevated
large diff (many lines added)
34
25b7aabaImprove KNN Graph Construction for Handling Exact Duplicates and Numerical Precision (#1119)
Elías Snorrason
2y ago+872 -83
98%Elevated
large diff (many lines added)
35
55409591Add existing issue managers to more tasks (#979)
Elías Snorrason
2y ago+896 -258
98%Elevated
large diff (many lines added)
36
dee32ad9Underperforming Group Issue Type (#838)
Ganesh Tata
2y ago+879 -9
98%Elevated
large diff (many lines added)
37
686cbf63Method to estimate label issues with limited memory via mini-batches (#615)
Jonas Mueller
3y ago+821 -23
98%Elevated
large diff (many lines added)
38
c32335c7Tutorial for multi-label classification (#517)
Aditya Thyagarajan
3y ago+574 -0
98%Elevated
large diff (many lines added)
39
b12d76b6Multilabel code restructuring with aggregation/scorer functions (#509)
Aditya Thyagarajan
3y ago+696 -332
98%Elevated
large diff (many lines added)
40
c50836cfAdd label quality scoring functions and user API to choose the method (#131)
Johnson Kuan
4y ago+651 -94
98%Elevated
large diff (many lines added)
41
6169fdcdAdd new documentation site
Wei Jing Lok
4y ago+720 -225
98%Elevated
large diff (many lines added)
42
86115263GNU GPL License
Curtis G. Northcutt
5y ago+674 -27
98%Elevated
large diff (many lines added)
43
0f558a69The RankPruning() class for learning with noisy labels.
Curtis Northcutt
7y ago+708 -246
98%Elevated
large diff (many lines added)
44
2a68fd72Improve knn graph handling and outlier detection in issue managers (#1155)
Elías Snorrason
2y ago+520 -249
97%Elevated
large diff (many lines added)
45
d9f589eeAdd table of issue type info and relevant column name descriptions (#1100)
Elías Snorrason
2y ago+512 -5
97%Elevated
large diff (many lines added)
46
71ba4b32Add a knn module (#1117)
Elías Snorrason
2y ago+667 -179
97%Elevated
large diff (many lines added)
47
e22ffd81Re-added tabular datalab tutorial
Matt Turk
2y ago+532 -0
97%Elevated
large diff (many lines added)
48
51de7776Multilabel Issue Manager for Classification (#929)
Ganesh Tata
2y ago+627 -91
97%Elevated
large diff (many lines added)
49
b93fdebfAdd compatibility for tensorflow and pytorch Dataset objects (#311)
Jonas Mueller
4y ago+637 -80
97%Elevated
large diff (many lines added)
50
b8e85284Added outlier detection tutorial into docs (#310)
Ulyana
4y ago+629 -0
97%Elevated
large diff (many lines added)
  • 64edc956Introduce Datalab (#614)
    Author
    Elías Snorrason
    When
    3y ago
    Lines
    +9.8K -20
    Risk
    100%Elevated
  • 31868642The glorious init commit.
    Author
    Curtis Northcutt
    When
    8y ago
    Lines
    +4.3K -0
    Risk
    100%Elevated
  • 460bb862replace w updated file 2
    Author
    Jonas Mueller
    When
    2y ago
    Lines
    +3.2K -0
    Risk
    100%Elevated
  • e82a6072Added fix to test data inspection/cleaning, changed wording, cleaned up notebook and added more on hyperparameter optimization section. This section still needs to be improved.
    Author
    Matt Turk
    When
    2y ago
    Lines
    +1.8K -1.0K
    Risk
    100%Elevated
  • 64488565Can ignore commented out code and also some code I pasted in from a previous version following the model eval on clean training + test data. Fixed section on using Datalab on training data to clean the data
    Author
    Matt Turk
    When
    2y ago
    Lines
    +2.4K -535
    Risk
    100%Elevated
  • b5c44b71Added WIP new CLOS train test split tutorial notebook
    Author
    Matt Turk
    When
    2y ago
    Lines
    +3.1K -0
    Risk
    100%Elevated
  • 5903c1b9Draft for new image tutorial
    Author
    Sanjana Garg
    When
    2y ago
    Lines
    +1.9K -0
    Risk
    100%Elevated
  • c8def3ecAddressed comments on PR
    Author
    Sanjana Garg
    When
    2y ago
    Lines
    +1.2K -214
    Risk
    99%Elevated
  • 9f302200Regression label quality scores (#572)
    Author
    Mayank Kumar
    When
    3y ago
    Lines
    +2.1K -0
    Risk
    99%Elevated
  • 35606fc4label error detection in semantic segmentation datasets (#677)
    Author
    Vedang Lad
    When
    3y ago
    Lines
    +1.8K -0
    Risk
    99%Elevated
  • da65a972identifying label errors in Object Detection data (#676)
    Author
    Ulyana
    When
    3y ago
    Lines
    +2.8K -2
    Risk
    99%Elevated
  • 9dfa0010Implementing get_ood_scores function (#338)
    Author
    Ulyana
    When
    3y ago
    Lines
    +1.4K -396
    Risk
    99%Elevated
  • 960c2b4aCL functionality for multiannotator data (#333)
    Author
    Hui Wen
    When
    3y ago
    Lines
    +2.1K -1
    Risk
    99%Elevated
  • 8f9f3f53Major API change. Introducing Cleanlab 2.0 (#128)
    Author
    Curtis G. Northcutt
    When
    4y ago
    Lines
    +2.3K -2.4K
    Risk
    99%Elevated
  • a0f24addUse existing Datalab to audit additional new data (#1049)
    Author
    ChG
    When
    2y ago
    Lines
    +1.3K -5
    Risk
    99%Elevated
  • 885fc46eChanged filename
    Author
    Sanjana Garg
    When
    2y ago
    Lines
    +862 -1.4K
    Risk
    99%Elevated
  • 3419ae5eAdded explanations to image tutorial
    Author
    Sanjana Garg
    When
    2y ago
    Lines
    +804 -736
    Risk
    99%Elevated
  • 3a09fb8eAdded fashion mnist image tutorial
    Author
    Sanjana Garg
    When
    2y ago
    Lines
    +1.1K -0
    Risk
    99%Elevated
  • f6146942Multiannotator Active Learning Support (#538)
    Author
    Hui Wen
    When
    3y ago
    Lines
    +1.2K -113
    Risk
    99%Elevated
  • e060f551supporting multilabel via one-vs-rest reductions (#483)
    Author
    Aditya Thyagarajan
    When
    3y ago
    Lines
    +856 -199
    Risk
    99%Elevated
  • 1bad2f82Adding functionality for cleanlab to find label errors in token classification datasets (#347)
    Author
    Eric Wang
    When
    3y ago
    Lines
    +1.3K -0
    Risk
    99%Elevated
  • 5b6d297bAdd audio tutorial to doc site (#165)
    Author
    Wei Jing
    When
    4y ago
    Lines
    +1.2K -1
    Risk
    99%Elevated
  • 11dff1faStandardize code style to Black (#107)
    Author
    Anish Athalye
    When
    4y ago
    Lines
    +1.4K -1.1K
    Risk
    99%Elevated
  • d3fd6280Add functionality for finding spurious correlations for image(#1140)
    Author
    Rahul Aditya
    When
    2y ago
    Lines
    +772 -0
    Risk
    98%Elevated
  • a7aab16eAdd notebook with miscellaneous Datalab workflows (#1125)
    Author
    Elías Snorrason
    When
    2y ago
    Lines
    +1.1K -0
    Risk
    98%Elevated
  • 83d4209cUpdated train and test datasets used, fixed bug with not dropping rows from training data that are exact duplicat with test set, updated seed usage to be proper, and fixed unit tests accordingly
    Author
    Matt Turk
    When
    2y ago
    Lines
    +729 -963
    Risk
    98%Elevated
  • f954aa86Fixed datasets and added sections on checking for near duplicates/non iid issues and filtered training data based on exact duplicates between training and test sets
    Author
    Matt Turk
    When
    2y ago
    Lines
    +667 -1.5K
    Risk
    98%Elevated
  • a6d13193Introduce regression support to Datalab (#796)
    Author
    OrdoAbChao
    When
    2y ago
    Lines
    +787 -231
    Risk
    98%Elevated
  • f8c1866cmove methods to multilabel_classification module (#657)
    Author
    Aditya Thyagarajan
    When
    3y ago
    Lines
    +1.1K -335
    Risk
    98%Elevated
  • d397cdebAdd in-depth tutorial [WIP] (#208)
    Author
    Wei Jing
    When
    4y ago
    Lines
    +970 -8
    Risk
    98%Elevated
  • 60be9f72Estimates and fully characterizes label noise
    Author
    Curtis Northcutt
    When
    7y ago
    Lines
    +772 -0
    Risk
    98%Elevated
  • 5bf6d689A PyTorch MNIST CNN wrapped in a partial sklearn template. Works with all algorithms.
    Author
    Curtis Northcutt
    When
    8y ago
    Lines
    +604 -0
    Risk
    98%Elevated
  • 69295dd6Updated tutorial hidden test thresholds, updated a few code blocks that were outdated with newest version of cleanlab package, and some wording in markdown
    Author
    Matt Turk
    When
    2y ago
    Lines
    +555 -486
    Risk
    98%Elevated
  • 25b7aabaImprove KNN Graph Construction for Handling Exact Duplicates and Numerical Precision (#1119)
    Author
    Elías Snorrason
    When
    2y ago
    Lines
    +872 -83
    Risk
    98%Elevated
  • 55409591Add existing issue managers to more tasks (#979)
    Author
    Elías Snorrason
    When
    2y ago
    Lines
    +896 -258
    Risk
    98%Elevated
  • dee32ad9Underperforming Group Issue Type (#838)
    Author
    Ganesh Tata
    When
    2y ago
    Lines
    +879 -9
    Risk
    98%Elevated
  • 686cbf63Method to estimate label issues with limited memory via mini-batches (#615)
    Author
    Jonas Mueller
    When
    3y ago
    Lines
    +821 -23
    Risk
    98%Elevated
  • c32335c7Tutorial for multi-label classification (#517)
    Author
    Aditya Thyagarajan
    When
    3y ago
    Lines
    +574 -0
    Risk
    98%Elevated
  • b12d76b6Multilabel code restructuring with aggregation/scorer functions (#509)
    Author
    Aditya Thyagarajan
    When
    3y ago
    Lines
    +696 -332
    Risk
    98%Elevated
  • c50836cfAdd label quality scoring functions and user API to choose the method (#131)
    Author
    Johnson Kuan
    When
    4y ago
    Lines
    +651 -94
    Risk
    98%Elevated
  • 6169fdcdAdd new documentation site
    Author
    Wei Jing Lok
    When
    4y ago
    Lines
    +720 -225
    Risk
    98%Elevated
  • 86115263GNU GPL License
    Author
    Curtis G. Northcutt
    When
    5y ago
    Lines
    +674 -27
    Risk
    98%Elevated
  • 0f558a69The RankPruning() class for learning with noisy labels.
    Author
    Curtis Northcutt
    When
    7y ago
    Lines
    +708 -246
    Risk
    98%Elevated
  • 2a68fd72Improve knn graph handling and outlier detection in issue managers (#1155)
    Author
    Elías Snorrason
    When
    2y ago
    Lines
    +520 -249
    Risk
    97%Elevated
  • d9f589eeAdd table of issue type info and relevant column name descriptions (#1100)
    Author
    Elías Snorrason
    When
    2y ago
    Lines
    +512 -5
    Risk
    97%Elevated
  • 71ba4b32Add a knn module (#1117)
    Author
    Elías Snorrason
    When
    2y ago
    Lines
    +667 -179
    Risk
    97%Elevated
  • e22ffd81Re-added tabular datalab tutorial
    Author
    Matt Turk
    When
    2y ago
    Lines
    +532 -0
    Risk
    97%Elevated
  • 51de7776Multilabel Issue Manager for Classification (#929)
    Author
    Ganesh Tata
    When
    2y ago
    Lines
    +627 -91
    Risk
    97%Elevated
  • b93fdebfAdd compatibility for tensorflow and pytorch Dataset objects (#311)
    Author
    Jonas Mueller
    When
    4y ago
    Lines
    +637 -80
    Risk
    97%Elevated
  • b8e85284Added outlier detection tutorial into docs (#310)
    Author
    Ulyana
    When
    4y ago
    Lines
    +629 -0
    Risk
    97%Elevated
Showing 50 of 1,726 commits

How the score behaves here

Change risk →

Two views of the same model: where the cuts fall, and what commit shape lands you above them.

Score distribution

Every scored commit, binned on the raw 0 to 10 score rather than the percentile. Percentile ranks are uniform by construction, so that axis has no shape to draw. The dashed lines are the tercile cuts behind each row's priority pill.

0290typical ↑elevated ↑0.05.010.0Change-risk score →
Below typical
Typical
Elevated

Size against diffusion

The 200 most recent commits, on their own recency sample rather than the feed above: that defaults to risk-sorted, so reusing it would plot only the top tercile and call it the spread. Big and scattered is what the model penalises. Click a dot to open it.

1101001,000Lines changed (log) →0.04.6Diffusion →
Below typical78
Typical64
Elevated58

Commit history for 2233admin/cleanlab

Repowise tracks change history across 155 files in 2233admin/cleanlab. No file in the repository changed in the last 90 days. Every commit is scored for change risk from its size, spread and the history of the files it touches.