Why a review budget must preserve tied scores
Distinguish every-K ranking from score thresholds, choose a reproducible tie policy, and interpret missing positives.
A fixed review count is a different decision
The original 2015 scikit-learn issue asks how precision, recall and F1 change when a researcher can review only N of 2000 held-out instances. The plot is public, but sample IDs, labels and scores are absent. The four comments link related proposals; the February 2023 closure is triage toward #7343, not evidence that a top-K API shipped.
Ties need an explicit order
A threshold can select a whole tied group at once. A fixed budget may end inside that group. Preserve source order for the least intrusive default, or use Unicode codepoint ID order for a declared secondary policy. Direction reverses scores, while the secondary policy stays ascending; locale sorting would change reproducibility.
- Keep all K rather than merging equal score thresholds.
- Keep original score spelling and the negative-zero flag separately from JSON numeric zero.
Undefined recall does not erase a budget
If the held-out set contains P=0 positives, recall=TP/P is undefined. Precision=TP/K is 0 and F1=2TP/(K+P) is 0 because K≥1. Retaining every budget and its exact numerator and denominator makes this state visible without inventing positives.
| Metric | Exact ratio | P=0 |
|---|---|---|
| precision | TP/K | 0 |
| recall | TP/P | Undefined |
| F1 | 2TP/(K+P) | 0 |
References
- Original every-budget classification question
Complete question, all four comments and six timeline items reviewed; plot only, no author sample values.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
Classification Rank BudgetRank complete labeled scores and inspect every top-K review budget with exact prefix counts.
Choose a positive class, score direction and tie policy, then see how a fixed review budget changes precision, recall and F1 across the entire held-out set.
Steps
- Paste complete CSV or choose one file, then select that active source.
- Set distinct one-based ID, label and score columns, positive/negative labels, direction and tie policy.
- Run, inspect any K, and save all six files with the full copied report.
Available options
- Active input
- Pasted CSV · One original CSV file
- Positive label
- 1
- Negative label
- 0
- Score direction
- Higher scores first · Lower scores first
- Tie policy
- Original source order · ID Unicode codepoint order
- ID column (1-based)
- 1
- Label column (1-based)
- 2
- Score column (1-based)
- 3
Capabilities and limits
- Strict UTF-8 CSV with one header and 1–1,000,000 complete records; unique nonempty string IDs, two explicit labels and finite decimal scores.
- 64 MiB source, 10,000,000 source cells, 12,000,000 curve cells and 100,000,000 combined work units; limits apply together.
- Files plus full text ≤512 MiB, typed result ≤512 MiB, their aggregate ≤768 MiB, wire and owned memory ≤1 GiB. A legal record count alone does not guarantee joint capacity.
- One 120-second deadline covers capture, read, validation, calculation, encoding, cleanup and first publication; a limit failure returns no partial curve.
- Score ties preserve every K. No positives gives undefined recall and F1=0 for every K. Preview is limited to 16 budgets and 4,000 codepoints; downloads and full copy are complete.