arXiv, February 2023
Does the evaluation stand up to evaluation? A first-principle approach to the evaluation of classifiers
Accuracy, F1 and MCC can recommend a classifier that loses money. We show that the only consistent way to compare classifiers is to weigh each outcome by what it's worth, and that even a rough guess at those values ranks classifiers better than the usual metrics.
What the usual metrics say
| Metric | A | B |
|---|---|---|
| Accuracy | 0.62 | 0.75 |
| F1 | 0.59 | 0.77 |
| MCC | 0.24 | 0.51 |
| Precision | 0.64 | 0.70 |
What each one earns per component

