References#
Methods#
Interquartile Mean (IQM)
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., & Bellemare, M. G. (2021). Deep Reinforcement Learning at the Edge of the Statistical Precipice. Advances in Neural Information Processing Systems, 34. https://arxiv.org/abs/2108.13264
ELO Ranking
Erickson, N., Purucker, L., Tschalzev, A., Holzmüller, D., Mutalik Desai, P., Salinas, D., & Hutter, F. (2025). TabArena: A Living Benchmark for Machine Learning on Tabular Data. Advances in Neural Information Processing Systems (Datasets and Benchmarks, Spotlight). https://arxiv.org/abs/2506.16791
Bayesian Pairwise Comparison
Benavoli, A., Corani, G., Demšar, J., & Zaffalon, M. (2017). Time for a Change: a Tutorial for Comparing Multiple Classifiers Through Bayesian Analysis. Journal of Machine Learning Research, 18(77), 1–36. https://jmlr.org/papers/v18/16-305.html
Frequentist Comparison
Demšar, J. (2006). Statistical Comparisons of Classifiers over Multiple Data Sets. Journal of Machine Learning Research, 7, 1–30. https://jmlr.org/papers/v7/demsar06a.html
Dolan-Moré Performance Profiles
Dolan, E. D., & Moré, J. J. (2002). Benchmarking Optimization Software with Performance Profiles. Mathematical Programming, 91(2), 201–213. https://doi.org/10.1007/s101070100263
Rank Sensitivity
Kendall, M. G. (1945). The Treatment of Ties in Ranking Problems. Biometrika, 33(3), 239–251. https://doi.org/10.1093/biomet/33.3.239
Libraries#
baycomp
Janez Demšar. baycomp: Bayesian comparison of classifiers. janezd/baycomp