Skip to main content
Back to top
Ctrl
+
K
Search
Ctrl
+
K
Contents
Overview
Configuration Guide
API Reference
evaluma
evaluma._version
evaluma.benchmark
evaluma.cli
evaluma.methods
evaluma.methods.aggregate
evaluma.methods.bayesian
evaluma.methods.elo
evaluma.methods.frequentist
evaluma.methods.improvability
evaluma.methods.iqm
evaluma.methods.profiles
evaluma.methods.rank_sensitivity
evaluma.metric_registry
evaluma.normalize
evaluma.plot
evaluma.results
Tutorials
Ranking Models Across a Benchmark: From Mean to IQM
Ranking Models by Head-to-Head Dominance: ELO from TabArena
Distance from the Best: Improvability Ranking
Frequentist Model Comparison: Friedman + Nemenyi / Wilcoxon + Holm
Evaluating a Benchmark from a Bayesian Perspective
Frequentist vs Bayesian Model Comparison
Performance Profiles
Rank Sensitivity: Do Model Rankings Hold Across Conditions?
References
Contributing
Repository
Open issue
Search
Error
Please activate JavaScript to enable the search functionality.
Ctrl
+
K