evaluma.methods.aggregate#
Attributes#
Functions#
|
Collapse a model × dataset matrix to one score per model. |
|
Compute a point-estimate descriptive ranking from a normalized score matrix. |
Module Contents#
- evaluma.methods.aggregate._AGG_MODES#
- evaluma.methods.aggregate._aggregate_scores(scores_matrix: pandas.DataFrame, agg='trimmed_mean') pandas.Series#
Collapse a model × dataset matrix to one score per model.
Shared primitive for the aggregate-ranking and rank-sensitivity paths so the two cannot drift in how they define a model’s summary score.
- Parameters:
scores_matrix – Normalized model × dataset score matrix.
agg – Aggregation mode — one of
"trimmed_mean","mean","median".
- Returns:
Per-model scores indexed by
scores_matrix.index.- Return type:
pd.Series
- Raises:
ValueError – If
aggis not one of the supported modes.
- evaluma.methods.aggregate.compute_aggregate(scores_matrix: pandas.DataFrame, agg='trimmed_mean') evaluma.results.AggregateResult#
Compute a point-estimate descriptive ranking from a normalized score matrix.
Note
This is a descriptive point estimate only (no CI). The trimmed-mean variant trims across datasets, not across seeds; with fewer than ~10 datasets the 25% trim is aggressive (e.g. 5 datasets → only 3 contribute). Treat results as exploratory. For a statistically grounded ranking with uncertainty, use
evaluma.methods.iqm.compute_iqm()(requires multiple seeds).- Parameters:
scores_matrix – Normalized model × dataset score matrix.
agg – Aggregation mode — one of
"trimmed_mean","mean","median".
- Returns:
Result with
.tablesorted descending byscore.- Return type:
- Raises:
ValueError – If
aggis not one of the supported modes.