Aggregate analysis

Capability Map

Explore a fixed full-coverage core through separate Evaluation Paradigm and Model Origin taxonomies, with backbone- and dataset-equal aggregation.

Fixed Overall capability core

The analysis uses the same complete method universe as the Standard Overall leaderboard: 12 configurations across 7 backbones, each evaluated on all 198 datasets.

Configuration scores are averaged within each backbone first. Backbones and datasets then receive equal weight, preventing classifier-rich backbones from receiving extra influence.

198

datasets

12

configurations

7

backbones

198 / 198

result coverage

Comparison taxonomy

Choose one independent analysis lens

Groups are mutually exclusive within each taxonomy. Evaluation Paradigm and Model Origin are calculated and ranked separately.

Property coverage

Choose a benchmark attribute

Every property has reviewed metadata for all 198 datasets; coverage remains visible for audit transparency.

Capability heatmap

Evaluation Paradigm × modality

198 of 198 datasets contribute to this property view. N is the dataset count in each condition.

Fixed core · 12 configurations · 7 backbones
Taxonomy groupUnivariate1 channelMultivariate2 or more channels
Native time-series ICL1 backbone · 1 configuration

81.2%

N=128

73.8%

N=70

Frozen representation + classifier4 backbones · 9 configurations

74.0%

N=128

68.6%

N=70

Tabular ICL adaptation2 backbones · 2 configurations

77.9%

N=128

69.7%

N=70

Aggregate evidence

Condition-level detail

Every row is a fixed-core aggregate; no dataset-level record is exposed or downloadable.

Hierarchical Accuracy

For each dataset, configuration accuracies are averaged within their shared backbone. Backbone scores are then averaged with equal backbone weight within each taxonomy group. The resulting group scores are averaged with equal dataset weight within the selected condition.

Best / Tied Best / Other

Best / Tied Best / Other counts whether a group is uniquely highest / tied for highest / below the highest within the active taxonomy on each dataset.

ConditionTaxonomy groupAccuracyAvg. Group RankBest / Tied / OtherDatasetsBackbonesConfigurations
UnivariateNative time-series ICL81.2%1.5069 / 2 / 5712811
UnivariateFrozen representation + classifier74.0%2.665 / 0 / 12312849
UnivariateTabular ICL adaptation77.9%1.8452 / 2 / 7412822
MultivariateNative time-series ICL73.8%1.4942 / 1 / 277011
MultivariateFrozen representation + classifier68.6%2.464 / 0 / 667049
MultivariateTabular ICL adaptation69.7%2.0623 / 1 / 467022

Reading guide

Methodology and privacy boundaries

Fixed result universe

The fixed core matches the Standard Overall leaderboard: 12 configurations across 7 backbones, each with complete 198 / 198 coverage.

Rank definition

Average Rank ranks the mutually exclusive groups within the active taxonomy on each dataset, using average ranks for exact ties.

Separate taxonomies

Evaluation Paradigm and Model Origin are separate taxonomies. Groups are mutually exclusive within each taxonomy and are never ranked across taxonomies.

Public data boundary

Only aggregate condition-level statistics are published. Dataset identities, source mappings, backbone names, and configuration names are excluded.