ContinuityBench Explorer

Explore aggregate rankings, item-level scores, retained judge evidence, and full conversation transcripts.

Rankings shown here are evaluation configurations—not universal measures of model quality.

19 models · 26 official variants |Methods & data |Construct-validity audit

About this view evaluation configuration, BC-Score formula, dispersion, missing data, judge setup
Dimension
Weights (composite)
Stressor families
Color domain
Rows are models, columns are stressors grouped by family. Click a model name, a stressor column, or a cell to inspect. Columns for families excluded by the current filter are dimmed but remain inspectable.
Construct-Validity Audit diagnostic pilot, n=5 models — not a leaderboard result

Scope. A diagnostic pilot on 5 models and 5 legacy items (ah_001ah_005). It is not a leaderboard result: none of its numbers enter any score or ranking shown elsewhere in this Explorer. It is also not a controlled version experiment — the legacy and official item sets were never matched on length, topic, authoring round or rubric metadata, so the difference cannot be attributed to item version.

The pilot asked whether a non-overlapping legacy item set could independently reproduce slice-level ranking instability. It could not. The five legacy items showed severe ceiling compression — all 25 model×item observations landed at or above 0.95 — so the rank correlation computed on them was rejected as evidence of ranking instability rather than reported. These legacy items are excluded from the official 26-variant leaderboard set and are never loaded by default.

0.0060
Legacy SD
0.0135
Legacy range
100%
Ceiling rate (25/25 ≥ 0.95)
5
Models at ranks #1, #9, #15, #16, #19

Under this degree of ceiling compression, the resulting rank correlation cannot support an interpretable estimate of how model ordering differs. Discriminability was also item-dependent: spreads among the three official Abstraction Hopping items ranged from 0.015 to 0.293, so this pilot does not establish a simple legacy-versus-official version effect. The audit therefore supports excluding the legacy items from the official leaderboard.