Explore aggregate rankings, item-level scores, retained judge evidence, and full conversation transcripts.
Rankings shown here are evaluation configurations—not universal measures of model quality.
19 models · 26 official variants |Methods & data |Construct-validity audit
Start here. This page exposes how the ContinuityBench ranking is constructed: every number on it is a function of an evaluation configuration, not a fixed property of a model. Drag the dimension weights, toggle stressor families, or switch Sort to Worst case and watch the order change — for example, GPT-5.3 Chat ranks #4 on mean score but #11 on worst case. The cell preloaded on the right shows why: a single Multi-Project Interleave item where its goal continuity collapses to 0.150 — click View conversation there to read the judged transcript turn by turn.
Scope. A diagnostic pilot on 5 models and 5 legacy items (ah_001–ah_005). It is not a leaderboard result: none of its numbers enter any score or ranking shown elsewhere in this Explorer. It is also not a controlled version experiment — the legacy and official item sets were never matched on length, topic, authoring round or rubric metadata, so the difference cannot be attributed to item version.
The pilot asked whether a non-overlapping legacy item set could independently reproduce slice-level ranking instability. It could not. The five legacy items showed severe ceiling compression — all 25 model×item observations landed at or above 0.95 — so the rank correlation computed on them was rejected as evidence of ranking instability rather than reported. These legacy items are excluded from the official 26-variant leaderboard set and are never loaded by default.
Under this degree of ceiling compression, the resulting rank correlation cannot support an interpretable estimate of how model ordering differs. Discriminability was also item-dependent: spreads among the three official Abstraction Hopping items ranged from 0.015 to 0.293, so this pilot does not establish a simple legacy-versus-official version effect. The audit therefore supports excluding the legacy items from the official leaderboard.
Full write-up, with observations, interpretations and open hypotheses stated separately: docs/construct_validity_audit.md