Compression BenchRankingMethodologyStatus
demo ranking · current pipeline · data as of 8 October 2026

Ranking

Every model scored so far, on one author's private prose and code, in nats per UTF-8 byte (lower is better). Post-trained models and base models are ranked separately. The arrows compare the two orders: does a post-trained model sit at the same place in its ranking as its base does in the base ranking? That is the question the recovery method exists to answer. The headline number for both is the re-based loss: the model after the fixed S1 fine-tune. This is a demo of the pipeline, not a launch: one corpus, one author, and the error bars from the methodology note apply.

Post-trained models

Recovered = the released model after the S1 fine-tune. Residual = recovered minus its matched base, where a public base exists. Rank vs base: this model's place among post-trained models (one representative per base, so a base with six descendants counts once) compared with its base's place among the bases. ▲ n = it ranks n places higher than its base does; ▼ n = n places lower; = same place. A method that recovers the base ordering perfectly would show "=" everywhere.

#ModelFamilyParams (B)Recovered lossNative lossResidualRank vs base

Base models

Pretrained checkpoints. Re-based = the base after the S1 fine-tune (the matched-base control); this is the published number because native loss conflates knowledge with register fit. Descendants: how many post-trained descendants were scored, and where the best of them sits in the post-trained ranking relative to this base's place here.

#ModelFamilyParams (B)Re-based lossNative lossS1 dropDescendants

Pending re-basing

Scored natively, matched S1 run queued. Listed by native loss; not comparable to the table above until re-based.

ModelFamilyParams (B)Native loss

Sources: reviewed aggregate JSON only (base-roster-v2 and extension, scorecard-v2 prose/code, recovery-expansion-v1, owner-v3 additions, blind-confirmation-v1). Qwen3 "base" checkpoints ship after an SFT stage; their native loss is far above trend and their re-based loss is the meaningful number. Private text and per-record outputs never leave the compute node.