Compression BenchRankingMethodologyStatus
Compression Bench · methodology note · 8 October 2026 · numbers rechecked against all 45 pairs 05:45 UTC

What the recovery residual is, and how much to trust a ranking

The central measurement question in this project: when we take a post-trained model, fine-tune it briefly on public text, and score it on private text, how close is that number to what the pretrained base would have scored, and is the error stable enough that rankings survive it? This note walks through what we have measured, in order, and what it means.

The setup

Four numbers exist for every model that has a public base. All are loss on the same private prose windows, in nats per UTF-8 byte.

ArmWhat it isExample: Llama-3.1-70B-Instruct
native baseThe released pretrained checkpoint, scored as-is.0.9365
matched baseThe same base after our fixed recovery fine-tune (S1: LoRA rank 256, 1,000 updates, 8.19M tokens of public web text). This is the control.0.9198
nativeThe post-trained release, scored as-is.1.08–1.3 for most instruct models
recoveredThe post-trained release after the identical S1 fine-tune. This is the number we rank on.0.9383

The residual is recovered minus matched base: here 0.9383 − 0.9198 = +0.0185. It is the part of the post-training damage that the recovery fine-tune does not remove, measured against a base that went through exactly the same fine-tune so that whatever the fine-tune teaches cancels. For a model with no public base, the residual is unobservable and becomes the error in our estimate of its base. Everything below is about the structure of that residual, measured on 45 post-trained models across 24 bases in four ancestral families (OLMo, Llama, Qwen, Gemma), plus one blind family (SmolLM3, two models).

One framing before the data. The same fine-tune applied to the base and to the descendant is a controlled experiment. If post-training only changed format (chat templates, refusal habits, answer style), the fine-tune would undo it and the residual would be zero. If post-training overwrote what the base knew, no amount of web text brings that back and the residual would be large. The residual is therefore a measurement of something real about the post-training, not just a nuisance term. The question is whether it is stable enough to see through.

Two ladders: native base loss is confounded by register

This was not in the original plan and matters for the base-model eval on its own.

The recovery fine-tune also moves base models, and by very different amounts. On 26 bases, the same 8M public tokens lower loss by 0.005 (OLMo-3-32B) to 0.053 (Qwen2.5-3B) for genuine pretrained checkpoints, and by 0.12–0.21 for the Qwen3 "Base" checkpoints that ship after an SFT stage. That is not noise. Our working interpretation, not yet tested, is that it is mostly register rather than knowledge: 8M tokens of ordinary web text should not teach a model much about the world, but the private prose runs to 2026 and the public pool is recent web text, so some information transfer is possible. The planned test is to repeat the fine-tune with a second public pool of a different register and see whether the drops and the residuals move. Register means how far the base's pretraining distribution sits from one person's conversational prose. A base trained on clean, filtered, partly synthetic text is surprised by messages and notes; a short exposure to messy web text removes the surprise.

Each base scored natively (open circle) and after the fixed recovery fine-tune (filled), joined by a line. The drop varies from 0.005 to 0.053 on genuine bases (0.12–0.21 on the post-SFT Qwen3 checkpoints). Kendall τ between the two ladders is 0.72; 46 of 325 pairs reverse. Rank reversals: natively Qwen2.5-3B scores worse than Qwen2.5-1.5B; re-based it lands between 1.5B and 7B where it belongs. Natively Gemma-3-27B is the worst of the 27–32B group; re-based it is the best and beats Qwen2.5-72B. Qwen3 "base" checkpoints ship after an SFT stage, which is why their native points are far off the trend.

Two rank reversals in the figure are the kind that decide whether a base-model eval is believable. Qwen2.5-3B scoring worse than Qwen2.5-1.5B natively is obviously wrong, and the re-based ladder fixes it. Gemma-3-27B going from worst to best among 27–32B bases is less obviously wrong either way, but Gemma 3's pretraining is known to include heavy distillation and instruction-shaped data, which is exactly what should produce a large register gap without a knowledge gap.

Consequence. The base-model ranking we publish should be the re-based ladder, not native loss, and every model on it, base or not, should go through the same fine-tune. Native loss conflates knowledge with register fit. This also makes a testable prediction for the 2026 multi-register corpus: with text from many registers, the native and re-based rankings should converge, because no single register dominates. If they do not, the register story is incomplete.

The residual is constant along a pipeline, not noisy

The worry: if two recovery runs on the same model, or two adjacent checkpoints in one lab's pipeline, gave residuals that differed by 0.01, then a 0.01 difference between two labs' models would be uninterpretable. The data say the opposite.

Residual (recovered − matched base) for every checkpoint in six published post-training pipelines. Within a pipeline, SFT → DPO → RL checkpoints land within 0.001 of each other (pooled sd 0.00025). OLMo-3-32B Think at RL steps 300, 750, 1500 and 2300 are indistinguishable from its SFT checkpoint. Between non-distilled pipelines on the same base the offsets differ by up to 0.013; including distillation, by up to 0.028 (Llama-70B), 0.034 (Qwen-32B) and 0.055 (Llama-8B).

Three facts from this figure:

So the residual is a property of (base, SFT recipe), deterministic to our measurement precision. For a model with a known base, it is simply measured. For a model without one, it is an unknown constant in the range set by its recipe class, and since we usually do not know the recipe either, the practical prior is the spread across all pipelines we have seen: 0.0005–0.03 for ordinary instruct and SFT chains (most between 0.004 and 0.02), 0.03–0.06 for distillation. The same-base sibling spread, 0.028 on Llama-70B, 0.034 on Qwen-32B and 0.055 on Llama-8B, is the honest size of the cross-lab error.

It is a dose meter for SFT

If the SFT stage sets the residual, the residual should grow with how much SFT a model received. AI2 published five checkpoints along OLMo-3-7B's Think SFT run. We recovered each one.

Residual against publisher SFT step for OLMo-3-7B Think (the final checkpoint's exact step is unpublished; plotted at the right edge). Over the four checkpoints with known steps the relationship is roughly linear in log(steps), about 0.014 nats/byte per e-fold of SFT; one run, one family, descriptive only. Native loss (unrecovered) on the same checkpoints rises from 1.013 to 1.144; recovery removes 63–81% of that, a share that falls as SFT exposure grows, and the part it cannot remove grows steadily. The fit is residual ≈ 0.0145 nats/byte per e-fold of SFT steps (R² 0.999).

This is the cleanest evidence that the residual is a real quantity. Nothing about the fine-tune changed between those five runs; only the SFT exposure of the input did. The same pattern appears across labs: Qwen2.5-Instruct, a light SFT pipeline, sits at a flat +0.008 from 0.5B to 72B; models that are known to be distilled or SFT-heavy (Llama 3.2 1B/3B, which Meta pruned and distilled from 8B/70B; the Qwen3 small checkpoints released after SFT; the R1-Distill series) are all at +0.03 to +0.06. We did not use any of that recipe information in the estimate. It simply agrees with it.

What happens with size

The concern: a 0.01 residual on an 8B model is small against the 0.04 gap between 8B and 70B, but at the frontier, where bases may sit within 0.01 of each other, the same residual would be the whole ranking.

Residual against parameter count for all 45 post-trained models with a public base, colored by pipeline type. Lines join same-lineage models at different sizes. The ±0.003 band is the worst-case recovery run noise.

So the direction is favourable: the residual does not grow with scale, and for the heavy pipelines it falls. But a frontier instruct model from a lab with a Llama-like pipeline would still carry about +0.015, and we would not know that number without its base. If two frontier bases truly differ by 0.01, the eval cannot order their post-trained releases across labs. It can say they are within 0.02 of each other, which is itself informative. And it can order releases within a lab to 0.0003, because the pipeline offset cancels.

Is the residual forgetting, or just deep register?

A residual of +0.03 could mean the model lost knowledge the base had, or that the chat prior is so ingrained that 8M tokens of web text cannot reset it even though the knowledge is intact. These have different implications. If it is forgetting, the recovered number is the right number: the post-trained model really is worse. If it is register, the recovered number overstates the damage and the base number would be the right one.

SimpleQA (long-tail factual recall, run by us on 1,500 questions with a fixed open judge) lets us look, because we have base and descendant scores for several lineages:

Base → descendantResidualSimpleQA baseSimpleQA desc.Δ points
Llama-3.1-70B → Tulu-3-70B+0.00424.724.2−0.5
Llama-3.1-70B → Llama-3.1-70B-Instruct+0.01924.722.8−1.9
Llama-3.1-70B → R1-Distill-Llama-70B+0.03224.710.0−14.7
Llama-3.1-8B → Tulu-3-8B+0.0097.68.2+0.6
Llama-3.1-8B → R1-Distill-Llama-8B+0.0647.62.8−4.8
Qwen2.5-32B → R1-Distill-Qwen-32B+0.0346.56.50.0
Qwen2.5-72B → Qwen2.5-72B-Instruct+0.0109.411.2+1.8
OLMo-2-7B → OLMo-2-7B-Instruct+0.0104.94.2−0.7
Qwen2.5-32B → Open-Reasoner-Zero-32B+0.0016.55.9−0.6

Reading this: residuals below about 0.02 mostly come with no measurable change in factual recall (SimpleQA noise at n=1,500 is roughly ±1 point). Across all 33 such pairs we have, 3 move by more than two standard errors, the largest being Gemma-3-27B-IT at −2.7 points with a residual of 0.014; across 41 pairs the Spearman between residual and SimpleQA change is −0.12 [−0.65, +0.29]. So small residuals are usually register the fine-tune could not fully reset, not lost knowledge, but not always. Large residuals from distillation are mixed: R1-Distill-Llama-70B lost 60% of its long-tail recall, R1-Distill-Qwen-32B lost none, at nearly the same residual. So the residual is not a pure forgetting meter. It is an upper bound on damage that sometimes is all register.

What this means for the number we publish. We cannot yet tell forgetting from deep register shift from the residual alone. Reporting recovered loss as the model's score treats the whole residual as real loss, which penalizes labs with heavier SFT relative to labs that release RL-only or base checkpoints, by an amount we can state per model. We show the residual next to the score rather than fold it in silently, and we treat a small residual as "probably register" only until a better recall test says otherwise.

What to trust, in numbers

ComparisonUncertaintyWhy
Two base models, same corpus, re-based ladder±0.001Corpus-sampling CI on ~1 MB of scored prose. Run noise of the re-basing fine-tune is ~0.0002.
Two releases from the same lab's pipeline (successive versions, SFT vs RL checkpoint)±0.001The pipeline offset cancels; within-pipeline spread is 0.0003.
A post-trained model vs its own public basemeasuredThe residual is observed directly; no estimate involved.
Post-trained models from different labs, base unknown±0.01 (ordinary SFT/instruct), ±0.03 (if possibly distilled)Unknown pipeline offset. The 90% interval from leave-one-family-out is ±0.019 on prose.
Same model, different corpus (prose vs code, or vs the 2026 corpus)unknown; prose/code agree at Spearman 0.70Not measurement noise: the ranking genuinely depends on the text. The 2026 corpus is the first independent draw.

The resolution metric, stated operationally: Δ95 = 0.027 is the 95th percentile of the absolute error in a predicted gap between two bases' losses, where the prediction comes from their post-trained descendants without seeing the bases. Pairs with a true gap under 0.02 are still misordered about a quarter of the time. Family gaps at fixed size are 0.01–0.03 and the 8B → 70B gap is 0.04, so this resolves tiers, not neighbours. Within a lab it resolves nearly everything.

Frontier models without a public base

The constancy of the pipeline offset suggests a way to make cross-lab comparisons at the frontier much tighter than ±0.01, without a base for every model. If a lab has released one base and one descendant, ever, the residual for that pair calibrates the lab's pipeline. Because the offset does not change along the pipeline and is flat or shrinking with size, it transfers to later releases from the same lab. DeepSeek V3-Base and V3 would calibrate DeepSeek; Kimi K2-Base and K2 would calibrate Moonshot; Qwen is already calibrated at seven sizes. The catch, visible in our own data: the offset moves between generations of the same lab (Qwen1.5 → 2 → 2.5 → 3, Llama 3.1 → 3.2 differ by 0.015–0.04), so a calibration transfers release-to-release on the same base generation, not across generations. Within that limit a later release could be reported with a bar near the within-pipeline spread rather than ±0.01. This is the concrete reason the DeepSeek V3 set is the first job for a second node: it is the only frontier-scale (base, descendant) pair that exists, and it anchors everything above 100B.

Labs that release no base at all (gpt-oss, GLM-4.5 and later, Kimi K3) would get the wide bar, with the residual prior set by whatever their pipeline most resembles, and the distillation flag raised if the native-to-recovered drop looks like the distilled models we have measured.

What we still do not know

  1. How much of a ranking is the corpus. Prose and code agree at 0.70. Until a third, multi-author corpus is scored, "the ranking" is "the ranking on this text".
  2. Whether the register story fully explains the native-vs-re-based gap. The 2026 multi-register corpus is the test: native and re-based rankings should converge on it.
  3. Whether a residual under 0.02 ever hides real damage. SimpleQA says rarely (3 of 33 pairs move beyond noise); a broader eval set (PopQA, TriviaQA) on the full set of pairs is running.
  4. Whether the residual keeps shrinking above 70B. We have no (base, descendant) pair larger than 72B. DeepSeek V3 is the one that exists.
  5. Whether the matched-base subtraction actually cancels what the fine-tune teaches. Every claim here assumes it does. The test: repeat the recovery with a second 8M-token public pool of a different register on six descendants spanning the residual range and their bases, and check whether residuals move by more than 0.003 or reorder. About 8–10 GPU-hours. Not yet run.
  6. Whether a longer or full-parameter recovery changes the residual. Recovery plateaus by update 600 of 1,000 (the two differ by 0.001), which says the residual is asymptotic for this recipe. A full fine-tune might reach further; it would also move the matched base further, and the difference is what matters. Not planned before launch.