The central measurement question in this project: when we take a post-trained model, fine-tune it briefly on public text, and score it on private text, how close is that number to what the pretrained base would have scored, and is the error stable enough that rankings survive it? This note walks through what we have measured, in order, and what it means.
Four numbers exist for every model that has a public base. All are loss on the same private prose windows, in nats per UTF-8 byte.
| Arm | What it is | Example: Llama-3.1-70B-Instruct |
|---|---|---|
| native base | The released pretrained checkpoint, scored as-is. | 0.9365 |
| matched base | The same base after our fixed recovery fine-tune (S1: LoRA rank 256, 1,000 updates, 8.19M tokens of public web text). This is the control. | 0.9198 |
| native | The post-trained release, scored as-is. | 1.08–1.3 for most instruct models |
| recovered | The post-trained release after the identical S1 fine-tune. This is the number we rank on. | 0.9383 |
The residual is recovered minus matched base: here 0.9383 − 0.9198 = +0.0185. It is the part of the post-training damage that the recovery fine-tune does not remove, measured against a base that went through exactly the same fine-tune so that whatever the fine-tune teaches cancels. For a model with no public base, the residual is unobservable and becomes the error in our estimate of its base. Everything below is about the structure of that residual, measured on 45 post-trained models across 24 bases in four ancestral families (OLMo, Llama, Qwen, Gemma), plus one blind family (SmolLM3, two models).
One framing before the data. The same fine-tune applied to the base and to the descendant is a controlled experiment. If post-training only changed format (chat templates, refusal habits, answer style), the fine-tune would undo it and the residual would be zero. If post-training overwrote what the base knew, no amount of web text brings that back and the residual would be large. The residual is therefore a measurement of something real about the post-training, not just a nuisance term. The question is whether it is stable enough to see through.
This was not in the original plan and matters for the base-model eval on its own.
The recovery fine-tune also moves base models, and by very different amounts. On 26 bases, the same 8M public tokens lower loss by 0.005 (OLMo-3-32B) to 0.053 (Qwen2.5-3B) for genuine pretrained checkpoints, and by 0.12–0.21 for the Qwen3 "Base" checkpoints that ship after an SFT stage. That is not noise. Our working interpretation, not yet tested, is that it is mostly register rather than knowledge: 8M tokens of ordinary web text should not teach a model much about the world, but the private prose runs to 2026 and the public pool is recent web text, so some information transfer is possible. The planned test is to repeat the fine-tune with a second public pool of a different register and see whether the drops and the residuals move. Register means how far the base's pretraining distribution sits from one person's conversational prose. A base trained on clean, filtered, partly synthetic text is surprised by messages and notes; a short exposure to messy web text removes the surprise.
Two rank reversals in the figure are the kind that decide whether a base-model eval is believable. Qwen2.5-3B scoring worse than Qwen2.5-1.5B natively is obviously wrong, and the re-based ladder fixes it. Gemma-3-27B going from worst to best among 27–32B bases is less obviously wrong either way, but Gemma 3's pretraining is known to include heavy distillation and instruction-shaped data, which is exactly what should produce a large register gap without a knowledge gap.
Consequence. The base-model ranking we publish should be the re-based ladder, not native loss, and every model on it, base or not, should go through the same fine-tune. Native loss conflates knowledge with register fit. This also makes a testable prediction for the 2026 multi-register corpus: with text from many registers, the native and re-based rankings should converge, because no single register dominates. If they do not, the register story is incomplete.
The worry: if two recovery runs on the same model, or two adjacent checkpoints in one lab's pipeline, gave residuals that differed by 0.01, then a 0.01 difference between two labs' models would be uninterpretable. The data say the opposite.
Three facts from this figure:
So the residual is a property of (base, SFT recipe), deterministic to our measurement precision. For a model with a known base, it is simply measured. For a model without one, it is an unknown constant in the range set by its recipe class, and since we usually do not know the recipe either, the practical prior is the spread across all pipelines we have seen: 0.0005–0.03 for ordinary instruct and SFT chains (most between 0.004 and 0.02), 0.03–0.06 for distillation. The same-base sibling spread, 0.028 on Llama-70B, 0.034 on Qwen-32B and 0.055 on Llama-8B, is the honest size of the cross-lab error.
If the SFT stage sets the residual, the residual should grow with how much SFT a model received. AI2 published five checkpoints along OLMo-3-7B's Think SFT run. We recovered each one.
This is the cleanest evidence that the residual is a real quantity. Nothing about the fine-tune changed between those five runs; only the SFT exposure of the input did. The same pattern appears across labs: Qwen2.5-Instruct, a light SFT pipeline, sits at a flat +0.008 from 0.5B to 72B; models that are known to be distilled or SFT-heavy (Llama 3.2 1B/3B, which Meta pruned and distilled from 8B/70B; the Qwen3 small checkpoints released after SFT; the R1-Distill series) are all at +0.03 to +0.06. We did not use any of that recipe information in the estimate. It simply agrees with it.
The concern: a 0.01 residual on an 8B model is small against the 0.04 gap between 8B and 70B, but at the frontier, where bases may sit within 0.01 of each other, the same residual would be the whole ranking.
So the direction is favourable: the residual does not grow with scale, and for the heavy pipelines it falls. But a frontier instruct model from a lab with a Llama-like pipeline would still carry about +0.015, and we would not know that number without its base. If two frontier bases truly differ by 0.01, the eval cannot order their post-trained releases across labs. It can say they are within 0.02 of each other, which is itself informative. And it can order releases within a lab to 0.0003, because the pipeline offset cancels.
A residual of +0.03 could mean the model lost knowledge the base had, or that the chat prior is so ingrained that 8M tokens of web text cannot reset it even though the knowledge is intact. These have different implications. If it is forgetting, the recovered number is the right number: the post-trained model really is worse. If it is register, the recovered number overstates the damage and the base number would be the right one.
SimpleQA (long-tail factual recall, run by us on 1,500 questions with a fixed open judge) lets us look, because we have base and descendant scores for several lineages:
| Base → descendant | Residual | SimpleQA base | SimpleQA desc. | Δ points |
|---|---|---|---|---|
| Llama-3.1-70B → Tulu-3-70B | +0.004 | 24.7 | 24.2 | −0.5 |
| Llama-3.1-70B → Llama-3.1-70B-Instruct | +0.019 | 24.7 | 22.8 | −1.9 |
| Llama-3.1-70B → R1-Distill-Llama-70B | +0.032 | 24.7 | 10.0 | −14.7 |
| Llama-3.1-8B → Tulu-3-8B | +0.009 | 7.6 | 8.2 | +0.6 |
| Llama-3.1-8B → R1-Distill-Llama-8B | +0.064 | 7.6 | 2.8 | −4.8 |
| Qwen2.5-32B → R1-Distill-Qwen-32B | +0.034 | 6.5 | 6.5 | 0.0 |
| Qwen2.5-72B → Qwen2.5-72B-Instruct | +0.010 | 9.4 | 11.2 | +1.8 |
| OLMo-2-7B → OLMo-2-7B-Instruct | +0.010 | 4.9 | 4.2 | −0.7 |
| Qwen2.5-32B → Open-Reasoner-Zero-32B | +0.001 | 6.5 | 5.9 | −0.6 |
Reading this: residuals below about 0.02 mostly come with no measurable change in factual recall (SimpleQA noise at n=1,500 is roughly ±1 point). Across all 33 such pairs we have, 3 move by more than two standard errors, the largest being Gemma-3-27B-IT at −2.7 points with a residual of 0.014; across 41 pairs the Spearman between residual and SimpleQA change is −0.12 [−0.65, +0.29]. So small residuals are usually register the fine-tune could not fully reset, not lost knowledge, but not always. Large residuals from distillation are mixed: R1-Distill-Llama-70B lost 60% of its long-tail recall, R1-Distill-Qwen-32B lost none, at nearly the same residual. So the residual is not a pure forgetting meter. It is an upper bound on damage that sometimes is all register.
What this means for the number we publish. We cannot yet tell forgetting from deep register shift from the residual alone. Reporting recovered loss as the model's score treats the whole residual as real loss, which penalizes labs with heavier SFT relative to labs that release RL-only or base checkpoints, by an amount we can state per model. We show the residual next to the score rather than fold it in silently, and we treat a small residual as "probably register" only until a better recall test says otherwise.
| Comparison | Uncertainty | Why |
|---|---|---|
| Two base models, same corpus, re-based ladder | ±0.001 | Corpus-sampling CI on ~1 MB of scored prose. Run noise of the re-basing fine-tune is ~0.0002. |
| Two releases from the same lab's pipeline (successive versions, SFT vs RL checkpoint) | ±0.001 | The pipeline offset cancels; within-pipeline spread is 0.0003. |
| A post-trained model vs its own public base | measured | The residual is observed directly; no estimate involved. |
| Post-trained models from different labs, base unknown | ±0.01 (ordinary SFT/instruct), ±0.03 (if possibly distilled) | Unknown pipeline offset. The 90% interval from leave-one-family-out is ±0.019 on prose. |
| Same model, different corpus (prose vs code, or vs the 2026 corpus) | unknown; prose/code agree at Spearman 0.70 | Not measurement noise: the ranking genuinely depends on the text. The 2026 corpus is the first independent draw. |
The resolution metric, stated operationally: Δ95 = 0.027 is the 95th percentile of the absolute error in a predicted gap between two bases' losses, where the prediction comes from their post-trained descendants without seeing the bases. Pairs with a true gap under 0.02 are still misordered about a quarter of the time. Family gaps at fixed size are 0.01–0.03 and the 8B → 70B gap is 0.04, so this resolves tiers, not neighbours. Within a lab it resolves nearly everything.
The constancy of the pipeline offset suggests a way to make cross-lab comparisons at the frontier much tighter than ±0.01, without a base for every model. If a lab has released one base and one descendant, ever, the residual for that pair calibrates the lab's pipeline. Because the offset does not change along the pipeline and is flat or shrinking with size, it transfers to later releases from the same lab. DeepSeek V3-Base and V3 would calibrate DeepSeek; Kimi K2-Base and K2 would calibrate Moonshot; Qwen is already calibrated at seven sizes. The catch, visible in our own data: the offset moves between generations of the same lab (Qwen1.5 → 2 → 2.5 → 3, Llama 3.1 → 3.2 differ by 0.015–0.04), so a calibration transfers release-to-release on the same base generation, not across generations. Within that limit a later release could be reported with a bar near the within-pipeline spread rather than ±0.01. This is the concrete reason the DeepSeek V3 set is the first job for a second node: it is the only frontier-scale (base, descendant) pair that exists, and it anchors everything above 100B.
Labs that release no base at all (gpt-oss, GLM-4.5 and later, Kimi K3) would get the wide bar, with the residual prior set by whatever their pipeline most resembles, and the distillation flag raised if the native-to-recovered drop looks like the distilled models we have measured.