Compression BenchRankingMethodologyStatus
Will DePue · brief for Mercor · updated 8 October 2026

Compression Bench

A language-model benchmark that cannot be trained on: we measure how well a model predicts private, human-written text that has never been on the internet, and we make it work for models that were only released after post-training.

What it is

The non-technical version.

Public benchmarks are losing their meaning. Their questions leak into training data, labs tune against them, and a score of 85 on MMLU now says more about the post-training recipe than about the underlying model. What people actually want to know about a new release is simpler: how good is the pretrained model underneath? How much of the world did it absorb, and how well does it generalize to text it has never seen?

Compression Bench answers that with the oldest measurement in language modelling: how well does the model predict the next byte of text it has never seen? The twist is the text. We use a private corpus of real human writing (prose and code) that has never been published, indexed or crawled, so no model has trained on it and no lab can target it. The score is loss in nats per byte: lower means the model finds the text less surprising, which means it understood more of the world the text came from.

Two things make this hard, and both have a working answer; neither is proven yet. First, most frontier models are released only after instruction tuning, RL or distillation, which wrecks their ability to predict raw text even though the pretrained knowledge is still inside. We have a fixed, cheap recipe that reverses most of that damage, so the eval works on released checkpoints without a public base. Second, a private corpus is only trustworthy if it is held out and diverse enough that a ranking means "model quality" rather than "fits this one author". That is where Mercor comes in.

The data today

Everything scored so far runs on one person's private text. That is the prototype, and its limits are the reason for the ask later on this page.

CorpusWhat it isSizeRole
Private proseOne author's personal writing from 2019 to 2026: messages and conversations, notes, documents and essays. Never published, indexed or shared with a model provider. A screen drops any document containing a URL or phrasing that suggests it was ever posted.dev 2,901 docs
≈1.0 MB scored
test 38,615 docs
≈25 MB
Every recipe decision and every number on this page uses the development split. The test split was opened once, to check that rankings hold across splits (Spearman 0.99).
Private codeThe same author's private repositories, pre-2025, with no AI-assisted commits. JavaScript 74%, TypeScript 16%, then HTML, C++, Python, CSS. Two repositories supply 94% of the tokens.1,373 records
1.40M tokens
5.4 MB
150K lines
Test-only: a second, independent corpus that no recipe decision has touched. It is what tells us prose and code rankings agree only at Spearman 0.70.
Public recovery textOrdinary public web text used by the S1 fine-tune, identical for every model, screened for overlap with the private corpora (no repeated spans found).8.19M tokens
per model
Training input for recovery only. It is never scored and never contributes to a ranking.
FineWeb sample2,000 documents of public web text.1.3M tokensPublic control: the baseline that private text has to beat as a predictor.

Scoring reads the private text in fixed 2,047-byte windows and reports nats per UTF-8 byte, so every tokenizer sees exactly the same bytes. Raw text, per-record losses and model weights stay on the compute node; only aggregate numbers leave it.

Three limits follow directly from this table. The prose is one author, so a ranking could reflect fit to that author's style and interests rather than model quality. The code is one author and two repositories, so the same applies. And the development prose is small, about a megabyte of scored bytes, which bounds how finely two models can be separated. None of these are fixable with more text from the same source.

How it works

The medium-technical version, for a research engineer.

The measurement

Every model is scored on the same fixed set of 2,047-byte windows cut from the private corpus. We compute total negative log-likelihood of the window in nats and divide by UTF-8 bytes, so models with different tokenizers are directly comparable (the window set is chosen so every tokenizer round-trips it losslessly). Prose and code are separate corpora and separate numbers. Scoring a model is one forward pass over the window set: minutes for a 7B model, well under an hour for 70B.

Recovery for post-trained models

Instruction tuning, SFT and distillation raise a model's raw-text loss far more than they change what it knows. Scoring a post-trained model directly therefore tells you about its chat template, not its pretraining. Our recovery recipe, S1, is a short LoRA fine-tune (rank 256, 1,000 updates, about 8M tokens of ordinary public web text, identical for every model) that teaches the model to be a plain language model again. The private text is never used in training. After recovery, the loss on private text is the model's recovered loss, and that is the number we rank on.

To know whether recovery works we need ground truth. For families that publish both the base and the post-trained checkpoints (OLMo, Llama, Qwen) we apply the same recipe to the base as well, and compare: the gap between the recovered descendant and the recovered base is the residual error of the method. We fit and evaluate everything leaving one family out, so each family's error is measured by a recipe never tuned on it.

0.027
Δ95 resolution, nats/byte, 33 post-trained models, each family held out
±0.003
run-to-run noise: re-running recovery with a different public-data order
0.001 → 0.06
residual vs. true base, from RL-only models to foreign-teacher distills

Δ95 is our resolution metric: the 95th percentile of the error in a predicted gap between two bases' losses, estimated from their post-trained descendants without seeing the bases. Base models in the current roster span about 0.93 to 1.19 nats/byte on prose, so 0.027 separates the roster into clear tiers; it is not yet fine enough to order two models from the same lab and size. Reducing it is the main engineering goal.

What recovery does and does not remove

The residual after recovery is not random. It tracks how far post-training moved the model: near zero for models trained with RL directly from the base (Open-Reasoner-Zero, OLMo-3 RL-Zero), 0.004 to 0.012 for ordinary SFT and DPO chains, 0.007 to 0.026 for direct instruction tuning and reasoning SFT, and 0.03 to 0.06 for models distilled from a different teacher (the R1-Distill series, the pruned-and-distilled Llama 3.2 small models). We read that as signal rather than error: heavier SFT and distillation may erode what the pretrained model knew. Our own long-tail recall check (SimpleQA) shows that for some distilled models and not others, so the residual is an upper bound on damage, not a measurement of it. The ranking is the recovered loss itself; no declared training-recipe information enters the estimate, because the models we most want to score do not disclose their recipes.

Recovered loss against the true (matched) base loss, private prose, 31 post-trained models from three families. The diagonal is perfect recovery; the shaded band is the ±0.003 run-to-run noise. Every model lands above the line by an amount that grows with how much post-training it had. The three Llama-3.1-70B descendants show the spread for one base: Tulu 3 (SFT/DPO/RL) sits 0.004 above, Llama-3.1-70B-Instruct 0.019, R1-Distill-Llama-70B 0.032.
The same residuals grouped by post-training recipe (recipe is a label here, never an input). Each dot is one model; the vertical band is the ±0.003 noise floor. RL directly from a base leaves the model essentially recoverable; ordinary SFT and DPO cost under 0.01; reasoning SFT 0.015 to 0.026; distillation from a different teacher 0.03 to 0.06. The spread within a group, not the group means, is what a better estimator has to shrink. The methodology note goes through what this residual is, how stable it is, and what it does and does not mean.
Native prose loss (nats per byte, lower is better) against parameter count for the 19 base checkpoints scored so far. Loss falls with size within every family, and families differ at fixed size. Qwen3 "Base" checkpoints sit far above the trend: they are released after a substantial SFT stage, and recovery brings them back onto it.

Where we are

Measured as of 8 October 2026 on one 8×B300 node. Development families are OLMo, Llama and Qwen; Gemma 3 was opened for diagnostics; SmolLM3 is the one blind family scored so far.

ComponentStateEvidence
Private corpus and scoring conditiondonefrozenOne author's prose (dev ≈1 MB scored, test ≈25 MB) and code (5.4 MB); fixed byte windows; scoring condition versioned and unchanged since 4 Oct
S1 recovery recipedonefrozenΔ95 0.027 prose / 0.027 code across 33 post-trained models, every family held out; 14 GPU-hours to score the whole registry
Noise floordoneRe-running recovery with a different public-data order moves a score by at most 0.003; a 600-update variant lands within 0.001 of the 1,000-update recipe at 40% less training time
Ranking stability on same-author textdoneTen bases ranked on development vs test splits agree at Spearman 0.99
Base rosterin progress54 bases in 11 families scored natively; 27 re-based so far, 28 queued
Blind confirmation on unseen familiespartialSmolLM3-3B run with the recipe frozen: errors +0.0008 prose / +0.0078 code, inside the intervals. Five other sealed families were incompatible with the frozen loader; no cross-family blind Δ95 yet
Predictive validity against capability evalsin progressOn 44 bases, private prose loss does not improve prediction of SimpleQA or MMLU-Pro beyond parameter count; private code loss carries size-independent signal (partial r −0.3 to −0.7) but not enough to lower held-out error yet. PopQA and TriviaQA are running; the 2026 corpus will give a second text
Author independencenot yetProse and code rank the same bases at Spearman 0.70. Both are one author. This is the open problem that multi-author data resolves

Where this stands. The eval is mostly working. Recovered loss sits close to the true base loss for lightly post-trained models and drifts away in proportion to how much SFT a model has had, which matches what we already believe about SFT. Across the base roster, loss falls cleanly with size inside every family and families separate at fixed size. What we have not shown yet is predictive capacity: whether these numbers forecast anything a lab cares about beyond parameter count, and how much a ranking depends on the text it was measured on. It is possible that SFT destroys too much for the recovered number to be useful on heavily post-trained models, and possible that the eval is simply less informative than it looks; the first fixed test (parameter count vs loss on SimpleQA and MMLU-Pro) has not shown a gain beyond size, and the better-targeted tests are still running. The work now is a better estimator, lower variance in the ranking, our own evals to validate against, and more diverse text.

What is next

Six bounded workstreams, started 6 October, each with a fixed deliverable.

  1. Roster expansion. Score 25+ additional public bases (Falcon, Apertus, StableLM, Pythia, older Llama and Qwen generations, small Gemma and OLMo variants) so that "does loss predict anything beyond size" can be tested across 12+ families and every size band.
  2. A 2026 held-out web corpus. Build a diverse corpus of human-written text published after every model's training cutoff (blogs, forums, court opinions, local news, preprints, fiction, new repositories), verified absent from Common Crawl and the Wayback Machine before 2026. This is our own first step toward author independence.
  3. Evals we run ourselves. SimpleQA, PopQA, TriviaQA, MMLU-Pro and GSM8K under one harness on every base and post-trained model, plus Open LLM Leaderboard v2 and LMArena scores for the same checkpoints.
  4. Meaning analysis. With the above: partial correlation of each loss corpus with each eval controlling for size; within-size-band rank agreement; private text versus 2026 web text versus FineWeb-Edu as predictors. The method is fixed before the data arrive.
  5. Reporting. Retire the class-based correction; report recovered loss with a single multiplicative base-equivalent label and a leave-one-family-out error bar.
  6. Blind confirmation. Run the frozen recipe on two sealed families and report the error without refitting anything.

Compute for all six is under 150 GPU-hours on the current node. The constraint for this phase is data, not GPUs.

Beyond this phase: a second node

Everything above fits on one 8×B300 node. The frontier-scale step does not. The ask we would make is for a second node of the same shape. With it we would, in order: run the DeepSeek V3-Base / R1-Zero / R1 / V3 set as a frontier-scale answer key, the only place where a 600B+ base and its post-trained descendants are both public (about 1,300 GPU-hours for the full set); then Kimi K2 across both nodes; then the base-less releases that are the real target of the recovery recipe, GLM, K3 and gpt-oss, where no public base exists and the recovered number is the only pretraining measurement anyone can make. One node can score models up to the 70B class comfortably; two nodes make the trillion-parameter class a routine run rather than a project.

What we would like Mercor to collect

Why multi-author held-out text is the thing that turns this from a method into a benchmark.

The private corpus proves the method but limits the claim. Every number above is "how well does this model predict Will's writing". To publish "how well does this model predict human text it has never seen", we need text from many unrelated authors, across registers and domains, that is verifiably unpublished, licensed for evaluation only, and never used for training by anyone. Mercor's expert network can reach authors who write real documents for a living and can attest, individually, that what they hand us is theirs and unpublished. That attestation is the asset; the bytes are secondary.

The one requirement that cannot be relaxed: no lab can ever have trained on it. Not in pretraining, not in SFT, not in RL, not as a reward-model or grader input, not as an eval that was later folded into training. The moment a document has been in any lab's training pipeline, its loss stops measuring what the model learned about the world and starts measuring whether it memorized that document, and the whole corpus is contaminated for that lab's models. This is the hard part for Mercor specifically: most of what the expert network produces is sold to labs as training data, so it is exactly the text we cannot use. We need the complement. Documents that have never touched the web and have never been delivered to any lab for any purpose.

Concretely, a document qualifies only if the author can attest that it has never been: posted or shared publicly; uploaded to or pasted into a chat model, coding assistant or any other LLM product; submitted to any lab, data vendor or annotation platform, including Mercor's own training-data pipelines; or included in any dataset licensed to anyone. And it must stay that way: the author licenses the text for evaluation only, and Mercor agrees not to resell, re-license or reuse the same documents, or near-duplicates of them, as training data for anyone, ever. A document that later leaks has to be retired from the corpus, which is why we want each one tagged with its own id and author so retirement is clean.

Pre-existing private writing (old internal docs, memos, code, letters) is better than text written to order. Text written on commission for us is fine as long as it meets the same rule, but it must be written by the expert directly, without a model in the loop, and never used by Mercor for anything else afterwards.

What counts as contamination

The question is always the same: could a model have seen the specific text, not just the subject. Three cases, from worst to fine.

  • The text itself, or pieces of it, exist somewhere a lab could train on. This is the problem case. It includes the obvious (the document was posted, sold or pasted into a chat model) and the less obvious: later documents that quote, copy-paste or closely paraphrase it. A carved-out memo whose follow-up memos, slides and emails are not carved out is leaking through them. So the carve-out unit has to be the thing that shares text: the whole repository rather than some files, the whole project or thread rather than one document, the whole client matter rather than one brief. Closed projects and archived material are safest because nothing will quote them next quarter.
  • Loose information about it is out there. The company is known, the project was announced, the field is public, the author has written publicly on the topic. This is fine. It is exactly the general knowledge the eval is supposed to reward.
  • Related but distant training work. An expert who wrote RL tasks or training data in the same domain, even for the same company, is fine as long as none of that work contains or was built on the carved-out text. Same domain is fine; same document is not.

Since authors cannot know what will be copied later, every document carries an id and author tag, and we re-score the corpus on each new model release looking for documents whose loss drops out of line with the rest. Those get retired. The published number is always tied to a corpus version.

Target strata

Eight strata, about 60 MB total after filtering, roughly 15–20M tokens. No stratum over 20% of bytes; no single author over 2% of any stratum. A third of every stratum, chosen by hash before anyone scores anything, is sealed and never looked at until a blind confirmation run.

StratumExample sources via the expert networkMBWhy it matters
Company code with historyPrivate repos from engineers at small firms, with commit messages and diffs; internal tools, not forks of open source12Code is our second-strongest signal; multi-author code tests whether code rankings transfer across codebases
Internal technical proseDesign docs, postmortems, RFCs, runbooks, architecture memos10Dense long-tail knowledge; the register where large models most clearly pull ahead
Expert professional analysisLawyers' memos, analysts' notes, de-identified clinical write-ups, engineers' failure reports9Tests nuance and domain reasoning rather than style; closest to "big-model smell"
Internal discussionChat and email threads, meeting notes, code-review comments, all participants consenting7Conversational, multi-speaker, abbreviated; a register absent from our current sets
Non-English proseNative-speaker writing in 4–6 languages: essays, work notes, letters8v1 is English-only; we need to know whether rankings hold across languages at all
Personal writing, many peopleJournals, letters, unpublished essays, newsletters sent only to friends6The direct control for our single-author prose set
Teaching and explanatory proseLecture notes, course handouts, tutoring write-ups never put online4Expository register; separates memorized explanations from understood ones
Fiction and creative proseUnpublished drafts, workshop pieces, screenplays4Style-heavy register where large models separate on nuance, not facts

Provenance per document

A short signed record, not a legal contract:

  • Authorship: the contributor wrote it themselves; co-authors named.
  • Exposure status: never posted publicly, submitted to a forum, pasted into an LLM product, or delivered to any lab, vendor or annotation platform as training, eval or grading data. If it was ever public or delivered anywhere, the earliest date that could have happened, so we can decide per model.
  • Date written (month and year), domain and language.
  • No LLM assistance: not drafted, rewritten or substantially edited by a model. Spell-check is fine.
  • Consent and license scope: evaluation only. Used to compute loss numbers; never trained on by us or anyone; never resold or re-licensed by Mercor; never quoted beyond short fragments in audits.
  • Sampling audit: we spot-check about 2% of documents by a brief author interview plus our own web-presence and AI-text checks. A failed check removes the document and flags the stratum, not the author.

What makes a document useful

Perplexity evaluation rewards natural text with real content. Useful: 2–50 KB per document; written for a reader with a purpose (a memo someone had to act on, a thread that resolved something, a chapter someone wanted finished); native register with abbreviations, hedges and mistakes intact; code with its commit messages and review comments rather than files alone.

Not useful, and filtered out: forms, templates, boilerplate, tables of numbers, logs, config dumps; translations, summaries or reprints of anything that exists elsewhere; AI-drafted or AI-polished text, which collapses toward what every model already predicts; PII beyond what the author consented to (third-party names replaced before we see it); many short fragments from one author. We want breadth across people, not depth in any one.

Why this is research, not procurement

A multi-author corpus is a measuring instrument we do not currently have. It lets us measure three things directly:

  • Author-specificity. How much of a model's score on the existing prose set is explained by that one author, with bootstrap intervals over authors rather than over bytes.
  • Cross-domain rank stability. Eight strata give eight rankings of the same roster. Where they agree we have a robust ordering; where they disagree we have found a real capability axis rather than a measurement artifact.
  • "Big-model smell." On expert and internal technical writing we expect large models to separate by recalled detail and nuance, not style. The strata split cleanly on that question, and we can quantify how much of the gap survives the recovery fine-tune.

The sealed third is what makes any of this publishable: everything we tune is done on the open two-thirds; the sealed third is scored once by a frozen pipeline and reported without adjustment. We would rather have 40 MB that passes every check than 100 MB that passes most of them.

Open research questions

Things the data would let us answer, in rough order of how much we care.

  1. Does private held-out loss predict long-tail knowledge and human preference beyond parameter count? If yes, the eval measures the thing labs actually buy pretraining compute for. If no, it is a compression score, which is still useful but a different product.
  2. How author-specific is a loss ranking? With eight or more independent author strata we can measure between-stratum rank variance directly and report a ranking with an interval that includes it.
  3. Does the recovery residual measure post-training damage? If distilled and heavily SFT'd models lose recoverable knowledge, that is a quantity labs would want to track per release, and it falls out of the same run.
  4. Can the eval be run on base-less frontier releases? The recipe already runs on any open checkpoint. With DeepSeek V3-Base/R1 as a frontier-scale answer key, we could publish numbers for gpt-oss, GLM, Kimi and similar releases that have no public base at all.