A language-model benchmark that cannot be trained on: we measure how well a model predicts private, human-written text that has never been on the internet, and we make it work for models that were only released after post-training.
The non-technical version.
Public benchmarks are losing their meaning. Their questions leak into training data, labs tune against them, and a score of 85 on MMLU now says more about the post-training recipe than about the underlying model. What people actually want to know about a new release is simpler: how good is the pretrained model underneath? How much of the world did it absorb, and how well does it generalize to text it has never seen?
Compression Bench answers that with the oldest measurement in language modelling: how well does the model predict the next byte of text it has never seen? The twist is the text. We use a private corpus of real human writing (prose and code) that has never been published, indexed or crawled, so no model has trained on it and no lab can target it. The score is loss in nats per byte: lower means the model finds the text less surprising, which means it understood more of the world the text came from.
Two things make this hard, and both have a working answer; neither is proven yet. First, most frontier models are released only after instruction tuning, RL or distillation, which wrecks their ability to predict raw text even though the pretrained knowledge is still inside. We have a fixed, cheap recipe that reverses most of that damage, so the eval works on released checkpoints without a public base. Second, a private corpus is only trustworthy if it is held out and diverse enough that a ranking means "model quality" rather than "fits this one author". That is where Mercor comes in.
Everything scored so far runs on one person's private text. That is the prototype, and its limits are the reason for the ask later on this page.
| Corpus | What it is | Size | Role |
|---|---|---|---|
| Private prose | One author's personal writing from 2019 to 2026: messages and conversations, notes, documents and essays. Never published, indexed or shared with a model provider. A screen drops any document containing a URL or phrasing that suggests it was ever posted. | dev 2,901 docs ≈1.0 MB scored test 38,615 docs ≈25 MB | Every recipe decision and every number on this page uses the development split. The test split was opened once, to check that rankings hold across splits (Spearman 0.99). |
| Private code | The same author's private repositories, pre-2025, with no AI-assisted commits. JavaScript 74%, TypeScript 16%, then HTML, C++, Python, CSS. Two repositories supply 94% of the tokens. | 1,373 records 1.40M tokens 5.4 MB 150K lines | Test-only: a second, independent corpus that no recipe decision has touched. It is what tells us prose and code rankings agree only at Spearman 0.70. |
| Public recovery text | Ordinary public web text used by the S1 fine-tune, identical for every model, screened for overlap with the private corpora (no repeated spans found). | 8.19M tokens per model | Training input for recovery only. It is never scored and never contributes to a ranking. |
| FineWeb sample | 2,000 documents of public web text. | 1.3M tokens | Public control: the baseline that private text has to beat as a predictor. |
Scoring reads the private text in fixed 2,047-byte windows and reports nats per UTF-8 byte, so every tokenizer sees exactly the same bytes. Raw text, per-record losses and model weights stay on the compute node; only aggregate numbers leave it.
Three limits follow directly from this table. The prose is one author, so a ranking could reflect fit to that author's style and interests rather than model quality. The code is one author and two repositories, so the same applies. And the development prose is small, about a megabyte of scored bytes, which bounds how finely two models can be separated. None of these are fixable with more text from the same source.
The medium-technical version, for a research engineer.
Every model is scored on the same fixed set of 2,047-byte windows cut from the private corpus. We compute total negative log-likelihood of the window in nats and divide by UTF-8 bytes, so models with different tokenizers are directly comparable (the window set is chosen so every tokenizer round-trips it losslessly). Prose and code are separate corpora and separate numbers. Scoring a model is one forward pass over the window set: minutes for a 7B model, well under an hour for 70B.
Instruction tuning, SFT and distillation raise a model's raw-text loss far more than they change what it knows. Scoring a post-trained model directly therefore tells you about its chat template, not its pretraining. Our recovery recipe, S1, is a short LoRA fine-tune (rank 256, 1,000 updates, about 8M tokens of ordinary public web text, identical for every model) that teaches the model to be a plain language model again. The private text is never used in training. After recovery, the loss on private text is the model's recovered loss, and that is the number we rank on.
To know whether recovery works we need ground truth. For families that publish both the base and the post-trained checkpoints (OLMo, Llama, Qwen) we apply the same recipe to the base as well, and compare: the gap between the recovered descendant and the recovered base is the residual error of the method. We fit and evaluate everything leaving one family out, so each family's error is measured by a recipe never tuned on it.
Δ95 is our resolution metric: the 95th percentile of the error in a predicted gap between two bases' losses, estimated from their post-trained descendants without seeing the bases. Base models in the current roster span about 0.93 to 1.19 nats/byte on prose, so 0.027 separates the roster into clear tiers; it is not yet fine enough to order two models from the same lab and size. Reducing it is the main engineering goal.
The residual after recovery is not random. It tracks how far post-training moved the model: near zero for models trained with RL directly from the base (Open-Reasoner-Zero, OLMo-3 RL-Zero), 0.004 to 0.012 for ordinary SFT and DPO chains, 0.007 to 0.026 for direct instruction tuning and reasoning SFT, and 0.03 to 0.06 for models distilled from a different teacher (the R1-Distill series, the pruned-and-distilled Llama 3.2 small models). We read that as signal rather than error: heavier SFT and distillation may erode what the pretrained model knew. Our own long-tail recall check (SimpleQA) shows that for some distilled models and not others, so the residual is an upper bound on damage, not a measurement of it. The ranking is the recovered loss itself; no declared training-recipe information enters the estimate, because the models we most want to score do not disclose their recipes.
Measured as of 8 October 2026 on one 8×B300 node. Development families are OLMo, Llama and Qwen; Gemma 3 was opened for diagnostics; SmolLM3 is the one blind family scored so far.
| Component | State | Evidence |
|---|---|---|
| Private corpus and scoring condition | donefrozen | One author's prose (dev ≈1 MB scored, test ≈25 MB) and code (5.4 MB); fixed byte windows; scoring condition versioned and unchanged since 4 Oct |
| S1 recovery recipe | donefrozen | Δ95 0.027 prose / 0.027 code across 33 post-trained models, every family held out; 14 GPU-hours to score the whole registry |
| Noise floor | done | Re-running recovery with a different public-data order moves a score by at most 0.003; a 600-update variant lands within 0.001 of the 1,000-update recipe at 40% less training time |
| Ranking stability on same-author text | done | Ten bases ranked on development vs test splits agree at Spearman 0.99 |
| Base roster | in progress | 54 bases in 11 families scored natively; 27 re-based so far, 28 queued |
| Blind confirmation on unseen families | partial | SmolLM3-3B run with the recipe frozen: errors +0.0008 prose / +0.0078 code, inside the intervals. Five other sealed families were incompatible with the frozen loader; no cross-family blind Δ95 yet |
| Predictive validity against capability evals | in progress | On 44 bases, private prose loss does not improve prediction of SimpleQA or MMLU-Pro beyond parameter count; private code loss carries size-independent signal (partial r −0.3 to −0.7) but not enough to lower held-out error yet. PopQA and TriviaQA are running; the 2026 corpus will give a second text |
| Author independence | not yet | Prose and code rank the same bases at Spearman 0.70. Both are one author. This is the open problem that multi-author data resolves |
Where this stands. The eval is mostly working. Recovered loss sits close to the true base loss for lightly post-trained models and drifts away in proportion to how much SFT a model has had, which matches what we already believe about SFT. Across the base roster, loss falls cleanly with size inside every family and families separate at fixed size. What we have not shown yet is predictive capacity: whether these numbers forecast anything a lab cares about beyond parameter count, and how much a ranking depends on the text it was measured on. It is possible that SFT destroys too much for the recovered number to be useful on heavily post-trained models, and possible that the eval is simply less informative than it looks; the first fixed test (parameter count vs loss on SimpleQA and MMLU-Pro) has not shown a gain beyond size, and the better-targeted tests are still running. The work now is a better estimator, lower variance in the ranking, our own evals to validate against, and more diverse text.
Six bounded workstreams, started 6 October, each with a fixed deliverable.
Compute for all six is under 150 GPU-hours on the current node. The constraint for this phase is data, not GPUs.
Everything above fits on one 8×B300 node. The frontier-scale step does not. The ask we would make is for a second node of the same shape. With it we would, in order: run the DeepSeek V3-Base / R1-Zero / R1 / V3 set as a frontier-scale answer key, the only place where a 600B+ base and its post-trained descendants are both public (about 1,300 GPU-hours for the full set); then Kimi K2 across both nodes; then the base-less releases that are the real target of the recovery recipe, GLM, K3 and gpt-oss, where no public base exists and the recovered number is the only pretraining measurement anyone can make. One node can score models up to the 70B class comfortably; two nodes make the trillion-parameter class a routine run rather than a project.
Why multi-author held-out text is the thing that turns this from a method into a benchmark.
The private corpus proves the method but limits the claim. Every number above is "how well does this model predict Will's writing". To publish "how well does this model predict human text it has never seen", we need text from many unrelated authors, across registers and domains, that is verifiably unpublished, licensed for evaluation only, and never used for training by anyone. Mercor's expert network can reach authors who write real documents for a living and can attest, individually, that what they hand us is theirs and unpublished. That attestation is the asset; the bytes are secondary.
The one requirement that cannot be relaxed: no lab can ever have trained on it. Not in pretraining, not in SFT, not in RL, not as a reward-model or grader input, not as an eval that was later folded into training. The moment a document has been in any lab's training pipeline, its loss stops measuring what the model learned about the world and starts measuring whether it memorized that document, and the whole corpus is contaminated for that lab's models. This is the hard part for Mercor specifically: most of what the expert network produces is sold to labs as training data, so it is exactly the text we cannot use. We need the complement. Documents that have never touched the web and have never been delivered to any lab for any purpose.
Concretely, a document qualifies only if the author can attest that it has never been: posted or shared publicly; uploaded to or pasted into a chat model, coding assistant or any other LLM product; submitted to any lab, data vendor or annotation platform, including Mercor's own training-data pipelines; or included in any dataset licensed to anyone. And it must stay that way: the author licenses the text for evaluation only, and Mercor agrees not to resell, re-license or reuse the same documents, or near-duplicates of them, as training data for anyone, ever. A document that later leaks has to be retired from the corpus, which is why we want each one tagged with its own id and author so retirement is clean.
Pre-existing private writing (old internal docs, memos, code, letters) is better than text written to order. Text written on commission for us is fine as long as it meets the same rule, but it must be written by the expert directly, without a model in the loop, and never used by Mercor for anything else afterwards.
The question is always the same: could a model have seen the specific text, not just the subject. Three cases, from worst to fine.
Since authors cannot know what will be copied later, every document carries an id and author tag, and we re-score the corpus on each new model release looking for documents whose loss drops out of line with the rest. Those get retired. The published number is always tied to a corpus version.
Eight strata, about 60 MB total after filtering, roughly 15–20M tokens. No stratum over 20% of bytes; no single author over 2% of any stratum. A third of every stratum, chosen by hash before anyone scores anything, is sealed and never looked at until a blind confirmation run.
| Stratum | Example sources via the expert network | MB | Why it matters |
|---|---|---|---|
| Company code with history | Private repos from engineers at small firms, with commit messages and diffs; internal tools, not forks of open source | 12 | Code is our second-strongest signal; multi-author code tests whether code rankings transfer across codebases |
| Internal technical prose | Design docs, postmortems, RFCs, runbooks, architecture memos | 10 | Dense long-tail knowledge; the register where large models most clearly pull ahead |
| Expert professional analysis | Lawyers' memos, analysts' notes, de-identified clinical write-ups, engineers' failure reports | 9 | Tests nuance and domain reasoning rather than style; closest to "big-model smell" |
| Internal discussion | Chat and email threads, meeting notes, code-review comments, all participants consenting | 7 | Conversational, multi-speaker, abbreviated; a register absent from our current sets |
| Non-English prose | Native-speaker writing in 4–6 languages: essays, work notes, letters | 8 | v1 is English-only; we need to know whether rankings hold across languages at all |
| Personal writing, many people | Journals, letters, unpublished essays, newsletters sent only to friends | 6 | The direct control for our single-author prose set |
| Teaching and explanatory prose | Lecture notes, course handouts, tutoring write-ups never put online | 4 | Expository register; separates memorized explanations from understood ones |
| Fiction and creative prose | Unpublished drafts, workshop pieces, screenplays | 4 | Style-heavy register where large models separate on nuance, not facts |
A short signed record, not a legal contract:
Perplexity evaluation rewards natural text with real content. Useful: 2–50 KB per document; written for a reader with a purpose (a memo someone had to act on, a thread that resolved something, a chapter someone wanted finished); native register with abbreviations, hedges and mistakes intact; code with its commit messages and review comments rather than files alone.
Not useful, and filtered out: forms, templates, boilerplate, tables of numbers, logs, config dumps; translations, summaries or reprints of anything that exists elsewhere; AI-drafted or AI-polished text, which collapses toward what every model already predicts; PII beyond what the author consented to (third-party names replaced before we see it); many short fragments from one author. We want breadth across people, not depth in any one.
A multi-author corpus is a measuring instrument we do not currently have. It lets us measure three things directly:
The sealed third is what makes any of this publishable: everything we tune is done on the open two-thirds; the sealed third is scored once by a frozen pipeline and reported without adjustment. We would rather have 40 MB that passes every check than 100 MB that passes most of them.
Things the data would let us answer, in rough order of how much we care.