Compression BenchRankingMethodologyStatus
status · updated 8 October 2026, 06:40 UTC

What is being worked on

Everything in flight, what it is for, and what done looks like. Updated when something changes, not on a schedule.

Right now (overnight 8 October)

Three agents on one GPU (the rest of the node is lent to another project) plus CPU. Results expected the morning of 8 October UTC.

WorkstreamWhat it producesWhyState
2026 held-out corpus v230–50 MB of human-written 2026 text across ≥10 source types (court opinions, single-author blogs and newsletters, hobby forums, fiction, explanatory writing, general-interest features; topical news capped), ≥10 MB of non-AI 2026 code, per-document provenance, a filled per-model cutoff table, frozen scoring windows, scores for all 54 bases, rank agreement with the private corpora, and a recency check on old-cutoff models.Every number so far is one author's text. This is the first independent draw; it tells us how much a ranking depends on the corpus. v1 was 60% news and had no code; it is being rebuilt.running
Evals v2SimpleQA, PopQA and TriviaQA across all 54 bases and ~55 post-trained models, after a validation pass (raw prompt/completion dumps per model kind, judge agreement on 200 items, fixed protocols). MMLU-Pro kept as a secondary column only.The meaning question: does private-text loss predict long-tail knowledge beyond parameter count. MMLU-Pro is a reasoning benchmark and was the wrong first target. The v1 SimpleQA 0-shot variant scored 0.0 on bases and is being replaced.running
Re-basing the remaining basesMatched S1 runs for the 28 bases that only have native scores, so the whole ladder is re-based.The re-based ladder is the published base ranking (owner decision 8 Oct): native loss conflates knowledge with register fit and reorders models by up to 0.05.queued behind corpus scoring and base evals
AnalysisLaunch table (all 103 scored models), re-based-ladder analysis, meaning analysis (size vs loss vs each eval, leave-one-family-out), and a recheck of every number on the methodology note.Turns the pile of result files into the pages on this site.running (CPU)
Site and methodology auditAn adversarial review of these pages: numerical spot-checks against the source JSON, claim-by-claim critique, overreach list, copy fixes.Before anyone outside reads this closely.running

Overnight results as they land

Time (UTC)FromResult
05:45analysisBeyond size, first pass (evals v1 data): no for prose. On 44 bases / 9 families, adding private prose loss to parameter count does not improve leave-one-family-out prediction of SimpleQA or MMLU-Pro (intervals include zero). Private code loss does carry size-independent signal: partial correlation given log-params is −0.3 to −0.7 with intervals excluding zero on SimpleQA, MMLU-Pro, BBH, GPQA and MATH-hard, but the MAE gain still includes zero. Among post-trained models, recovered prose predicts SimpleQA better than native prose (MAE 4.5 vs 6.6) but not better than size (4.6). Rerun when evals v2 lands.
05:45analysisRe-based ladder: Kendall τ 0.72 vs native across 26 bases; 46 of 325 pairs reverse, mostly the post-SFT Qwen3 checkpoints and Gemma 3. Re-basing makes the ladder more size-ordered (Spearman with size −0.77 → −0.95). Launch table: 113 rows, 85 ranked, 28 bases pending re-basing (queued).
05:45analysisMethodology note corrected on eight points after recomputation over all 45 pairs (drop range, recovery share, between-pipeline spread, recipe-class R² 48% on 45 pairs vs 78% on the original 31, a mislabeled RL-Zero model, the SimpleQA "no change under 0.02" claim softened: 3 of 33 pairs move beyond noise).

Done and frozen

ItemResult
Scoring conditionFixed 2,047-byte UTF-8 windows, lossless across every roster tokenizer, nats per byte; unchanged since 4 Oct.
S1 recovery recipeLoRA rank 256, 1,000 updates, 8.19M public web tokens, identical for every model; plateaus by update 600. Frozen.
Base roster54 bases, 11 families, scored natively on prose and code (1.5 GPU-hours). 26 have matched S1 runs.
Post-trained roster49 models with a public base recovered and compared to their matched base; residuals 0.0005–0.064.
Residual structureConstant along a lab's pipeline (sd 0.0003 across SFT/DPO/RL), log-linear in SFT steps, shrinks with size for heavy pipelines, large for distillation. See the methodology note.
ResolutionΔ95 = 0.027 nats/byte (prose and code) across labs without a base; ±0.001 within a lab or between bases.
Blind familySmolLM3-3B: frozen-estimator errors +0.0008 prose / +0.0078 code, inside the intervals. One family only; five other sealed families were incompatible with the frozen loader and remain unscored.
First meaning testMMLU-Pro on 19 bases: loss does not beat parameter count. Treated as the wrong benchmark for this question; superseded by evals v2.

Next, in order

  1. Read the overnight results and decide whether private loss adds anything beyond size on long-tail knowledge evals. If yes, the site says so with the numbers. If no, the claim stays "an interesting score whose meaning we do not know yet".
  2. Recovery-pool control (9 October, owner-approved): re-run S1 with a second 8M-token public pool of a different register on six descendants spanning the residual range and their four bases; if residuals move by more than 0.003 or reorder, the matched-base subtraction is not a cancellation and the calibration and register sections of the methodology note are withdrawn. ~8–10 GPU-hours.
  3. Launch page: the re-based ladder with error bars, the 2026 corpus column, the methodology note, the corpus description. Target: this week.
  4. Multi-author held-out text with Mercor: the thing that turns a one-author score into a benchmark. See the overview page.
  5. Frontier (base, descendant) pairs: DeepSeek V3-Base/V3 to calibrate the pipeline offset above 100B; needs a second node.
  6. Mistral as a second blind family, once the loader supports its tokenizer pool. Not before launch.

Stopped

Further recovery-recipe tuning (saturated: 600 vs 1,000 updates differ by 0.001); additional estimator variants (class-based corrections rejected because deployed models do not declare recipes; β and power-law corrections tested and worse); readiness gates and seals as process; public model-card score auditing. The 2026 corpus v1 (news-heavy) is superseded by v2.