Everything in flight, what it is for, and what done looks like. Updated when something changes, not on a schedule.
Three agents on one GPU (the rest of the node is lent to another project) plus CPU. Results expected the morning of 8 October UTC.
| Workstream | What it produces | Why | State |
|---|---|---|---|
| 2026 held-out corpus v2 | 30–50 MB of human-written 2026 text across ≥10 source types (court opinions, single-author blogs and newsletters, hobby forums, fiction, explanatory writing, general-interest features; topical news capped), ≥10 MB of non-AI 2026 code, per-document provenance, a filled per-model cutoff table, frozen scoring windows, scores for all 54 bases, rank agreement with the private corpora, and a recency check on old-cutoff models. | Every number so far is one author's text. This is the first independent draw; it tells us how much a ranking depends on the corpus. v1 was 60% news and had no code; it is being rebuilt. | running |
| Evals v2 | SimpleQA, PopQA and TriviaQA across all 54 bases and ~55 post-trained models, after a validation pass (raw prompt/completion dumps per model kind, judge agreement on 200 items, fixed protocols). MMLU-Pro kept as a secondary column only. | The meaning question: does private-text loss predict long-tail knowledge beyond parameter count. MMLU-Pro is a reasoning benchmark and was the wrong first target. The v1 SimpleQA 0-shot variant scored 0.0 on bases and is being replaced. | running |
| Re-basing the remaining bases | Matched S1 runs for the 28 bases that only have native scores, so the whole ladder is re-based. | The re-based ladder is the published base ranking (owner decision 8 Oct): native loss conflates knowledge with register fit and reorders models by up to 0.05. | queued behind corpus scoring and base evals |
| Analysis | Launch table (all 103 scored models), re-based-ladder analysis, meaning analysis (size vs loss vs each eval, leave-one-family-out), and a recheck of every number on the methodology note. | Turns the pile of result files into the pages on this site. | running (CPU) |
| Site and methodology audit | An adversarial review of these pages: numerical spot-checks against the source JSON, claim-by-claim critique, overreach list, copy fixes. | Before anyone outside reads this closely. | running |
| Time (UTC) | From | Result |
|---|---|---|
| 05:45 | analysis | Beyond size, first pass (evals v1 data): no for prose. On 44 bases / 9 families, adding private prose loss to parameter count does not improve leave-one-family-out prediction of SimpleQA or MMLU-Pro (intervals include zero). Private code loss does carry size-independent signal: partial correlation given log-params is −0.3 to −0.7 with intervals excluding zero on SimpleQA, MMLU-Pro, BBH, GPQA and MATH-hard, but the MAE gain still includes zero. Among post-trained models, recovered prose predicts SimpleQA better than native prose (MAE 4.5 vs 6.6) but not better than size (4.6). Rerun when evals v2 lands. |
| 05:45 | analysis | Re-based ladder: Kendall τ 0.72 vs native across 26 bases; 46 of 325 pairs reverse, mostly the post-SFT Qwen3 checkpoints and Gemma 3. Re-basing makes the ladder more size-ordered (Spearman with size −0.77 → −0.95). Launch table: 113 rows, 85 ranked, 28 bases pending re-basing (queued). |
| 05:45 | analysis | Methodology note corrected on eight points after recomputation over all 45 pairs (drop range, recovery share, between-pipeline spread, recipe-class R² 48% on 45 pairs vs 78% on the original 31, a mislabeled RL-Zero model, the SimpleQA "no change under 0.02" claim softened: 3 of 33 pairs move beyond noise). |
| Item | Result |
|---|---|
| Scoring condition | Fixed 2,047-byte UTF-8 windows, lossless across every roster tokenizer, nats per byte; unchanged since 4 Oct. |
| S1 recovery recipe | LoRA rank 256, 1,000 updates, 8.19M public web tokens, identical for every model; plateaus by update 600. Frozen. |
| Base roster | 54 bases, 11 families, scored natively on prose and code (1.5 GPU-hours). 26 have matched S1 runs. |
| Post-trained roster | 49 models with a public base recovered and compared to their matched base; residuals 0.0005–0.064. |
| Residual structure | Constant along a lab's pipeline (sd 0.0003 across SFT/DPO/RL), log-linear in SFT steps, shrinks with size for heavy pipelines, large for distillation. See the methodology note. |
| Resolution | Δ95 = 0.027 nats/byte (prose and code) across labs without a base; ±0.001 within a lab or between bases. |
| Blind family | SmolLM3-3B: frozen-estimator errors +0.0008 prose / +0.0078 code, inside the intervals. One family only; five other sealed families were incompatible with the frozen loader and remain unscored. |
| First meaning test | MMLU-Pro on 19 bases: loss does not beat parameter count. Treated as the wrong benchmark for this question; superseded by evals v2. |
Further recovery-recipe tuning (saturated: 600 vs 1,000 updates differ by 0.001); additional estimator variants (class-based corrections rejected because deployed models do not declare recipes; β and power-law corrections tested and worse); readiness gates and seals as process; public model-card score auditing. The 2026 corpus v1 (news-heavy) is superseded by v2.