The winning ridge, stress-tested against longer training

The arch2 symbolic-regression run ended with a formula whose core is a ridge solve in the : solve (Gpre[S,S] + λI) x = 1 on the trained subset S, predict Gpre[:,S] x for everything else, with λ = λ0·tr(Gpre[S,S])/|S| and a fleet-tuned λ0 = 40. Follow-up runs at longer horizons showed the formula's falling as training went on. This piece turns the one knob the formula has: sweep λ0 over ten orders of magnitude, at every horizon, and see what was the constant's fault and what wasn't. Everything here uses the winner's pattern term only — the rank-1 interaction term is switched off throughout — and only training runs that actually reduce loss.

§1At the horizon it shipped for, 40 was exactly right

First the anchor. The held-out eval that picked the winner scores predictors on 16-step of gemma-3-1b across four optimizer families. Re-scoring the pattern term through the unmodified pipeline at thirteen values of λ0:

Held-out LDS vs ridge strength: rises from 0.35 at λ0=0.01 to 0.617 at λ0=40, then plateaus flat to λ0=10000.
Fig. 1 · Pattern-term LDS on the 16-step held-out 1B grid vs ridge strength. Thin blue lines: the four optimizer families; thick line: their mean — the shipped score. The dashed line marks the fleet's λ0 = 40.

λ0 = 40 is the exact argmax of the held-out score (0.617) — but the optimum is a one-sided plateau, and the solve itself is worth just +0.003 over its own λ→∞ limit.

Everything from λ0 ≈ 20 to 104 scores 0.614–0.617, while small λ falls off a cliff (0.35 at λ0 = 0.01). At sixteen steps the inversion barely matters — as λ→∞ the solution degenerates to a plain column sum of the Gram, and that already gets you 0.614. The fleet's constant sits at the start of the safe plateau: correctly tuned, for a dial that is nearly free at this horizon.

§2One line per λ, across 1200-step runs that learn

The stress test: 1200-step LoRA fine-tunes of gemma-3-1b on random 32-doc subsets — plain SGD and AdamW, learning rates calibrated so the runs genuinely descend (net probe-loss change −0.07 and −0.10 nats respectively), with loss snapshots at τ = 16, 64, 256, 1024 steps. Scoring is the eval's own protocol ( from a held-back half of runs, LDS on the other half; nval = 13 SGD / 16 AdamW). One line per λ0, goodness of fit against number of training steps:

LDS vs training steps, one line per lambda. SGD: all lambda lines flat around 0.2-0.3 at every horizon. AdamW: all lambda lines fall together from 0.31 at 16 steps to about zero at 256+ steps; the no-solve first-order line sits above all of them at late steps.
Fig. 2 · LDS vs training steps, one line per ridge strength (teal ramp, dark→pale = weak→strong ridge; orange = the shipped λ0 = 40). Grey dashed: .

Across eight orders of magnitude of λ the lines form one tight band: for these runs, ridge strength is nearly irrelevant — what changes with horizon is whether the geometry predicts at all.

Two entirely different fates, neither of which the dial controls. SGD stays predictable at every horizon — LDS ≈ 0.21–0.29 from 16 steps to 1024, drifting mildly upward for weaker ridges (best fixed retune λ0 ≈ 0.03: +0.054 at τ = 1024, paired bootstrap CI [−0.015, +0.115], P(Δ>0) = 0.94 — suggestive, not conclusive). AdamW collapses for every λ: 0.31 at 16 steps, ≈ 0 by 256, and the best value over the whole grid at τ ≥ 256 is 0.02.

§3What fails is the solve, not the constant

LDS vs horizon for shipped lambda, retuned lambda, and the no-solve first-order prediction, with bootstrap confidence whiskers. For AdamW the first-order line is above both ridge lines at 256 and 1024 steps.
Fig. 3 · The shipped ridge, the best fixed retune, and the solve-free first-order prediction across horizon. Whiskers: paired-bootstrap 95% CI (val runs resampled, nboot = 2000) on the retuned line.

For AdamW at τ = 1024 the ranking flips outright: the first-order prediction beats the best ridge in the sweep by Δ = 0.067, and the paired bootstrap puts that difference at 95% significance ([+0.003, +0.123] in favour of no-solve, P = 0.019). Once a preconditioned optimizer has run long enough to leave the regime where w0-gradient geometry describes the trajectory, inverting that geometry harder cannot help — the failure is the linearization, not the regularization.

armpredictorτ=16τ=64τ=256τ=1024
SGDridge @40 (shipped)+0.205+0.266+0.260+0.226
ridge @0.0316 (retuned)+0.210+0.289+0.290+0.280
first-order (no solve)+0.230+0.298+0.296+0.275
AdamWridge @40 (shipped)+0.308+0.130−0.079+0.006
ridge @0.001 (≈ best on grid)+0.322+0.148−0.047+0.018
first-order (no solve)+0.356+0.215+0.017+0.085

§4A data-hygiene footnote

The first pass of this sweep showed the SGD arm decaying too — 0.25 at 16 steps down to 0.03 at equilibrium — which would have made a tidier story. It was an artifact: the run pool contained four late top-up runs executed with a different config (3× the learning rate, ⅓ the weight-decay tether), three of them landing in the scored half. Filtering each arm to its modal config restored the flat curves above and reconciled the numbers with the committed baseline. Worth remembering whenever runs accumulate across fleet waves: the dedupe that keeps the latest file is not the dedupe that keeps the same experiment.

§5Why you'd expect the dial to want to shrink

Kernel-regression theory says gradient flow trained for time t is equivalent to ridge regression with λ ∝ 1/t (Ali, Kolter & Tibshirani 2019), and the same implicit damping appears inside modern unrolled-attribution methods (SOURCE, Bae et al. 2024: per-segment damping ≈ 1/(η̄K)). A constant tuned at 16 steps should be ~64× too strong by 1024 — and the SGD arm's best-λ does drift down with horizon (0.3 at τ = 16 → 0.03 later), in units where λ0 counts multiples of the subset Gram's mean eigenvalue (EK-FAC's damping convention is 0.1 in the same units; the shipped value is 40). But for these learning runs the drift is worth a few hundredths of LDS at most. The practical summary for the formula: the constant was never the problem; horizon is a regime change, not a recalibration.

Caveats: 13–16 scored runs per arm make single-cell LDS noisy (the bootstrap CIs in Fig. 3 are the honest read); the retuned λ is chosen on the same runs it is scored on, so treat retuned-vs-shipped gaps as upper bounds; and the two arms here carry small parameter noise and a decoupled weight-decay tether so that 1200-step runs stay bounded — the τ-snapshots are quenched loss changes measured against the same pre-run baseline as the 16-step eval. That last caveat turned out to matter: §6 removes it.

§6Postscript: on clean quenches, the optimum does move

Added 2026-07-14. The sections above left a confound: the only long-horizon runs carried injected noise and a decay-to-init tether, and there λ barely mattered. So we re-ran the question on true deterministic quenches — the held-out eval's own standard cell (identical 32-doc subsets, per-family learning rates, seeds, fp32 harness; τ=16 snapshots reproduce the original runs digit-for-digit), extended 16 → 1024 steps with probes at every power of two. 96 runs — SGD, Adam, Muon, LoRA-Adam — on a 15×H100 fleet, ~$220.

Four panels, one per optimizer family: LDS vs training steps with one line per ridge strength. SGD's lines fan apart after 64 steps with small-lambda lines on top; Adam, Muon and LoRA-Adam collapse to zero or below by 128-512 steps for every lambda; the no-solve line goes strongly negative late for SGD and Adam.
Fig. 4 · Clean quenches: LDS vs training steps, one line per λ₀ (dark→pale teal = weak→strong; orange = shipped λ₀ = 40; grey dashed = no-solve first-order). Unlike the noisy arms of Fig. 2, the SGD λ-lines fan apart beyond 64 steps.

On clean SGD quenches the ridge optimum slides down with horizon exactly as kernel-regression theory predicts — the same λ₀ = 0.316 that significantly hurts at 16 steps (P = 0.003) buys +0.30 of LDS at 256 steps (P = 0.997).

Left: LDS vs ridge strength for SGD, one line per horizon; the argmax star moves from the right plateau at 16 steps to progressively smaller lambda at 128 through 1024 steps. Right: optimal lambda vs training steps on log-log axes for SGD and Muon, falling between 1/tau and tau^-2 guide lines.
Fig. 5 · Left: the SGD LDS-vs-λ₀ curve per horizon; ★ = argmax. At τ = 16 the optimum is the familiar right-hand plateau; from τ = 128 an interior optimum appears and marches left — λ₀* ≈ 0.3 at 128, 0.1 at 256, 0.01 at 512. Right: λ₀*(τ) on log-log axes, falling between the gradient-flow 1/τ scaling and τ⁻² (steeper than 1/τ, consistent with response saturation on top of the time-rescaling).

Three things the clean runs settle. First, the horizon-dependent ridge is real — §2's "λ is nearly irrelevant" was a property of the noisy/tethered dynamics, not of long training as such. On clean SGD, re-tuning rescues late prediction outright: LDS 0.076 → 0.379 at τ = 256 (Δ = +0.303 [+0.072, +0.491], P = 0.997), 0.034 → 0.261 at τ = 512, with the paired sign-flip against τ = 16 making the moving optimum unambiguous. Second, at long horizon the solve is no longer decoration. The first-order response goes anti-predictive once out-of-subset transfer reverses into interference (mean out-of-S Δloss flips from −0.06 nats at τ = 16 to +0.1…+3.5 nats by τ = 256–1024): fdt scores −0.05 at τ = 256 while the retuned solve holds +0.38 (Δ = +0.43, P = 1.000). The 16-step picture — solve worth +0.003 over a column sum — inverts completely. Third, Adam's collapse is the linearization, full stop. With no noise and no tether, Adam still dies at τ ≥ 128 for every λ across twelve decades (best 0.09, then ≤ 0; LoRA-Adam identical), having dug ~2.3× deeper into the subset by that point than SGD. Muon sits in between: predictable through τ = 256 — though there via the plain first-order form, not the solve (P = 0.023 in fdt's favour) — with its argmax drifting 10 → 0.03 before it too collapses by τ = 512.

The practical summary for the winning formula sharpens: its fixed λ₀ = 40 is correct at the horizon it shipped for, over-damped ~100–4000× beyond a couple hundred SGD steps (where a 1/(η·τ)-style schedule would keep most of the signal), and no λ schedule saves the preconditioned optimizers — their fix has to come from re-linearizing (TracIn-style multi-checkpoint Grams, or SOURCE-style segment curvature), not from the damping dial.

Source · latent-learning repo, experiments/training_prediction/equilibrated/ (ridge_lambda_sweep.py, ridge_sweep_bootstrap.py, heldout_lambda_sweep.py; results in ridge_sweep_20260713_090849/, commits 651e132…4da37c3).
Long quenches (§6) · experiments/training_prediction/tp_long_quench.py + long_quench/ (SPEC, fleet script, sweep_20260714/, commits cbe83d9…5f665e8) · 96 runs, 15×H100, 2026-07-13/14.
Data · arcadia-impact/training-prediction-1b-heldout-20260708 (held-out grid, 1200-step runs under equilibrated/, clean long quenches under long_quench/).
Cut from this page · the same sweep over the thread's thermal/Langevin (SGLD) arms, which heat rather than learn; those results live in the repo's RESULTS.md alongside the sweep outputs, and the equilibrated-FT story itself is §12 of the run page.