What was measured
When training runs through compressed pipeline cuts while sampling runs dense, the trainer and the sampler disagree about the policy's own log probabilities. That disagreement, in nats, is the gap this page prices. The run that collapsed carried about 17.9 nats of it while a dense control carries 0.0002. Each probe here trains for exactly one step and reads the gap at step 1, before any parameter update, so it prices the compression scheme itself rather than any drift it later causes.
One probe per arm, about 15 minutes each. Seven single-cut probes, one per pipeline cut, with that cut at 95 percent sparsity and the other six dense. Four multi-cut patterns and the winning layout on top of them, one of which, the five-cut pattern, is read from the first step of a full-length 16k run rather than a short probe. Probes ran at a 3072-token generation length because the step-1 gap turned out to be length invariant: all seven cuts compressed reads 17.23 at short generations and 17.23 at the 16k limit, and the winner reads 12.43 short against 12.37 at 16k. Short cheap probes price long-context deployments.
The gap rises with depth
Priced alone, every cut is expensive, and deeper cuts are reliably more expensive. The price climbs from 6.44 nats at the shallowest cut to 9.95 at the deepest, and the order never breaks.
A · The solo price of each cut
One-step probes. One cut compressed at 95 percent sparsity, the other six dense.
The middle band is not special. An earlier variant left the two middle cuts dense and compressed the other five, and its gap came in at 16.88 against 17.23 for compressing everything, a saving of about 2 percent. The bars explain why that protection bought almost nothing. The two protected cuts price no differently than their neighbours, and the deep cuts keep charging.
Composition saturates
The seven solo prices sum to 56.5 nats. Compressed together the seven cuts measure 17.23. One compressed cut already costs more than a third of the full gap, and a spread pattern of three cuts is within 7 percent of compressing everything.
B · Gap against number of compressed cuts
Each dot is one probe. The green dot is the winning layout.
The saturation means most of the gap is already present once a few cuts are compressed, so removing one or two buys little. The exception is placement. The three shallowest cuts together read 12.43 while a spread arrangement of the same count reads 16.07.
The winning layout
The cheapest arrangement found compresses the three shallow cuts and keeps the four deep cuts dense. It is also exactly what the depth rule predicts: pay only the low prices.
C · The layout that priced lowest
The 36-layer pipeline with its seven cuts, shallow on the left.
Compress shallow, keep deep dense. The gap tracks the computation behind the cut, and composition saturates so quickly that placement is the only lever with real range.
Every probe, measured
D · All twelve probe arms
| Arm | Compressed cuts | Gap at step 1 | Share of the all-seven gap |
|---|---|---|---|
| Single, cut 1 | cut 1 (14% depth) | 6.44 nats | 37% |
| Single, cut 2 | cut 2 (28% depth) | 7.04 nats | 41% |
| Single, cut 3 | cut 3 (42% depth) | 7.11 nats | 41% |
| Single, cut 4 | cut 4 (56% depth) | 7.90 nats | 46% |
| Single, cut 5 | cut 5 (67% depth) | 8.66 nats | 50% |
| Single, cut 6 | cut 6 (78% depth) | 9.38 nats | 54% |
| Single, cut 7 | cut 7 (89% depth) | 9.95 nats | 58% |
| Three spread | cuts 2, 4, 6 | 16.07 nats | 93% |
| Four alternating | cuts 1, 3, 5, 7 | 17.05 nats | 99% |
| Five, middle dense | cuts 1, 2, 5, 6, 7 | 16.88 nats | 98% |
| All seven | cuts 1 to 7 | 17.23 nats | 100% |
| Winner, three shallowest | cuts 1, 2, 3 | 12.43 nats (12.37 at 16k) | 72% |
What this changes
A depth rule replaces the band intuition. The gap grows with how much computation sits behind the cut, not with whether the cut lands in some special middle band. That is why protecting the two middle cuts bought a 2 percent saving while protecting the four deep ones bought 28 percent.
The pattern lever has a ceiling. Within this compression scheme at 95 percent sparsity, no arrangement of three or more compressed cuts brings the gap anywhere near dense. The winner lands at 12.4 against 17.2 for compressing everything, still far above the 0.0002 of a dense control. But it is below 14.4, the level the smaller model carried safely for its full 600-step run.
The line is now being tested. A 600-step training run at the 16k generation length with the winning layout is underway, to test whether starting below that line changes the collapse story.
Two honest caveats. The probes price the step-1 gap only. They say nothing about training dynamics, about how the gap feeds drift over hundreds of steps, or about where a run finally fails. And the depth rule is measured on one model at one sparsity, so its shape at other budgets is unknown.