Com-Eff RLVR
Cloudflare Pages-ready static site · GitHub repo target: com-eff-RLVR

Research hub

Communication-efficient RLVR

The two-circuit GRPO research hub: weekly status, curated reports, and the full run record.

Most important tab

Weekly Update

The current readout for people who only have time for one page.

2026-07-24

Two-circuit pipeline-parallel RL at longer horizons

The exact-k PRF fast circuit completed the full long run without collapse while holding a fixed, strongly compressed boundary payload. The new report uses full W&B trajectories to explain why a fresh random mask survives while a reused PowerSGD Q subspace can turn ordinary policy movement into a directional failure.

Long run heldExact-k PRF completed the full horizon with no reward or entropy cliff.
Validation gainedMATH validation improved early and stayed inside the learned band.
Budget heldThe fixed compressed payload remained unchanged at every pipeline boundary.
No collapse
through the full exact-k PRF run
Slow KL rise
codec-view drift stayed controlled
Fixed payload
mask budget held
Latest update: randomness removes the persistent first-order direction. The remaining task is to reduce variance and add a dense-view drift controller.

Curated evidence

Essential Reports

Only the reports that carry the current decision chain are copied into this repo.

2026-07-08 ANALYSIS

Same Rank-1, Different Projectability - What Breaks from GSM8K to Big-Math

Decomposes the two tasks. The rank-1 spatial structure only dents (EVR₁ ~90-96% on both), but forward linear projectability collapses (op ratio 0.90 → 1.00). Reconciled by two timescales: the cumulative direction is stable on both, while raw per-step updates align only on GSM8K (0.86 vs 0.15). Rank-1 is necessary but not sufficient for projection.

2026-07-08 ANALYSIS

Rank-1 Trajectories on a Harder Task - Big-Math vs GSM8K

Re-runs the rank-1 analysis on Big-Math (EXP-58). Rank-1 structure replicates, but the single line's durability degrades on the harder task: ~6× more of the move lands off the line (25% vs 4.5%) and the direction rotates further over the run (|cos| 0.46 vs 0.64).

2026-07-08 VERDICT

Big-Math signed-EMA (α, β) sweep - how much correction, how fresh an anchor?

Seven communication-efficient runs varying the merger's correction weight α and anchor-memory β, plus a dense control. α=0.25 with a fresh anchor (β=0.00) is the stable winner; only the stalest anchor (β=0.90) collapses. Dense stays the ceiling.

2026-07-07 LIVE RESULT

Big-Math Training Reward - Does Weight Projection Help?

Live training on Big-Math: no compressed variant beats dense, and plain stale signed EMA every step (no projection) beats every projection variant. A GSM8K appendix shows the coherent-trajectory contrast that explains why.

2026-07-04 PASS

MOAT Projector Verdict - ANCHOR/FAST from GSM8K to Big-Math

Integration rollup for the projector lanes. Fixed damped-linear wins on GSM8K, but is not adopted as a universal default because Big-Math collapses to do-nothing.

2026-07-05 LIMITED

Math-only Projector Ablation

EXP-61 confirms the harder Math trace is only marginally projectable at the freshest anchor and do-nothing-optimal at operational staleness.

2026-07-03 ANALYSIS

Rank-1 Trajectory Family on the EXP-57 Trace

RELEX-style rank-1 structure replicates, but the paper form is the wrong objective for short-horizon staleness repair. Use rank-1 as position state, not blind extrapolation.

2026-07-04 LIVE PATH

LA-WARMUP-V2 - PP-GRPO at 20/20 Without Collapse

Live training result for communication-efficient pipeline-parallel GRPO at longer cadence, useful as the bridge from offline projection geometry to end-to-end stability.