Com-Eff RLVR
← Back to the hub

Raw evidence, one page per issue

Experiment Runs

Auto-published by the research harness at /close: each card is one closed GitHub issue's full record — verdict, per-criterion results, provenance, and artifact pointers. The curated Reports and Weekly tabs are written by hand FROM these pages; this tab is the unedited feed.

2026-07-29 PRF EXACT-K IS NOW THE DEFAULT

#93 · Twelve arms, re-scored on how steadily they train

The program was scored against a bar that leads with capability, and capability turned out to be a tie: a7 and a9 land on 0.6713 to the digit, the three 600-step runs finish within 0.006 of each other. Re-ranked on gap stationarity, gradient behaviour and non-collapse, twelve arms were run to beat the incumbent PRF exact-k and none did. FRLR is 2.55x better at step 100, crosses at 424, and ends 2.12x worse with its optimizer drifting 9.25x. Includes the two counterexamples that discipline the reading and a nine-entry corrections ledger.

2026-07-28 COMPLETE, 29 SECTIONS

#93 · The full program: four papers, twelve arms, 14 runs, and a codec that survives 600 steps

The complete record, sections 0 to 29: what four papers changed, the operator's questions answered and corrected, round A's five arms, the finding that codec-view drift is not drift, the theory that predicted FRLR's crossover before the run existed, the PPO-ratio bug the operator spotted in a WandB plot and its staging fix, and the 600-step durability run that resolved the prediction. Ends with the stability re-scoring of all twelve arms.

2026-07-24 HEALTHY @337

#90 · Mid-run diagnosis: the KL climb is a codec-view artifact; run healthy; mismatch roadmap

Live 600-step PRF exact-k run audited at step 337: val 0.451 to 0.663 and holding, wedge flat at 14.5 nats, kl_loss climb decelerating; zero clipfrac is by construction. Answers the bypass_mode, anchor-budget, systematic-bias and ALP (arXiv 2603.19470) questions with measured numbers.

2026-07-23 PASS

#89 · Autoresearch: drive PRF codec to dropout-benign KL at fixed 95% compression

The open-ended search ran 8 one-lever codec candidates against the fresh prf-constant incumbent at matched horizon and rejected all 8 under the strict 5-part gate; the loop exited exactly at the money gate (hard budget e

2026-07-10 PASS

#64 · Blockwise-freeze GRPO: train only the middle block (L11–15) for 75 steps on GSM8K + Big-Math

**C(block) = (S_frozen − S_base)/(S_dense − S_base)** — 2-seed replication, train ONLY decoder block L11–15 (frozen) vs full-parameter (dense), vanilla GRPO comm-eff OFF, Qwen2.5-1.5B-Instruct:

2026-07-10 PASS

#63 · Run comm-eff signed_ema vs dense on DeepScaleR RLVR, R1-Distill-1.5B, AIME val

**On this truncated read, comm-eff `signed_ema` (β_anc=0.50) holds the dense line** — within noise on the AIME headline and tightly on the low-variance training-reward curve. Not a full-length verdict (see caveats).

2026-07-08 PASS

#62 · Add RLVR-paper models + math datasets to comm-eff GRPO (additive, reward/prompt-validated, smoke-tested)

All 4 stage gates GREEN (none skipped):

Artifacts. Small per-run downloads live in this repo's local-only artifacts/<run-id>/ folder (gitignored — never deployed). Bulk artifacts (checkpoints, weight dumps, full metrics) live in Cloudflare R2 under autonomous-harness-rlvr-compression/<run-id>/; each run page links its prefix.