← All runs · 89-prf-codec-autoresearch-benign-kl

Autoresearch: drive PRF codec to dropout-benign KL at fixed 95% compression

PASSclosed 2026-07-23kind: experiment

Small artifacts (local only, not deployed): artifacts/89-prf-codec-autoresearch-benign-kl/

Close-out verdict (issue record)

VERDICT: PASS (clean symmetric negative, the plan's own falsifiable branch)

The open-ended search ran 8 one-lever codec candidates against the fresh prf-constant incumbent at matched horizon and rejected all 8 under the strict 5-part gate; the loop exited exactly at the money gate (hard budget enforcement at 40 GPU-hr). The trajectory, mechanism findings, and recommended config are shipped.

HEADLINE RESULTS (criterion | observed | target):

RECOMMENDED CONFIG (gate-legal): compression_type=prf_mask, p=0.95, rescale_mode=constant, exact_k=true, mask_recompute=true, mask_reference=true, pp_size=8. GOAL-NEAREST (documented, not promoted): FRLR 48+28+1 at the same 77-value budget.

NOT RUN (named unticked criteria): the single MATH confirmation val (operator-cancelled 2026-07-23) and the winner-projection-off pair (budget teardown + operator close-out; the lever is characterized in the #88 close-out).

COST: ~41.5 box-hr x $3.97/hr ~ $165 on Vast 45458698 (team, external attach), torn down 2026-07-23T03:05Z reason budget-exceeded. Zero env incidents across 62 monitor cycles.

PR (open, NOT merged per operator): https://github.com/shamanez/verl/pull/32

Report page: https://com-eff-rlvr.pages.dev/runs/89-prf-codec-autoresearch-benign-kl.html

WandB: https://wandb.ai/shamanework-pl/89-prf-codec-autoresearch-benign-kl

Analyst verdict (verdict.md)

Verdict — 89-prf-codec-autoresearch-benign-kl

VERDICT: PASS

Autoresearch search loop ran to the money gate and shipped a **symmetric clean

negative** with a full 8-candidate trajectory and a gate-legal recommended

config. The plan's Hypothesis section explicitly declares this a PASS: "if no

candidate tried within the GPU-hour budget beats the constant-rescale incumbent

under that gate ... the incumbent stays PRF constant. Either way we ship the

trajectory and the final recommended config." No candidate reached the goal

(drive codec-view entropy toward the dropout ~0.13 floor while holding

actor/kl_loss at or below the incumbent); the one gate-legal accept (exact-k)

is within noise. Every number below is greppable from runs/89-prf-codec-autoresearch-benign-kl/.

Judged against the plan cache (.claude/state/plan-cache/89.md) as amended by

the operator on the issue (2026-07-21..23). Per-trial + Stage-4 validation and

the winner-projection-off pair were operator-CANCELLED; the 30-step horizon and

the step-21 read of candidate 8 are operator-sanctioned. These are amendments to

the predicate, not analyst-unmeasured criteria.

---

Reference frame (Stage 2, fresh controls, this project) — metrics/train_<cell>.log @ step 40

metric (WandB key)dense-controldropout p=0.1prf-constant-incumbent
actor/entropy0.180730.137027.8158
actor/kl_loss0.00330280.123730.11658
actor/ppo_kl-2.95e-060.090155-2.27e-04
actor/comm_eff/mask_ration/an/a0.94995
critic/score/mean0.645510.604250.63159
training/rollout_probs_diff_mean0.00225820.00182770.90265
actor/perf/max_memory_allocated_gb45.80645.756110.48

Reads confirm the three-divergence structure: PRF-constant reproduces

codec-view entropy ~7.82 (structural, from the h*mask/(1-p) 20x amplification)

with ppo_kl at machine-zero (within-step mask bit-identity intact) and a slow

kl_loss climb (0.032@5 to 0.117@40), while dropout is entropy ~0.14 with a flat

kl_loss ~0.12 plateau. Peak GPU mem 110.48 GB after the first anchor refresh

(two-circuit clone at ~step 10) — at the ~110 GB gate, well inside 141 GB. No OOM.

Incumbent matched-step-30 reference (trials ran to 30, so all candidates are

judged here): `entropy=7.8134 kl_loss=0.064933 ppo_kl=0.0029998 mask=0.94995

score=0.59399 rpd=0.89082`.

---

Incumbent trajectory — 8 candidates vs incumbent @ matched step 30 (candidate 8 @ 21)

Amended 5-part gate: KEEP iff (actor/entropy reduced) AND (actor/kl_loss no

higher) AND (actor/ppo_kl < ~1e-2) AND (mask_ratio in [0.94,0.96]) AND

(critic/score/mean early slope not degraded). Entropy compared intra-PRF

(apples to apples per plan).

#candidate (lever)entropyΔentkl_lossΔklppo_klmaskscoreverdictone-line diagnosis
prf-constant-incumbent7.81340.0649330.00300.949950.59399incumbent20x constant rescale; structural entropy, machine-zero ppo_kl
1prf-rms-match (rescale_mode=rms_match)5.8606-1.952.4678+2.40 (38x)0.00980.949950.59961REJECTremoving the variance blow-up drops entropy but re-aligns the surviving 5% across steps → coherent residual → KL explodes 38x
2prf-exact-k (exact_k=true)7.8128-0.00060.061633-0.00330.00330.949950.59326ACCEPT (gate-legal, marginal)fixing the per-token keep count to round((1-p)H) is within-noise on every axis — count jitter is irrelevant; kept only because it is harmless + a hair lower KL
3prf-nonuniform-p (p_by_boundary mean 0.95)7.8030-0.0100.073971+0.00900.00270.949990.60229REJECTreallocating the byte budget across depth at constant average moves entropy negligibly and slightly WORSENS KL — depth reallocation is low-value
4prf-antithetic (antithetic=true)7.8135+0.00010.060319-0.0046 (-7.1%)0.00220.949950.60693REJECT (entropy not reduced)complementary cross-step draw leaves entropy untouched (structural) and buys only ~7% KL reduction — temporal decorrelation is a weak lever
5prf-frlr r32k44 (frlr rank32 k44, rescale=none)2.9141-4.900.32856+0.264 (5.1x)0.00500.949870.58984REJECTNEW lever (seed queue exhausted): replace 20x amplification with a fast-refresh rank-32 residual correction. Crushes entropy 7.81→2.91 (amplification DOES drive the inflation) but re-injects coherent residual energy → KL 0.33
6prf-frlr r16k60 (rank16 k60)3.4259-4.390.45316+0.388 (7x)0.00650.949870.57251REJECT (KL higher; ppo_kl 0.0109@20)shrinking correction rank 32→16 raises residual energy → higher KL; the r+k≈76 budget trades rank for worse alignment
7prf-frlr slowq (rank32 k44, slow Q: 14 vs 203 refreshes)5.8597-1.951.0532+0.988 (16x)0.00490.949870.61035REJECT (KL higher, unstable)slowing Q refresh makes the basis stale → KL both higher AND sawtooths (1.46→2.85→0.74→1.05); slope = Q adaptivity
8prf-frlr r48k28 (rank48 k28) @ step 213.27433-4.50.05050636+0.010 vs incum@21 (0.04044), climbing0.00440.949870.48755REJECT (incomplete + KL climbing; goal-nearest)highest correction rank held KL lowest of the FRLR family (0.003 through step 9 via fast-Q, 0.0505 by step 21) while crushing entropy to 3.27 — closest any candidate came to the goal, but KL had already crossed above the incumbent @21 and was rising ~0.023/step; never reached the step-30 horizon (cut by budget teardown)

All ppo_kl < 1e-2 everywhere (within-step mask bit-identity preserved across

old/train/reference forwards on every codec) except a single transient

0.010888 @ step 20 on r16k60. mask_ratio in [0.94, 0.96] on every cell.

Diagnose-and-propose chain (loop never terminated on a reject): seed queue

rms-match → exact-k → nonuniform-p → antithetic exhausted with no goal hit; the

antithetic/nonuniform nulls plus the rms-match KL blow-up pointed at the

residual (not the count/schedule), so the loop proposed the NEW FRLR lever

(fast-refresh low-rank residual correction) and swept its rank/refresh budget

r32k44 → r16k60 → slowq → r48k28 until the money gate.

---

Recommended config

Shipped (gate-legal): PRF constant + exact-k. The only lever that passes the

amended gate without cost; exact-k marginally lowers KL (0.0649→0.0616 @30) at

zero entropy or capability cost.

COMM_EFF_ENABLED=true
COMM_EFF_COMPRESSION_TYPE=prf_mask
COMM_EFF_MASK_ENABLED=true
COMM_EFF_MASK_P=0.95
COMM_EFF_MASK_RESCALE_MODE=constant
COMM_EFF_MASK_EXACT_K=true
COMM_EFF_MASK_RECOMPUTE=true
COMM_EFF_MASK_REFERENCE=true
COMM_EFF_MASK_PP_SIZE=8
COMM_EFF_ANCHOR_OWNS_Q=false
COMM_EFF_POWERSGD_FAST_Q_BOOTSTRAP=false
# anchor.lookahead_anchor=true (projection-off pair operator-CANCELLED; see #88 close-out)

Goal-nearest, documented (NOT recommended): FRLR r48k28

(rescale_mode=none frlr=true frlr_rank=48 frlr_k=28). Measured tradeoff:

entropy 7.81→3.27 (the largest inflation cut of any candidate at low-ish KL) but

KL 0.003→0.0505 by step 21 and still climbing past the incumbent; unmeasured at

the step-30 horizon. It is the direction to reopen if the search budget is

re-approved (larger rank + faster Q to hold the residual alignment), not a

shippable config today.

---

Mechanism findings (all grounded in the trajectory reads above)

1. Codec-view entropy is structural inflation, and it IS reducible. The ~7.82

entropy is the h*mask/(1-p) 20x variance amplification; touching the rescale

(rms_match 5.86, FRLR 2.91–3.43) moves it directly. So the amplification is a

real driver of the entropy inflation — refuting the "amplification is

irrelevant" half of the null.

2. But entropy and the reference-KL climb are COUPLED through the residual.

Every lever that cut entropy (rms_match, all FRLR) re-introduced coherence in

the discarded/rescaled energy and paid for it in kl_loss (2.47, 0.33, 0.45,

1.05). KL level = residual energy × alignment asymmetry; KL slope = Q

adaptivity (fast-Q r48k28 held 0.003 to step 9; stale/slow-Q sawtooths). You

cannot cheaply trade one for the other at this horizon — the benign-KL,

low-entropy corner was not reachable within budget.

3. Count jitter is irrelevant (exact-k within noise on every axis).

4. Depth reallocation is low-value (nonuniform-p: negligible entropy move,

slightly worse KL).

5. Temporal decorrelation is weak (antithetic: entropy untouched, only ~7.1%

KL reduction: (0.064933−0.060319)/0.064933).

Net: the constant-rescale incumbent stays the recommended PRF codec; the

amplification is reducible but not without re-inflating the residual KL, i.e. the

residual is coherence/anchor-coupled (findings Section 8), consistent with the

plan's declared clean-negative conclusion.

---

Success-criteria checklist (amended items marked)

incumbent = canonical byte-identical PRF ⇒ off-path parity) ✓; Stage 2

(3 controls to step 40, no OOM, anchor clone fits @110.48 GB, reference frame

logged) ✓; Stage 3 (8-candidate loop, exited at money gate) ✓; Stage 4

AMENDED-CANCELLED (val + projection pair cancelled by operator; recommended

config + trajectory still shipped).

(30, operator-sanctioned via step_target_fallback); mask_ratio in

[0.94,0.96] on every cell including the one accept (exact-k 0.94995).

(val-core/DigitalLearningGmbH/MATH-lighteval/acc/mean@1) was operator-CANCELLED

(no MATH val pass anywhere; capability guard = free critic/score/mean slope,

which tracks dense: incumbent 0.34→0.63 vs dense 0.39→0.65).

operator-CANCELLED (budget teardown + close-out-now); the lever is characterized

in the #88 close-out. lookahead_anchor=true held on every cell.

the ledger cap at ~03:05Z 2026-07-23 (teardown_reason=budget-exceeded); operator

amendment 3 declares this the plan's money-gate exit. Every rejection has a

recorded diagnose-and-propose entry (seed queue → FRLR family).

EXPECTED (rollout uncompressed; dense/dropout 0.0018–0.0027). Not a stop trigger.

ppo_mini 256, n=8, prompt/response 1024/3072, lr 1e-6, kl_loss_coef 0.001,

cadence/delay 20/20); the sole sanctioned exception (projection-off pair) was

cancelled, so no non-codec variance was ever introduced.

---

Budget / teardown record

---

Notes

synced metrics/train_<cell>.log are the on-box metric TAIL only; the launcher

python3 -m verl.trainer.main_ppo set -x trace was not synced, so there is no

resolved_cmd.txt. Ground-truth codec + non-codec config was recovered instead

from each cell's own Hydra config-echo (CommEffMaskConfig dump) →

resolved_params.txt. No plan-vs-ran divergence found: the incumbent echo

matches the run.json canonical PRF cell exactly (p=0.95, rescale_mode=constant,

pp_size=8, exact_k/antithetic false, p_by_boundary=[]).

uncompressed; the codec touches only the trainer forward) — explicitly not a

failure and not a stop trigger.

this run. Capability was guarded during the search by the free

critic/score/mean early slope only, which held at dense levels on every cell.

instance was destroyed mid-cell at the budget cap. Its step-1..21 curve is read

from monitor-detail.log (cycle58–cycle62 SPLITMETRIC blocks); judged at step

21 per operator decision.

run's per-issue WandB project (it hardcodes verl_compression_research +

resume='must' with the cell name as run id) and errors out. It parsed the

local log correctly (score=0.631592 == step-40), confirming every final-step row

is already local, so no curve is missing and backfill is not blocking.

clean negative is fully measured, so STOP (budget-exhausted-UNMEASURED) does not

apply — this is the planned money-gate exit with a complete trajectory.

Provenance

resolved_params.txt (what actually ran)
# Resolved parameters for 89-prf-codec-autoresearch-benign-kl
# SOURCE: capture_resolved_config.py could NOT run -- the synced per-cell logs
# (metrics/train_<cell>.log) are the on-box METRIC TAIL only; the launcher
# `set -x` `python3 -m verl.trainer.main_ppo <args>` trace was NOT synced.
# => RESOLVED_CONFIG_MISSING (flagged in verdict Notes; no resolved_cmd.txt).
# The values below are GROUND TRUTH extracted from each cell's own Hydra
# config-echo (CommEffMaskConfig dump) inside metrics/train_<cell>.log, plus
# run.json cell overrides. Greppable per cell via:
#   grep "'rescale_mode'\|'exact_k'\|'antithetic'\|'frlr" metrics/train_<cell>.log

## Non-codec surface (identical across ALL cells; incumbent echo)
train_batch_size=512
ppo_mini_batch_size=256
rollout.n=8
max_prompt_length=1024
max_response_length=3072
actor.optim.lr=1e-06
actor.kl_loss_coef=0.001
ref.low_var_kl=0.001
anchor.cadence=20
anchor.delay=20
anchor.lookahead_anchor=true   # held true on every cell; projection-off pair CANCELLED
total_training_steps=40 (controls+incumbent) / 30 (trials, step_target_fallback)
VAL_BEFORE_TRAIN=False ; TEST_FREQ=-1   # per-trial val OFF (operator amendment)
model=Qwen/Qwen2.5-Math-1.5B ; data=$HOME/data/math (MATH parquet)

## Canonical PRF codec (incumbent echo: prf-constant-incumbent)
comm_eff.enabled=true
comm_eff.compression_type=prf_mask
comm_eff.mask.enabled=true
comm_eff.mask.p=0.95
comm_eff.mask.rescale_mode=constant
comm_eff.mask.exact_k=false
comm_eff.mask.antithetic=false
comm_eff.mask.p_by_boundary=[]
comm_eff.mask.mask_recompute=true
comm_eff.mask.mask_reference=true
comm_eff.mask.pp_size=8
comm_eff.anchor_owns_q=false
comm_eff.powersgd.fast_q_bootstrap=false

## Per-cell codec lever DELTA vs canonical (from each cell's config-echo)
dense-control            : comm_eff.enabled=false (dropout 0)
dropout-p10-control      : comm_eff.enabled=false ; model.override_config.attention_dropout=0.1
prf-constant-incumbent   : (canonical, rescale_mode=constant)
prf-rms-match            : rescale_mode=rms_match
prf-exact-k              : rescale_mode=constant ; exact_k=true
prf-nonuniform-p         : rescale_mode=constant ; p_by_boundary=<7-vec mean 0.95>
prf-antithetic           : rescale_mode=constant ; antithetic=true
prf-frlr (r32k44)        : rescale_mode=none ; frlr=true frlr_rank=32 frlr_k=44 frlr_unbiased=false
prf-frlr-r16k60          : rescale_mode=none ; frlr=true frlr_rank=16 frlr_k=60 frlr_unbiased=false
prf-frlr-slowq           : rescale_mode=none ; frlr=true frlr_rank=32 frlr_k=44 (slow Q: 14 refreshes vs 203)
prf-frlr-r48k28 (@21)    : rescale_mode=none ; frlr=true frlr_rank=48 frlr_k=28 (config from monitor-detail.log; on-box log lost with instance)