Back to research hub

Paper discussions

Frontier training reports, read for implementation.

Concise research-team briefings that separate disclosed training stages from inference, compare the important design choices, and surface what we can use in our own work.

Latest discussions

Post-training pipelines, specialist teachers, distillation, reasoning effort, model merging, and trainer efficiency.

30 July 2026 Post-training

Kimi K3, GLM-5.2, and DeepSeek-V4: how the post-training recipes differ

A two-page comparison focused on the stages after SFT, including Kimi's 9 domain-effort teachers, its MOPD dense token-wise loss signal, GLM's sequential reasoning, agentic, and general RL stages, and DeepSeek's full-vocabulary consolidation.

SpecialistsDomain and reasoning effort as separate training axes.
DistillationTeacher logits consolidate specialists into one student.
EfficiencyQAT, rollout scheduling, teacher loading, and recovery.
Open the two-page briefing
30 July 2026 Figures

Kimi K3 Post-Training Figures

A one-page, figure-first walkthrough of Kimi K3's nine specialist teachers, teacher routing, unified student, and MOPD token-wise signal.

Open the one-page figures