Paper discussions
Frontier training reports, read for implementation.
Concise research-team briefings that separate disclosed training stages from inference, compare the important design choices, and surface what we can use in our own work.
Latest discussions
Post-training pipelines, specialist teachers, distillation, reasoning effort, model merging, and trainer efficiency.
30 July 2026
Post-training
Kimi K3, GLM-5.2, and DeepSeek-V4: how the post-training recipes differ
A two-page comparison focused on the stages after SFT, including Kimi's 9 domain-effort teachers, its MOPD dense token-wise loss signal, GLM's sequential reasoning, agentic, and general RL stages, and DeepSeek's full-vocabulary consolidation.
SpecialistsDomain and reasoning effort as separate training axes.
DistillationTeacher logits consolidate specialists into one student.
EfficiencyQAT, rollout scheduling, teacher loading, and recovery.
30 July 2026
Figures
Kimi K3 Post-Training Figures
A one-page, figure-first walkthrough of Kimi K3's nine specialist teachers, teacher routing, unified student, and MOPD token-wise signal.
Open the one-page figures