技巧精选

Sebastian Raschka 第二讲:手把手用 GRPO 从零训练推理模型

精选理由

Raschka 的 RLVR 系列第二课,用 GRPO 亲手训推理模型,讲透 clipped ratio、KL 损失和 reward hacking,代码全给,适合想上手 RL 后训练的人。

Sebastian Raschka 发布了《Reasoning From Scratch》系列 RLVR 第二讲视频,全程约 1 小时 50 分钟。课程围绕 GRPO 训练展开,覆盖 clipped policy ratio、KL loss 项、entropy 追踪和 format reward 的实现细节。视频演示在 MATH-500 上评估 checkpoint 的完整流程,并讲解如何诊断训练不稳定和 reward hacking 问题。所有训练脚本、日志绘制和熵计算的 PyTorch 代码均随课提供。

原文 · Sebastian Raschka

Reasoning From Scratch: Reinforcement Learning with Verifiable Rewards (RLVR) round 2. Covering clipped policy ratios, KL loss term, format rewards, and other GRPO tips & tricks.

00:00 Introduction and recap 01:52 Interpreting basic GRPO training metrics 06:34 Planned improvements to GRPO 08:58 Running longer training jobs with Python scripts 13:39 Running the baseline GRPO training script 17:29 Loading and plotting training logs 19:29 Diagnosing unstable training 23:55 Evaluating checkpoints on MATH-500 26:26 Downloading existing checkpoints 30:09 Tracking advantage statistics 34:53 Understanding entropy 40:32 Computing entropy in PyTorch 44:17 Interpreting entropy values 48:58 Adding entropy tracking to GRPO 53:36 Analyzing advantage and entropy metrics 56:18 Stabilizing GRPO with clipped policy ratios 1:03:27 Implementing the clipped policy loss 1:09:39 Analyzing clipped policy training results 1:11:25 KL divergence and reward hacking 1:15:12 Adding a KL loss term 1:20:34 Limitations of the simplified KL loss 1:23:04 Format rewards and think tags 1:25:47 Adding special tokens to the tokenizer 1:30:29 Implementing the format reward 1:35:56 Analyzing format reward training 1:38:25 Rewarding format only for correct answers 1:40:48 Further GRPO improvements from research 1:45:43 Next steps and distillation