技巧精选

Sebastian Raschka 发布第6期教程:从零实现 RLVR 与 GRPO 训练推理模型

精选理由

Raschka 又出新教程了,这次手把手带你从零写 GRPO,把 DeepSeek-R1 那套推理训练方法完整实现一遍,还带 MATH-500 评测结果。

Sebastian Raschka 发布"Reasoning from scratch"系列第6期视频教程,时长约1小时25分,讲解并实现 RLVR(可验证奖励强化学习)与 GRPO 算法。教程用 DeepSeek-R1 的训练思路解释推理模型与普通模型的区别,对比 GRPO 与 PPO 的差异。实操部分涵盖加载预训练模型、用 MATH 数据集训练、实现序列 log 概率、计算 GRPO loss,最后在 MATH-500 上评估 checkpoint 并讨论显存需求。

原文 · Sebastian Raschka

Reasoning from scratch, round number 6! An introduction (and implementation) of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO).

00:00 Introduction 01:54 What makes a reasoning model different? 04:25 Reasoning traces and model capability 08:29 Accuracy and format rewards 11:34 Aha moments and DeepSeek-R1 training 14:41 Reasoning effort and answer length 18:38 RLHF and RLVR 23:04 GRPO vs. PPO 26:40 GRPO explained with a cooking analogy 31:43 The KL term and simplified GRPO 35:04 Loading the pretrained model 36:07 Loading the MATH training data 39:26 Sampling model responses 46:30 Computing verifiable rewards 49:55 Computing advantages 51:54 Token and sequence log probabilities 55:29 Implementing sequence log probabilities 57:37 Fixing the inference-mode error 1:02:24 Computing the GRPO loss 1:04:37 Putting the GRPO step together 1:09:19 The GRPO training loop 1:12:57 Training settings, logging, and checkpoints 1:17:24 Running training and inspecting outputs 1:19:28 Loading and evaluating checkpoints 1:22:33 MATH-500 results and training stability 1:24:05 Memory requirements and next steps