1 min readfrom Machine Learning

Deep Dive on RL and OPD for Training LLMs [D]

Hi everyone, if you have been reading the tech reports of Kimi, DS, Qwen and GLM, you will realize how much on policy distillation and GRPO style algorithms power the frontier. I thought it will be quite beneficial to do a deep dive explaining the maths and code behind this algorithms and how they connect to pretraining and supervised fine tuning.

I have published a deep dive on this topics here

Hope you enjoy it and it helps you understand training of LLMs better. Happy to answer questions on this

https://youtu.be/MaZWafi4gYY?is=8jLkAp\_Fe86abUVP

submitted by /u/johnolafenwa
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#LLMs
#RL
#OPD
#Policy Distillation
#GRPO
#Pretraining
#Supervised Fine Tuning
#Kimi
#Qwen
#GLM
#DS
#Algorithms
#Training
#Tech Reports
#Maths
#Code
#Frontier
#MachineLearning
#Deep Dive
#On Policy