•1 min read•from Machine Learning
Deep Dive on RL and OPD for Training LLMs [D]
Hi everyone, if you have been reading the tech reports of Kimi, DS, Qwen and GLM, you will realize how much on policy distillation and GRPO style algorithms power the frontier. I thought it will be quite beneficial to do a deep dive explaining the maths and code behind this algorithms and how they connect to pretraining and supervised fine tuning.
I have published a deep dive on this topics here
Hope you enjoy it and it helps you understand training of LLMs better. Happy to answer questions on this
[link] [comments]
Want to read more?
Check out the full article on the original site
Tagged with
#LLMs
#RL
#OPD
#Policy Distillation
#GRPO
#Pretraining
#Supervised Fine Tuning
#Kimi
#Qwen
#GLM
#DS
#Algorithms
#Training
#Tech Reports
#Maths
#Code
#Frontier
#MachineLearning
#Deep Dive
#On Policy