1 min readfrom Data Science

Defaulting to Adam without understanding will cost you. Don't "just throw adam at it"

Defaulting to Adam without understanding will cost you. Don't "just throw adam at it"
Defaulting to Adam without understanding will cost you. Don't "just throw adam at it"

Work in RL has caused me to rethink adam. It leads to extremely wonky behavior and hard to explain "burstiness" in the loss values that makes me want to rip my hair out,

It still works, but needs to be coaxed into it.

This article re-covers the mathematical intuitions behind adam, and where it fails spectacularly. If you're someone who works in RL, or trains deep transformers, it's a must read

Check it out. Do you agree?

submitted by /u/Nice-Dragonfly-4823
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#Adam
#RL
#Reinforcement Learning
#Deep Transformers
#Optimization
#Loss Values
#Mathematical Intuitions
#Burstiness
#Deep Learning
#Training
#Neural Networks
#Algorithm
#Gradient Descent
#Data Science
#Model Training
#Wonky Behavior
#Coaxing
#Failure Modes
#Parameter Tuning