•1 min read•from Data Science
Defaulting to Adam without understanding will cost you. Don't "just throw adam at it"

| Work in RL has caused me to rethink adam. It leads to extremely wonky behavior and hard to explain "burstiness" in the loss values that makes me want to rip my hair out, It still works, but needs to be coaxed into it. This article re-covers the mathematical intuitions behind adam, and where it fails spectacularly. If you're someone who works in RL, or trains deep transformers, it's a must read Check it out. Do you agree? [link] [comments] |
Want to read more?
Check out the full article on the original site
Tagged with
#Adam
#RL
#Reinforcement Learning
#Deep Transformers
#Optimization
#Loss Values
#Mathematical Intuitions
#Burstiness
#Deep Learning
#Training
#Neural Networks
#Algorithm
#Gradient Descent
#Data Science
#Model Training
#Wonky Behavior
#Coaxing
#Failure Modes
#Parameter Tuning