Shows that L2 regularization and weight decay are equivalent for SGD (after rescaling) but not for Adam, and proposes AdamW, which decouples weight decay from the adaptive gradient step (ICLR 2019).
Weight Decay & AdamW: Decoupled Regularization
Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.
Intuition
The idea in plain words, before the symbols.
Regularization often starts with a simple preference: all else equal, prefer smaller weights.
There are two closely related ways to express that idea:
- add an penalty to the loss,
- directly shrink parameters a little bit every step.
For plain SGD, those give the same update once the penalty strength is matched to the learning rate. That is why "L2 regularization" and "weight decay" are often used as synonyms.
For Adam, they are not the same. Adam rescales coordinates using running estimates of gradient magnitude, so an term added to the gradient gets rescaled too. That means different parameters can experience very different effective regularization strengths.
AdamW (Loshchilov and Hutter, 2019) decouples weight decay from the adaptive gradient step: it takes the Adam step on the loss alone and shrinks the parameters separately.
Study prompts for this section: Intuition
Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.
Mathematics
Definitions, assumptions and the derivation.
Let be the task loss and the regularization strength.
With an penalty, the regularized objective is
so the gradient becomes
For SGD, this leads to
which is exactly a weight-decay step.
For Adam, the adaptive preconditioner changes things. AdamW writes the update as
The shrinkage term is not divided by , so every coordinate decays at the same relative rate instead of a rate set by its gradient history. (Here and are the bias-corrected averages defined on the Adam page.)
Conventions for differ, so compare values only within one. Loshchilov and Hutter (2019) write weight decay per step, , which for plain SGD equals an penalty with coefficient (their Proposition 1); in their AdamW (Algorithm 2) the decay is scaled by a schedule multiplier, not by the learning rate. PyTorch's AdamW multiplies by the learning rate, as written above.
One step with fixed per-coordinate factors
To see the difference without the moving averages, freeze Adam's per-coordinate factors for one step as a diagonal matrix ; in Adam, would hold . Each method's step is a task part plus a regularization part, both computed from the same pre-update :
Take , , , and . The task part is . The part is and the decoupled part is : the penalty pulls four times harder where and four times more weakly where . The new weights are with the penalty and with decoupled decay. In the second coordinate the part points towards zero, yet the weight grows from to , because the task part is larger: a pull towards zero is not the whole step.
Changing only from to makes the first coordinate agree ( under both methods) while the second still differs. By the difference formula, with the same the two methods take the same step for every and only when . A uniform can be matched by using in the penalty, but no single coefficient matches a whose diagonal entries differ; this is Loshchilov and Hutter's Proposition 2, and their Proposition 3 interprets decoupled decay under a fixed as a rescaled penalty. In Adam, changes at every step, and an term added to the gradient also enters both moving averages.
The first Adam step
From zero averages and with , Adam's first step is in each coordinate, where is the gradient it receives (see the Adam page). With the penalty ; in the example , with the same signs as , so the step is , exactly as with no penalty at all. AdamW takes the same Adam step and subtracts , giving . At this step the penalty disappears into Adam's normalization unless it reverses a gradient's sign; later it does not vanish, but it is still divided by .
Study prompts for this section: Mathematics
Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.
Code
The same calculation as runnable code, in the notation of the derivation.
import numpy as np
theta = np.array([1.0, 1.0])
g = np.array([0.1, 0.1])
vhat = np.array([1e-4, 1.0]) # very different squared-gradient averages
lr = 1e-2
wd = 0.1
eps = 1e-8
adam_with_l2 = theta - lr * ((g + wd * theta) / (np.sqrt(vhat) + eps))
adamw = theta - lr * (g / (np.sqrt(vhat) + eps)) - lr * wd * theta
print("Adam + L2 :", np.round(adam_with_l2, 4))
print("AdamW :", np.round(adamw, 4))
In the first coordinate, the decay part of the "Adam + L2" step is 100 times larger than AdamW's ( instead of ), because Adam divides it by the small sqrt(vhat) . In the second coordinate the two agree.
The second block works the example above in exact fractions with Python's standard library: one step with fixed factors, the change of , , and Adam's first step with each kind of decay.
from fractions import Fraction as F
theta = (F(2), F(2))
g = (F(1), F(-1))
eta, lam = F(1, 10), F(1, 5)
def one_step(P):
"""One step with fixed per-coordinate factors P: task part, L2 part, decoupled part."""
task = tuple(-eta * p * gi for p, gi in zip(P, g))
l2 = tuple(-eta * p * lam * t for p, t in zip(P, theta)) # the penalty's gradient goes through P
decoupled = tuple(-eta * lam * t for t in theta) # applied directly
new_l2 = tuple(t + a + b for t, a, b in zip(theta, task, l2))
new_decoupled = tuple(t + a + b for t, a, b in zip(theta, task, decoupled))
return new_l2, new_decoupled
new_l2, new_decoupled = one_step((F(4), F(1, 4)))
assert new_l2 == (F(36, 25), F(403, 200)) # coordinate 2 grows, although its L2 part points inward
assert new_decoupled == (F(39, 25), F(397, 200))
changed_l2, changed_decoupled = one_step((F(1), F(1, 4)))
assert changed_l2[0] == changed_decoupled[0] == F(93, 50)
assert one_step((F(1), F(1)))[0] == one_step((F(1), F(1)))[1] # P = I: the same step
def sign(x):
return (x > 0) - (x < 0)
# Adam's first step from zero averages (epsilon = 0) moves each coordinate by -eta * sign(input).
adam_l2 = tuple(-eta * sign(gi + lam * t) for gi, t in zip(g, theta))
adam_w = tuple(-eta * sign(gi) - eta * lam * t for gi, t in zip(g, theta))
assert adam_l2 == tuple(-eta * sign(gi) for gi in g) # the penalty has no effect at this step
assert adam_w == (F(-7, 50), F(3, 50))
print("one step, L2 penalty:", [str(x) for x in new_l2])
print("one step, decoupled decay:", [str(x) for x in new_decoupled])
print("Adam first step, L2 penalty:", [str(x) for x in adam_l2])
print("AdamW first step:", [str(x) for x in adam_w])
Study prompts for this section: Code
Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.
Demo
Predict where the two penalties differ, work one step of each, then change one factor and run Adam's first step.
Where do the two penalties differ?
Observation guide
Optional and ungraded. Opening a guide saves nothing.
Choose what to inspect in Weight Decay & AdamW: Decoupled Regularization, then open its guide.
Predict where the two penalties differ, then work one step: the task, and decoupled parts, the two new weights, a change of one factor, and finally Adam's own first step. Each step can be checked, and its solution shown after one try or the hint. The demo on the Adam page has an AdamW switch that adds the decoupled decay to a whole run.
Study prompts for this section: Demo
Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.
Concept: Weight Decay & AdamW: Decoupled Regularization
What is the smallest example of Weight Decay & AdamW you can work through by hand, and what does it show?
DetailsOptimization
Reference id
concept:optimization/weight-decay-adamwAfter the first pass
Go deeper
Optional material for a second reading. Sources are always shown; the other panels open from their headings or from these links, and closing one keeps what you wrote or revealed.What to look for in each sectionChoose a section and what you expect to see in it, then read short advice.
Section by section
Choose what to look for in each section
Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.
Start with the picture, metaphor, or geometric mechanism.
Choose what you expect to see in this section of Weight Decay & AdamW: Decoupled Regularization.
Questions to ask of the pictureChoose a part of the picture and what you expect, then read short advice.
Looking at the picture
Questions to ask of the picture
Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.
Which part of the picture should you look at first?
Pick the part of the picture you expect to explain Weight Decay & AdamW: Decoupled Regularization.
Sources
References for this page
A reference shows where an idea comes from; it does not vouch for every step on this page.Defines the Adam update whose per-coordinate scaling AdamW keeps (ICLR 2015).
Textbook treatment of the L2 penalty, its gradient and the per-step shrinkage it causes under gradient descent.
Documents the decay step theta <- theta - lr * weight_decay * theta applied outside the adaptive step. Documentation only; no PyTorch code runs here.
What each claim rests on1 claim, each with its sources, where it appears on this page, and its caveats.
Claims and sources
Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.
The checks were made by this site, not by an independent reviewer.
Loshchilov and Hutter show the equivalence for SGD, its failure for adaptive methods (Proposition 2) and define AdamW; PyTorch documents the decoupled decay step. The page's one-step comparison with fixed fa...
One step with factors fixed by hand, plus Adam's first step from zero. It shows how the two penalties enter the update, not which one trains or generalizes better.
Read Loshchilov and Hutter in the arXiv source: Proposition 1 (SGD equivalence with lambda/alpha), Algorithm 2 (L2 enters g_t before the moments; decay sits outside the adaptive fraction) and Propositions 2 and 3. Recomputed the page's one-step example and Adam's first step in exact fractions.
Explain it without the pageWrite your own explanation before asking for feedback. Your draft and any hints stay while this page is open.
Practice · Weight Decay & AdamW: Decoupled Regularization
Try the idea in your own words
Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.
Concept · Selected for practice
Weight Decay & AdamW: Decoupled Regularization
What it rests on: Sources: Decoupled Weight Decay Regularization; Adam: A Method for Stochastic Optimization; Deep Learning
Context and links
Optimization
Explain the mechanism
For Weight Decay & AdamW: Decoupled Regularization: What is the smallest example of Weight Decay & AdamW you can work through by hand, and what does it show? Explain your answer, including what changes, why, and which assumption matters.
A rough first thought is enough. Your draft stays when you change tasks.
Your draft stays on this page and is cleared when you leave or reload.
A little help · Explain
Open one hint at a time. These are suggestions, not your answer or a grade.
0 of 3 hints shown for this question.
Clearing your answer keeps this help record. Outside help is your own report; this page cannot check it.
Where am I stuck? (optional)
Take your draft to a feedback conversation
No AI feedback runs here. You can copy a prompt to use elsewhere; nothing is sent automatically. Review the text before sharing, and leave out private information.
Write an attempt before copying a feedback prompt.
A draft, or an AI reply to it, is not a test of what you have learned. To check that, try a different case later without help.
Ask about this pagePick one item on this page and copy a prompt about it into an AI assistant you already use. Notes you write here stay in this browser.
assert new_l2 == (F(36, 25), F(403, 200)) # coordinate 2 grows, although its L2 part points inwardYour question
Choose what your question is about
Pick the idea, equation, source, code, claim, misconception or demo state you want to ask about. The prompt you copy includes it, so the answer can stay on that item.Open the draft below to save one note and next action in this browser.
Weight Decay & AdamW: Decoupled Regularization
What is the smallest example of Weight Decay & AdamW you can work through by hand, and what does it show?
Fixed prompts generated from this item. They are not comments from people or independent reviews.
The 4 cited sources must support this exact item, not just the surrounding topic.
Connect the definition to one equation, piece of code or demo before widening the discussion.
Change one input, then check that the same relationship holds in the mathematics, the code and the demo.
You can state the mechanism in your own words
Local action draftNo local draft saved yetOpen when you are ready to write one next action
This draft stays in this browser and is attached to this item only.
- What the 4 cited sources say about this exact item
- The definition, its prerequisites, and a contrasting concept
- The equation or code that makes the concept concrete
- One demo state that shows what stays the same as the inputs change
- You can state the mechanism in your own words
- You can name the prerequisite that would clear up the confusion
- You can predict how the result changes when one input changes
Reference id
concept/concept-notebook/optimization/weight-decay-adamw
concept:optimization/weight-decay-adamw