Weight Decay & AdamW: Decoupled Regularization

Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.

Intermediate · Undergraduate mathematics · about 13 minutes

Reading map and next steps

Intuition

The idea in plain words, before the symbols.

Regularization often starts with a simple preference: all else equal, prefer smaller weights.

There are two closely related ways to express that idea:

  • add an L2L_2 penalty to the loss,
  • directly shrink parameters a little bit every step.

For plain SGD, those give the same update once the penalty strength is matched to the learning rate. That is why "L2 regularization" and "weight decay" are often used as synonyms.

For Adam, they are not the same. Adam rescales coordinates using running estimates of gradient magnitude, so an L2L_2 term added to the gradient gets rescaled too. That means different parameters can experience very different effective regularization strengths.

AdamW (Loshchilov and Hutter, 2019) decouples weight decay from the adaptive gradient step: it takes the Adam step on the loss alone and shrinks the parameters separately.

Study prompts for this section: Intuition
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Mathematics

Definitions, assumptions and the derivation.

Let L(θ)L(\theta) be the task loss and λ>0\lambda > 0 the regularization strength.

With an L2L_2 penalty, the regularized objective is

Lreg(θ)=L(θ)+λ2∥θ∥22,L_{reg}(\theta) = L(\theta) + \frac{\lambda}{2} \lVert \theta \rVert_2^2,

so the gradient becomes

∇Lreg(θ)=∇L(θ)+λθ.\nabla L_{reg}(\theta) = \nabla L(\theta) + \lambda \theta.

For SGD, this leads to

θt+1=θt−η∇L(θt)−ηλθt=(1−ηλ)θt−η∇L(θt),\theta_{t+1} = \theta_t - \eta \nabla L(\theta_t) - \eta \lambda \theta_t = (1 - \eta \lambda)\theta_t - \eta \nabla L(\theta_t),

which is exactly a weight-decay step.

For Adam, the adaptive preconditioner changes things. AdamW writes the update as

mt=β1mt−1+(1−β1)gt,m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t,
vt=β2vt−1+(1−β2)gt2,v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2,
θt+1=θt−ηm^tv^t+ϵ−ηλθt.\theta_{t+1} = \theta_t - \eta \frac{\hat m_t}{\sqrt{\hat v_t} + \epsilon} - \eta \lambda \theta_t.

The shrinkage term −ηλθt-\eta \lambda \theta_t is not divided by v^t+ϵ\sqrt{\hat v_t} + \epsilon, so every coordinate decays at the same relative rate ηλ\eta\lambda instead of a rate set by its gradient history. (Here m^t\hat m_t and v^t\hat v_t are the bias-corrected averages defined on the Adam page.)

Conventions for λ\lambda differ, so compare values only within one. Loshchilov and Hutter (2019) write weight decay per step, θt+1=(1−λ)θt−α∇ft(θt)\theta_{t+1}=(1-\lambda)\theta_t-\alpha\nabla f_t(\theta_t), which for plain SGD equals an L2L_2 penalty with coefficient λ/α\lambda/\alpha (their Proposition 1); in their AdamW (Algorithm 2) the decay is scaled by a schedule multiplier, not by the learning rate. PyTorch's AdamW multiplies λ\lambda by the learning rate, as written above.

One step with fixed per-coordinate factors

To see the difference without the moving averages, freeze Adam's per-coordinate factors for one step as a diagonal matrix PP; in Adam, PP would hold 1/(v^t+ϵ)1/(\sqrt{\hat v_t}+\epsilon). Each method's step is a task part plus a regularization part, both computed from the same pre-update θ\theta:

ΔL2=−ηPg−ηP(λθ),ΔW=−ηPg−ηλθ,ΔL2−ΔW=−ηλ (P−I) θ.\Delta_{L_2}=-\eta P g-\eta P(\lambda\theta), \qquad \Delta_{\mathrm W}=-\eta P g-\eta\lambda\theta, \qquad \Delta_{L_2}-\Delta_{\mathrm W}=-\eta\lambda\,(P-I)\,\theta .

Take θ=(2,2)\theta=(2,2), g=(1,−1)g=(1,-1), η=1/10\eta=1/10, λ=1/5\lambda=1/5 and P=diag⁡(4,1/4)P=\operatorname{diag}(4,1/4). The task part is (−2/5, 1/40)(-2/5,\,1/40). The L2L_2 part is (−4/25, −1/100)(-4/25,\,-1/100) and the decoupled part is (−1/25, −1/25)(-1/25,\,-1/25): the penalty pulls four times harder where P=4P=4 and four times more weakly where P=1/4P=1/4. The new weights are (36/25, 403/200)=(1.44, 2.015)(36/25,\,403/200)=(1.44,\,2.015) with the L2L_2 penalty and (39/25, 397/200)=(1.56, 1.985)(39/25,\,397/200)=(1.56,\,1.985) with decoupled decay. In the second coordinate the L2L_2 part points towards zero, yet the weight grows from 22 to 2.0152.015, because the task part is larger: a pull towards zero is not the whole step.

Changing only P11P_{11} from 44 to 11 makes the first coordinate agree (93/5093/50 under both methods) while the second still differs. By the difference formula, with the same λ\lambda the two methods take the same step for every θ\theta and gg only when P=IP=I. A uniform P=sIP=sI can be matched by using λ/s\lambda/s in the penalty, but no single coefficient matches a PP whose diagonal entries differ; this is Loshchilov and Hutter's Proposition 2, and their Proposition 3 interprets decoupled decay under a fixed PP as a rescaled L2L_2 penalty. In Adam, PP changes at every step, and an L2L_2 term added to the gradient also enters both moving averages.

The first Adam step

From zero averages and with ϵ=0\epsilon=0, Adam's first step is −η u/∣u∣-\eta\,u/|u| in each coordinate, where uu is the gradient it receives (see the Adam page). With the L2L_2 penalty u=g+λθu=g+\lambda\theta; in the example u=(7/5, −3/5)u=(7/5,\,-3/5), with the same signs as gg, so the step is (−1/10, 1/10)(-1/10,\,1/10), exactly as with no penalty at all. AdamW takes the same Adam step and subtracts ηλθ=(1/25, 1/25)\eta\lambda\theta=(1/25,\,1/25), giving (−7/50, 3/50)(-7/50,\,3/50). At this step the penalty disappears into Adam's normalization unless it reverses a gradient's sign; later it does not vanish, but it is still divided by v^t\sqrt{\hat v_t}.

Study prompts for this section: Mathematics
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Code

The same calculation as runnable code, in the notation of the derivation.

import numpy as np

theta = np.array([1.0, 1.0])
g = np.array([0.1, 0.1])
vhat = np.array([1e-4, 1.0])  # very different squared-gradient averages
lr = 1e-2
wd = 0.1
eps = 1e-8

adam_with_l2 = theta - lr * ((g + wd * theta) / (np.sqrt(vhat) + eps))
adamw = theta - lr * (g / (np.sqrt(vhat) + eps)) - lr * wd * theta

print("Adam + L2 :", np.round(adam_with_l2, 4))
print("AdamW     :", np.round(adamw, 4))

In the first coordinate, the decay part of the "Adam + L2" step is 100 times larger than AdamW's (0.10.1 instead of 0.0010.001), because Adam divides it by the small sqrt(vhat) =0.01=0.01. In the second coordinate the two agree.

The second block works the example above in exact fractions with Python's standard library: one step with fixed factors, the change of P11P_{11}, P=IP=I, and Adam's first step with each kind of decay.

from fractions import Fraction as F

theta = (F(2), F(2))
g = (F(1), F(-1))
eta, lam = F(1, 10), F(1, 5)


def one_step(P):
    """One step with fixed per-coordinate factors P: task part, L2 part, decoupled part."""
    task = tuple(-eta * p * gi for p, gi in zip(P, g))
    l2 = tuple(-eta * p * lam * t for p, t in zip(P, theta))      # the penalty's gradient goes through P
    decoupled = tuple(-eta * lam * t for t in theta)                # applied directly
    new_l2 = tuple(t + a + b for t, a, b in zip(theta, task, l2))
    new_decoupled = tuple(t + a + b for t, a, b in zip(theta, task, decoupled))
    return new_l2, new_decoupled


new_l2, new_decoupled = one_step((F(4), F(1, 4)))
assert new_l2 == (F(36, 25), F(403, 200))         # coordinate 2 grows, although its L2 part points inward
assert new_decoupled == (F(39, 25), F(397, 200))

changed_l2, changed_decoupled = one_step((F(1), F(1, 4)))
assert changed_l2[0] == changed_decoupled[0] == F(93, 50)
assert one_step((F(1), F(1)))[0] == one_step((F(1), F(1)))[1]   # P = I: the same step


def sign(x):
    return (x > 0) - (x < 0)


# Adam's first step from zero averages (epsilon = 0) moves each coordinate by -eta * sign(input).
adam_l2 = tuple(-eta * sign(gi + lam * t) for gi, t in zip(g, theta))
adam_w = tuple(-eta * sign(gi) - eta * lam * t for gi, t in zip(g, theta))
assert adam_l2 == tuple(-eta * sign(gi) for gi in g)    # the penalty has no effect at this step
assert adam_w == (F(-7, 50), F(3, 50))

print("one step, L2 penalty:", [str(x) for x in new_l2])
print("one step, decoupled decay:", [str(x) for x in new_decoupled])
print("Adam first step, L2 penalty:", [str(x) for x in adam_l2])
print("AdamW first step:", [str(x) for x in adam_w])
Study prompts for this section: Code
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Demo

Predict where the two penalties differ, work one step of each, then change one factor and run Adam's first step.

ExplorePredict which coordinate the L2 penalty pulls in harder before you work the step.An optional observation guide follows the demo.

Where do the two penalties differ?

Loading the interactive demo…

Observation guide

Optional and ungraded. Opening a guide saves nothing.

Choose what to inspect in Weight Decay & AdamW: Decoupled Regularization, then open its guide.

Predict where the two penalties differ, then work one step: the task, L2L_2 and decoupled parts, the two new weights, a change of one factor, and finally Adam's own first step. Each step can be checked, and its solution shown after one try or the hint. The demo on the Adam page has an AdamW switch that adds the decoupled decay to a whole run.

Study prompts for this section: Demo
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

In focus

Concept: Weight Decay & AdamW: Decoupled Regularization

What is the smallest example of Weight Decay & AdamW you can work through by hand, and what does it show?

BeforeAdam OptimizerNextScaling Laws & Emergent Abilities
DetailsOptimization

After the first pass

Go deeper

Optional material for a second reading. Sources are always shown; the other panels open from their headings or from these links, and closing one keeps what you wrote or revealed.
What to look for in each sectionChoose a section and what you expect to see in it, then read short advice.

Section by section

Choose what to look for in each section

Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.

Choose what to look for (no demo yet)01 / Intuition
What to look for

Start with the picture, metaphor, or geometric mechanism.

Choose first

Choose what you expect to see in this section of Weight Decay & AdamW: Decoupled Regularization.

Questions to ask of the pictureChoose a part of the picture and what you expect, then read short advice.

Looking at the picture

Questions to ask of the picture

Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.

4 of 4 sections writtenNo live demo yet
Question

Which part of the picture should you look at first?

Choose first

Pick the part of the picture you expect to explain Weight Decay & AdamW: Decoupled Regularization.

Sources

References for this page

A reference shows where an idea comes from; it does not vouch for every step on this page.
paper · 2019Decoupled Weight Decay RegularizationLoshchilov and HutterSee Section 2, Propositions 1 to 3 and Algorithm 2
Why it is cited

Shows that L2 regularization and weight decay are equivalent for SGD (after rescaling) but not for Adam, and proposes AdamW, which decouples weight decay from the adaptive gradient step (ICLR 2019).

Open source
paper · 2015Adam: A Method for Stochastic OptimizationKingma and Ba
Why it is cited

Defines the Adam update whose per-coordinate scaling AdamW keeps (ICLR 2015).

Open source
book · 2016Deep LearningGoodfellow, Bengio, and CourvilleSee Section 7.1.1, L2 Parameter Regularization
Why it is cited

Textbook treatment of the L2 penalty, its gradient and the per-step shrinkage it causes under gradient descent.

Open source
documentation · 2026PyTorch 2.14 torch.optim.AdamWPyTorch
Why it is cited

Documents the decay step theta <- theta - lr * weight_decay * theta applied outside the adaptive step. Documentation only; no PyTorch code runs here.

Open source
What each claim rests on1 claim, each with its sources, where it appears on this page, and its caveats.

Claims and sources

Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.

Checked against the cited passages

The checks were made by this site, not by an independent reviewer.

With an L2 penalty, Adam's per-coordinate factors also scale the penalty's gradient, so weights with large past gradients are decayed less; decoupled weight decay shrinks every weight by the same fraction. With the same coefficient, the two one-step updates agree for every weight only when the factors are all 1.
What the sources support

Loshchilov and Hutter show the equivalence for SGD, its failure for adaptive methods (Proposition 2) and define AdamW; PyTorch documents the decoupled decay step. The page's one-step comparison with fixed fa...

Where to see it on this page
Equation 3
θt+1=θt−η∇L(θt)−ηλθt=(1−ηλ)θt−η∇L(θt),\theta_{t+1} = \theta_t - \eta \nabla L(\theta_t) - \eta \lambda \theta_t = (1 - \eta \lambda)\theta_t - \eta \nabla L(\theta_t),
Caveat

One step with factors fixed by hand, plus Adam's first step from zero. It shows how the two penalties enter the update, not which one trains or generalizes better.

StatusChecked against the cited passages.

Read Loshchilov and Hutter in the arXiv source: Proposition 1 (SGD equivalence with lambda/alpha), Algorithm 2 (L2 enters g_t before the moments; decay sits outside the adaptive fraction) and Propositions 2 and 3. Recomputed the page's one-step example and Adam's first step in exact fractions.

Explain it without the pageWrite your own explanation before asking for feedback. Your draft and any hints stay while this page is open.

Practice · Weight Decay & AdamW: Decoupled Regularization

Try the idea in your own words

Why shrinking weights is not the same as adding an L2 penalty inside Adam, and how AdamW restores the intended regularization behavior.

Concept · Selected for practice

Weight Decay & AdamW: Decoupled Regularization

What it rests on: Sources: Decoupled Weight Decay Regularization; Adam: A Method for Stochastic Optimization; Deep Learning

Context and links
Choose a task

Explain the mechanism

For Weight Decay & AdamW: Decoupled Regularization: What is the smallest example of Weight Decay & AdamW you can work through by hand, and what does it show? Explain your answer, including what changes, why, and which assumption matters.

No answer yet

A rough first thought is enough. Your draft stays when you change tasks.

Your draft stays on this page and is cleared when you leave or reload.

A little help · Explain

Open one hint at a time. These are suggestions, not your answer or a grade.

0 of 3 hints shown for this question.

    Clearing your answer keeps this help record. Outside help is your own report; this page cannot check it.

    Where am I stuck? (optional)
    Your own description, not an automatic diagnosis

    Choose one, or leave this unspecified. Select it again to clear it.

    Take your draft to a feedback conversation

    No AI feedback runs here. You can copy a prompt to use elsewhere; nothing is sent automatically. Review the text before sharing, and leave out private information.

    Write an attempt before copying a feedback prompt.

    A draft, or an AI reply to it, is not a test of what you have learned. To check that, try a different case later without help.

    Ask about this pagePick one item on this page and copy a prompt about it into an AI assistant you already use. Notes you write here stay in this browser.
    Ask about this pageClose
    ConceptWeight Decay & AdamW: Decoupled RegularizationSources: Decoupled Weight Decay Regularization; Adam: A Method for Stochastic Optimization; Deep Learningassert new_l2 == (F(36, 25), F(403, 200)) # coordinate 2 grows, although its L2 part points inward

    Your question

    Choose what your question is about

    Pick the idea, equation, source, code, claim, misconception or demo state you want to ask about. The prompt you copy includes it, so the answer can stay on that item.
    Next local actionNo local draft saved yet

    Open the draft below to save one note and next action in this browser.

    conceptOptimization

    Weight Decay & AdamW: Decoupled Regularization

    Starting question

    What is the smallest example of Weight Decay & AdamW you can work through by hand, and what does it show?

    SourcesCheck the 4 cited sources listed under Sources.Notes can be saved for this item
    Points of view on this item

    Fixed prompts generated from this item. They are not comments from people or independent reviews.

    Learner’s first stepAsk what would make "Weight Decay & AdamW: Decoupled Regularization" feel predictable rather than familiar.
    Assumption

    The 4 cited sources must support this exact item, not just the surrounding topic.

    Source-checking summary

    Connect the definition to one equation, piece of code or demo before widening the discussion.

    Proposed experiment

    Change one input, then check that the same relationship holds in the mathematics, the code and the demo.

    Next action

    You can state the mechanism in your own words

    Evidence4 checks
    ObservationChecking for a saved observation
    ActionReady for one action
    PromptLearner prompt ready to copy
    Go to this item
    01ObservationChecking what this browser has saved
    02EvidenceChecking for a saved observation
    03SourcesCheck the 4 cited sources listed under Sources.
    04Next stepSave one next action
    Local action draftNo local draft saved yetOpen when you are ready to write one next action
    Local action draft

    This draft stays in this browser and is attached to this item only.

    No local draft saved.
    Evidence to inspect
    • What the 4 cited sources say about this exact item
    • The definition, its prerequisites, and a contrasting concept
    • The equation or code that makes the concept concrete
    • One demo state that shows what stays the same as the inputs change
    What would resolve this
    • You can state the mechanism in your own words
    • You can name the prerequisite that would clear up the confusion
    • You can predict how the result changes when one input changes
    Reference idconcept/concept-notebook/optimization/weight-decay-adamw concept:optimization/weight-decay-adamw