Gradient Clipping & Explosion Prevention

How to stop one extreme batch from blowing up training: clip the gradient norm, bound the update, and keep the gradient's direction.

Intermediate · Undergraduate mathematics · about 13 minutes

Reading map and next steps

Intuition

The idea in plain words, before the symbols.

Most batches are boring. A few are not.

On those bad batches, the gradient can spike because activations, logits, or long chains of Jacobians produce an unusually large backpropagated signal. If you apply that gradient naively, one step can throw the model far away from the regime where training had been stable.

Gradient clipping (Pascanu et al., 2013) is a simple safeguard:

  • let normal gradients pass untouched,
  • when a gradient is too large, scale it back before taking the optimizer step.

The point is not to "fix" the optimization problem permanently. The point is to stop rare outliers from destroying the run.

Norm clipping is usually the right mental model: keep the direction of the gradient, but cap its magnitude.

One caution, worked through below: clipping caps the gradient the optimizer receives. Only for plain gradient descent does that alone cap the step; an adaptive optimizer rescales what it receives, and may average it with earlier gradients, before it moves the weights.

Study prompts for this section: Intuition
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Mathematics

Definitions, assumptions and the derivation.

Let g∈Rdg \in \mathbb{R}^d be the gradient and let c>0c > 0 be the clipping threshold.

Global norm clipping defines

g~={gif ∥g∥2≤c,c∥g∥2gif ∥g∥2>c.\tilde g = \begin{cases} g & \text{if } \lVert g \rVert_2 \le c, \\\\ \frac{c}{\lVert g \rVert_2} g & \text{if } \lVert g \rVert_2 > c. \end{cases}

Equivalently,

g~=min⁡(1,c∥g∥2)g.\tilde g = \min\left(1, \frac{c}{\lVert g \rVert_2}\right) g.

If your optimizer step is Δθ=−ηg~\Delta \theta = -\eta \tilde g, then whenever clipping activates,

∥Δθ∥2=ηc.\lVert \Delta \theta \rVert_2 = \eta c.

So clipping puts a hard cap on the update size.

Why can gradients explode in the first place? In a deep network, gradients are products of Jacobians:

∂L∂h(1)=∂L∂h(L)∏ℓ=2L∂h(ℓ)∂h(ℓ−1).\frac{\partial L}{\partial h^{(1)}} = \frac{\partial L}{\partial h^{(L)}} \prod_{\ell=2}^{L} \frac{\partial h^{(\ell)}}{\partial h^{(\ell-1)}}.

If those Jacobians repeatedly amplify norms, backpropagated signals can grow exponentially with depth or sequence length.

Which vector is bounded

Clipping acts on the gradient before the optimizer uses it. The simplest thing an optimizer can then do is rescale each coordinate of the clipped gradient by a fixed factor. Write those factors as a matrix PP, so that

q=Pg~,Δθ=−η q,∥Δθ∥2≤η ∥P∥2 min⁡(∥g∥2, c),q=P\tilde g, \qquad \Delta\theta=-\eta\,q, \qquad \lVert\Delta\theta\rVert_2\le\eta\,\lVert P\rVert_2\,\min\left(\lVert g\rVert_2,\,c\right),

where ∥P∥2\lVert P\rVert_2 is the largest factor by which PP can stretch a vector. Plain gradient descent is P=IP=I, and the bound becomes the equality above whenever clipping is active.

Take g=(6,8)g=(6,8), c=5c=5 and η=1/10\eta=1/10. Then ∥g∥2=10\lVert g\rVert_2=10, the scale is 1/21/2 and g~=(3,4)\tilde g=(3,4), of length exactly 55. Plain gradient descent steps by (−3/10,−2/5)(-3/10,-2/5), of length 1/2=ηc1/2=\eta c, where without clipping it would have stepped by (−3/5,−4/5)(-3/5,-4/5), of length 11. Now let P=diag⁡(2,1/2)P=\operatorname{diag}(2,1/2). Then q=(6,2)q=(6,2) and Δθ=(−3/5,−1/5)\Delta\theta=(-3/5,-1/5), of length 10/5≈0.632\sqrt{10}/5\approx0.632. The clip is still correct, but the step is longer than ηc\eta c and no longer points along −g-g; it stays within the bound η∥P∥2c=1\eta\lVert P\rVert_2c=1. A uniform scaling P=sIP=sI would keep the direction and multiply the length by ss. When the gradient's norm equals cc exactly, the scale is 11 whether an implementation rescales at ∥g∥2>c\lVert g\rVert_2>c or at ∥g∥2≥c\lVert g\rVert_2\ge c.

Adam does more than rescale. Its step divides a running average of past gradients by the root of a running average of their squares, so it depends on the clipped gradients of earlier steps too: with a zero gradient at one step, Adam still moves whenever its average of earlier gradients is not zero, while the bound above would allow no movement. Clipping therefore bounds what enters Adam's averages, not Adam's step.

In PyTorch, torch.nn.utils.clip_grad_norm_ computes one norm over the gradients of all the parameters given to it, as if they were concatenated into a single vector, and rescales them in place; the optimizer step then runs on the clipped gradients.

Study prompts for this section: Mathematics
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Code

The same calculation as runnable code, in the notation of the derivation.

import numpy as np

g = np.array([6.0, 8.0, 0.0])  # norm = 10
c = 5.0
lr = 1e-2

scale = min(1.0, c / np.linalg.norm(g))
g_clip = scale * g

print("raw grad norm    :", np.linalg.norm(g))
print("clipped grad norm:", np.linalg.norm(g_clip))
print("raw step norm    :", np.linalg.norm(lr * g))
print("clipped step norm:", np.linalg.norm(lr * g_clip))

This keeps the update direction the same but limits the step size to lr * c.

The second block works the example above in exact fractions, using only Python's standard library: the clip, the plain step of length ηc\eta c, and the step after the per-coordinate factors, which is longer and turned.

from fractions import Fraction as F

g = (F(6), F(8))
c, eta = F(5), F(1, 10)


def norm_squared(v):
    return sum(x * x for x in v)


scale = min(F(1), c / 10)            # ||g|| = 10 exactly, since 6^2 + 8^2 = 10^2
g_clip = tuple(scale * x for x in g)
assert norm_squared(g) == 100 and g_clip == (3, 4) and norm_squared(g_clip) == c * c

plain_step = tuple(-eta * x for x in g_clip)
assert plain_step == (F(-3, 10), F(-2, 5)) and norm_squared(plain_step) == (eta * c) ** 2

P = (F(2), F(1, 2))                  # per-coordinate factors applied after the clip
step = tuple(-eta * p * x for p, x in zip(P, g_clip))
assert step == (F(-3, 5), F(-1, 5))
assert norm_squared(step) == F(2, 5)               # length sqrt(10)/5, more than eta * c = 1/2
assert step[0] * g[1] - step[1] * g[0] != 0        # no longer parallel to g
assert norm_squared(step) <= (eta * max(P) * c) ** 2   # within eta * ||P|| * c = 1

print("clipped gradient:", [str(x) for x in g_clip])
print("plain step:", [str(x) for x in plain_step], "squared length", norm_squared(plain_step))
print("scaled step:", [str(x) for x in step], "squared length", norm_squared(step))
Study prompts for this section: Code
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Demo

Predict what clipping bounds, clip one gradient, then change only what the optimizer does after the clip.

ExplorePredict which length clipping fixes at the threshold before you work the example.An optional observation guide follows the demo.

Which vector does clipping bound?

Loading the interactive demo…

Observation guide

Optional and ungraded. Opening a guide saves nothing.

Choose what to inspect in Gradient Clipping & Explosion Prevention, then open its guide.

Predict which length clipping fixes, then work the example: clip g=(6,8)g=(6,8) at c=5c=5, take a plain gradient-descent step, and finally scale the coordinates after the clip by P=diag⁡(2,1/2)P=\operatorname{diag}(2,1/2) to see the step turn and lengthen. Each step can be checked, and its solution shown after one try or the hint.

Study prompts for this section: Demo
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

In focus

Concept: Gradient Clipping & Explosion Prevention

What is the smallest example of Gradient Clipping & Explosion Prevention you can work through by hand, and what does it show?

BeforeBackpropagationNextAdam Optimizer
DetailsOptimization

After the first pass

Go deeper

Optional material for a second reading. Sources are always shown; the other panels open from their headings or from these links, and closing one keeps what you wrote or revealed.
What to look for in each sectionChoose a section and what you expect to see in it, then read short advice.

Section by section

Choose what to look for in each section

How to stop one extreme batch from blowing up training: clip the gradient norm, bound the update, and keep the gradient's direction.

Choose what to look for (no demo yet)01 / Intuition
What to look for

Start with the picture, metaphor, or geometric mechanism.

Choose first

Choose what you expect to see in this section of Gradient Clipping & Explosion Prevention.

Questions to ask of the pictureChoose a part of the picture and what you expect, then read short advice.

Looking at the picture

Questions to ask of the picture

How to stop one extreme batch from blowing up training: clip the gradient norm, bound the update, and keep the gradient's direction.

4 of 4 sections writtenNo live demo yet
Question

Which part of the picture should you look at first?

Choose first

Pick the part of the picture you expect to explain Gradient Clipping & Explosion Prevention.

Sources

References for this page

A reference shows where an idea comes from; it does not vouch for every step on this page.
paper · 2013On the difficulty of training recurrent neural networksPascanu, Mikolov, and BengioSee Section 3.2, Algorithm 1
Why it is cited

Analyzes exploding and vanishing gradients in recurrent networks and proposes rescaling the gradient whenever its norm exceeds a threshold (ICML 2013).

Open source
book · 2016Deep LearningGoodfellow, Bengio, and CourvilleSee Section 10.11.1, Clipping Gradients
Why it is cited

Textbook treatment of norm clipping and element-wise clipping for recurrent networks, and of why a bounded step helps near steep regions of the loss.

Open source
documentation · 2026PyTorch 2.14 torch.nn.utils.clip_grad_norm_PyTorch
Why it is cited

Computes one total norm over the gradients of all the given parameters, as if they were one vector, rescales the gradients in place and returns that norm. Documentation only; no PyTorch code runs here.

Open source
What each claim rests on1 claim, each with its sources, where it appears on this page, and its caveats.

Claims and sources

How to stop one extreme batch from blowing up training: clip the gradient norm, bound the update, and keep the gradient's direction.

Checked against the cited passages

The checks were made by this site, not by an independent reviewer.

Global norm clipping shortens a gradient longer than c to length c without changing its direction, and leaves shorter gradients unchanged. When clipping is active, a plain gradient-descent step has length eta*c; if the optimizer rescales coordinates after the clip, the step can turn and be longer, up to eta*||P||*c.
What the sources support

Pascanu et al. define norm clipping as rescaling the gradient to the threshold; PyTorch documents clipping one total norm in place before the optimizer step. The step-length equality and bound follow from th...

Where to see it on this page
Equation 3
∥Δθ∥2=ηc.\lVert \Delta \theta \rVert_2 = \eta c.
Caveat

A worked example with a fixed per-coordinate rescaling; Adam also averages past gradients, so its step depends on earlier clipped gradients too. It shows which vector is bounded, not how clipping affects tra...

StatusChecked against the cited passages.

Pascanu et al. Section 3.2, Algorithm 1 rescales the gradient to the threshold when its norm exceeds it; the PyTorch 2.14 documentation for clip_grad_norm_ describes the total norm over concatenated gradients and the in-place rescaling. The plain-step equality, the operator-norm bound and the worked example (g = (6, 8), c = 5, eta = 1/10, P = diag(2, 1/2)) were derived and recomputed in exact fractions.

Explain it without the pageWrite your own explanation before asking for feedback. Your draft and any hints stay while this page is open.

Practice · Gradient Clipping & Explosion Prevention

Try the idea in your own words

How to stop one extreme batch from blowing up training: clip the gradient norm, bound the update, and keep the gradient's direction.

Concept · Selected for practice

Gradient Clipping & Explosion Prevention

What it rests on: Sources: On the difficulty of training recurrent neural networks; Deep Learning; PyTorch 2.14 torch.nn.utils.clip_grad_norm_

Context and links
Choose a task

Explain the mechanism

For Gradient Clipping & Explosion Prevention: What is the smallest example of Gradient Clipping & Explosion Prevention you can work through by hand, and what does it show? Explain your answer, including what changes, why, and which assumption matters.

No answer yet

A rough first thought is enough. Your draft stays when you change tasks.

Your draft stays on this page and is cleared when you leave or reload.

A little help · Explain

Open one hint at a time. These are suggestions, not your answer or a grade.

0 of 3 hints shown for this question.

    Clearing your answer keeps this help record. Outside help is your own report; this page cannot check it.

    Where am I stuck? (optional)
    Your own description, not an automatic diagnosis

    Choose one, or leave this unspecified. Select it again to clear it.

    Take your draft to a feedback conversation

    No AI feedback runs here. You can copy a prompt to use elsewhere; nothing is sent automatically. Review the text before sharing, and leave out private information.

    Write an attempt before copying a feedback prompt.

    A draft, or an AI reply to it, is not a test of what you have learned. To check that, try a different case later without help.

    Ask about this pagePick one item on this page and copy a prompt about it into an AI assistant you already use. Notes you write here stay in this browser.
    Ask about this pageClose
    ConceptGradient Clipping & Explosion PreventionSources: On the difficulty of training recurrent neural networks; Deep Learning; PyTorch 2.14 torch.nn.utils.clip_grad_norm_assert norm_squared(g) == 100 and g_clip == (3, 4) and norm_squared(g_clip) == c * c

    Your question

    Choose what your question is about

    Pick the idea, equation, source, code, claim, misconception or demo state you want to ask about. The prompt you copy includes it, so the answer can stay on that item.
    Next local actionNo local draft saved yet

    Open the draft below to save one note and next action in this browser.

    conceptOptimization

    Gradient Clipping & Explosion Prevention

    Starting question

    What is the smallest example of Gradient Clipping & Explosion Prevention you can work through by hand, and what does it show?

    SourcesCheck the 3 cited sources listed under Sources.Notes can be saved for this item
    Points of view on this item

    Fixed prompts generated from this item. They are not comments from people or independent reviews.

    Learner’s first stepAsk what would make "Gradient Clipping & Explosion Prevention" feel predictable rather than familiar.
    Assumption

    The 3 cited sources must support this exact item, not just the surrounding topic.

    Source-checking summary

    Connect the definition to one equation, piece of code or demo before widening the discussion.

    Proposed experiment

    Change one input, then check that the same relationship holds in the mathematics, the code and the demo.

    Next action

    You can state the mechanism in your own words

    Evidence4 checks
    ObservationChecking for a saved observation
    ActionReady for one action
    PromptLearner prompt ready to copy
    Go to this item
    01ObservationChecking what this browser has saved
    02EvidenceChecking for a saved observation
    03SourcesCheck the 3 cited sources listed under Sources.
    04Next stepSave one next action
    Local action draftNo local draft saved yetOpen when you are ready to write one next action
    Local action draft

    This draft stays in this browser and is attached to this item only.

    No local draft saved.
    Evidence to inspect
    • What the 3 cited sources say about this exact item
    • The definition, its prerequisites, and a contrasting concept
    • The equation or code that makes the concept concrete
    • One demo state that shows what stays the same as the inputs change
    What would resolve this
    • You can state the mechanism in your own words
    • You can name the prerequisite that would clear up the confusion
    • You can predict how the result changes when one input changes
    Reference idconcept/concept-notebook/optimization/gradient-clipping concept:optimization/gradient-clipping