Residual Connections & Skip Connections

Why deep networks can keep useful features alive: each layer learns a correction to the identity instead of rewriting the whole representation from scratch.

Intermediate · Undergraduate mathematics · about 14 minutes

Reading map and next steps

Intuition

The idea in plain words, before the symbols.

Without a skip connection, a deep layer has to reinvent the whole representation it receives.

That is a hard optimization problem. If the best thing a layer could do is "mostly keep what already works, and add a small correction," a plain stack has no easy way to express that.

Residual connections (He et al., 2016) make that easy. Instead of asking a layer to learn a full map x↦yx \mapsto y, we ask it to learn a change:

y=x+F(x).y = x + F(x).

Now the safe default is clear:

  • if the layer has nothing useful to add, it can make F(x)≈0F(x) \approx 0,
  • if it has a useful feature, it writes that feature on top of the existing state,
  • information always has a direct path forward through depth.

In transformers, this leads to the picture of a residual stream: attention and MLP blocks both read the current state, write an update, and pass the shared stream onward.

Study prompts for this section: Intuition
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Demo

Send one vector forward and one gradient backward through a residual block, then change only the branch's slope.

PredictAnswer each problem before you open the figure.Each solution appears after you check or ask for a hint.

Follow one vector forward and one gradient backward through y=x+a xy=x+a\,x: predict whether the branch can cancel the skip, work both directions at a=12a=\tfrac12, then change only aa to −1-1. The figure lets you try other slopes and compare the block with a plain layer y=F(x)y=F(x).

Can the branch cancel the skip?

Loading the interactive demo…

Observation guide

Optional and ungraded. Opening a guide saves nothing.

Choose what to inspect in Residual Connections & Skip Connections, then open its guide.

Study prompts for this section: Demo
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Mathematics

Definitions, assumptions and the derivation.

Let x∈Rdx \in \mathbb{R}^d be the input to a block and let F:Rd→RdF: \mathbb{R}^d \to \mathbb{R}^d be the learned update. A residual block outputs

y=x+F(x).y = x + F(x).

The Jacobian of this mapping is

∂y∂x=I+JF(x),\frac{\partial y}{\partial x} = I + J_F(x),

where II is the identity matrix and JF(x)J_F(x) is the Jacobian of FF.

That identity term matters for optimization. Across many layers,

h(ℓ+1)=h(ℓ)+F(ℓ)(h(ℓ)),h^{(\ell+1)} = h^{(\ell)} + F^{(\ell)}(h^{(\ell)}),

so backpropagation multiplies matrices of the form I+JF(ℓ)I + J_{F^{(\ell)}}. Even if the learned part is small or noisy, there is still a direct gradient path through the identity.

A direct path is not a guarantee that the total is nonzero. The identity and the branch meet at a plus sign, forward and backward, and the branch can cancel the identity exactly. Take x=(2,−1)x=(2,-1) and the linear branch F(x)=a xF(x)=a\,x, so ∂y/∂x=(1+a)I\partial y/\partial x=(1+a)I, and let the gradient arriving at yy be g=(2,−3)g=(2,-3). At a=12a=\tfrac12,

y=(3,−32),∂L∂x=32 g=(3,−92),y=\left(3,-\tfrac32\right),\qquad \frac{\partial L}{\partial x}=\tfrac32\,g=\left(3,-\tfrac92\right),

the sum of the skip's gg and the branch's 12g\tfrac12 g. At a=−1a=-1 the branch outputs −x-x and its Jacobian is −I-I:

y=x−x=(0,0),∂y∂x=I−I=0,∂L∂x=(0,0).y=x-x=(0,0),\qquad \frac{\partial y}{\partial x}=I-I=0,\qquad \frac{\partial L}{\partial x}=(0,0).

One linear example cannot say how often trained networks come near such a cancellation. It refutes only the stronger claim that a skip guarantees a nonzero gradient. The identity-mapping analysis of He et al. (2016b, §2) makes the same point in words: the gradient is unlikely to cancel across a whole mini-batch, which is not the same as impossible.

For a pre-norm transformer layer, a common pattern is

h′=h+Attn⁡(LN⁡(h)),h' = h + \operatorname{Attn}(\operatorname{LN}(h)),
hnext=h′+MLP⁡(LN⁡(h′)).h_{next} = h' + \operatorname{MLP}(\operatorname{LN}(h')).

This says each sublayer writes an update into a shared stream rather than replacing the stream entirely.

Study prompts for this section: Mathematics
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Code

The same calculation as runnable code, in the notation of the derivation.

import numpy as np

rs = np.random.RandomState(0)
d = 128
x0 = rs.randn(d)

def run(depth=40, residual=True, alpha=0.2, seed=1):
    weights = np.random.RandomState(seed)  # the same matrices for both runs
    h = x0.copy()
    norms = []
    for _ in range(depth):
        W = weights.randn(d, d) / np.sqrt(d)
        update = np.tanh(W @ h)
        h = h + alpha * update if residual else update
        norms.append(np.linalg.norm(h))
    return norms

plain = run(residual=False)
resid = run(residual=True)

print("plain layer-1/layer-40:", round(plain[0], 3), round(plain[-1], 3))
print("resid layer-1/layer-40:", round(resid[0], 3), round(resid[-1], 3))

Both runs use the same 40 random matrices. The plain stack replaces its state with tanh⁡(Wh)\tanh(Wh) at every layer, and its norm shrinks from about 7.3 to 1.2. The residual stack keeps its state and adds 0.2tanh⁡(Wh)0.2\tanh(Wh) to it, so it stays near the input's scale (its norm is 11.9), going from about 11.8 to 15.4. The two runs differ in the skip and in that 0.2 scale on the update.

Residual connections do not remove every instability, but they give optimization a much safer default: keep the current representation unless there is a good reason to change it.

Study prompts for this section: Code
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

In focus

Concept: Residual Connections & Skip Connections

What is the smallest example of Residual Connections & Skip Connections you can work through by hand, and what does it show?

BeforeScaled Dot-Product Attention & Transformer LayersNextLayer Normalization & RMSNorm
DetailsAttention & Transformers

After the first pass

Go deeper

Optional material for a second reading. Sources are always shown; the other panels open from their headings or from these links, and closing one keeps what you wrote or revealed.
What to look for in each sectionChoose a section and what you expect to see in it, then read short advice.

Section by section

Choose what to look for in each section

Why deep networks can keep useful features alive: each layer learns a correction to the identity instead of rewriting the whole representation from scratch.

Choose what to look for (no demo yet)01 / Intuition
What to look for

Start with the picture, metaphor, or geometric mechanism.

Choose first

Choose what you expect to see in this section of Residual Connections & Skip Connections.

Questions to ask of the pictureChoose a part of the picture and what you expect, then read short advice.

Looking at the picture

Questions to ask of the picture

Why deep networks can keep useful features alive: each layer learns a correction to the identity instead of rewriting the whole representation from scratch.

4 of 4 sections writtenNo live demo yet
Question

Which part of the picture should you look at first?

Choose first

Pick the part of the picture you expect to explain Residual Connections & Skip Connections.

Sources

References for this page

A reference shows where an idea comes from; it does not vouch for every step on this page.
paper · 2016Deep Residual Learning for Image RecognitionHe, Zhang, Ren, and Sun
Why it is cited

CVPR 2016. Sections 3.1–3.2 define the residual block y = F(x) + x with an identity shortcut added element by element; Section 4 finds deep residual networks easier to optimize than plain ones.

Open source
paper · 2016Identity Mappings in Deep Residual NetworksHe, Zhang, Ren, and Sun
Why it is cited

ECCV 2016. Section 2 shows the identity path adds a direct term to the gradient and says cancellation across a mini-batch is unlikely, not impossible.

Open source
book · 2024Deep Learning: Foundations and ConceptsBishop and Bishop
Why it is cited

Springer. Chapter 12 (Transformers) builds the transformer layer from attention and feed-forward sublayers, each wrapped in a residual connection.

Open source
What each claim rests on1 claim, each with its sources, where it appears on this page, and its caveats.

Claims and sources

Why deep networks can keep useful features alive: each layer learns a correction to the identity instead of rewriting the whole representation from scratch.

Checked against the cited passages

The checks were made by this site, not by an independent reviewer.

An identity skip makes a block compute y = x + F(x), whose Jacobian is I + J_F: the skip adds a direct term forward and backward, and the branch can cancel either total.
What the sources support

He et al. (CVPR 2016) define y = F(x) + x with an element-wise identity shortcut. He et al. (ECCV 2016, §2) show the direct gradient term and call cancellation unlikely, not impossible.

Where to see it on this page
Equation 2
∂y∂x=I+JF(x),\frac{\partial y}{\partial x} = I + J_F(x),
Caveat

The branch F(x) = a·x and its slopes are chosen for exact arithmetic. One linear example refutes a guarantee; it says nothing about how trained networks behave.

StatusChecked against the cited passages.

Both papers were read for the residual form, the element-wise identity shortcut and the gradient argument after Eqn. 5 of the identity-mappings paper. The forward and backward values at a = 1/2 and a = −1 were recomputed independently with exact fractions.

Explain it without the pageWrite your own explanation before asking for feedback. Your draft and any hints stay while this page is open.

Practice · Residual Connections & Skip Connections

Try the idea in your own words

Why deep networks can keep useful features alive: each layer learns a correction to the identity instead of rewriting the whole representation from scratch.

Concept · Selected for practice

Residual Connections & Skip Connections

What it rests on: Sources: Deep Residual Learning for Image Recognition; Identity Mappings in Deep Residual Networks; Deep Learning: Foundations and Concepts

Context and links
Choose a task

Explain the mechanism

For Residual Connections & Skip Connections: What is the smallest example of Residual Connections & Skip Connections you can work through by hand, and what does it show? Explain your answer, including what changes, why, and which assumption matters.

No answer yet

A rough first thought is enough. Your draft stays when you change tasks.

Your draft stays on this page and is cleared when you leave or reload.

A little help · Explain

Open one hint at a time. These are suggestions, not your answer or a grade.

0 of 3 hints shown for this question.

    Clearing your answer keeps this help record. Outside help is your own report; this page cannot check it.

    Where am I stuck? (optional)
    Your own description, not an automatic diagnosis

    Choose one, or leave this unspecified. Select it again to clear it.

    Take your draft to a feedback conversation

    No AI feedback runs here. You can copy a prompt to use elsewhere; nothing is sent automatically. Review the text before sharing, and leave out private information.

    Write an attempt before copying a feedback prompt.

    A draft, or an AI reply to it, is not a test of what you have learned. To check that, try a different case later without help.

    Ask about this pagePick one item on this page and copy a prompt about it into an AI assistant you already use. Notes you write here stay in this browser.
    Ask about this pageClose
    ConceptResidual Connections & Skip ConnectionsSources: Deep Residual Learning for Image Recognition; Identity Mappings in Deep Residual Networks; Deep Learning: Foundations and Conceptsrs = np.random.RandomState(0)

    Your question

    Choose what your question is about

    Pick the idea, equation, source, code, claim, misconception or demo state you want to ask about. The prompt you copy includes it, so the answer can stay on that item.
    Next local actionNo local draft saved yet

    Open the draft below to save one note and next action in this browser.

    conceptAttention & Transformers

    Residual Connections & Skip Connections

    Starting question

    What is the smallest example of Residual Connections & Skip Connections you can work through by hand, and what does it show?

    SourcesCheck the 3 cited sources listed under Sources.Notes can be saved for this item
    Points of view on this item

    Fixed prompts generated from this item. They are not comments from people or independent reviews.

    Learner’s first stepAsk what would make "Residual Connections & Skip Connections" feel predictable rather than familiar.
    Assumption

    The 3 cited sources must support this exact item, not just the surrounding topic.

    Source-checking summary

    Connect the definition to one equation, piece of code or demo before widening the discussion.

    Proposed experiment

    Change one input, then check that the same relationship holds in the mathematics, the code and the demo.

    Next action

    You can state the mechanism in your own words

    Evidence4 checks
    ObservationChecking for a saved observation
    ActionReady for one action
    PromptLearner prompt ready to copy
    Go to this item
    01ObservationChecking what this browser has saved
    02EvidenceChecking for a saved observation
    03SourcesCheck the 3 cited sources listed under Sources.
    04Next stepSave one next action
    Local action draftNo local draft saved yetOpen when you are ready to write one next action
    Local action draft

    This draft stays in this browser and is attached to this item only.

    No local draft saved.
    Evidence to inspect
    • What the 3 cited sources say about this exact item
    • The definition, its prerequisites, and a contrasting concept
    • The equation or code that makes the concept concrete
    • One demo state that shows what stays the same as the inputs change
    What would resolve this
    • You can state the mechanism in your own words
    • You can name the prerequisite that would clear up the confusion
    • You can predict how the result changes when one input changes
    Reference idconcept/concept-notebook/attention-transformers/residual-connections concept:attention-transformers/residual-connections