SwiGLU: Gated MLP Blocks in Transformers

Why modern transformer MLPs often use a learned multiplicative gate: one projection proposes a token-local write while another projection controls how much of each channel reaches the residual stream.

Advanced · Graduate mathematics · about 16 minutes

Reading map and next steps

Intuition

The idea in plain words, before the symbols.

Main sources: Shazeer (2020), "GLU Variants Improve Transformer", Chowdhery et al. (2023), "PaLM: Scaling Language Modeling with Pathways", Touvron et al. (2023), "LLaMA: Open and Efficient Foundation Language Models", and Vaswani et al. (2017), "Attention Is All You Need".

Attention is the token-mixing part of a transformer block: each token reads from other tokens. The MLP or feedforward sublayer is the token-local writing part: each token takes its own residual vector, expands it into hidden channels, applies a nonlinearity, and writes a correction back.

A plain ReLU or GELU feedforward block asks each hidden channel one question: "how active is this feature?" SwiGLU asks two questions. One projection proposes a value. A second projection produces a gate. The final hidden channel is their product after passing the gate logit through SiLU.

That small change matters conceptually. The MLP is no longer just "linear, activation, linear." It becomes a bank of conditional writes: feature channels can be suppressed, passed, amplified or reversed in sign, depending on another learned view of the same token state.

This page is only about the channel-level mechanism. It is not a benchmark claim that SwiGLU always beats GELU, and the common two-thirds hidden-width rule is a parameter-budget comparison, not a law of nature.

Study prompts for this section: Intuition
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Mathematics

Definitions, assumptions and the derivation.

Use row-vector notation for one token. Let

x∈Rdmodelx \in \mathbb R^{d_{\text{model}}}

be the residual stream vector for that token. A standard two-matrix feedforward block with hidden width dffd_{\text{ff}} has roughly

2dmodeldff2d_{\text{model}}d_{\text{ff}}

matrix parameters, ignoring biases.

SwiGLU uses two input projections and one output projection. Choose a gated hidden width dff′d_{\text{ff}}^{\prime} and define

Wv,Wg∈Rdmodel×dff′,Wo∈Rdff′×dmodel.W_v, W_g \in \mathbb R^{d_{\text{model}}\times d_{\text{ff}}^{\prime}}, \qquad W_o \in \mathbb R^{d_{\text{ff}}^{\prime}\times d_{\text{model}}}.

The value and gate logits are

v=xWv∈Rdff′,g=xWg∈Rdff′.v=xW_v \in \mathbb R^{d_{\text{ff}}^{\prime}}, \qquad g=xW_g \in \mathbb R^{d_{\text{ff}}^{\prime}}.

SiLU, also called Swish, is

SiLU⁡(z)=zσ(z),σ(z)=11+e−z.\operatorname{SiLU}(z)=z\sigma(z), \qquad \sigma(z)=\frac{1}{1+e^{-z}}.

The gated hidden vector is the elementwise product

h=v⊙SiLU⁡(g).h=v\odot \operatorname{SiLU}(g).

The token-local MLP write is

y=hWo=(v⊙SiLU⁡(g))Wo=(xWv⊙SiLU⁡(xWg))Wo.y=hW_o =\bigl(v\odot \operatorname{SiLU}(g)\bigr)W_o =\bigl(xW_v\odot \operatorname{SiLU}(xW_g)\bigr)W_o.

The parameter count is approximately

3dmodeldff′.3d_{\text{model}}d_{\text{ff}}^{\prime}.

To compare it to a two-matrix block with expansion width dffd_{\text{ff}}, set

dff′≈23dff.d_{\text{ff}}^{\prime}\approx \frac{2}{3}d_{\text{ff}}.

Then

3dmodeldff′≈3dmodel(23dff)=2dmodeldff.3d_{\text{model}}d_{\text{ff}}^{\prime} \approx 3d_{\text{model}}\left(\frac{2}{3}d_{\text{ff}}\right) =2d_{\text{model}}d_{\text{ff}}.

So the usual comparison is not "SwiGLU has a free extra matrix." It is "SwiGLU spends a similar budget differently: two narrower input projections create a learned multiplicative gate."

One channel, worked by hand

Hold one hidden channel's value at vi=4v_i=4 and change only its gate logit. At gi=ln⁡3g_i=\ln 3, e−ln⁡3=1/3e^{-\ln 3}=1/3, so σ(ln⁡3)=3/4\sigma(\ln 3)=3/4, SiLU⁡(ln⁡3)=34ln⁡3\operatorname{SiLU}(\ln 3)=\tfrac34\ln 3 and hi=3ln⁡3≈3.296h_i=3\ln 3\approx 3.296. At gi=−ln⁡3g_i=-\ln 3 the sigmoid is still positive, σ(−ln⁡3)=1/4\sigma(-\ln 3)=1/4, but SiLU also multiplies by the negative logit: SiLU⁡(−ln⁡3)=−14ln⁡3\operatorname{SiLU}(-\ln 3)=-\tfrac14\ln 3 and hi=−ln⁡3≈−1.099h_i=-\ln 3\approx-1.099. The gate has reversed the sign of the value. It cannot push it far the other way: for g<0g<0, SiLU⁡(g)\operatorname{SiLU}(g) lies between about −0.278-0.278 (its lowest point, at g≈−1.278g\approx-1.278) and 00. The sign of one channel before WoW_o does not fix the sign of yy, which WoW_o mixes from every channel.

The budget rule is exact in a small case. A bias-free plain block with dmodel=3d_{\text{model}}=3 and dff=6d_{\text{ff}}=6 has 3⋅6+6⋅3=363\cdot6+6\cdot3=36 matrix entries. A bias-free SwiGLU block with dff′=4=23⋅6d_{\text{ff}}^{\prime}=4=\tfrac23\cdot 6 has 3⋅4+3⋅4+4⋅3=363\cdot4+3\cdot4+4\cdot3=36. This is how Shazeer (2020, §2) keeps the parameter count and computation fixed when comparing variants; it counts matrix entries, not speed, activation memory or quality.

Study prompts for this section: Mathematics
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Code

The same calculation as runnable code, in the notation of the derivation.

This code follows the math step by step. It uses one token vector, two input projections, a SiLU gate, an elementwise product, and an output projection. The final assertion checks the common two-thirds budget comparison.

import numpy as np

def silu(z):
    return z / (1.0 + np.exp(-z))

d_model = 6
d_ff = 12
d_ff_prime = round((2.0 / 3.0) * d_ff)

x = np.array([[0.7, -0.4, 0.2, 1.1, -0.3, 0.5]])  # shape: (1, d_model)

rng = np.random.default_rng(7)
W_v = rng.normal(0.0, 0.25, size=(d_model, d_ff_prime))
W_g = rng.normal(0.0, 0.25, size=(d_model, d_ff_prime))
W_o = rng.normal(0.0, 0.25, size=(d_ff_prime, d_model))

v = x @ W_v                      # shape: (1, d_ff_prime)
g = x @ W_g                      # shape: (1, d_ff_prime)
gate = silu(g)                   # shape: (1, d_ff_prime)
h = v * gate                     # elementwise product
y = h @ W_o                      # shape: (1, d_model)

relu_ffn_params = 2 * d_model * d_ff
swiglu_params = 3 * d_model * d_ff_prime
ratio = swiglu_params / relu_ffn_params

assert v.shape == (1, d_ff_prime)
assert g.shape == (1, d_ff_prime)
assert gate.shape == (1, d_ff_prime)
assert h.shape == (1, d_ff_prime)
assert y.shape == (1, d_model)
assert 0.95 <= ratio <= 1.05

channel = 3
print({
    "v_i": float(v[0, channel]),
    "g_i": float(g[0, channel]),
    "SiLU(g_i)": float(gate[0, channel]),
    "product": float(h[0, channel]),
    "budget_ratio": ratio,
})

The printed channel is the scalar version of the interactive demo. One number proposes a hidden-channel write, another number controls the gate, and their product decides what reaches the output projection.

Study prompts for this section: Code
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Demo

Change an input, predict the result, and see which quantities respond.

ExploreChange one control at a time and name what responds.An optional observation guide follows the demo.

Explore SwiGLU: Gated MLP Blocks in Transformers

Loading the interactive demo…

Observation guide

Optional and ungraded. Opening a guide saves nothing.

Choose what to inspect in SwiGLU: Gated MLP Blocks in Transformers, then open its guide.

Use the Gated MLP Write demo to inspect one synthetic hidden channel. Before reveal, you can see the value projection viv_i and the gate logit gig_i, but not SiLU⁡(gi)\operatorname{SiLU}(g_i) or the product.

Predict whether the gate reverses, suppresses, passes, or amplifies the value. Then reveal the gate coefficient, the product viSiLU⁡(gi)v_i\operatorname{SiLU}(g_i), and a small toy selected-channel contribution through the output projection. The full MLP write would sum many such channel contributions; this demo isolates one. The point is not to memorize a curve but to see why a gated MLP can decide, channel by channel, what to write.

Study prompts for this section: Demo
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

When the gate turns a value around

Hold one channel's value fixed, flip the sign of its gate logit, then count every matrix in two equal budgets.

Loading the exercise…
In focus

Concept: SwiGLU: Gated MLP Blocks in Transformers

What is the smallest example of SwiGLU you can work through by hand, and what does it show?

BeforeScaled Dot-Product Attention & Transformer LayersNextSparse Mixture of Experts: Routing, Load Balancing & Expert Parallelism
DetailsAttention & Transformers

After the first pass

Go deeper

Optional material for a second reading. Sources are always shown; the other panels open from their headings or from these links, and closing one keeps what you wrote or revealed.
What to look for in each sectionChoose a section and what you expect to see in it, then read short advice.

Section by section

Choose what to look for in each section

Why modern transformer MLPs often use a learned multiplicative gate: one projection proposes a token-local write while another projection controls how much of each channel reaches the residual stream.

Choose what to look for (no demo yet)01 / Intuition
What to look for

Start with the picture, metaphor, or geometric mechanism.

Choose first

Choose what you expect to see in this section of SwiGLU: Gated MLP Blocks in Transformers.

Questions to ask of the pictureChoose a part of the picture and what you expect, then read short advice.

Looking at the picture

Questions to ask of the picture

Why modern transformer MLPs often use a learned multiplicative gate: one projection proposes a token-local write while another projection controls how much of each channel reaches the residual stream.

4 of 4 sections writtenNo live demo yet
Question

Which part of the picture should you look at first?

Choose first

Pick the part of the picture you expect to explain SwiGLU: Gated MLP Blocks in Transformers.

Sources

References for this page

A reference shows where an idea comes from; it does not vouch for every step on this page.
paper · 2020GLU Variants Improve TransformerShazeer
Why it is cited

arXiv preprint. Section 2 defines SwiGLU as (Swish(xW) ⊗ xV)W₂ with three weight matrices and reduces the hidden width to two thirds of the original to keep parameters and computation constant.

Open source
paper · 2017Attention Is All You NeedVaswani et al.
Why it is cited

NeurIPS 2017. Section 3.3 defines the position-wise feed-forward block: two linear maps with a ReLU between them.

Open source
paper · 2023PaLM: Scaling Language Modeling with PathwaysChowdhery et al.
Why it is cited

JMLR 2023. Section 2 lists SwiGLU as the activation in PaLM's feed-forward blocks.

Open source
paper · 2023LLaMA: Open and Efficient Foundation Language ModelsTouvron et al.
Why it is cited

arXiv preprint. Section 2.2 replaces ReLU with SwiGLU and uses a hidden width of two thirds of 4d.

Open source
What each claim rests on1 claim, each with its sources, where it appears on this page, and its caveats.

Claims and sources

Why modern transformer MLPs often use a learned multiplicative gate: one projection proposes a token-local write while another projection controls how much of each channel reaches the residual stream.

Checked against the cited passages

The checks were made by this site, not by an independent reviewer.

A SwiGLU block multiplies a value projection by a SiLU-gated projection before the output projection, so a negative gate logit reverses a channel's sign; two-thirds hidden width matches a plain block's matrix parameter count.
What the sources support

Shazeer (2020, §2) defines SwiGLU with three weight matrices and reduces the hidden width to two thirds to keep parameters and computation constant. The sign reversal follows from SiLU(g) = g·σ(g).

Where to see it on this page
Equation 2
2dmodeldff2d_{\text{model}}d_{\text{ff}}
Caveat

v = 4, g = ±ln 3 and the 36-entry blocks are chosen for exact arithmetic. One channel before W_o, not the block's output; the count says nothing about speed or quality.

StatusChecked against the cited passages.

Section 2 of Shazeer (2020) was read for Equation 6, the three-matrix form and the two-thirds width. The ±ln 3 values, SiLU's lowest point and the 36-versus-36 counts were recomputed independently.

Explain it without the pageWrite your own explanation before asking for feedback. Your draft and any hints stay while this page is open.

Practice · SwiGLU: Gated MLP Blocks in Transformers

Try the idea in your own words

Why modern transformer MLPs often use a learned multiplicative gate: one projection proposes a token-local write while another projection controls how much of each channel reaches the residual stream.

Concept · Selected for practice

SwiGLU: Gated MLP Blocks in Transformers

What it rests on: Sources: GLU Variants Improve Transformer; Attention Is All You Need; PaLM: Scaling Language Modeling with Pathways

Context and links
Choose a task

Explain the mechanism

For SwiGLU: Gated MLP Blocks in Transformers: What is the smallest example of SwiGLU you can work through by hand, and what does it show? Explain your answer, including what changes, why, and which assumption matters.

No answer yet

A rough first thought is enough. Your draft stays when you change tasks.

Your draft stays on this page and is cleared when you leave or reload.

A little help · Explain

Open one hint at a time. These are suggestions, not your answer or a grade.

0 of 3 hints shown for this question.

    Clearing your answer keeps this help record. Outside help is your own report; this page cannot check it.

    Where am I stuck? (optional)
    Your own description, not an automatic diagnosis

    Choose one, or leave this unspecified. Select it again to clear it.

    Take your draft to a feedback conversation

    No AI feedback runs here. You can copy a prompt to use elsewhere; nothing is sent automatically. Review the text before sharing, and leave out private information.

    Write an attempt before copying a feedback prompt.

    A draft, or an AI reply to it, is not a test of what you have learned. To check that, try a different case later without help.

    Ask about this pagePick one item on this page and copy a prompt about it into an AI assistant you already use. Notes you write here stay in this browser.
    Ask about this pageClose
    ConceptSwiGLU: Gated MLP Blocks in TransformersSources: GLU Variants Improve Transformer; Attention Is All You Need; PaLM: Scaling Language Modeling with PathwaysThis code follows the math step by step. It uses one token vector, two input projections, a SiLU gate, an elementwise product, and an output projec...

    Your question

    Choose what your question is about

    Pick the idea, equation, source, code, claim, misconception or demo state you want to ask about. The prompt you copy includes it, so the answer can stay on that item.
    Next local actionNo local draft saved yet

    Open the draft below to save one note and next action in this browser.

    conceptAttention & Transformers

    SwiGLU: Gated MLP Blocks in Transformers

    Starting question

    What is the smallest example of SwiGLU you can work through by hand, and what does it show?

    SourcesCheck the 4 cited sources listed under Sources.Notes can be saved for this item
    Points of view on this item

    Fixed prompts generated from this item. They are not comments from people or independent reviews.

    Learner’s first stepAsk what would make "SwiGLU: Gated MLP Blocks in Transformers" feel predictable rather than familiar.
    Assumption

    The 4 cited sources must support this exact item, not just the surrounding topic.

    Source-checking summary

    Connect the definition to one equation, piece of code or demo before widening the discussion.

    Proposed experiment

    Change one input, then check that the same relationship holds in the mathematics, the code and the demo.

    Next action

    You can state the mechanism in your own words

    Evidence4 checks
    ObservationChecking for a saved observation
    ActionReady for one action
    PromptLearner prompt ready to copy
    Go to this item
    01ObservationChecking what this browser has saved
    02EvidenceChecking for a saved observation
    03SourcesCheck the 4 cited sources listed under Sources.
    04Next stepSave one next action
    Local action draftNo local draft saved yetOpen when you are ready to write one next action
    Local action draft

    This draft stays in this browser and is attached to this item only.

    No local draft saved.
    Evidence to inspect
    • What the 4 cited sources say about this exact item
    • The definition, its prerequisites, and a contrasting concept
    • The equation or code that makes the concept concrete
    • One demo state that shows what stays the same as the inputs change
    What would resolve this
    • You can state the mechanism in your own words
    • You can name the prerequisite that would clear up the confusion
    • You can predict how the result changes when one input changes
    Reference idconcept/concept-notebook/attention-transformers/swiglu concept:attention-transformers/swiglu