Positional Encoding

How transformers represent order: sinusoidal encodings, learned embeddings, and relative-position methods like RoPE.

Introductory · Undergraduate mathematics · about 12 minutes

Reading map and next steps

Intuition

The idea in plain words, before the symbols.

Self-attention compares tokens by content (queries vs keys). But content alone does not tell you where a token is in the sequence.

If you shuffle the tokens in a sequence and keep their embeddings the same, plain attention has no built-in way to notice the shuffle. In other words, attention is permutation-equivariant unless we inject order information.

Positional encodings are the mechanism that turns "a bag of token vectors" into "an ordered sequence". There are many variants:

  • Absolute position (sinusoidal or learned): add a position vector to each token.
  • Relative position (RoPE, ALiBi, etc.): make attention scores depend on relative offsets.

RoPE is one modern relative-position method, but it is easier to appreciate once you understand the basic goal: the model needs a stable coordinate system for time/position.

Study prompts for this section: Intuition
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Demo

Predict what a shared shift does, check it on one pair, then move one position and find a collision.

PredictAnswer each problem before you open the figure.Each solution appears after you check or ask for a hint.

Work one sine–cosine pair by hand: predict what a shared shift does, compute two positions and their shifted copies, then move one position from 11 to 55 and explain why nothing changes. The figure lets you pick any two positions and compare the single pair with a full eight-dimensional encoding.

One sine–cosine pair, by hand

Loading the interactive demo…

Observation guide

Optional and ungraded. Opening a guide saves nothing.

Choose what to inspect in Positional Encoding, then open its guide.

Study prompts for this section: Demo
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Mathematics

Definitions, assumptions and the derivation.

The classic sinusoidal encoding (Vaswani et al., 2017, §3.5) defines, for model dimension dd and position mm, with base 10,000:

PE(m,2i)=sin⁡(m ωi),PE(m,2i+1)=cos⁡(m ωi),\mathrm{PE}(m,2i) = \sin\left(m\,\omega_i\right),\qquad \mathrm{PE}(m,2i+1) = \cos\left(m\,\omega_i\right),

with frequencies

ωi=base−2i/d,i∈{0,…,d2−1}.\omega_i = \mathrm{base}^{-2i/d},\qquad i\in\{0,\dots,\tfrac d2-1\}.

Two useful facts.

Shift structure. The dot product between the encodings of positions mm and nn depends only on their offset:

PE(m)⋅PE(n)=∑icos⁡(ωi (m−n)).\mathrm{PE}(m)\cdot \mathrm{PE}(n) = \sum_i \cos\left(\omega_i\,(m-n)\right).

Each term comes from one sine–cosine pair and the cosine difference identity, sin⁡(mω)sin⁡(nω)+cos⁡(mω)cos⁡(nω)=cos⁡((m−n)ω)\sin(m\omega)\sin(n\omega)+\cos(m\omega)\cos(n\omega)=\cos\bigl((m-n)\omega\bigr), so moving both positions by the same amount leaves the dot product unchanged. This is a statement about the position vectors alone. Once they are added to token embeddings and passed through learned query and key projections, an attention score also depends on the tokens.

Additive absolute encoding. The simplest way to use PE is to add it to token embeddings:

hm=em+PE(m).h_m = e_m + \mathrm{PE}(m).

Relative-position methods (like RoPE) instead change the attention computation so that the score between a query at position mm and a key at position nn depends on the offset n−mn-m directly.

One pair, worked by hand

Take a single pair s(m)=(sin⁡(mω),cos⁡(mω))s(m)=\bigl(\sin(m\omega),\cos(m\omega)\bigr) and set ω=π/2\omega=\pi/2, a quarter turn per position. This frequency is chosen so the arithmetic is exact; it is not one of the published frequencies. Every s(m)s(m) is then one of four points on the unit circle:

s(0)=(0,1),s(1)=(1,0),s(2)=(0,−1),s(3)=(−1,0).s(0)=(0,1),\qquad s(1)=(1,0),\qquad s(2)=(0,-1),\qquad s(3)=(-1,0).

Positions 00 and 11 give s(0)⋅s(1)=0s(0)\cdot s(1)=0. Moving both two places later gives s(2)⋅s(3)=0s(2)\cdot s(3)=0 again: the offset is still −1-1. Now keep m=0m=0 and move nn to 55. Four quarter turns make a full turn, so s(5)=s(1)s(5)=s(1) and s(0)⋅s(5)=0s(0)\cdot s(5)=0. This pair cannot tell distance 11 from distance 55: one frequency repeats, so its dot product is not a measure that falls as positions move apart.

The collision belongs to the single pair. The full encoding stacks pairs that turn at different speeds, and the slow ones have not come round yet: with d=8d=8, PE(0)⋅PE(1)≈3.535\mathrm{PE}(0)\cdot\mathrm{PE}(1)\approx 3.535 but PE(0)⋅PE(5)≈3.160\mathrm{PE}(0)\cdot\mathrm{PE}(5)\approx 3.160. The demo works these numbers one step at a time.

Study prompts for this section: Mathematics
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

Code

The same calculation as runnable code, in the notation of the derivation.

import numpy as np

def sinusoidal_pe(T, d, base=10000.0):
    assert d % 2 == 0
    pos = np.arange(T)[:, None]
    i = np.arange(d // 2)[None, :]
    w = base ** (-2 * i / d)
    angles = pos * w
    pe = np.zeros((T, d))
    pe[:, 0::2] = np.sin(angles)
    pe[:, 1::2] = np.cos(angles)
    return pe

pe = sinusoidal_pe(T=32, d=16)
dot = lambda m, n: float(pe[m] @ pe[n])

for delta in [0, 1, 5, 10]:
    a = dot(0, delta)
    b = dot(7, 7 + delta)  # same relative offset
    print("delta =", delta, "dot =", round(a, 3), "dot (shifted) =", round(b, 3))
Study prompts for this section: Code
Study prompts

Each button copies a prompt about this section to your clipboard, for an AI assistant you already use. Nothing is sent by this site.

In focus

Concept: Positional Encoding

What is the smallest example of Positional Encoding you can work through by hand, and what does it show?

BeforeScaled Dot-Product Attention & Transformer LayersNextRotary Position Embeddings (RoPE)
DetailsAttention & Transformers

After the first pass

Go deeper

Optional material for a second reading. Sources are always shown; the other panels open from their headings or from these links, and closing one keeps what you wrote or revealed.
What to look for in each sectionChoose a section and what you expect to see in it, then read short advice.

Section by section

Choose what to look for in each section

How transformers represent order: sinusoidal encodings, learned embeddings, and relative-position methods like RoPE.

Choose what to look for (no demo yet)01 / Intuition
What to look for

Start with the picture, metaphor, or geometric mechanism.

Choose first

Choose what you expect to see in this section of Positional Encoding.

Questions to ask of the pictureChoose a part of the picture and what you expect, then read short advice.

Looking at the picture

Questions to ask of the picture

How transformers represent order: sinusoidal encodings, learned embeddings, and relative-position methods like RoPE.

4 of 4 sections writtenNo live demo yet
Question

Which part of the picture should you look at first?

Choose first

Pick the part of the picture you expect to explain Positional Encoding.

Sources

References for this page

A reference shows where an idea comes from; it does not vouch for every step on this page.
paper · 2017Attention Is All You NeedVaswani et al.
Why it is cited

NeurIPS 2017. Section 3.5 adds sine and cosine encodings of geometrically spaced wavelengths to the input embeddings, chosen so that a fixed offset acts on them as a linear function.

Open source
book · 2024Deep Learning: Foundations and ConceptsBishop and Bishop
Why it is cited

Springer. Chapter 12 (Transformers) treats positional encoding, including the sinusoidal form, as a textbook topic.

Open source
What each claim rests on1 claim, each with its sources, where it appears on this page, and its caveats.

Claims and sources

How transformers represent order: sinusoidal encodings, learned embeddings, and relative-position methods like RoPE.

Checked against the cited passages

The checks were made by this site, not by an independent reviewer.

The original Transformer adds a sinusoidal position vector to each token embedding, and the dot product of two such position vectors depends only on the offset between the positions.
What the sources support

Vaswani et al. (§3.5) add sine and cosine encodings to the input embeddings and chose them so a fixed offset acts linearly. The offset-only dot product follows from the cosine difference identity, derived on...

Where to see it on this page
Equation 1
PE(m,2i)=sin⁡(m ωi),PE(m,2i+1)=cos⁡(m ωi),\mathrm{PE}(m,2i) = \sin\left(m\,\omega_i\right),\qquad \mathrm{PE}(m,2i+1) = \cos\left(m\,\omega_i\right),
Equation 3
PE(m)⋅PE(n)=∑icos⁡(ωi (m−n)).\mathrm{PE}(m)\cdot \mathrm{PE}(n) = \sum_i \cos\left(\omega_i\,(m-n)\right).
Caveat

The quarter-turn pair and its positions are chosen for exact arithmetic. Its collision holds for that one pair, not the full encoding, and says nothing about attention scores.

StatusChecked against the cited passages.

Section 3.5 of Vaswani et al. was read for the addition to embeddings, the sinusoidal formula and the fixed-offset motivation. The pair identity was derived again, and the positions 0, 1, 2, 3 and 5 and the eight-dimensional comparison were recomputed independently.

Explain it without the pageWrite your own explanation before asking for feedback. Your draft and any hints stay while this page is open.

Practice · Positional Encoding

Try the idea in your own words

How transformers represent order: sinusoidal encodings, learned embeddings, and relative-position methods like RoPE.

Concept · Selected for practice

Positional Encoding

What it rests on: Sources: Attention Is All You Need; Deep Learning: Foundations and Concepts

Context and links
Choose a task

Explain the mechanism

For Positional Encoding: What is the smallest example of Positional Encoding you can work through by hand, and what does it show? Explain your answer, including what changes, why, and which assumption matters.

No answer yet

A rough first thought is enough. Your draft stays when you change tasks.

Your draft stays on this page and is cleared when you leave or reload.

A little help · Explain

Open one hint at a time. These are suggestions, not your answer or a grade.

0 of 3 hints shown for this question.

    Clearing your answer keeps this help record. Outside help is your own report; this page cannot check it.

    Where am I stuck? (optional)
    Your own description, not an automatic diagnosis

    Choose one, or leave this unspecified. Select it again to clear it.

    Take your draft to a feedback conversation

    No AI feedback runs here. You can copy a prompt to use elsewhere; nothing is sent automatically. Review the text before sharing, and leave out private information.

    Write an attempt before copying a feedback prompt.

    A draft, or an AI reply to it, is not a test of what you have learned. To check that, try a different case later without help.

    Ask about this pagePick one item on this page and copy a prompt about it into an AI assistant you already use. Notes you write here stay in this browser.
    Ask about this pageClose
    ConceptPositional EncodingSources: Attention Is All You Need; Deep Learning: Foundations and Conceptsassert d % 2 == 0

    Your question

    Choose what your question is about

    Pick the idea, equation, source, code, claim, misconception or demo state you want to ask about. The prompt you copy includes it, so the answer can stay on that item.
    Next local actionNo local draft saved yet

    Open the draft below to save one note and next action in this browser.

    conceptAttention & Transformers

    Positional Encoding

    Starting question

    What is the smallest example of Positional Encoding you can work through by hand, and what does it show?

    SourcesCheck the 2 cited sources listed under Sources.Notes can be saved for this item
    Points of view on this item

    Fixed prompts generated from this item. They are not comments from people or independent reviews.

    Learner’s first stepAsk what would make "Positional Encoding" feel predictable rather than familiar.
    Assumption

    The 2 cited sources must support this exact item, not just the surrounding topic.

    Source-checking summary

    Connect the definition to one equation, piece of code or demo before widening the discussion.

    Proposed experiment

    Change one input, then check that the same relationship holds in the mathematics, the code and the demo.

    Next action

    You can state the mechanism in your own words

    Evidence4 checks
    ObservationChecking for a saved observation
    ActionReady for one action
    PromptLearner prompt ready to copy
    Go to this item
    01ObservationChecking what this browser has saved
    02EvidenceChecking for a saved observation
    03SourcesCheck the 2 cited sources listed under Sources.
    04Next stepSave one next action
    Local action draftNo local draft saved yetOpen when you are ready to write one next action
    Local action draft

    This draft stays in this browser and is attached to this item only.

    No local draft saved.
    Evidence to inspect
    • What the 2 cited sources say about this exact item
    • The definition, its prerequisites, and a contrasting concept
    • The equation or code that makes the concept concrete
    • One demo state that shows what stays the same as the inputs change
    What would resolve this
    • You can state the mechanism in your own words
    • You can name the prerequisite that would clear up the confusion
    • You can predict how the result changes when one input changes
    Reference idconcept/concept-notebook/attention-transformers/positional-encoding concept:attention-transformers/positional-encoding