Sources. This article primarily draws on Jianlin Su’s “Transformer Upgrade Path 1: Tracing Sinusoidal Positional Encoding to Its Source”, together with the Transformer and RoFormer papers.
Thirty-second summary. A sin/cos pair can be viewed as a hand rotating around the unit circle. Every step forward by one token rotates the hand by a fixed angle, so the distance between two positions becomes a phase difference between two hands. One hand repeats periodically, so the full encoding places many hands with different speeds side by side. The most elegant feature of sinusoidal positional encoding is not that it “assigns different numbers to positions,” but that it turns translation along a sequence into rotation in a representation space.
We will repeatedly use the following notation:
| Symbol | Meaning |
|---|---|
| Absolute position indices of two tokens | |
| Relative displacement between the two positions | |
| Total dimensionality of the position vector, usually even | |
| Index of the -th sin/cos dimension pair | |
| Radians rotated by the -th hand for each token step | |
| Number of token steps required for the -th hand to complete one revolution | |
| Base of the geometric frequency schedule; the classic Transformer uses | |
| Full position vector at position |
1. What Does Self-Attention Lack Without Positional Encoding?
Consider a minimal example: “cat chases dog” and “dog chases cat” contain the same three tokens, but changing their order changes the meaning completely. If a model knows only token content and receives no position-dependent signal, it lacks a coordinate system for distinguishing these two sequences.
More precisely, remove the attention mask, position bias, and positional encoding. Let
denote a sequence of token representations, and let be any permutation matrix. Pure self-attention satisfies
This property is called permutation equivariance, not permutation invariance: if the inputs are permuted, the outputs are permuted in the same way. The model can compare token content, but it has no independent signal for “which position,” “before which token,” or “how many steps apart.”
A decoder’s causal mask already supplies the directional constraint that a token cannot attend to the future, but that constraint is not the coordinate encoding discussed here. Our question is: how can we assign a vector to every position so that the model can distinguish absolute positions and conveniently use relative distances?
The classic Transformer takes the most direct approach and adds a position vector to the token at position :
Using different values at different positions breaks the original permutation symmetry. But making every position different is only the minimum requirement.
2. Why Is Making Every Position “Different” Not Enough?
Compare the two position pairs
Their absolute coordinates are completely different, but both relative displacements equal . If distinctness were the only goal, random vectors, one-hot encodings, or a learned embedding table would all suffice. A more structured objective is to make it easy for the model to discover that these two pairs share the same relative relationship.
To isolate the geometry supplied by the position vectors, define a simplified positional interaction kernel:
We would like this kernel to depend only on relative displacement:
Then, for any global shift ,
In other words, two position pairs separated by the same distance have the same pure positional similarity.
The boundary of this argument should be explicit: is a simplified object used to isolate positional structure. It is not the complete attention score produced by additive positional encoding. The real score also mixes token content and projection matrices; Section 7 expands it in full.
Moreover, is a clear and tractable sufficient condition, not a definition that every positional encoding must obey. Our next task is to construct a simple family of vectors that satisfies it.
3. One Clock: How Do Sin/Cos Encode Relative Position?
Start with only two dimensions and place position on the unit circle:
The parameter means “how many radians the clock hand rotates for each token step.” Therefore:
- Position has angle .
- Position has angle .
- Position has angle .
- When the position index advances from to , the hand rotates by an additional . If , it completes exactly one revolution, so the one-revolution period is .
The quantity that changes is the position index ; is the fixed “rotation per step” of this hand. Thus is the total angle accumulated by the time the hand reaches position . “Completing one revolution” means that the accumulated angle increases by and the hand returns to the same point on the unit circle. If is not an integer, exactly one revolution does not correspond to an integer number of token steps.
For example, let :
| Position index | Total angle | Corresponding 2D vector |
|---|---|---|
| 0 | ||
| 1 | ||
| 2 | ||
| 3 |
The pairs and are both separated by two tokens. The angle between each pair of hands is also , so both inner products equal .
In general, the inner product between two position vectors is
The second step is simply the cosine subtraction identity. Therefore,
The right-hand side no longer depends on and separately; it depends only on the relative displacement . The unit circle establishes the correspondence
The canonical formula orders each coordinate pair as , whereas the derivation above uses . Swapping the two coordinates changes neither the inner product nor the rotation structure.
This derivation does not imply that sin/cos is the unique solution. It shows that if we seek a two-dimensional representation with fixed length that rotates uniformly with position, sin/cos provides the most natural and simplest family of coordinates.
4. Why Is One Clock Not Enough?
A clock repeats periodically. If , then
Hence,
Positions and land on the same point of the unit circle. At integer positions, this two-dimensional hand has only distinct phases. Once a sequence contains more than positions, the phases repeat, so this hand alone cannot distinguish positions separated by tokens. This does not mean that the entire Transformer can hold only tokens: the limitation belongs to this deliberately chosen single frequency, while the full encoding combines many hands with different periods.
A -dimensional sinusoidal positional encoding places clocks with different speeds side by side:
When one clock returns to its initial phase, the others will usually not reset at the same time. Multiple frequencies therefore provide two benefits:
- They reduce positional ambiguity caused by a single period.
- They resolve displacement at several distance scales.
Larger values rotate quickly, so adjacent tokens create noticeable phase changes and local displacement is easier to resolve. Smaller values rotate slowly and retain gradually changing structure over a longer interval.
Multiple frequencies are not an unconditional guarantee against collisions at arbitrary lengths or under finite numerical precision. Distance information is carried by the joint phase pattern across all dimensions, and the model must still learn how to use that pattern during training.
5. Why Space the Frequencies Geometrically?
The general form of a geometric frequency schedule is
Here is the base that controls the span of the full frequency set. The classic Transformer chooses , giving
The corresponding periods are
For , we can see directly how the clocks become progressively slower:
| Frequency pair | Period (approx.) | Intuition | |
|---|---|---|---|
| 0 | Rapid change, local emphasis | ||
| 1 | Shorter scale | ||
| 2 | Longer scale | ||
| 3 | Slow change, global emphasis |
This is the main value of geometric spacing: a finite number of dimensions can cover logarithmic scales relatively evenly. With an arithmetic frequency schedule, many dimensions may cluster in a narrow absolute range. Geometric spacing instead acts like a multiscale ruler that moves from local detail toward global structure.
Two questions should remain separate:
- Why sin/cos? Rotations and the angle-subtraction identity make relative displacement appear naturally in positional interactions.
- Why the base 10000? It is an engineering choice governing frequency range and resolution, not a constant uniquely forced by the preceding identity.
The Transformer paper also gives a direct motivation: for any fixed offset , can be represented as a linear function of . A fixed function can also generate encodings at positions longer than those seen during training. The latter means only that “the formula can be evaluated”; it does not automatically imply reliable length extrapolation by the model.
6. Distance Behavior of the Multi-Frequency Position Kernel: Oscillatory Decorrelation
First, write the two complete position vectors side by side:
Their inner product multiplies sin by sin and cos by cos within each frequency pair, then sums over all pairs:
The last step applies the cosine subtraction identity to every dimension pair. The result says that the full-vector inner product is the sum of all clocks’ “votes” on their phase differences.
Let and divide by the number of clocks, , to obtain the normalized position kernel
The normalization gives
because when , every clock is perfectly aligned and every cosine term equals .
When is small, the phases at different frequencies remain relatively synchronized, so many cosine terms reinforce one another. As distance grows, the phases spread apart and positive and negative terms begin to cancel. Over commonly used distance ranges, this often produces oscillatory decorrelation: distant positions tend to have lower average similarity, but the curve repeatedly rises and falls along the way.
Figure 1 makes this process explicit by drawing the three hands at position and position in every row. Because the position kernel depends only on , we can set as the reference without loss of generality; all three hands on the left then point to 12 o’clock. On the right, the angle between each matching color and the reference direction is determined by . The cosine of that angle is the inner-product contribution of the corresponding frequency pair.
An important point is that every two-dimensional block remains on the unit circle, so the length of the full position vector is constant:
As distance grows, the positional encoding itself does not “shrink to zero.” What may become smaller is the normalized inner product between two position vectors. In other words, the average alignment between their directions changes, not their vector lengths.
This trend should not be mistaken for a strict theorem:
With finitely many frequencies, is a finite cosine sum. It is not guaranteed to decrease monotonically with distance, nor can we generally claim that it converges strictly to zero as .
To study the overall trend of many discrete frequencies, treat the frequency index as a continuous variable on . From , we obtain , and the discrete average can be approximated by
Here is the frequency-schedule base defined in Section 5, with the classic value . The integral can be read as an average over a continuous distribution of clocks with different speeds. As grows, the integrand oscillates between positive and negative values more rapidly as a function of , so its average is more easily cancelled.
The purpose of this figure is to turn “positive and negative contributions from many frequencies gradually cancel” into a visible continuous curve. It is not evidence that a particular finite-dimensional positional encoding must eventually converge to zero. The schedule is not the only possible choice either; it is an engineering compromise among local resolution, long-distance variation, and frequency coverage.
Section takeaway: what is often called “long-range decay” is more precisely oscillatory decorrelation. As grows, phases at different frequencies are more likely to spread out, and the normalized inner product tends to be smaller over commonly used distance ranges. A finite cosine sum, however, is neither guaranteed to decrease monotonically nor guaranteed to converge to zero. What changes is the average similarity between two positions, not the length of either position vector.
7. How Does Real Attention Differ from This Simplified Model?
So far, we have studied . Now place it back inside real attention.
Ignoring the scaling constant and bias, suppose the Query and Key are computed from tokens after their position vectors have been added:
Let
The attention score then expands exactly as
Therefore, using sinusoidal positional encoding does not automatically make the real score a function of alone. The earlier identity
first isolates the pure positional interaction, then examines the simplest case to reveal the structure supplied by the encoding itself.
Directly expanding the attention score also exposes the boundary clearly. When is a general matrix, the pure positional term need not depend only on ; the other three terms additionally mix content and position. The kernel in Section 6 is therefore a simplified tool for analyzing encoding geometry, not an equivalent description of complete attention behavior.
A more accurate conclusion is:
Sinusoidal positional encoding gives the model a relative phase structure that it can exploit, but it does not dictate how a trained model must use that structure.
8. What Does This Explanation Establish—and What Does It Not?
It establishes or directly demonstrates that:
- A sin/cos pair forms two-dimensional rotation coordinates.
- The inner product within one frequency pair is exactly .
- A fixed positional offset corresponds to a rotation independent of the absolute position.
- Multiple frequencies extend this structure across several distance scales.
- Geometric frequencies produce an oscillatory decorrelation trend under the continuous approximation.
It does not establish that:
- Sinusoidal PE is the only positional encoding.
- Sinusoidal PE is the optimal positional encoding.
- The base is theoretically mandatory or optimal.
- A finite-dimensional inner product decreases strictly monotonically with distance.
- The real attention score of additive sinusoidal PE depends only on relative position.
- Being able to evaluate encodings at longer positions guarantees reliable length extrapolation.
The real value of this analysis is not to prove that “Google’s formula is irreplaceable.” It demonstrates a reusable way of thinking: first specify the structure we want a positional representation to have, then find a simple explicit construction, and finally inspect the additional properties and assumptions that come with it.
9. Extension: Why Does This Idea Lead Naturally to RoPE?
The principle that “relative displacement becomes phase difference” can enter attention through a different mechanism. This is the natural starting point for understanding RoPE.
Traditional sinusoidal PE adds a position vector to a token representation:
RoPE instead applies the position-dependent rotation directly to content-derived Queries and Keys. For one two-dimensional frequency block, let denote the rotation at position :
Their inner product is
Full RoPE uses two-dimensional rotation blocks with different frequencies across dimension pairs. Position enters the Q/K inner product explicitly through , while the complete score still depends on the content vectors and .
The two methods therefore share the mathematical core that relative displacement corresponds to phase difference, but their mechanisms differ:
- Sinusoidal PE adds an absolute position vector at the input, after which position and content pass through the projections together.
- RoPE rotates content directly in Q/K space, making relative rotation explicit in the score.
10. Quick Reference
| Question | Shortest answer |
|---|---|
| Why do we need positional encoding? | Position-free self-attention is permutation equivariant and lacks an independent coordinate for sequence order. |
| Why is making every position different not enough? | We also want the model to recognize easily that and share the same relative relationship. |
| Why use sin/cos? | Unit-circle rotations make the inner product depend only on position difference through the angle-subtraction identity. |
| Why pair every two dimensions? | Two dimensions are exactly enough to represent the two coordinates of one rotating hand. |
| Why use many frequencies? | A single frequency repeats periodically; multiple frequencies reduce ambiguity and provide several scales. |
| Why use geometric frequencies? | They cover logarithmic scales relatively evenly with a finite number of dimensions. |
| Is oscillatory decorrelation strictly monotonic? | No. A finite-dimensional inner product is an oscillatory cosine sum and only shows a decorrelation trend over commonly used ranges. |
| Is theoretically mandatory? | No. It is an engineering choice governing frequency range and resolution. |
| Does real attention depend only on relative distance? | No. With additive PE, the score contains four kinds of content-position interaction. |
| How is this related to RoPE? | Both use phase differences; RoPE places the rotation directly inside the Q/K inner product. |
The single most important sentence to remember is:
Sinusoidal positional encoding does not use trigonometric functions merely to “number positions.” It uses many clocks with different speeds to represent relative displacement as a multiscale phase difference.
References and Image Sources
- Vaswani, A. et al. (2017). Attention Is All You Need, §3.5.
- Jianlin Su (2021). “Transformer Upgrade Path 1: Tracing Sinusoidal Positional Encoding to Its Source”. The motivating questions, relative-position inner product, and continuous-frequency approximation in this article begin from that post. The pedagogical sequence has been reorganized, and a direct attention-score decomposition replaces the Taylor expansion.
- Su, J. et al. (2021). RoFormer: Enhanced Transformer with Rotary Position Embedding.
- Figure 1 is an author-created mechanism diagram. Figure 2 is taken from Jianlin Su’s post above and reproduced unaltered under the site’s CC BY-NC-ND 2.5 CN notice; open the image to view the source file.