🌏 中文版
This guide follows the 3/27 materials of NTU Machine Learning 2026 Spring by Hung-yi Lee. It is part 9 of the Reading NTU Hung-yi Lee Machine Learning 2026 Spring series. The two previous lectures asked why generation is slow: Flash Attention deals with memory traffic, KV Cache deals with repeated computation, and HW3 measured both on a GPU. This lecture asks a different question. Agents routinely feed models hundreds of thousands of tokens. How does the model know where each token sits? And why does it break on inputs longer than anything it saw in training?
The course schedule titles this row "inside the model: how models handle very long inputs". The official materials are the slides pos.pdf (64 pages, also as pptx) and the video How does a Transformer know the order of input tokens? Absolute, Relative, RoPE, and no Positional Embedding (in Mandarin). Access level is A3: slides and recording are public, and this lecture has no quiz or leaderboard.
The problem: "you hit me" vs. "I hit you"
Slides 2–3 make the point quickly. Feed tokens A B C D into self-attention and look at D. If you reorder the first three as C B A, D's weighted sum comes out exactly the same. Yet "you hit me" and "I hit you" mean opposite things. The Transformer needs position information from somewhere.
Slide 4 splits the lecture into five parts, and this guide follows the same order:
- Absolute Positional Embedding
- Relative Positional Embedding
- RoPE
- Train short, test long
- No Positional Embedding?!
Absolute positions: sinusoidal and its blind spot
The first approach gives each position a vector p and adds it to the token vector x (slide 6). The sinusoidal version from Attention Is All You Need builds p from sines and cosines at different frequencies. The slides use a clock: high-frequency dimensions are the second hand and separate neighbouring positions; low-frequency dimensions are the hour hand and separate distant ones (slides 9–11).
Slides 13–14 state what we actually need: relative position matters. Whether "the cat ate the fish" appears at the start of the text or at token 101, the link between "cat" and "fish" should be the same. If a long stretch separates them, the link should be much weaker.
Slides 16–22 expand the attention score after adding positions. The dot product of x+p breaks into terms that depend only on content, terms that mix content and position, and terms that depend only on position. The trouble with sinusoidal embeddings is that none of these terms clearly tracks relative position. Slide 22 writes down the goal: the position term should depend only on relative position.
Expand: the sinusoidal formula
From the original paper, for position pos and dimension pair i:
- PE(pos, 2i) = sin(pos / 10000^(2i/d))
- PE(pos, 2i+1) = cos(pos / 10000^(2i/d))
Small i means high frequency (the second hand) and large i means low frequency (the hour hand). The numbers 6.3, 628.3 and 54410.1 on slide 10 are how many positions it takes different dimensions to complete one full turn.
Relative positions: ALiBi and T5
If relative position is what matters, skip p and modify the attention score directly.
- ALiBi (slide 24) subtracts b×(m−n) from the score, where m−n is the distance between two tokens. The slide's comment: "the farther apart, the smaller the attention. That's it!" b is set by hand, with a different value for each attention head.
- T5 (slide 26) also adds a distance-based bias to the score, but that bias is a trainable parameter.
RoPE: position as rotation
Slide 27 names RoPE as the method used by Llama, Qwen and Gemma, and stresses two advantages: it leaves the attention computation unchanged, and it works with the KV cache.
The intuition: treat every two dimensions of the query and key as an arrow on a plane, and rotate the arrow by nθ for a token at position n. When two tokens take a dot product, only the angle difference (m−n)θ survives. So "cat" and "fish" get the same score at positions 1 and 3 as at positions 101 and 103 (the example on slide 32; the rotation diagrams are on slides 29–31).
Expand: how the rotation angles are set
Each dimension pair gets a base angle θ_i = 1 / 10000^(2i/d), for i = 0, 1, …, d/2−1 (slide 49). Small-i dimensions rotate quickly and large-i dimensions rotate slowly. It is the same second-hand / hour-hand structure as the sinusoidal embedding, with rotation in place of addition.
Does RoPE decay with distance?
A common claim is that RoPE makes attention decay with distance. Slides 39–40 cite Round and Round We Go! to push back: RoPE does not simply decay. The paper studies Gemma 7B and argues that decay is unlikely to be why RoPE works. The slides include a sample Colab so you can plot it yourself.
Slide 40 adds that this "is not necessarily a bad thing": a query can learn to line up with keys at a specific distance. The example is "my cat" and "his dog" (in Chinese, "我 的 貓" and "他 的 狗"), where the noun consistently attends to the word two positions back.
Train short, test long
Slide 42 sets the scene: training sequences are all shorter than 1M tokens, but at test time the input is longer than 1M. Slide 45 explains why RoPE fails. In training, a key rotates at most Nθ. At test time it rotates 2Nθ, pointing in a direction the model has never seen.
Slide 41 lists Aman Arora's How LLMs Scaled from 512 to 2M Context as the reference for this section. The fixes, in order:
| Method | What it does (per the slides) | Cost |
|---|---|---|
| Position Interpolation (slides 46–47) | Squeezes 1, 2, 3, 4 into 0.5, 1, 1.5, 2, scaling every position back into the training range; kaiokendev did the same around the same time | The slides note it "still requires fine-tuning the model" |
| Frequency-based (slides 48–49) | High-frequency dimensions have already made many full turns within the training length, so going past N is harmless; leave them alone. Low-frequency dimensions never finished one turn in training, so only they need compressing | You must decide which dimensions count as high frequency |
| NTK-Aware Scaling (slides 50–51) | Multiplies dimension pair i by f(L, i) = (1/L)^(2i/(d−2)): the highest-frequency θ_0 is untouched, the lowest-frequency dimension is scaled by 1/L, with a smooth transition in between. A community chart shows LLaMA 7B running well past 2048 tokens without fine-tuning | The source is a Reddit post, not a paper |
| YaRN (slide 52) | "Yet another RoPE extensioN method", the paper version of the frequency-based idea | Still a position-scaling approach |
| Dynamic Scaling (slides 53–54) | A fixed compression ratio makes "long sequences work but short sequences get worse"; instead, compress only once a sequence exceeds the training length | The ratio changes with length |
| LongRoPE (slide 56) | Frequency-based plus dynamic, with per-dimension scaling factors found by evolutionary search | Requires a search |
No Positional Embedding?!
The last section overturns the opening premise. Slide 59 shows two decoders fed "cat eats fish?" and "fish eats cat?". Each token can only see itself and earlier tokens. The final token sees the same set of words either way, but the middle tokens see different prefixes, and stacking layers makes the final output differ. The slide concludes: "so there's no need to add a Positional Embedding!"
Two papers back this up:
- NoPE (slide 61) compares length generalization in decoder-only Transformers with APE, T5 relative bias, ALiBi, Rotary, and no positional encoding at all. NoPE did best on their reasoning and math tasks.
- DroPE (slides 62–63; the YaRN chart on slide 52 also comes from this paper) argues that positional embeddings help training converge but are also what stops models from generalizing to longer sequences. Dropping them after pretraining, followed by a short recalibration, gives zero-shot context extension.
My reading: the role of the positional embedding shifts from "the model can't understand order without it" to "training wheels". That also explains why all the scaling tricks above hit limits. They are still wrestling with an explicit position signal.
Back to the model: what this lecture answered
Back to the opening question. Most modern LLMs encode position with RoPE because it puts relative distance into the dot product and works with the KV cache. Stretching context from a few thousand tokens to a million does not come from retraining. It comes from rescaling RoPE's angles at inference time, with little or no fine-tuning. When a model card mentions "rope scaling", "YaRN" or "128K context", it is talking about the second half of this lecture.
Try this: open the config.json of an open-weight model you use and find rope_theta and rope_scaling. If rope_scaling is set, match it against the table above and work out how many times longer the advertised context is than the training length.
What this guide can and cannot confirm
Confirmed: the structure of the 64 slides, each slide's title, the labels on the figures and the cited sources. Every paper title and abstract was checked on arXiv, and the video title and uploader were checked via YouTube oEmbed.
Not confirmed: I did not transcribe the video, so examples and numbers the lecturer only said aloud are not included. NTK-Aware and Dynamic Scaling come from Reddit posts. The slides only reuse their charts, and I did not independently verify those numbers.
Further reading
- The Stanford CME295 guide on Transformer tricks, which also walks from positional encoding to RoPE
- The CMU 11-785 guide on Transformer architectures
Series navigation: Series overview | Previous: HW3: LLM Fast Inference | Next: HW4: Training a Transformer
References
- NTU Machine Learning 2026 Spring course page (Hung-yi Lee) (in Mandarin)
- pos.pdf (Positional Embedding slides)
- Video: How does a Transformer know the order of input tokens? Absolute, Relative, RoPE, and no Positional Embedding (in Mandarin)
- Attention Is All You Need (arXiv 1706.03762)
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation (ALiBi, arXiv 2108.12409)
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5, arXiv 1910.10683)
- RoFormer: Enhanced Transformer with Rotary Position Embedding (arXiv 2104.09864)
- Round and Round We Go! What makes Rotary Positional Encodings useful? (arXiv 2410.06205)
- Extending Context Window of Large Language Models via Positional Interpolation (arXiv 2306.15595)
- kaiokendev: Extending Context is Hard
- YaRN: Efficient Context Window Extension of Large Language Models (arXiv 2309.00071)
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens (arXiv 2402.13753)
- The Impact of Positional Encoding on Length Generalization in Transformers (NoPE, arXiv 2305.19466)
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings (DroPE, arXiv 2512.12167)
- Aman Arora: How LLMs Scaled from 512 to 2M Context: A Technical Deep Dive
- NTK-Aware Scaled RoPE (r/LocalLLaMA post)
- Dynamically Scaled RoPE (r/LocalLLaMA post)
Loading...