A guide to the BERT and its Family unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). It starts from "I record the record": Word2Vec and GloVe give both records the same vector. ELMo fixes this with a bidirectional LSTM language model. With Transformers, pretraining splits into three roads: encoders (BERT: MLM plus NSP, strong at understanding, weak at generation), encoder-decoders (T5's span corruption, BART's five noise types), and decoders (GPT: pure next-token prediction). The last part covers GPT-3's in-context learning and scaling laws, and why decoders became today's dominant backbone.
A guide to the decoding and evaluation unit in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The first half covers how to pick a word once the model outputs a probability distribution: greedy decoding can't take back a mistake, beam search keeps several candidates but favors short outputs, and top-k / top-p trade determinism for diversity (missing from the slides; the professor covers it verbally in class). The second half covers scoring generated text: BLEU's modified precision and brevity penalty, ROUGE-N and ROUGE-L, perplexity, and what GLUE, SQuAD 2.0, MTEB, and MMLU each measure.
A guide to the GPT-2 / T5 tutorial in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). One task, LCSTS Chinese summarization, is solved twice. Decoder-only GPT-2 is written in native PyTorch: you join article and summary into one sequence, switch to left padding, and set padding labels to −100. Encoder-decoder mT5 uses Seq2SeqTrainer: no left padding needed, and DataCollatorForSeq2Seq handles the −100 for you. Both segment with jieba and score word-level ROUGE.
Hung-Yu Kao's Fall 2025 W8 slides walk from GPT-1 to GPT-3, explain how the Sparse Transformer behind GPT-3 cuts attention cost, and then use InstructGPT to show the gap between continuing text and following instructions. The maximum likelihood objective can't tell a fabricated fact from a slightly wrong synonym, so three extra stages are added: SFT learns how humans write, a reward model learns how humans grade, and PPO optimizes against that grade while a KL penalty keeps the model from drifting too far. The lecture closes with Llama-2: separate safety and helpfulness reward models, context distillation, and GQA for faster inference.
A guide to the Hugging Face BERT tutorial and HW3 in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The tutorial walks through binary IMDb sentiment classification: AutoTokenizer, the input_ids / token_type_ids / attention_mask fields, AutoModelForSequenceClassification, and Trainer. HW3 applies the same tools to SemEval 2014 Task 1: one bert-base-uncased with two heads, one regressing a 1–5 relatedness score and one classifying entailment into three classes. You add the two losses and write the training loop yourself, because Trainer is not allowed.
HW1 tests word vectors on the 19,544 Google Analogy questions (8,869 semantic, 10,675 syntactic). You first answer them with pretrained glove-wiki-gigaword-100 loaded through Gensim, then train your own Word2Vec on a 20% sample of a pre-cleaned Wikipedia dump, and plot t-SNE for the family subcategory both times. Seven TODOs are worth 55%, the report 45%. Fall 2026 keeps the same TODOs but asks for an .ipynb with outputs.
Week 1 of Hung-Yu Kao's NLP course at NTHU is a 91-page deck, W1_NLP_brief. It opens with 'Watch for kids', five readings of the telescope sentence, and a Chinese tongue-twister about eleven uncles to show why language is hard. Then, from an information-retrieval angle, it builds one pipeline: inverted index, tokenization, stemming, TF-IDF, BM25. The second half hits that pipeline's dead ends (synonyms, polysemy, vocabulary mismatch), moves to SVD-based LSA, and closes with a preview of dense vectors through Skip-gram, GloVe, and FastText.
Hung-Yu Kao's Natural Language Processing at National Tsing Hua University is a graduate-level flagship course in the TAICA alliance. The syllabus caps it at 1,200 students, it is taught in Mandarin, and it runs from TF-IDF and word vectors to RLHF, PEFT, and RAG. For Fall 2025, the slides, 32 class recordings, and 4 assignments with starter notebooks are all on GitHub, which rates A3. Solutions, grading, and the term-project spec are not public. Fall 2026 is in progress and only goes up to W3, so it rates A2. Grading changed to 75% assignments plus a 25% in-person midterm, and a Reasoning/Agent unit was added.
This 34-slide TA session answers a practical question. Pasting data into the ChatGPT web page one row at a time is slow and hits hourly limits, so research and homework should use the API. The notebook runs one SemEval 2014 entailment example through Gemini, Claude, and OpenAI in turn: prompts live in prompts.yaml, output is forced into JSON, then few-shot and token counting. The material is from 2024. The slide cover says 2024/11/21, and the notebook uses gemini-1.5-pro, gpt-4o, and claude-3-5-sonnet-20241022. That Claude model was retired on 2025-10-28, and Google's old Gemini SDK reached end of support on 2025-11-30.
Hung-Yu Kao's Fall 2025 PEFT slides open with a budget: full fine-tuning of Llama 2-7B in 16-bit needs about 56GB of GPU memory, while training only 0.2M parameters brings it down to about 17GB, because gradients and optimizer states nearly vanish. Intrinsic dimensionality then explains why tuning a small slice is enough: the longer a model is pretrained and the larger it is, the fewer effective dimensions fine-tuning needs. Methods fall into additive (Adapters, Prompt Tuning), selective (BitFit), reparametrization (LoRA), and hybrid (MAM Adapters, S4). The second half runs from GPT-2's task descriptions and verbalizers to the trade-offs between prefix tuning and soft prompt tuning.
HW2 treats expressions like "14*(43+20)=882" as character sequences and asks a two-layer LSTM to generate the answer one character at a time after it sees "=". The training set has 2,369,250 rows and the eval set 263,250, with every number in 0–49. Six TODOs run from building a vocabulary and batching with loss only after "=" to a generator, teacher-forced training, and exact-match evaluation. The W4 PyTorch TA session is the toolbox for it.
Each of the two RAG TA sessions builds one version. The first installs Ollama on Colab to run llama3.2:1b and wires up a minimal RAG with LangChain's Chroma, MMR, and retrieval chain. The second uses LangChain only for data prep and writes the rest by hand: chunking, text and vector stores, hybrid BM25 + cosine retrieval merged with RRF, then generation with Llama-3.2-1B-Instruct. HW4 applies the first session's skeleton to 150 cat facts and 150 GPT-5-generated QA pairs. The generator must be Llama3.2-1b and the embedding model jina-embeddings-v2-base-en, and you report recall@1, recall@5, and exact match. Code is 45% of the grade and the report 55%; the report analyzes how prompts, data format, document order, and counterfactual information change the results.
The second half of W11_RAG.pdf starts at the "From Retrievers to QA" slide and turns a retriever plus a reader into a full QA system. It begins with ORQA and REALM (2019–2020), where BERT is the reader, then covers the first paper named RAG and REPLUG, which keeps the LLM frozen. A single table then sorts seven recent fixes into three groups: rewrite the query (Query Rewriting, HyDE), make the generator robust to noise (RetRobust, RAFT, RAAT), and decide when to retrieve (FLARE, Self-RAG). It ends with noise types, four abilities an LLM needs inside RAG, and generative retrieval (GR) with reliable response generation (RRG). On the recording side, the W11 Thursday lecture stops at the RAG paper; I could not find a recording that covers the later slides.
An LLM will confidently answer that Oppenheimer was born in 1967 (the year he died). RAG retrieves first and generates second. The first 60 pages of Hung-Yu Kao's Fall 2025 RAG deck are all about finding the right material. Sparse vectors (bag-of-words, TF-IDF, BM25) are cheap and dependable; dense vectors catch paraphrases. BERT's [CLS] isn't a good sentence vector as-is, hence Sentence-BERT pooling and bi-encoders. A cross-encoder is accurate but needs nearly 50 million passes for 10,000 sentences, while a bi-encoder needs 20,000. SimCSE uses dropout as data augmentation, DPR beats BM25 with only 1,000 training examples, and GTR shows that scaling up a dual encoder improves out-of-domain retrieval.
"Look over there" is 3 tokens, the Chinese version 4, the Japanese version 12. When output length does not track input length, the classifier trick of adding an FFN at the end fails. This 33-slide deck starts from encoder-decoder models, works through vanishing gradients in RNNs and the three LSTM gates, and ends with attention fixing both long-range memory and parallelism, including why scores are divided by √d.
A guide to the Sub-word Tokenization unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). Splitting on white space only works for Western languages and cannot handle unseen words or German-style compounds. BPE starts from characters and repeatedly merges the most frequent adjacent pair into the vocabulary, so each merge adds one entry; its weakness is that the greedy split is not always the best one. The Unigram LM picks splits by probability and can even sample different ones. The slides give vocabulary sizes of 30522 for BERT, 50257 for GPT-2/GPT-3 and 32,128 for T5.
In week 14 of Fall 2025, Hung-Yu Kao closed the course with two slide decks. Course_summary sorts the semester into three columns (NLP Fundamentals, NLP Models, NLP Advances), adds two columns for the TA labs, and lists five directions for further study. The second deck is his notes on Denny Zhou's (Google DeepMind) April 2025 Stanford talk. Its claim: pretrained models can already reason, and decoding is what brings it out. CoT decoding, self-consistency and retrieval + reasoning add up to four inequalities. The Fall 2026 W16 'Reasoning / Agent' unit has not been released yet.
Fall 2025 was graded 70% assignments + 30% term project. Projects were done in groups of 3–4 and split into Proposal 6%, Progress 6%, Poster 6% and Report 12%, with no GPUs provided. The repo has no project spec, only the syllabus structure, an end-of-term reminder and the W15–W16 recordings. Fall 2026 switches to 75% assignments (4 of them) + a 25% in-person midterm in W14. The schedule drops the presentation weeks, adds a Reasoning/Agent unit, and brings in an AI-TA for grading support and TAICA compute credits. The official materials don't say why.
A guide to the Transformer unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). The slides start from two RNN problems: words interact only across O(N) steps, and time steps cannot run in parallel. Self-attention lets every word look at every other word and computes it all in one matrix multiplication, fixing both at once. The cost is that the model can no longer tell word order, so sinusoidal positional encoding is added. The lecture then assembles multi-head attention, Add & Norm, feed forward, cross-attention, masked attention and teacher forcing, and closes with GPT-2, ViT and four variants to show how far the architecture went.
Week 2 of Hung-Yu Kao's NLP course at NTHU is a 62-page deck about predicting the next word. The first part covers statistical language models with a bigram count table, add-one smoothing, and perplexity, then sparse vectors through a PPMI example with cherry and digital. The middle returns to Word2Vec's negative sampling and uses the Chinese word for 'apple' to show why contextualized embeddings are needed. The last part derives RNNs from three weaknesses of feedforward networks and shows how they handle NER, sentence classification, and stacked and bidirectional variants.
An RNN translator has to squeeze the whole source sentence into one vector, and long sentences don't fit. Attention lets the decoder look back at every input position each time it produces a word, score each one, and take a weighted sum. Call the scorer the query, the thing being scored the key, and the thing being averaged the value, and you have dot-product attention. The Transformer goes one step further: the input attends to itself, recurrence disappears, and in return you get parallelism and a constant path length. The price is that position information has to be added back by hand.
A static word vector gives "apple" one embedding, whether it means the fruit or the company. The BERT lecture in ADL Fall 2025 starts from that polysemy problem. TagLM feeds language-model features into a tagger, ELMo builds contextual embeddings from a deep bidirectional LSTM, and BERT swaps the LSTM for a Transformer, pre-trained with Masked LM and Next Sentence Prediction; downstream, you add a classifier or tagger on the top layer and fine-tune. The optional BERT Variants slides go one ring further out: Transformer-XL for longer context, XLNet's permutation LM to get both AR and AE benefits, RoBERTa's better data and training recipe, SpanBERT's span masking, and mBERT and XLM for many languages. This is the direct prerequisite for HW1, which uses bert-base-chinese for extractive QA.
Applied Deep Learning (ADL) Fall 2025, taught by Yun-Nung (Vivian) Chen in NTU's CSIE department, is a deep learning course built around NLP. It runs from neural network basics through Transformers, BERT, pretraining and prompting, post-training, LoRA, RAG, generation and evaluation, alignment issues, and language agents. Lectures L0–L11 come with slide PDFs and segmented videos, and the playlist adds videos for L12–L14. On the assignment side only the HW1 spec is public; HW2, HW3, and the final project have explainer videos only. That makes it A2.
The first self-study deck of ADL Fall 2025 describes machine learning as finding a function from data, and deep learning as a production line of simple functions where the machine learns what every station does. It uses speech and vision to contrast deep and shallow models, credits big data and GPUs for the post-2010 breakthroughs, and uses the universality theorem to ask why networks should be deep rather than fat. The most practical part comes last: the output domain decides the learning task, and the architecture should fit the properties of the input domain.
Lecture 9 of ADL Fall 2025 answers two questions. The model gives you a probability distribution at every step, so how do you pick a word from it? And once you have a sentence, how do you judge it? The slides start with teacher forcing and exposure bias to show the gap between training and generation, then compare greedy, beam search, sampling, top-k, and nucleus sampling, and file temperature and the penalties under 'control' rather than decoding algorithms. The evaluation half covers BLEU, ROUGE, perplexity, and LLM-Eval, then explains why you would use RL to optimize whole-sentence quality directly.
When an LLM is too big to fine-tune in full, the LLM Adaptation slides of NTU ADL Fall 2025 offer three ways to change only a small part of it: insert small Adapter modules into the Transformer, represent the weight update with low-rank matrices (LoRA), or learn only a prefix or soft prompt (prompt tuning). The slides conclude that no single method fits every task. For HW2, the only public information is its title, "LLM Tuning and Prompt Tuning for Classical Chinese Translation"; the data, baseline, and grading have no written spec.
A pre-trained model can continue text, but that does not mean it follows instructions. The Post-Training slides of NTU ADL Fall 2025 fix this in two steps. Instruction tuning (FLAN, T0) teaches the model to read task descriptions. RLHF then pulls its outputs toward human preference. Three limits of instruction tuning connect the two steps, a reward model and pairwise comparisons solve two practical RL problems, and InstructGPT's SFT → reward model → PPO pipeline ties it all together. ChatGPT runs the same pipeline on multi-turn dialogue.
Lecture 6 of ADL Fall 2025 sorts pre-trained models into three families: encoders (the BERT family, bidirectional context), decoders (the GPT series, good at generation), and encoder-decoders (BART and T5, pre-trained with denoising). It then names two practical obstacles of the pre-trained-model era: downstream labeled data is scarce, and models keep growing until one copy per task no longer fits. The slides' answer is prompt learning. GPT-3's in-context learning shows a model can do a task without updating parameters; hand-written hard prompts (template plus verbalizer, LM-BFF) then give way to soft prompts optimized as vectors (P-Tuning, Prefix-Tuning, Prompt Tuning); and Liu et al.'s prompting typology closes the lecture.
LLMs cannot memorize long-tail facts, their knowledge goes stale, and they cannot see private documents. The RAG slides of NTU ADL Fall 2025 open with an LLM hallucinating about the lecturer herself, then split RAG into indexing, retrieval, and generation: sparse (TF-IDF, BM25) and dense (DPR, Contriever) retrieval, how dense retrievers are trained, pre- and post-retrieval techniques including pointwise and pairwise reranking. A closing roadmap organizes RAG, RETRO, FLARE, Search-R1, and others by what, how, and when to retrieve. For HW3, only the title is public: "Retriever & Reranker Training for RAG".
The first real lecture of ADL Fall 2025 has one through-line: a language model predicts the next word. It starts from one-hot vectors and co-occurrence matrices, moves through the zero-probability problem of n-grams and the smoothing that a neural LM gets for free, and ends at the RNN LM, which folds all previous words into a hidden state. BPTT and vanishing/exploding gradients are the training cost, and LSTM and GRU patch it with gating. The lecture closes by splitting applications into sequence input versus sequence output, which separates tagging from encoder-decoder models.
With whole words as units, any unseen word becomes UNK. With single characters, meaning is hard to reassemble. This 22-page ADL deck explains the mainstream compromise, subwords, and the most common way to build them, BPE. The core is a tiny corpus of 4 words and 16 occurrences: start from characters, merge the most frequent adjacent pair each round, and after 9 merges you have units like newest</w> and low</w>, which then segment the unseen words lowest and powest. It ends with a GPT-3 tokenizer screenshot where the Chinese version of a sentence takes more than twice as many tokens as the English.
The Analysis methods unit of CS224U (Spring 2023) starts by grading three families of methods on a three-column scorecard. Probing is strong at characterizing representations but can't support causal claims. Integrated gradients only gives you a scalar about each representation, but it satisfies the sensitivity axiom, so it does come with a causal guarantee. This post covers slides 1–40, videos 33–35, and feature_attribution.ipynb, including where the notebook breaks in today's environment.
CS224U's fourth unit opens with one question: what can behavioral testing prove, and what can't it? It can never give a guarantee, and when a model fails you first have to ask whether the model or the dataset is at fault. BERT scored 2.2% on negated NLI examples, then 90% after fine-tuning on a small set of them. The unit then covers SQuAD distractor sentences, Breaking NLI, ANLI's human-and-model adversarial collection, and ends with DynaSent's two rounds.
Causal abstraction rests on one operation: take the internal state a model computes for a source input at some location, swap it into the same location for a base input, and check whether the output changes the way your hypothesized high-level program says it should. In CS224U's iit_equality.ipynb, a network with 0.99 test accuracy scores only 0.50 and 0.54 on this check. After IIT training, its counterfactual accuracy is 1.00. DAS replaces guessing which neurons match which variable with learning a rotation matrix.
A few COGS generalization splits score 0 for nearly every model. CS224U uses its own ReCOGS work to explain why: the zeros on the recursion splits are mostly a length-generalization problem, and the zeros on the prepositional-phrase split come from training data that only ever put PPs in certain variables and positions. Assignment 3, hw_recogs.ipynb, uses 135K ReCOGS training pairs. It first has you find Charlie and Lina, two names whose train and test roles are exact opposites, then shows a trained model stumbling on them.
CS224U Spring 2023 tells the story of the Transformer families through BERT's four known limitations. RoBERTa addresses the first (optimization was only partly explored). ELECTRA addresses the second and third (the [MASK] mismatch, and only about 15% of tokens giving a learning signal per batch). XLNet addresses the fourth (the assumption that masked tokens are independent of each other). GPT changes the objective and the mask, T5 and BART change the architecture and how inputs are corrupted, and distillation changes model size. The course's 2023 view: autoregressive architectures have taken over, but bidirectional models may still have the edge for representation.
The first three parts of CS224U's Spring 2023 contextual representations unit start with examples like "break" and "crane" to show why static word vectors were never going to be enough. They then build a Transformer block step by step on the three words "The Rock rules." Only attention connects the columns; every other step runs on each column independently. Finally, two questions sort three positional encoding schemes: Do you have to fix the set of positions ahead of time? Does the scheme get in the way of generalizing to new positions? Absolute encoding fails both, sinusoidal encoding passes the first, and the relative encoding of Shaw et al. (2018) passes both.
The second half of CS224U's 'NLP methods and metrics' unit skips metric formulas. It asks whether your experiment holds up. Naturalistic or crowdsourced data, adversarial or common cases: the course answers 'both' each time. Lock the test set away. Pick baselines when you write the hypothesis. Compare two models with confidence intervals, Wilcoxon, or McNemar, and run several random initializations. The slides, three videos, and two notebooks are all public. Kawin Ethayarajh's guest session 'Real-world NLP assessments' has no public slides or video.
In Spring 2023, CS224U slipped two talks by members of its own teaching team into the Transformer unit. Lisa Li presented Diffusion-LM: instead of generating left to right one word at a time, it denoises a sequence of Gaussian vectors into word vectors. It loses to autoregressive models on both training and decoding efficiency, and in exchange lets a classifier's gradient steer the output at every step. Sidd Karamcheti showed how to cut GPT-2 Small's single-GPU training clock from 99.63 days to 3.37 days by stacking data parallelism, mixed precision, and ZeRO. The first talk survives only as slides; the second has slides and two recordings.
CS224U's first assignment, hw_sentiment.ipynb, is ternary sentiment classification: you develop on two rounds of DynaSent plus SST-3, and the bake-off test set mixes in mystery sentences from undisclosed sources. The original-system question is worth 3 of the 9 homework points, and it has exactly one rule: never touch the three public test sets during development. Run as-is today, the first data-loading cell breaks because Hugging Face datasets 4.0 dropped trust_remote_code.
CS224U's second assignment, hw_openqa.ipynb, asks you to answer questions that come with no passage, using only a frozen language model and a frozen ColBERT retriever. The Spring 2023 version was written for DSP; in January 2024 the repo switched to DSPy and pinned dspy-ai==2.4.13. Before you start you need an OpenAI API key, a ColBERTv2 checkpoint of about 406 MB, and a 600 MB prebuilt index. The notebook's first setup call, dspy.OpenAI, no longer exists in DSPy 3.4.
The Spring 2023 edition of CS224U defines in-context learning as a frozen language model performing a task only by conditioning on the prompt, and warns that the second condition of few-shot learning (no examples of the behavior seen in training) is almost impossible to verify. Potts's 38-page deck runs from GPT-2's TL;DR trick through choosing demonstrations, chain of thought, self-consistency, and DSP, and ends with four recommendations: build dev/test sets first, learn your target model's instruction format, and treat prompt writing as AI system design. Mina Lee's guest lecture asks the reverse question: who should learn to read prompts, people or models?
The Spring 2023 edition of CS224U spends a whole unit on retrieval, because OpenQA gives you only the question and you have to find the evidence yourself, while large language models fabricate sources. The slides by Potts and Omar Khattab go from TF-IDF and BM25 through Success@K, MRR, and average precision to four neural IR designs (cross-encoder, DPR, ColBERT, SPLADE) and how each trades expressiveness against scale, and they end by asking you to count latency and cost as metrics too.
The first CS224U lecture of Spring 2023 asks "Which U.S. states border no U.S. states?" of every system from Chat-80 (1980) to text-davinci-001. The answers show that the progress is real. The lecture then questions whether that progress counts as understanding, using Levesque's "cheap tricks," models that invent links, and benchmarks that saturate within a year or two. That splits the course map in two: the first half teaches you to build systems with Transformers and retrieval-augmented in-context learning, and the second half teaches you to test them with harder benchmarks, behavioral evaluation, and causal explanation methods.
The first two deliverables of the CS224U final project are a literature review and an experiment protocol. The lit review covers 5, 7, or 9 papers depending on team size, under five suggested sections. The protocol has seven required sections, and its core is a hypothesis you can state. The course supplies a six-step paper-search loop, a rule that AI-assistant output must be quoted, and a worked example: a student's final project that became a Findings of EMNLP paper. The Gradescope format and rubric slides, and past exemplary papers, are behind a login.
The CS224U slides compute two numbers from one three-class confusion matrix: accuracy 0.81 and macro F1 0.43. One says the system is good; the other says it gets the two small classes almost entirely wrong. The unit's claim is that different metrics encode different values, and it goes through the bounds, values, and weaknesses of accuracy, the three F-score averages, perplexity, word error rate, and BLEU. Final projects are graded on whether the metrics fit, not on how high the scores are.
CS224U's 'Presenting your research' lecture has four parts: the course-specific rules for the final paper, how to write an NLP paper, how conference submission works, and how to give a talk. Three things matter most. The final paper must include Known project limitations and an Authorship statement. Write as a Shieber-style 'rational reconstruction,' not a chronological tour of your dead ends. At submission, your title largely decides reviewer bidding. The slides, four videos, and projects.md are public; past example papers need a Stanford login.
Models don't understand text — they only understand numbers. Embeddings map each token to a vector of several hundred dimensions, where semantically similar words end up close together in vector space. This is the shared foundation behind search, RAG, and classification.
Every time a model predicts the next token, it assigns a probability to every candidate word. A loss function measures how far that probability distribution is from the correct answer — the further off, the higher the loss, the more the model knows it got it wrong. Cross-entropy is the standard formula; perplexity is its human-readable translation.
Models charge by tokens, not characters. The BPE algorithm starts from individual bytes and repeatedly merges the most frequent adjacent pair to build a vocabulary. English 'understanding' might be 1-2 tokens, but Chinese '理解' could take 2-3 — same meaning, higher cost.
The core of the Transformer is self-attention: for each token, the model computes how relevant every other token is, then takes a weighted sum. This lets the model reach across distance to figure out that 'it' refers to 'cat' not 'mat' — and is the foundation for how it handles long documents.
2021 marked NLP’s shift from fine-tuning an entire model to adapting only a small fraction of its parameters. Prefix-Tuning at ACL, LoRA on arXiv, and Prompt Tuning at EMNLP all appeared that year; ACL Rolling Review launched; and the Findings track established itself as a second publication channel.
2022 marked NLP’s shift from demonstrating model capabilities toward aligning and controlling them. InstructGPT brought RLHF into the mainstream, Chain-of-Thought showed that prompts could unlock reasoning, and Flan 2022 matured instruction-tuning methodology. ACL and NAACL adopted ARR as their sole review path, exposing infrastructure and reviewer-load problems. ChatGPT launched at year-end and rewrote the rules of NLP research.
2023 was the first full academic year after ChatGPT, and LLMs rewrote the NLP conference agenda. ACL's Best Papers examined humor understanding and the propagation of political bias; an EMNLP Best Paper explained in-context learning through information flow; and the HackAPrompt competition paper also won an EMNLP Best Paper award, signaling that security research had entered the mainstream. The year's largest shift was from asking how to make models more accurate to asking how we can tell when a model is misleading us.
NLP conferences redefined themselves under LLM dominance in 2024. ACL made open science its annual theme, and four of its seven Best Papers probed fundamental limits of language models. EMNLP turned toward multilingual and cross-cultural work, with Best Papers spanning speech representations and gradient interpretability. ACL and EMNLP received more than 10,000 submissions combined, but the deeper anxiety was what remains of NLP when LLMs can perform nearly every traditional NLP task.
NLP conference submissions nearly doubled in 2025: ACL received 8,360 papers and EMNLP 8,174. China-based first authors exceeded 51% at ACL, and DeepSeek's Native Sparse Attention won Best Paper. The deeper story was an identity crisis: an ACL president said 'ACL is not an AI conference,' a quantitative study asked 'Has ACL Lost Its Crown?', and EMNLP faced questions about what still distinguished it from ACL or NAACL.
CS124 Winter 2026 opens by mapping a ten-week path from tokenization and classification to retrieval, speech, networks, and LLMs, while PA0 establishes the Jupyter environment used throughout the quarter.
Week 2 builds three layers: a token vocabulary with BPE, sequence comparison with dynamic-programming edit distance, and probability approximation with n-grams; PA1 turns regex and BPE into executable work.
Week 4 builds candidates with an inverted index, ranks them with tf-idf and cosine similarity, and then connects retrieved evidence to generation; PA3 exposes RAG's inspectable retrieval half.
Week 5's public materials support the distributional hypothesis, word embeddings, and cosine similarity; the paired Social NLP lecture is unrecorded and restricted, so concrete audit methods are labeled as author extensions.
Week 6 uses public neural-network slides for weighted sums, nonlinearities, loss, and backpropagation, then a public LLM/Transformer deck labeled 2025 for decoder-only architecture without treating it as the 2026 live transcript.
Week 8 sends text through TTS and back through STT, requiring error classification, formatting-loss analysis, and accent stress tests, while Lab 4 prepares Git collaboration for the team agent project.
Lecture 3 decomposes neural-network training into computation graphs, local derivatives, and the chain rule: the forward pass computes a result; backprop accumulates gradients from the output so every parameter knows how to move.
Lecture 11 divides evaluation into what to test, how to measure it, and when the result stops being trustworthy. Benchmarks saturate or leak, prompts change scores, and an LLM judge remains a biased model.
Lecture 6 completes the Transformer picture with encoders, decoders, and cross-attention, then breaks the final project into formats, assessment, research topics, and data. A viable topic needs one explicit baseline and metric.
Winter 2026 Lecture 1 divides NLP into four eras: early exploration, symbolic systems, statistical machine learning, and deep/self-supervised learning. The point is not the dates but how each era redefined the language problem.
Lecture 4 defines a language model as a next-word probability distribution, then uses an RNN to compress an arbitrarily long prefix. It also exposes recurrence's central cost: information and gradients travel one time step at a time.
Lecture 5 moves from the long-range and sequential bottlenecks of RNNs to self-attention and the Transformer. It shortens information paths and enables parallel computation, at the price of quadratic attention and separately encoded position.
Lecture 2 moves from word2vec's prediction task, objective, and gradients to count-based vectors and evaluation. Meaning becomes a high-dimensional position learned from context, not a label retrieved from a dictionary.
CS288 moves from n-grams to RAG, reasoning, and agents through 18 public slide units and three assignments; Berkeley-only recordings make this an A3 materials route, not a public video course.
CS124 is the first course in Stanford's NLP branch. Its textbook is Jurafsky's own Speech and Language Processing, free online, and all nine assignment repos are public. But a banner sits on the course homepage: it will not be taught at all in AY 2026–27. And the chapter numbers the syllabus points at no longer match the August 2026 textbook.
CS224N has kept every course website since 2000 online. In Winter 2019, Transformers were lecture 14, taught by a guest. In Winter 2026 they are lecture 5, and every lecture after that assumes you already know them. The machine translation assignment is gone; assignment 3 now has you code a decoder-only Transformer from scratch, with pytest suites that run on your laptop.
CS224U's teaching material isn't a slide deck — it's an Apache-2.0 GitHub repo holding the lecture notebooks, all three assignments, and the grading document for the final project. But the on-campus course has skipped three straight academic years since Spring 2023, and ExploreCourses briefly put it back on the books for Spring 2026-27, then dropped that section again by 29 September 2026. The official description still lists relation extraction and semantic parsing; the 2023 syllabus covers neither. And the data-loading cell in the first assignment breaks in a fresh environment today, on a Hugging Face compatibility change.
The dividing line in LLM interviews is whether you've actually used these things. High-frequency topics: BPE tokenization logic and multilingual challenges, pretraining objectives (CLM vs MLM), three levels of fine-tuning (full/LoRA/prompt tuning), RLHF workflow and failure modes, prompting as engineering practice, and the difficulty of LLM evaluation with current methods.
In a climbing RAG system, 'recommend the next route' (progression) and 'recommend a similar route' (similarity) were conflated by a single hasSimilarRouteIntent() function, causing recommendation quality to collapse. The fix is a two-stage intent classification with a Regex Fast Path + LLM Fallback.