Skip to content
All tags

#nlp

79 posts

NTHU NLP Guide 8: ELMo, BERT, T5, BART, GPT and the Three Roads to Pretraining

A guide to the BERT and its Family unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). It starts from "I record the record": Word2Vec and GloVe give both records the same vector. ELMo fixes this with a bidirectional LSTM language model. With Transformers, pretraining splits into three roads: encoders (BERT: MLM plus NSP, strong at understanding, weak at generation), encoder-decoders (T5's span corruption, BART's five noise types), and decoders (GPT: pure next-token prediction). The last part covers GPT-3's in-context learning and scaling laws, and why decoders became today's dominant backbone.

NTHU NLP Guide 10: Why the Same Model Sounds Stiff or Rambles — Decoding Strategies and NLG Evaluation

A guide to the decoding and evaluation unit in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The first half covers how to pick a word once the model outputs a probability distribution: greedy decoding can't take back a mistake, beam search keeps several candidates but favors short outputs, and top-k / top-p trade determinism for diversity (missing from the slides; the professor covers it verbally in class). The second half covers scoring generated text: BLEU's modified precision and brevity penalty, ROUGE-N and ROUGE-L, perplexity, and what GLUE, SQuAD 2.0, MTEB, and MMLU each measure.

NTHU NLP Guide 11: GPT-2 vs. T5 for Chinese Summarization — Left Padding, −100, and ROUGE

A guide to the GPT-2 / T5 tutorial in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). One task, LCSTS Chinese summarization, is solved twice. Decoder-only GPT-2 is written in native PyTorch: you join article and summary into one sequence, switch to left padding, and set padding labels to −100. Encoder-decoder mT5 uses Seq2SeqTrainer: no left padding needed, and DataCollatorForSeq2Seq handles the −100 for you. Both segment with jieba and score word-level ROUGE.

Reading NTHU Kao's NLP: GPT-3, InstructGPT, and RLHF — How a Model That Continues Text Becomes an Assistant That Follows Instructions

Hung-Yu Kao's Fall 2025 W8 slides walk from GPT-1 to GPT-3, explain how the Sparse Transformer behind GPT-3 cuts attention cost, and then use InstructGPT to show the gap between continuing text and following instructions. The maximum likelihood objective can't tell a fabricated fact from a slightly wrong synonym, so three extra stages are added: SFT learns how humans write, a reward model learns how humans grade, and PPO optimizes against that grade while a KL penalty keeps the model from drifting too far. The lecture closes with Llama-2: separate safety and helpfulness reward models, context distillation, and GQA for faster inference.

NTHU NLP Guide 9: One BERT That Scores and Classifies at Once — The Hugging Face Tutorial and HW3 Multi-Output Learning

A guide to the Hugging Face BERT tutorial and HW3 in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The tutorial walks through binary IMDb sentiment classification: AutoTokenizer, the input_ids / token_type_ids / attention_mask fields, AutoModelForSequenceClassification, and Trainer. HW3 applies the same tools to SemEval 2014 Task 1: one bert-base-uncased with two heads, one regressing a 1–5 relatedness score and one classifying entailment into three classes. You add the two losses and write the training loop yourself, because Trainer is not allowed.

NTHU NLP HW1: Testing Word Vectors on Google Analogy — Pretrained GloVe vs. Word2Vec Trained on 20% of Wikipedia

HW1 tests word vectors on the 19,544 Google Analogy questions (8,869 semantic, 10,675 syntactic). You first answer them with pretrained glove-wiki-gigaword-100 loaded through Gensim, then train your own Word2Vec on a 20% sample of a pre-cleaned Wikipedia dump, and plot t-SNE for the family subcategory both times. Seven TODOs are worth 55%, the report 45%. Fall 2026 keeps the same TODOs but asks for an .ipynb with outputs.

NTHU Hung-Yu Kao NLP, Week 1: Why Language Is Hard, and How Text Became Numbers Before LLMs

Week 1 of Hung-Yu Kao's NLP course at NTHU is a 91-page deck, W1_NLP_brief. It opens with 'Watch for kids', five readings of the telescope sentence, and a Chinese tongue-twister about eleven uncles to show why language is hard. Then, from an information-retrieval angle, it builds one pipeline: inverted index, tokenization, stemming, TF-IDF, BM25. The second half hits that pipeline's dead ends (synonyms, polysemy, vocabulary mismatch), moves to SVD-based LSA, and closes with a preview of dense vectors through Skip-gram, GloVe, and FastText.

Reading NTHU Hung-Yu Kao's Natural Language Processing: What Outsiders Can Get from a 1,200-Seat TAICA Course

Hung-Yu Kao's Natural Language Processing at National Tsing Hua University is a graduate-level flagship course in the TAICA alliance. The syllabus caps it at 1,200 students, it is taught in Mandarin, and it runs from TF-IDF and word vectors to RLHF, PEFT, and RAG. For Fall 2025, the slides, 32 class recordings, and 4 assignments with starter notebooks are all on GitHub, which rates A3. Solutions, grading, and the term-project spec are not public. Fall 2026 is in progress and only goes up to W3, so it rates A2. Grading changed to 75% assignments plus a 25% in-person midterm, and a Reasoning/Agent unit was added.

NTHU NLP LLM API Lab: NLI Classification with Gemini, OpenAI, and Claude, Using prompts.yaml, JSON Output, Few-Shot, and Token Counts

This 34-slide TA session answers a practical question. Pasting data into the ChatGPT web page one row at a time is slow and hits hourly limits, so research and homework should use the API. The notebook runs one SemEval 2014 entailment example through Gemini, Claude, and OpenAI in turn: prompts live in prompts.yaml, output is forced into JSON, then few-shot and token counting. The material is from 2024. The slide cover says 2024/11/21, and the notebook uses gemini-1.5-pro, gpt-4o, and claude-3-5-sonnet-20241022. That Claude model was retired on 2025-10-28, and Google's old Gemini SDK reached end of support on 2025-11-30.

Reading NTHU Kao's NLP: Parameter-Efficient Fine-Tuning — Fine-Tuning Large Models Without an A100 Cluster

Hung-Yu Kao's Fall 2025 PEFT slides open with a budget: full fine-tuning of Llama 2-7B in 16-bit needs about 56GB of GPU memory, while training only 0.2M parameters brings it down to about 17GB, because gradients and optimizer states nearly vanish. Intrinsic dimensionality then explains why tuning a small slice is enough: the longer a model is pretrained and the larger it is, the fewer effective dimensions fine-tuning needs. Methods fall into additive (Adapters, Prompt Tuning), selective (BitFit), reparametrization (LoRA), and hybrid (MAM Adapters, S4). The second half runs from GPT-2's task descriptions and verbalizers to the trade-offs between prefix tuning and soft prompt tuning.

NTHU NLP HW2: Arithmetic as a Language — PyTorch TA Session, a Two-Layer LSTM, and Teacher Forcing

HW2 treats expressions like "14*(43+20)=882" as character sequences and asks a two-layer LSTM to generate the answer one character at a time after it sees "=". The training set has 2,369,250 rows and the eval set 263,250, with every number in 0–49. Six TODOs run from building a vocabulary and batching with loss only after "=" to a generator, teacher-forced training, and exact-match evaluation. The W4 PyTorch TA session is the toolbox for it.

NTHU NLP RAG Labs + HW4: Building a Cat-Facts RAG Two Ways, with LangChain and by Hand

Each of the two RAG TA sessions builds one version. The first installs Ollama on Colab to run llama3.2:1b and wires up a minimal RAG with LangChain's Chroma, MMR, and retrieval chain. The second uses LangChain only for data prep and writes the rest by hand: chunking, text and vector stores, hybrid BM25 + cosine retrieval merged with RRF, then generation with Llama-3.2-1B-Instruct. HW4 applies the first session's skeleton to 150 cat facts and 150 GPT-5-generated QA pairs. The generator must be Llama3.2-1b and the embedding model jina-embeddings-v2-base-en, and you report recall@1, recall@5, and exact match. Code is 45% of the grade and the report 55%; the report analyzes how prompts, data format, document order, and counterfactual information change the results.

NTHU NLP RAG, Part 2: Connecting the Retriever to a Reader, from ORQA and REALM to Self-RAG

The second half of W11_RAG.pdf starts at the "From Retrievers to QA" slide and turns a retriever plus a reader into a full QA system. It begins with ORQA and REALM (2019–2020), where BERT is the reader, then covers the first paper named RAG and REPLUG, which keeps the LLM frozen. A single table then sorts seven recent fixes into three groups: rewrite the query (Query Rewriting, HyDE), make the generator robust to noise (RetRobust, RAFT, RAAT), and decide when to retrieve (FLARE, Self-RAG). It ends with noise types, four abilities an LLM needs inside RAG, and generative retrieval (GR) with reliable response generation (RRG). On the recording side, the W11 Thursday lecture stops at the RAG paper; I could not find a recording that covers the later slides.

Reading NTHU Kao's NLP: RAG (Part 1) — Hallucination and Retrievers, from BM25 to DPR and GTR

An LLM will confidently answer that Oppenheimer was born in 1967 (the year he died). RAG retrieves first and generates second. The first 60 pages of Hung-Yu Kao's Fall 2025 RAG deck are all about finding the right material. Sparse vectors (bag-of-words, TF-IDF, BM25) are cheap and dependable; dense vectors catch paraphrases. BERT's [CLS] isn't a good sentence vector as-is, hence Sentence-BERT pooling and bi-encoders. A cross-encoder is accurate but needs nearly 50 million passes for 10,000 sentences, while a bi-encoder needs 20,000. SimCSE uses dropout as data augmentation, DPR beats BM25 with only 1,000 training examples, and GTR shows that scaling up a dual encoder improves out-of-domain retrieval.

NTHU NLP W3: When Input and Output Lengths Differ — Seq2seq, LSTM, and Attention

"Look over there" is 3 tokens, the Chinese version 4, the Japanese version 12. When output length does not track input length, the classifier trick of adding an FFN at the end fails. This 33-slide deck starts from encoder-decoder models, works through vanishing gradients in RNNs and the three LSTM gates, and ends with attention fixing both long-range memory and parallelism, including why scores are divided by √d.

NTHU NLP Guide 7: Sub-word Tokenization, or Why a Model's Vocabulary Is Made of Word Pieces

A guide to the Sub-word Tokenization unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). Splitting on white space only works for Western languages and cannot handle unseen words or German-style compounds. BPE starts from characters and repeatedly merges the most frequent adjacent pair into the vocabulary, so each merge adds one entry; its weakness is that the greedy split is not always the best one. The Unigram LM picks splits by probability and can even sample different ones. The slides give vocabulary sizes of 30522 for BERT, 50257 for GPT-2/GPT-3 and 32,128 for T5.

NTHU NLP Course Summary and LLM Reasoning Notes: A Three-Column Map of the Semester, Then Denny Zhou's Question of Whether Pretrained Models Can Reason

In week 14 of Fall 2025, Hung-Yu Kao closed the course with two slide decks. Course_summary sorts the semester into three columns (NLP Fundamentals, NLP Models, NLP Advances), adds two columns for the TA labs, and lists five directions for further study. The second deck is his notes on Denny Zhou's (Google DeepMind) April 2025 Stanford talk. Its claim: pretrained models can already reason, and decoding is what brings it out. CoT decoding, self-consistency and retrieval + reasoning add up to four inequalities. The Fall 2026 W16 'Reasoning / Agent' unit has not been released yet.

NTHU NLP Term Project and the Fall 2026 Redesign: A 30% Group Project Becomes an In-Person W14 Midterm, Plus Reasoning/Agentic AI, an AI-TA and TAICA Compute

Fall 2025 was graded 70% assignments + 30% term project. Projects were done in groups of 3–4 and split into Proposal 6%, Progress 6%, Poster 6% and Report 12%, with no GPUs provided. The repo has no project spec, only the syllabus structure, an end-of-term reminder and the W15–W16 recordings. Fall 2026 switches to 75% assignments (4 of them) + a 25% in-person midterm in W14. The schedule drops the presentation weeks, adds a Reasoning/Agent unit, and brings in an AI-TA for grading support and TAICA compute credits. The official materials don't say why.

NTHU NLP Guide 6: Without the RNN, How Does a Transformer Know How Words Relate and Where They Sit?

A guide to the Transformer unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). The slides start from two RNN problems: words interact only across O(N) steps, and time steps cannot run in parallel. Self-attention lets every word look at every other word and computes it all in one matrix multiplication, fixing both at once. The cost is that the model can no longer tell word order, so sinusoidal positional encoding is added. The lecture then assembles multi-head attention, Add & Norm, feed forward, cross-attention, masked attention and teacher forcing, and closes with GPT-2, ViT and four variants to show how far the architecture went.

NTHU Hung-Yu Kao NLP, Week 2: Word Embeddings and Language Models, from Counting N-grams to RNNs

Week 2 of Hung-Yu Kao's NLP course at NTHU is a 62-page deck about predicting the next word. The first part covers statistical language models with a bigram count table, add-one smoothing, and perplexity, then sparse vectors through a PPMI example with cherry and digital. The middle returns to Word2Vec's negative sampling and uses the Chinese word for 'apple' to show why contextualized embeddings are needed. The last part derives RNNs from three weaknesses of feedforward networks and shows how they handle NER, sentence classification, and stacked and bidirectional variants.

NTU ADL Lecture 4: Attention and the Transformer

An RNN translator has to squeeze the whole source sentence into one vector, and long sentences don't fit. Attention lets the decoder look back at every input position each time it produces a word, score each one, and take a weighted sum. Call the scorer the query, the thing being scored the key, and the thing being averaged the value, and you have dot-product attention. The Transformer goes one step further: the input attends to itself, recurrence disappears, and in return you get parallelism and a constant path length. The price is that position information has to be added back by hand.

Reading NTU ADL 2025 Fall: BERT and Its Family — From the Polysemy Problem to XLNet, RoBERTa, and mBERT

A static word vector gives "apple" one embedding, whether it means the fruit or the company. The BERT lecture in ADL Fall 2025 starts from that polysemy problem. TagLM feeds language-model features into a tagger, ELMo builds contextual embeddings from a deep bidirectional LSTM, and BERT swaps the LSTM for a Transformer, pre-trained with Masked LM and Next Sentence Prediction; downstream, you add a classifier or tagger on the top layer and fine-tune. The optional BERT Variants slides go one ring further out: Transformer-XL for longer context, XLNet's permutation LM to get both AR and AE benefits, RoBERTa's better data and training recipe, SpanBERT's span masking, and mBERT and XLM for many languages. This is the direct prerequisite for HW1, which uses bert-base-chinese for extractive QA.

Reading NTU Yun-Nung Chen's Applied Deep Learning 2025 Fall: Course Map, A2 Rating, and How to Read It

Applied Deep Learning (ADL) Fall 2025, taught by Yun-Nung (Vivian) Chen in NTU's CSIE department, is a deep learning course built around NLP. It runs from neural network basics through Transformers, BERT, pretraining and prompting, post-training, LoRA, RAG, generation and evaluation, alignment issues, and language agents. Lectures L0–L11 come with slide PDFs and segmented videos, and the playlist adds videos for L12–L14. On the assignment side only the HW1 spec is public; HW2, HW3, and the final project have explainer videos only. That makes it A2.

NTU ADL 2025 Lecture 1: What Machine Learning and Deep Learning Are

The first self-study deck of ADL Fall 2025 describes machine learning as finding a function from data, and deep learning as a production line of simple functions where the machine learns what every station does. It uses speech and vision to contrast deep and shallow models, credits big data and GPUs for the post-2010 breakthroughs, and uses the universality theorem to ask why networks should be deep rather than fat. The most practical part comes last: the output domain decides the learning task, and the architecture should fit the properties of the input domain.

NTU ADL 2025 Lecture 9: Decoding, Generation Control, and Evaluation for NLG

Lecture 9 of ADL Fall 2025 answers two questions. The model gives you a probability distribution at every step, so how do you pick a word from it? And once you have a sentence, how do you judge it? The slides start with teacher forcing and exposure bias to show the gap between training and generation, then compare greedy, beam search, sampling, top-k, and nucleus sampling, and file temperature and the penalties under 'control' rather than decoding algorithms. The evaluation half covers BLEU, ROUGE, perplexity, and LLM-Eval, then explains why you would use RL to optimize whole-sentence quality directly.

NTU ADL 2025 Lecture 7.5: PEFT — Adapter, LoRA, Prompt Tuning, and HW2

When an LLM is too big to fine-tune in full, the LLM Adaptation slides of NTU ADL Fall 2025 offer three ways to change only a small part of it: insert small Adapter modules into the Transformer, represent the weight update with low-rank matrices (LoRA), or learn only a prefix or soft prompt (prompt tuning). The slides conclude that no single method fits every task. For HW2, the only public information is its title, "LLM Tuning and Prompt Tuning for Classical Chinese Translation"; the data, baseline, and grading have no written spec.

NTU ADL 2025 Lecture 7: Post-Training — Instruction Tuning, RLHF, and InstructGPT

A pre-trained model can continue text, but that does not mean it follows instructions. The Post-Training slides of NTU ADL Fall 2025 fix this in two steps. Instruction tuning (FLAN, T0) teaches the model to read task descriptions. RLHF then pulls its outputs toward human preference. Three limits of instruction tuning connect the two steps, a reward model and pairwise comparisons solve two practical RL problems, and InstructGPT's SFT → reward model → PPO pipeline ties it all together. ChatGPT runs the same pipeline on multi-turn dialogue.

Reading NTU ADL 2025 Fall: Three Pre-training Families and Prompt Learning — From BERT, GPT, and T5 to Prompts Only Machines Understand

Lecture 6 of ADL Fall 2025 sorts pre-trained models into three families: encoders (the BERT family, bidirectional context), decoders (the GPT series, good at generation), and encoder-decoders (BART and T5, pre-trained with denoising). It then names two practical obstacles of the pre-trained-model era: downstream labeled data is scarce, and models keep growing until one copy per task no longer fits. The slides' answer is prompt learning. GPT-3's in-context learning shows a model can do a task without updating parameters; hand-written hard prompts (template plus verbalizer, LM-BFF) then give way to soft prompts optimized as vectors (P-Tuning, Prefix-Tuning, Prompt Tuning); and Liu et al.'s prompting typology closes the lecture.

NTU ADL 2025 Lecture 8: RAG — From Retrieval and Reranking to Search-R1, Plus HW3

LLMs cannot memorize long-tail facts, their knowledge goes stale, and they cannot see private documents. The RAG slides of NTU ADL Fall 2025 open with an LLM hallucinating about the lecturer herself, then split RAG into indexing, retrieval, and generation: sparse (TF-IDF, BM25) and dense (DPR, Contriever) retrieval, how dense retrievers are trained, pre- and post-retrieval techniques including pointwise and pairwise reranking. A closing roadmap organizes RAG, RETRO, FLARE, Search-R1, and others by what, how, and when to retrieve. For HW3, only the title is public: "Retriever & Reranker Training for RAG".

NTU ADL Lecture 3: Word Representations, Language Models, and RNNs

The first real lecture of ADL Fall 2025 has one through-line: a language model predicts the next word. It starts from one-hot vectors and co-occurrence matrices, moves through the zero-probability problem of n-grams and the smoothing that a neural LM gets for free, and ends at the RNN LM, which folds all previous words into a hidden state. BPTT and vanishing/exploding gradients are the training cost, and LSTM and GRU patch it with gating. The lecture closes by splitting applications into sequence input versus sequence output, which separates tagging from encoder-decoder models.

NTU ADL Lecture 5: Tokenization and BPE, or Where the Vocabulary Comes From

With whole words as units, any unseen word becomes UNK. With single characters, meaning is hard to reassemble. This 22-page ADL deck explains the mainstream compromise, subwords, and the most common way to build them, BPE. The core is a tiny corpus of 4 words and 16 occurrences: start from characters, merge the most frequent adjacent pair each round, and after 9 merges you have units like newest</w> and low</w>, which then segment the unseen words lowest and powest. It ends with a GPT-3 tokenizer screenshot where the Chinese version of a sentence takes more than twice as many tokens as the English.

CS224U Analysis Methods I: Probing Shows You Representations, Feature Attribution Gives You Causal Guarantees

The Analysis methods unit of CS224U (Spring 2023) starts by grading three families of methods on a three-column scorecard. Probing is strong at characterizing representations but can't support causal claims. Integrated gradients only gives you a scalar about each representation, but it satisfies the sensitivity axiom, so it does come with a causal guarantee. This post covers slides 1–40, videos 33–35, and feature_attribution.ipynb, including where the notebook breaks in today's environment.

CS224U Behavioral Evaluation: Analytical Considerations, Adversarial Tests, ANLI, and DynaSent

CS224U's fourth unit opens with one question: what can behavioral testing prove, and what can't it? It can never give a guarantee, and when a model fails you first have to ask whether the model or the dataset is at fault. BERT scored 2.2% on negated NLI examples, then 90% after fine-tuning on a small set of them. The unit then covers SQuAD distractor sentences, Breaking NLI, ANLI's human-and-model adversarial collection, and ends with DynaSent's two rounds.

CS224U Analysis Methods II: From "The Information Is There" to "The Model Uses It" with Interchange Interventions — Causal Abstraction, IIT, and DAS

Causal abstraction rests on one operation: take the internal state a model computes for a source input at some location, swap it into the same location for a base input, and check whether the output changes the way your hypothesized high-level program says it should. In CS224U's iit_equality.ipynb, a network with 0.99 test accuracy scores only 0.50 and 0.54 on this check. After IIT training, its counterfactual accuracy is 1.00. DAS replaces guessing which neurons match which variable with learning a rotation matrix.

CS224U Compositional Generalization: COGS, ReCOGS, and Assignment 3

A few COGS generalization splits score 0 for nearly every model. CS224U uses its own ReCOGS work to explain why: the zeros on the recursion splits are mostly a length-generalization problem, and the zeros on the prepositional-phrase split come from training data that only ever put PPs in certain variables and positions. Assignment 3, hw_recogs.ipynb, uses 135K ReCOGS training pairs. It first has you find Charlie and Lina, two names whose train and test roles are exact opposites, then shows a trained model stumbling on them.

CS224U Contextual Representations II: What GPT, BERT, RoBERTa, ELECTRA, T5, BART, and Distillation Each Change in Pretraining

CS224U Spring 2023 tells the story of the Transformer families through BERT's four known limitations. RoBERTa addresses the first (optimization was only partly explored). ELECTRA addresses the second and third (the [MASK] mismatch, and only about 15% of tokens giving a learning signal per batch). XLNet addresses the fourth (the assumption that masked tokens are independent of each other). GPT changes the objective and the mask, T5 and BART change the architecture and how inputs are corrupted, and distillation changes model size. The course's 2023 view: autoregressive architectures have taken over, but bidirectional models may still have the edge for representation.

CS224U Contextual Representations I: "Break" Has Eight Meanings, and How the Transformer Lets Each Word Read Its Context

The first three parts of CS224U's Spring 2023 contextual representations unit start with examples like "break" and "crane" to show why static word vectors were never going to be enough. They then build a Transformer block step by step on the three words "The Rock rules." Only attention connects the columns; every other step runs on each column independently. Finally, two questions sort three positional encoding schemes: Do you have to fix the set of positions ahead of time? Does the scheme get in the way of generalizing to new positions? Absolute encoding fails both, sinusoidal encoding passes the first, and the relative encoding of Shaw et al. (2018) passes both.

CS224U Methods and Metrics II: Datasets, Data Splits, and Comparing Models

The second half of CS224U's 'NLP methods and metrics' unit skips metric formulas. It asks whether your experiment holds up. Naturalistic or crowdsourced data, adversarial or common cases: the course answers 'both' each time. Lock the test set away. Pick baselines when you write the hypothesis. Compare two models with confidence intervals, Wilcoxon, or McNemar, and run several random initializations. The slides, three videos, and two notebooks are all public. Kawin Ethayarajh's guest session 'Real-world NLP assessments' has no public slides or video.

Reading Stanford CS224U, Part 17: Two Extension Lectures — Generating Text with Diffusion, and Turning LLM Training Folklore into Intuition

In Spring 2023, CS224U slipped two talks by members of its own teaching team into the Transformer unit. Lisa Li presented Diffusion-LM: instead of generating left to right one word at a time, it denoises a sequence of Gaussian vectors into word vectors. It loses to autoregressive models on both training and decoding efficiency, and in exchange lets a classifier's gradient steer the output at every step. Sidd Karamcheti showed how to cut GPT-2 Small's single-GPU training clock from 99.63 days to 3.37 days by stacking data parallelism, mixed precision, and ZeRO. The first talk survives only as slides; the second has slides and two recordings.

CS224U Homework 1: Multi-Domain Sentiment and the Bake-Off

CS224U's first assignment, hw_sentiment.ipynb, is ternary sentiment classification: you develop on two rounds of DynaSent plus SST-3, and the bake-off test set mixes in mystery sentences from undisclosed sources. The original-system question is worth 3 of the 9 homework points, and it has exactly one rule: never touch the three public test sets during development. Run as-is today, the first data-loading cell breaks because Hugging Face datasets 4.0 dropped trust_remote_code.

CS224U Assignment 2: Few-Shot OpenQA with DSPy

CS224U's second assignment, hw_openqa.ipynb, asks you to answer questions that come with no passage, using only a frozen language model and a frozen ColBERT retriever. The Spring 2023 version was written for DSP; in January 2024 the repo switched to DSPy and pinned dspy-ai==2.4.13. Before you start you need an OpenAI API key, a ColBERTv2 checkpoint of about 406 MB, and a 600 MB prebuilt index. The notebook's first setup call, dspy.OpenAI, no longer exists in DSPy 3.4.

CS224U In-Context Learning: Origins, Core Concepts, and Suggested Methods

The Spring 2023 edition of CS224U defines in-context learning as a frozen language model performing a task only by conditioning on the prompt, and warns that the second condition of few-shot learning (no examples of the behavior seen in training) is almost impossible to verify. Potts's 38-page deck runs from GPT-2's TL;DR trick through choosing demonstrations, chain of thought, self-consistency, and DSP, and ends with four recommendations: build dev/test sets first, learn your target model's instruction format, and treat prompt writing as AI system design. Mina Lee's guest lecture asks the reverse question: who should learn to read prompts, people or models?

CS224U Information Retrieval: From Classical IR and IR Metrics to Neural IR

The Spring 2023 edition of CS224U spends a whole unit on retrieval, because OpenQA gives you only the question and you have to find the evidence yourself, while large language models fabricate sources. The slides by Potts and Omar Khattab go from TF-IDF and BM25 through Success@K, MRR, and average precision to four neural IR designs (cross-encoder, DPR, ColBERT, SPLADE) and how each trades expressiveness against scale, and they end by asking you to count latency and cost as metrics too.

CS224U Opening Lecture: One Question Asked for Forty Years, and How a 2023 NLU Course Defines Understanding

The first CS224U lecture of Spring 2023 asks "Which U.S. states border no U.S. states?" of every system from Chat-80 (1980) to text-davinci-001. The answers show that the progress is real. The lecture then questions whether that progress counts as understanding, using Levesque's "cheap tricks," models that invent links, and benchmarks that saturate within a year or two. That splits the course map in two: the first half teaches you to build systems with Transformers and retrieval-augmented in-context learning, and the second half teaches you to test them with harder benchmarks, behavioral evaluation, and causal explanation methods.

CS224U Final Project Workflow: Lit Review and Experiment Protocol

The first two deliverables of the CS224U final project are a literature review and an experiment protocol. The lit review covers 5, 7, or 9 papers depending on team size, under five suggested sections. The protocol has seven required sections, and its core is a hypothesis you can state. The course supplies a six-step paper-search loop, a rule that AI-assistant output must be quoted, and a worked example: a student's final project that became a Findings of EMNLP paper. The Gradescope format and rubric slides, and past exemplary papers, are behind a login.

CS224U Methods and Metrics I: A Classifier with 0.81 Accuracy and 0.43 Macro F1 — What Classifier and Generation Metrics Each Encode

The CS224U slides compute two numbers from one three-class confusion matrix: accuracy 0.81 and macro F1 0.43. One says the system is good; the other says it gets the two small classes almost entirely wrong. The unit's claim is that different metrics encode different values, and it goes through the bounds, values, and weaknesses of accuracy, the three F-score averages, perplexity, word error rate, and BLEU. Final projects are graded on whether the metrics fit, not on how high the scores are.

CS224U: Writing NLP Papers, Submitting, and Giving Talks

CS224U's 'Presenting your research' lecture has four parts: the course-specific rules for the final paper, how to write an NLP paper, how conference submission works, and how to give a talk. Three things matter most. The final paper must include Known project limitations and an Authorship statement. Write as a Shieber-style 'rational reconstruction,' not a chronological tour of your dead ends. At submission, your title largely decides reviewer bidding. The slides, four videos, and projects.md are public; past example papers need a Stanford login.

Harvard CS50 AI Week 6: Language — N-gram Language Models, TF-IDF QA, Parser & Attention

Week 6 processes natural language: N-gram conditional probability & smoothing, CFG syntax parsing with CYK, TF-IDF vector retrieval, attention mechanism & Transformer basics. Projects: Parser (syntactic generation) and Questions (TF-IDF QA system).

Embeddings: How Models Turn Words Into Computable Vectors

Models don't understand text — they only understand numbers. Embeddings map each token to a vector of several hundred dimensions, where semantically similar words end up close together in vector space. This is the shared foundation behind search, RAG, and classification.

How a Model Knows It's Wrong: Loss Functions and Cross-Entropy

Every time a model predicts the next token, it assigns a probability to every candidate word. A loss function measures how far that probability distribution is from the correct answer — the further off, the higher the loss, the more the model knows it got it wrong. Cross-entropy is the standard formula; perplexity is its human-readable translation.

Tokenization: The BPE Algorithm, and Why Chinese Costs More Than English

Models charge by tokens, not characters. The BPE algorithm starts from individual bytes and repeatedly merges the most frequent adjacent pair to build a vocabulary. English 'understanding' might be 1-2 tokens, but Chinese '理解' could take 2-3 — same meaning, higher cost.

Transformers and Attention: How Models Decide Which Words to Look At

The core of the Transformer is self-attention: for each token, the model computes how relevant every other token is, then takes a weighted sum. This lets the model reach across distance to figure out that 'it' refers to 'cat' not 'mat' — and is the foundation for how it handles long documents.

2021 AI Conference Guide: Natural Language Processing

2021 marked NLP’s shift from fine-tuning an entire model to adapting only a small fraction of its parameters. Prefix-Tuning at ACL, LoRA on arXiv, and Prompt Tuning at EMNLP all appeared that year; ACL Rolling Review launched; and the Findings track established itself as a second publication channel.

2022 AI Conference Guide: Natural Language Processing

2022 marked NLP’s shift from demonstrating model capabilities toward aligning and controlling them. InstructGPT brought RLHF into the mainstream, Chain-of-Thought showed that prompts could unlock reasoning, and Flan 2022 matured instruction-tuning methodology. ACL and NAACL adopted ARR as their sole review path, exposing infrastructure and reviewer-load problems. ChatGPT launched at year-end and rewrote the rules of NLP research.

A Guide to the Top AI Conferences of 2023: Natural Language Processing

2023 was the first full academic year after ChatGPT, and LLMs rewrote the NLP conference agenda. ACL's Best Papers examined humor understanding and the propagation of political bias; an EMNLP Best Paper explained in-context learning through information flow; and the HackAPrompt competition paper also won an EMNLP Best Paper award, signaling that security research had entered the mainstream. The year's largest shift was from asking how to make models more accurate to asking how we can tell when a model is misleading us.

2024 AI Conference Review: Natural Language Processing

NLP conferences redefined themselves under LLM dominance in 2024. ACL made open science its annual theme, and four of its seven Best Papers probed fundamental limits of language models. EMNLP turned toward multilingual and cross-cultural work, with Best Papers spanning speech representations and gradient interpretability. ACL and EMNLP received more than 10,000 submissions combined, but the deeper anxiety was what remains of NLP when LLMs can perform nearly every traditional NLP task.

2025 AI Conference Review: Natural Language Processing

NLP conference submissions nearly doubled in 2025: ACL received 8,360 papers and EMNLP 8,174. China-based first authors exceeded 51% at ACL, and DeepSeek's Native Sparse Attention won Best Paper. The deeper story was an identity crisis: an ACL president said 'ACL is not an AI conference,' a quantitative study asked 'Has ACL Lost Its Crown?', and EMNLP faced questions about what still distinguished it from ACL or NAACL.

CS124 Week 1 Introduction and Setup: Turning Language Problems into Computable Components

CS124 Winter 2026 opens by mapping a ten-week path from tokenization and classification to retrieval, speech, networks, and LLMs, while PA0 establishes the Jupyter environment used throughout the quarter.

CS124 Week 2 Words, Tokens, Edit Distance, and N-grams: Decide What the Model Sees First

Week 2 builds three layers: a token vocabulary with BPE, sequence comparison with dynamic-programming edit distance, and probability approximation with n-grams; PA1 turns regex and BPE into executable work.

CS124 Week 3 Logistic Regression and Text Classification: From Features to Probability and Loss

Week 3 connects text features, sigmoid probabilities, cross-entropy loss, and gradient descent, producing a classifier whose feature contributions remain inspectable.

CS124 Week 4 Information Retrieval: The Indexing and Ranking Layer Beneath RAG

Week 4 builds candidates with an inverted index, ranks them with tf-idf and cosine similarity, and then connects retrieved evidence to generation; PA3 exposes RAG's inspectable retrieval half.

CS124 Week 5 Embeddings and Social NLP: Context Vectors and the Public-Evidence Boundary

Week 5's public materials support the distributional hypothesis, word embeddings, and cosine similarity; the paired Social NLP lecture is unrecorded and restricted, so concrete audit methods are labeled as author extensions.

CS124 Week 6 Neural Networks and LLMs: From Units and Backpropagation to Decoder-Only Models

Week 6 uses public neural-network slides for weighted sums, nonlinearities, loss, and backpropagation, then a public LLM/Transformer deck labeled 2025 for decoder-only architecture without treating it as the 2026 live transcript.

CS124 Week 8 Speech and the PA7/Git Lab: Auditing Information Loss in a TTS-to-STT Pipeline

Week 8 sends text through TTS and back through STT, requiring error classification, formatting-loss analysis, and accent stress tests, while Lab 4 prepares Git collaboration for the team agent project.

CS224N Lecture 3: Matrix Calculus and Backpropagation

Lecture 3 decomposes neural-network training into computation graphs, local derivatives, and the chain rule: the forward pass computes a result; backprop accumulates gradients from the output so every parameter knows how to move.

CS224N Lecture 11: Why LLM Benchmarks Expire

Lecture 11 divides evaluation into what to test, how to measure it, and when the result stops being trustworthy. Benchmarks saturate or leak, prompts change scores, and an LLM judge remains a biased model.

CS224N Lecture 6: Turn a Final Project into a Testable Question

Lecture 6 completes the Transformer picture with encoders, decoders, and cross-attention, then breaks the final project into formats, assessment, research topics, and data. A viable topic needs one explicit baseline and metric.

CS224N Lecture 1: Four Paradigm Shifts in NLP

Winter 2026 Lecture 1 divides NLP into four eras: early exploration, symbolic systems, statistical machine learning, and deep/self-supervised learning. The point is not the dates but how each era redefined the language problem.

CS224N Lecture 4: Language Models, RNNs, and Vanishing Gradients

Lecture 4 defines a language model as a next-word probability distribution, then uses an RNN to compress an arbitrarily long prefix. It also exposes recurrence's central cost: information and gradients travel one time step at a time.

CS224N Lecture 5: From Recurrence to the Transformer

Lecture 5 moves from the long-range and sequential bottlenecks of RNNs to self-attention and the Transformer. It shortens information paths and enables parallel computation, at the price of quadratic attention and separately encoded position.

CS224N Lecture 2: How word2vec Turns Meaning into Vectors

Lecture 2 moves from word2vec's prediction task, objective, and gradients to count-based vectors and evaluation. Meaning becomes a high-dimensional position learned from context, not a label retrieved from a dictionary.

Berkeley CS288 Part 1: From N-grams and Word Representations to Text Classification

The first four units make text countable, representable, and classifiable; A1 then moves from n-grams and perceptrons to an NBOW MLP.

Berkeley CS288 Spring 2026: 18 Slide Units, Three Assignments, and the Limits of Self-Study

CS288 moves from n-grams to RAG, reasoning, and agents through 18 public slide units and three assignments; Berkeley-only recordings make this an A3 materials route, not a public video course.

Berkeley CS288 Part 2: Sequence Models, Seq2Seq, and Transformers

Units 05–07 move from recurrent state to encoder-decoder models, then rewrite the information path with attention and Transformer blocks.

Stanford CS124: Numbered 100, Four Prerequisites Written Into the Catalog, and Not Offered at All Next Year

CS124 is the first course in Stanford's NLP branch. Its textbook is Jurafsky's own Speech and Language Processing, free online, and all nine assignment repos are public. But a banner sits on the course homepage: it will not be taught at all in AY 2026–27. And the chapter numbers the syllabus points at no longer match the August 2026 textbook.

Stanford CS224N: Open the 2019 Syllabus and Transformers Are Still Lecture 14

CS224N has kept every course website since 2000 online. In Winter 2019, Transformers were lecture 14, taught by a guest. In Winter 2026 they are lecture 5, and every lecture after that assumes you already know them. The machine translation assignment is gone; assignment 3 now has you code a decoder-only Transformer from scratch, with pytest suites that run on your laptop.

Stanford CS224U: The Course Site Stopped in Spring 2023, but You Can Clone the Whole Thing

CS224U's teaching material isn't a slide deck — it's an Apache-2.0 GitHub repo holding the lecture notebooks, all three assignments, and the grading document for the final project. But the on-campus course has skipped three straight academic years since Spring 2023, and ExploreCourses briefly put it back on the books for Spring 2026-27, then dropped that section again by 29 September 2026. The official description still lists relation extraction and semantic parsing; the 2023 syllabus covers neither. And the data-loading cell in the first assignment breaks in a fresh environment today, on a Hugging Face compatibility change.

NLP & LLM Interview Guide: From Tokenization to RLHF

The dividing line in LLM interviews is whether you've actually used these things. High-frequency topics: BPE tokenization logic and multilingual challenges, pretraining objectives (CLM vs MLM), three levels of fine-tuning (full/LoRA/prompt tuning), RLHF workflow and failure modes, prompting as engineering practice, and the difficulty of LLM evaluation with current methods.

"Recommend the next route" and "Recommend something similar" are not the same thing — Intent Disambiguation in RAG Recommendation Systems

In a climbing RAG system, 'recommend the next route' (progression) and 'recommend a similar route' (similarity) were conflated by a single hasSimilarRouteIntent() function, causing recommendation quality to collapse. The fix is a two-stage intent classification with a Regex Fast Path + LLM Fallback.