Skip to content
All tags

#bert

6 posts

CMU 10-423 L5: CNNs, Encoder-only Transformers, and ViT — Why a Generative AI Course Starts Images with Understanding

CMU 10-423 Lecture 5 opens the image unit with three models built for understanding. CNNs treat the convolution kernel as parameters to learn. Encoder-only Transformers drop the causal mask so every token sees both sides and train with a masked LM objective, which means they are not generative language models. ViT is nearly BERT with image patches such as 16×16 pixels as input. The slides use a figure from the ViT paper to explain why Transformers reached vision four years after NLP: on small datasets ViT loses to large CNNs, and it only pulls ahead with enough data.

NTHU NLP Guide 8: ELMo, BERT, T5, BART, GPT and the Three Roads to Pretraining

A guide to the BERT and its Family unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). It starts from "I record the record": Word2Vec and GloVe give both records the same vector. ELMo fixes this with a bidirectional LSTM language model. With Transformers, pretraining splits into three roads: encoders (BERT: MLM plus NSP, strong at understanding, weak at generation), encoder-decoders (T5's span corruption, BART's five noise types), and decoders (GPT: pure next-token prediction). The last part covers GPT-3's in-context learning and scaling laws, and why decoders became today's dominant backbone.

NTHU NLP Guide 9: One BERT That Scores and Classifies at Once — The Hugging Face Tutorial and HW3 Multi-Output Learning

A guide to the Hugging Face BERT tutorial and HW3 in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The tutorial walks through binary IMDb sentiment classification: AutoTokenizer, the input_ids / token_type_ids / attention_mask fields, AutoModelForSequenceClassification, and Trainer. HW3 applies the same tools to SemEval 2014 Task 1: one bert-base-uncased with two heads, one regressing a 1–5 relatedness score and one classifying entailment into three classes. You add the two losses and write the training loop yourself, because Trainer is not allowed.

Reading NTU ADL 2025 Fall: BERT and Its Family — From the Polysemy Problem to XLNet, RoBERTa, and mBERT

A static word vector gives "apple" one embedding, whether it means the fruit or the company. The BERT lecture in ADL Fall 2025 starts from that polysemy problem. TagLM feeds language-model features into a tagger, ELMo builds contextual embeddings from a deep bidirectional LSTM, and BERT swaps the LSTM for a Transformer, pre-trained with Masked LM and Next Sentence Prediction; downstream, you add a classifier or tagger on the top layer and fine-tune. The optional BERT Variants slides go one ring further out: Transformer-XL for longer context, XLNet's permutation LM to get both AR and AE benefits, RoBERTa's better data and training recipe, SpanBERT's span masking, and mBERT and XLM for many languages. This is the direct prerequisite for HW1, which uses bert-base-chinese for extractive QA.

Reading NTU ADL 2025 Fall: HW1 Chinese Extractive QA — Finding the Answer Span Among Four Paragraphs

HW1 in ADL Fall 2025 gives a question and four Chinese paragraphs. The model first picks the relevant paragraph (paragraph selection, framed as four-way multiple choice), then marks the answer's start and end inside it (span selection), scored by Exact Match. The spec slides point you straight at Hugging Face's run_swag_no_trainer.py and run_qa_no_trainer.py. The simple baseline uses bert-base-chinese, length 512, effective batch size 2, and learning rate 3e-5, and both stages together take under three hours on an 8GB RTX 3070. The Kaggle leaderboard closed 9/29, and code plus report were due 10/1 on NTU COOL. Outside readers cannot get the Kaggle data or grading, but the task design, baseline settings, and five report questions are all usable for practice.

CME295 Lecture 2: How One Transformer Grew into BERT, GPT, and a Zoo of Attention Variants

CME295 Lecture 2 takes the original Transformer apart and refits it: position information moves from "added to the embedding" to RoPE's "rotate Q and K inside attention"; attention gets cheaper with sliding windows and MQA/GQA; models split into encoder-only, encoder-decoder, and decoder-only families; and the second half dissects BERT's MLM (15% of tokens) and NSP pretraining. The 2026 edition folds all of this into a single "Large Language Models" lecture, and BERT is no longer a syllabus item.