🌏 中文版
This is post 3 of the Reading NTHU Hung-Yu Kao Natural Language Processing series. It follows Word Embeddings and Language Models, which covered how word vectors are trained. This post puts them to a test: do word vectors really learn that king is to queen as man is to woman, and how far behind is a model you train yourself?
The sources are the Fall 2025 Assignment 1 folder in the IKMLab course repo: the handout NLP_HW1_word_emb.pdf, the starter main.ipynb, the processed questions-words.csv, and the TA's walkthrough video (titled "Week 2 Thu. - Assignment 1" and listed in the W2 row of the 2025 schedule). The video is in Mandarin. Access level is A3: the handout, starter code, and data are public. Solutions and grading scripts live on NTU COOL, which outside readers cannot reach.
What the assignment tests
The handout phrases analogy as "A is to B as C is to D." With word vectors you compute B − A + C, find the nearest word, and check whether it is D.
Slide 2 writes "King + Queen − Man ≈ Woman", which is a typo. The hint in the starter code has it right: word_b + word_c - word_a should be close to word_d. For the question king queen man woman, that means queen − king + man ≈ woman.
The data is the Google Word Analogy set (Mikolov et al., 2013). I counted the questions-words.csv in the repo:
| Category | Subcategories | Questions | Example |
|---|---|---|---|
| Semantic | 5 | 8,869 | family: king queen man woman |
| Syntactic | 9 | 10,675 | gram1-adjective-to-adverb: infrequent infrequently cheerful cheerfully |
| Total | 14 | 19,544 |
Subcategory sizes vary a lot: capital-world has 4,524 questions, family only 506. Keep that in mind when you compare which categories a model handles well.
The three parts of the starter notebook
The work happens on Colab, and main.ipynb has three parts.
Part I: preprocessing. You download the raw questions-words.txt with wget. Each line holds four words, and lines starting with : are subcategory headers. A comment notes that the first five headers are semantic and the other nine syntactic. TODO1 turns this into a DataFrame with Question, Category, and SubCategory columns. The TAs ship the processed CSV, but the handout says you still have to write the TODO1 code.
Part II: pretrained embeddings. You load glove-wiki-gigaword-100 through the Gensim downloader. A comment says you may swap in other pretrained models listed by Gensim. TODO2 predicts every answer and keeps the gold labels. The evaluation cell is already written. It prints accuracy for both categories and all 14 subcategories, counting a prediction as correct only if it matches the gold word exactly. TODO3 plots the words in the family subcategory with t-SNE.
Part III: train your own embeddings. Downloading the raw Wikipedia dump takes a long time, and cleaning it with Gensim's WikiCorpus takes longer. So the TAs pre-cleaned it into 11 .txt.gz files that you fetch with gdown. The notebook says each file has 562,365 lines, one article per line (except the last file). Cleaning used WikiCorpus defaults, so single-character words were dropped.
Then:
- TODO4: sample 20% of the articles.
- TODO5: train your own Word2Vec on the sample.
- TODO6 and TODO7: repeat TODO2 and TODO3 with your own model.
Slide 28 suggests preprocessing steps: drop non-English words, remove stop words, lemmatize (rocks → rock), tokenize better than whitespace splitting, and keep only frequent words in the vocabulary. The same slide warns that not every trick helps, so you have to test them.
Grading
Code is worth 55%, split across the seven TODOs:
| TODO | Task | Weight |
|---|---|---|
| 1 | Convert analogy data to a DataFrame | 5% |
| 2 | Answer with pretrained embeddings | 10% |
| 3 | family t-SNE for pretrained embeddings | 5% |
| 4 | Sample 20% of Wikipedia articles | 5% |
| 5 | Train your own embeddings on the sample | 10% |
| 6 | Answer with your own embeddings | 10% |
| 7 | family t-SNE for your own embeddings | 10% |
The report is worth 45%:
- Which embedding model, which preprocessing steps, which hyperparameters (5%)
- Performance when TODO4 samples 5%, 10%, and 20% (10%)
- Performance by category or subcategory when trained on a different corpus (15%): present results, describe your corpus and how it differs from Wikipedia in size, topic, and structure, and explain why performance rose or fell, 5% each
- Pick a few words, retrieve the five most similar words for each, and describe what you see (10%)
- Anything else that strengthens the report (5%)
The second question forces you to plot corpus size against accuracy. The third makes you leave Wikipedia and find your own corpus.
Submission rules
The 2025 version asks for three files zipped and uploaded to NTU COOL:
- Code: the
.pydownloaded from Colab, namedNLP_HW1_school_studentID.py - Packages:
requirements.txt(the example isgensim==4.3.3) - Report: a
.docxfollowing the template
The report must state the runtime environment and Python version. If you use generative AI, say so in both code comments and the report. Link any code you borrowed from the web. Wrong filenames, a missing requirements.txt, or edits to the code template (only data loading may change) cost 5 points each. If your code or report is highly similar to another student's, both lose 100 points. You get three weeks.
What changed in Fall 2026
The 2026 assignment page has only HW1 so far, with a new walkthrough video. I diffed the 2026 handout against 2025 line by line and compared the two notebooks:
- Same tasks: TODO1–7 and their weights, the dataset, the pretrained model, and the Wikipedia files are unchanged.
- New submission format: code is now an
.ipynbthat must keep its outputs for the 20% Wikipedia run. Without outputs, no points go to anything based on results or plots. The report moves from Word into a report cell inside the notebook. - Tweaked report questions: the other-corpus question now says "except wiki", and the extra experiment code must sit in the notebook or the question scores 0. The similar-words question now asks for at least five words. A new "Generative AI Usage" field costs 10 points if left empty.
So if you study with the 2025 materials, you are doing the same assignment as the 2026 class.
Before you start
- Finish Part II before touching Wikipedia. It only needs the GloVe download, and within minutes you have per-category accuracy. That is your baseline for every later comparison.
- Run the full pipeline at 5% first. The report needs 5%, 10%, and 20% anyway. Starting small surfaces preprocessing bugs without waiting for a 20% training run each time.
- Handle out-of-vocabulary words. GloVe's vocabulary differs from yours. How your code handles a question with a missing word directly affects accuracy, so explain it in the report.
- Read t-SNE plots for relations, not clusters. The family questions are pairs like king/queen and man/woman. Good embeddings make the offsets between pairs point in similar directions.
Further reading
- Previous in this series: Word Embeddings and Language Models (n-gram → RNN)
- Next in this series: Seq2seq, LSTM, and Attention
- The same topic in an English-language course: CS224N: Word Vectors
- Back to the series overview
References
- IKMLab/NTHU_Natural_Language_Processing (GitHub repo)
- 2025 assignment index
- 2025 HW1 handout NLP_HW1_word_emb.pdf
- 2025 HW1 starter main.ipynb
- 2025 HW1 walkthrough video (in Mandarin)
- 2026 HW1 handout
- 2026 HW1 walkthrough video (in Mandarin)
- Mikolov et al., 2013, Efficient Estimation of Word Representations in Vector Space
- Gensim Word2Vec docs and pretrained model list
- Gensim WikiCorpus docs
Loading...