Skip to content

NTHU NLP HW1: Testing Word Vectors on Google Analogy — Pretrained GloVe vs. Word2Vec Trained on 20% of Wikipedia

Sep 30, 20261 min
TL;DRHW1 tests word vectors on the 19,544 Google Analogy questions (8,869 semantic, 10,675 syntactic). You first answer them with pretrained glove-wiki-gigaword-100 loaded through Gensim, then train your own Word2Vec on a 20% sample of a pre-cleaned Wikipedia dump, and plot t-SNE for the family subcategory both times. Seven TODOs are worth 55%, the report 45%. Fall 2026 keeps the same TODOs but asks for an .ipynb with outputs.

🌏 中文版

This is post 3 of the Reading NTHU Hung-Yu Kao Natural Language Processing series. It follows Word Embeddings and Language Models, which covered how word vectors are trained. This post puts them to a test: do word vectors really learn that king is to queen as man is to woman, and how far behind is a model you train yourself?

The sources are the Fall 2025 Assignment 1 folder in the IKMLab course repo: the handout NLP_HW1_word_emb.pdf, the starter main.ipynb, the processed questions-words.csv, and the TA's walkthrough video (titled "Week 2 Thu. - Assignment 1" and listed in the W2 row of the 2025 schedule). The video is in Mandarin. Access level is A3: the handout, starter code, and data are public. Solutions and grading scripts live on NTU COOL, which outside readers cannot reach.

What the assignment tests

The handout phrases analogy as "A is to B as C is to D." With word vectors you compute B − A + C, find the nearest word, and check whether it is D.

Slide 2 writes "King + Queen − Man ≈ Woman", which is a typo. The hint in the starter code has it right: word_b + word_c - word_a should be close to word_d. For the question king queen man woman, that means queen − king + man ≈ woman.

The data is the Google Word Analogy set (Mikolov et al., 2013). I counted the questions-words.csv in the repo:

CategorySubcategoriesQuestionsExample
Semantic58,869family: king queen man woman
Syntactic910,675gram1-adjective-to-adverb: infrequent infrequently cheerful cheerfully
Total1419,544

Subcategory sizes vary a lot: capital-world has 4,524 questions, family only 506. Keep that in mind when you compare which categories a model handles well.

The three parts of the starter notebook

The work happens on Colab, and main.ipynb has three parts.

Part I: preprocessing. You download the raw questions-words.txt with wget. Each line holds four words, and lines starting with : are subcategory headers. A comment notes that the first five headers are semantic and the other nine syntactic. TODO1 turns this into a DataFrame with Question, Category, and SubCategory columns. The TAs ship the processed CSV, but the handout says you still have to write the TODO1 code.

Part II: pretrained embeddings. You load glove-wiki-gigaword-100 through the Gensim downloader. A comment says you may swap in other pretrained models listed by Gensim. TODO2 predicts every answer and keeps the gold labels. The evaluation cell is already written. It prints accuracy for both categories and all 14 subcategories, counting a prediction as correct only if it matches the gold word exactly. TODO3 plots the words in the family subcategory with t-SNE.

Part III: train your own embeddings. Downloading the raw Wikipedia dump takes a long time, and cleaning it with Gensim's WikiCorpus takes longer. So the TAs pre-cleaned it into 11 .txt.gz files that you fetch with gdown. The notebook says each file has 562,365 lines, one article per line (except the last file). Cleaning used WikiCorpus defaults, so single-character words were dropped.

Then:

  • TODO4: sample 20% of the articles.
  • TODO5: train your own Word2Vec on the sample.
  • TODO6 and TODO7: repeat TODO2 and TODO3 with your own model.

Slide 28 suggests preprocessing steps: drop non-English words, remove stop words, lemmatize (rocks → rock), tokenize better than whitespace splitting, and keep only frequent words in the vocabulary. The same slide warns that not every trick helps, so you have to test them.

Grading

Code is worth 55%, split across the seven TODOs:

TODOTaskWeight
1Convert analogy data to a DataFrame5%
2Answer with pretrained embeddings10%
3family t-SNE for pretrained embeddings5%
4Sample 20% of Wikipedia articles5%
5Train your own embeddings on the sample10%
6Answer with your own embeddings10%
7family t-SNE for your own embeddings10%

The report is worth 45%:

  • Which embedding model, which preprocessing steps, which hyperparameters (5%)
  • Performance when TODO4 samples 5%, 10%, and 20% (10%)
  • Performance by category or subcategory when trained on a different corpus (15%): present results, describe your corpus and how it differs from Wikipedia in size, topic, and structure, and explain why performance rose or fell, 5% each
  • Pick a few words, retrieve the five most similar words for each, and describe what you see (10%)
  • Anything else that strengthens the report (5%)

The second question forces you to plot corpus size against accuracy. The third makes you leave Wikipedia and find your own corpus.

Submission rules

The 2025 version asks for three files zipped and uploaded to NTU COOL:

  • Code: the .py downloaded from Colab, named NLP_HW1_school_studentID.py
  • Packages: requirements.txt (the example is gensim==4.3.3)
  • Report: a .docx following the template

The report must state the runtime environment and Python version. If you use generative AI, say so in both code comments and the report. Link any code you borrowed from the web. Wrong filenames, a missing requirements.txt, or edits to the code template (only data loading may change) cost 5 points each. If your code or report is highly similar to another student's, both lose 100 points. You get three weeks.

What changed in Fall 2026

The 2026 assignment page has only HW1 so far, with a new walkthrough video. I diffed the 2026 handout against 2025 line by line and compared the two notebooks:

  • Same tasks: TODO1–7 and their weights, the dataset, the pretrained model, and the Wikipedia files are unchanged.
  • New submission format: code is now an .ipynb that must keep its outputs for the 20% Wikipedia run. Without outputs, no points go to anything based on results or plots. The report moves from Word into a report cell inside the notebook.
  • Tweaked report questions: the other-corpus question now says "except wiki", and the extra experiment code must sit in the notebook or the question scores 0. The similar-words question now asks for at least five words. A new "Generative AI Usage" field costs 10 points if left empty.

So if you study with the 2025 materials, you are doing the same assignment as the 2026 class.

Before you start

  1. Finish Part II before touching Wikipedia. It only needs the GloVe download, and within minutes you have per-category accuracy. That is your baseline for every later comparison.
  2. Run the full pipeline at 5% first. The report needs 5%, 10%, and 20% anyway. Starting small surfaces preprocessing bugs without waiting for a 20% training run each time.
  3. Handle out-of-vocabulary words. GloVe's vocabulary differs from yours. How your code handles a question with a missing word directly affects accuracy, so explain it in the report.
  4. Read t-SNE plots for relations, not clusters. The family questions are pairs like king/queen and man/woman. Good embeddings make the offsets between pairs point in similar directions.

Further reading

References