🌏 中文版
This is post 5 of the Reading NTHU Hung-Yu Kao Natural Language Processing series. The previous post covered RNNs, LSTMs, and vanishing gradients. This one is hands-on: show an LSTM a few million arithmetic expressions. Can it learn arithmetic?
The post draws on two sets of material in the IKMLab course repo:
- The W4 Tue TA session: the 62-slide deck pytorch_tutorial_NTHU_NLP.pdf, with a recording titled "Week 4 Tue.[助教課]" (助教課 means TA session; in Mandarin).
- Assignment 2: the folder holds the handout NLP_HW2_arithmetic.pdf, the starter main.ipynb,
arithmetic_train.csv, andarithmetic_eval.csv. There is also a walkthrough video titled "Week 5 Thu. - Assignment 2"; the 2025 schedule lists HW2 in the W5 row.
Access level is A3: the handout, starter code, and full data are public. Solutions and grading scripts are on NTU COOL. Fall 2026 has not released HW2 yet, so everything here is the 2025 version.
The TA session: a toolbox for the assignment
The session starts with setup (Anaconda, conda commands, installing PyTorch for your CUDA version), then uses y = ax² + b to introduce the model, the loss, and the optimizer. Below are the parts that matter for HW2.
Tensors. Several slides cover creating and manipulating tensors, including the difference between view and reshape: view fails on non-contiguous tensors, while reshape copies them first. One slide normalizes a 100×100 tensor both ways: 0.127 seconds element by element, 0.000088 seconds as a matrix operation.
nn.Module. A model is a torch.nn.Module that must define __init__ and forward. The slides stress putting submodules in nn.ModuleList rather than a plain Python list. state_dict(), parameters(), train(), and eval() all recurse into submodules, so layers stored in a list are never registered and never trained.
Autograd. A small example traces step by step how backward() walks the computation graph in reverse and uses dependency counts to order the work. Three ways to turn gradients off are listed: requires_grad = False, with torch.no_grad():, and tensor.detach().
The five-step training loop. Slide 45 maps each step to one line:
optimizer.zero_grad() # 1. clear gradients
output = model(**batch) # 2. feed data to the model
loss = loss_fn(output, ground_truth) # 3. compute loss
loss.backward() # 4. compute gradients
optimizer.step() # 5. update parameters
Dataset and DataLoader. Dataset.__getitem__ returns one sample. The DataLoader groups a batch into a tuple and hands it to a collate function, which turns it into tensors. TODO3 in HW2 is exactly this.
RNN data flow and teacher forcing. Slides 54–57 use the sentence "I love AI !". One-hot vectors pass through an embedding layer into 768-dimensional vectors, then into the RNN. The slides then compare two ways to train:
- Generative training: each step feeds the model's own previous prediction. If the model guesses "eat" early, it ends up learning P(AI | I eat). The error carries forward and training becomes unstable.
- Teacher forcing: each step feeds the ground truth, so the model learns P(love | I), P(AI | I love), and so on. The slides conclude this is more stable.
The last few slides load BERT from Hugging Face, which belongs to post 9.
HW2: what the task looks like
The handout states the idea up front: treat arithmetic expressions as a language, train a sequence generation model with RNNs or LSTMs, and reflect on how much the model really understands about arithmetic.
The dataset:
| Item | Details |
|---|---|
| Training set | 2,369,250 rows |
| Eval set | 263,250 rows |
| Question form | A (+/-/) B (+/-/) C = ?, every number in [0, 50) |
| Operators | +, -, *, parentheses |
| Format | Two CSV columns: src (e.g. 14*(43+20)=) and tgt (e.g. 882) |
I downloaded both CSVs and the row counts match the handout. The handout says each item has "2~3 numbers". In practice nearly every training row has three; only 6,784 have two. Answers can be negative too, as in the eval row 30-(48+13)=,-31.
The model reads "1+1=" one character at a time, generates "2" after it sees "=", then emits <eos> and stops.
Starter code and the six TODOs
main.ipynb already defines the model. CharRNN is an embedding layer, two nn.LSTM layers, and a two-layer fully connected head (ReLU in between) that outputs a probability for each character. The loss is cross entropy and the optimizer is Adam. You fill in:
| TODO | Task | Weight |
|---|---|---|
| 1 | Build the vocabulary: char_to_id and id_to_char, including <pad> and <eos> | 5% |
| 2 | Preprocess each expression into model input and output, ending with <eos> | 5% |
| 3 | Data batching: write the Dataset and DataLoader | 5% |
| 4 | Generation: write the generator, predicting one character at a time until <eos> | 10% |
| 5 | Training: teacher forcing, on GPU | 10% |
| 6 | Evaluation: generate full answers with the generator and compute exact match on the eval set | 10% |
The weights follow the summary table on slide 27, which totals 45%. The TODO2 heading on slide 19 says "10%", which disagrees with that table.
The handout and notebook flag several sticking points:
- Compute loss only after "=". For
1+2-3=0, the notebook's target is/ / / / / 0 <eos>, where each "/" becomes<pad>and the loss ignores<pad>. The model does not need to predict the next character of the expression itself. - Take the last position when generating. Feed the current sequence in full each time and use the last element of the output as the next-token prediction.
- Evaluate with the generator. The handout requires generating each complete answer and comparing it with the gold answer, not scoring teacher-forced outputs.
- Gradient clipping is already in place. The training loop calls
torch.nn.utils.clip_grad_value_to clamp gradients to ±1, which ties back to the exploding gradients from the previous post.
The notebook's hyperparameter table lists batch size 64, embedding and hidden size 256, learning rate 0.001, and 10 epochs, but the code cell below it sets epochs = 2. The two disagree, so state in your report what you actually used.
Grading and report questions
The total has three parts: code 45%, accuracy 10% (higher accuracy earns more; attach a screenshot at the end of the report), and the report 45%. The report questions are worth reading because they all probe whether the model understands arithmetic:
- List your training hyperparameters: learning rate, batch size, hidden size, epochs, and so on (5%)
- What happens to answer quality if you use an RNN or GRU instead of an LSTM, and why? (10%)
- What happens if training uses three-digit numbers but evaluation uses two-digit numbers? (10%)
- If 20% of the training answers are wrong, how does that affect the output? Give examples (10%)
- Why is gradient clipping needed during training? (5%)
- Anything else that strengthens the report (5%)
The handout asks for results as text rather than only images, to make grading easier. Submission rules match HW1: a .py, a requirements.txt, and a .docx, zipped and uploaded to NTU COOL within three weeks. Generative AI use must be disclosed, and plagiarism costs both students 100 points.
Before you start
- Run the whole pipeline on a small slice first. The training set has over two million rows. Take a few tens of thousands, confirm the loss drops and the generator stops, then scale up.
- Write the generator before training. An untrained model can already run
model.generator('1+1='); it just outputs garbage. The starter generator defaults tomax_len=200, so a model that never learned to emit<eos>runs to 200 characters on every question, and evaluating over two hundred thousand eval rows becomes very slow. Get the stopping condition right before you train. - Print the wrong answers. Exact match is a single number. Whether errors cluster in multiplication or addition, large numbers or negatives, is the material for any claim about whether the model understands arithmetic.
- Actually run the RNN/GRU question. Swapping
nn.LSTMfornn.GRUornn.RNNin the starter code is a two-line change, and measured numbers make a stronger answer than reasoning alone.
Further reading
- Previous in this series: Seq2seq, LSTM, and Attention
- Next in this series: Transformers and Self-Attention
- RNNs in an English-language course: CMU 11-785: RNNs, part 1, CMU 11-785: RNNs, part 2
- Back to the series overview
References
- pytorch_tutorial_NTHU_NLP.pdf (W4 TA session slides)
- Fall 2025 W4 Tue TA session recording (in Mandarin)
- 2025 HW2 handout NLP_HW2_arithmetic.pdf
- 2025 HW2 starter main.ipynb
- 2025 HW2 folder (with arithmetic_train.csv and arithmetic_eval.csv)
- 2025 HW2 walkthrough video (in Mandarin)
- 2025 schedule README
- PyTorch official site
Loading...