Skip to content

Hung-yi Lee ML 2026 HW8: Spending More Inference Compute — What Voting, Self-Certainty, and DeepConf Each Buy in Accuracy

Sep 30, 20261 min
TL;DRHW8 involves no coding and no code submission. The TAs provide a finished Colab that runs Llama-3.2-1B-Instruct on the first 100 GSM8K questions and compares direct inference, Self-Consistency, Self-Certainty, and DeepConf (Confidence), sampling 16 reasoning traces per method. You read three papers, run the notebook, and answer 20 questions on NTU COOL: 18 about the papers and 2 about the Colab results. The prerequisite is Lecture 7 (Reasoning) of Lee's 2025 course. All questions are printed in hw8.pdf in Chinese and English, and the Colab is publicly downloadable. Only the COOL quiz and grades need an NTU account.

🌏 中文版

This post covers HW8 of NTU Hung-yi Lee's Machine Learning 2026 Spring. It is part 17 of the series Reading NTU Hung-yi Lee Machine Learning 2026 Spring. Official materials: the slides hw8.pdf, the assignment Colab (34 cells), and the TA's walkthrough video. The course page lists it as released 5/15 and due 2026/06/04 23:59 (UTC+8), with TAs 江履方, 陳品睿, 尹廷安, and 林育正. Grades were due by 2026/06/07.

Access rating: A3 minus grading. The slides print all 20 questions in both Chinese and English, and the Colab can be downloaded and run by anyone. The only things you can't get are the NTU COOL quiz itself and the grades. This assignment doesn't use JudgeBoi.

Prerequisite: 2025 Lecture 7 on Reasoning

Page 3 of hw8.pdf asks you to watch Machine Learning in the Age of Generative AI (2025), Lecture 7: How do LLMs like DeepSeek-R1 "think deeply" (Reasoning)? (in Mandarin) first. This term has no new lecture that covers it.

The closest lectures this term are Self-Correction and Self-Improving AI (Part 1). Part 1's certainty-based loss uses how concentrated the output distribution is as a signal to update parameters. HW8's Self-Certainty and DeepConf use the same kind of signal, but to pick an answer, leaving the parameters untouched.

The task: five concepts, three methods

The slides' goal is to learn several test-time scaling methods, run them on an open-source LLM, and look at how accuracy differs. They recommend these papers for an overview:

The Colab: already written, you just run it

The notebook says up front that the code is complete. You only need to run it on a Colab T4, and you should not modify any cells. Treat it as a tutorial. The fixed settings:

ItemSetting
Modelunsloth/Llama-3.2-1B-Instruct, loaded with vLLM and wrapped in deepconf's DeepThinkLLM
DataFirst 100 questions of the GSM8K test split
Samples per method16 reasoning traces
Samplingtemperature 0.7, top-p 0.95, up to 384 new tokens
Answer formatThe model is asked for \boxed{answer}; the extractor also accepts GSM8K's #### answer or the last number

Each GSM8K item is a word problem whose reference solution ends with an answer in the form #### 10. The slides' Metric page is explicit: an answer counts only if the value and the format are both right. Wrong working with a lucky final number, or #### written as ##!!, changes the outcome.

Baseline: direct inference

One trace at temperature 0, answer extracted directly. It's the control for asking whether test-time scaling beats a single generation at all.

Method 1: Self-Consistency

Sample N traces y₁…y_N, extract an answer a_i from each, and return the most frequent:

â = argmax_a Σ_i 𝟙[a_i = a]

This needs no logprobs, so it works with any API.

Method 2: Self-Certainty

Score each trace y_k, where V is the vocabulary size and n is the number of generated tokens:

Self-Certainty(y_k) = −(1 / nV) · Σ_i Σ_j log( V · p(j | x, y_{k,<i}) )

The intuition: the further the distribution is from uniform (the more concentrated it is), the higher the score. Traces are then ranked by score, and the trace at rank r gets weight (N − r + 1)^p in a Borda-style weighted vote. With p = 0 this reduces to plain majority voting. The Colab sets p = 0.3.

One approximation to note: the full vocabulary costs too much VRAM, so the Colab takes only the top 50 logprobs from vLLM and spreads the remaining probability mass evenly over the other tokens.

Method 3: Confidence (DeepConf)

This one isn't hand-coded. The notebook calls DeepThinkLLM.deepthink() from the official deepconf package in offline mode. As the notebook explains:

  • At each position, compute the entropy H_{k,i}, then convert it to a normalized confidence c_{k,i} = 1 − H_{k,i} / log K, where K is the number of token probabilities used. Lower entropy means higher confidence.
  • Aggregate a trace's c values into a trace-level confidence C(y_k). Variants include mean, tail, and bottom-window aggregation.
  • Run confidence-weighted voting: each answer sums the confidence of the traces supporting it, and the largest total wins.

The Colab prefers bottom-window, then tail, then mean confidence-weighted voting, followed by the top-10% filtered variants, and falls back to majority voting only as a last resort.

At the end the notebook prints an accuracy table (accuracy, number correct, average runtime in seconds, average number of valid answers) and plots an accuracy bar chart, correct-vs-wrong counts, and per-question runtime. It reminds you that the accuracy gained from sampling many traces usually costs more inference time.

The quiz: 20 questions, all on NTU COOL

PartQuestionsPoints
Part 1: Paper reading180.5 each
Part 2: Coding20.5 each

The total is 10 points. No code submission; the COOL quiz has unlimited attempts and keeps your best score; no late submissions.

Based on the questions printed in hw8.pdf, the paper questions cover:

  • Self-Consistency (Q1–Q3): the correct order of steps, when it works poorly, and its core idea.
  • Self-Certainty (Q4–Q6): the core concept, experimental findings (including why Borda voting combines certainty ranking with answer frequency), and how it differs from AvgLogP / negative perplexity.
  • DeepConf (Q7–Q9): motivation; the Token, Average Trace, Bottom 10% Group, Tail, and Lowest Group confidence measures; and online thinking with early stopping in DeepConf-low / high.
  • CoT and comparisons (Q10–Q15): how CoT relates to Self-Consistency, when CoT doesn't help, sensible implementation and evaluation practices, a hand calculation of majority vs. confidence-weighted voting, and Best-of-N.
  • Beam search (Q16–Q18): how it differs from Self-Consistency, a hand expansion with beam size 2, and its limits on reasoning tasks.

The two Part 2 questions: screenshot the Colab's accuracy table (Q19), then use it to say which method has the lowest accuracy (Q20).

This post doesn't give answers. Q13 and Q17 are hand calculations you can do with the formulas above.

What outside readers can't get

  • NTU COOL: the quiz needs an NTU account, so you can't see scores or official answers.
  • Walkthrough video: it has no captions on YouTube, so this post does not draw on the video.

Everything else is available. Something you can do tonight: copy the Colab, run it on a T4, and write down each method's accuracy and average runtime. Work out how many times more time each extra percentage point of accuracy costs. Then quiz yourself with the 18 paper questions from hw8.pdf, and go back to the paper for any you can't answer.

Going deeper

  • Papers: the DeepConf arXiv paper is the source for Q7–Q9, and its definitions of the confidence measures are the best use of your time. The Self-Certainty arXiv paper explains why it uses the full distribution rather than just the chosen token.
  • Modify it yourself: the assignment forbids editing the notebook, but outside the assignment you can set N_SAMPLES and each method's budget to 4, 8, and 32 instead of 16, plot accuracy against sample count, and see which method saturates first.
  • Related reading: CME295 LLM Reasoning covers reasoning models and inference-time compute. BrowseConf applies confidence-driven test-time scaling to browsing agents. For RL fundamentals, see Berkeley CS285 policy and value methods.

Series navigation: Series overview | Previous: HW7: Model Merging | Next: Self-Improving AI (Part 2): improving the harness and learning to learn

References