Skip to content

CS189 Spring 2026 HW1 Guide: Linear Algebra / Calculus / Probability Warm-up + Fashion Coding

Sep 29, 20261 min
TL;DRCS189 Spring 2026 HW1 has three pieces: ten written math warm-ups (a linear system, limits of matrix powers via eigendecomposition, SVD, a matrix that flips an image, partial derivatives, a chain rule over a recursion, and four probability problems including Bayes for cancer screening), plus two public Modal notebooks on Fashion-MNIST. Part 1 drills pandas / Plotly / K-means / an MLP / matrix-based image augmentation / tensor puzzles; Part 2 does price regression, MAE / MSE / R², confusion matrices, and finally a secret test set whose images have been rotated. Due Feb 20; no official solutions.

🌏 中文版

This guide follows CS189 Spring 2026 (Jennifer Listgarten / Alex Dimakis). HW1 was posted in the week of Lec 3 (Jan 27) alongside Discussion 1, and it is due Friday, Feb 20, 11:59 PM, the day after the Lec 7–10 linear regression block ends.

HW1 is titled "AGI, Everywhere, All at Once," but the content is practical. It checks two things:

  1. Whether your linear algebra, calculus, and probability are strong enough for the course.
  2. Whether you can take an image dataset from loading through exploration and training to debugging with pandas, Plotly, scikit-learn, and PyTorch tensors.

This post goes problem by problem: what each one tests and which lecture has the tools. It does not include solutions. HW1 has no public official solutions, and pasting answers would not help a self-learner anyway.

Official materials and scope

The schedule lists three links under HW1:

  • Part 1: Written: a Drive folder with hw1.pdf and the LaTeX template hw1_student.tex
  • Part 1: Coding: the Modal notebook fashion_pt_1, titled "Homework 1.1"
  • Part 2: Coding: the Modal notebook fashion_pt_2, titled "Homework 1.2"

The syllabus says every homework has two parts: Part 1 is a shorter Warmup and Part 2 is the Main Homework, released together with the same deadline. Homeworks have both public and hidden autograder tests. The public ones are sanity checks, such as confirming you entered a number and not a word; the hidden ones check correctness.

There are three Gradescope assignments: the written PDF goes to "HW1 Write-Up," and each notebook is zipped and submitted to "HW1.1 Coding" and "HW1.2 Coding."

Access level: the problem PDF, the LaTeX template, and both notebooks open without a login, so the problems themselves are A3 (see the global AI/CS course map). You cannot get Gradescope, the hidden tests, official solutions, or Ed. Outside readers can work every problem but have to verify correctness themselves.

Written part: ten math warm-ups

There are ten written problems worth 46 points in total according to the PDF, in three topics. You also have to copy out and sign an honor code statement first.

Linear algebra (problems 1–4)

#ProblemWhat it testsWhere the tools are
1System of equations (3 pts)How many solutions a 3×3 linear system has, with reasoningRank and linear dependence, the same idea as "when is XᵀX invertible" in Lec 8
2Asymptotic powers of 2×2 matrices (6 pts)lim Mⁿ for three 2×2 matricesThe hint is eigendecomposition M = PDP⁻¹, so Mⁿ = PDⁿP⁻¹; compare the eigenvalues' magnitudes with 1
3Singular Value Decomposition (4 pts)Find the unit vector x that maximizes ‖Ax‖In the SVD, A maps each right singular vector to a left singular vector scaled by its singular value; which one gets stretched most?
4Image Flipping (6 pts)Flatten a 3×3 image to 1×9 and flip it vertically with a matrix T; then generalize to N×N and to horizontal flipsPermutation matrices. The PDF says this is a warm-up for coding problem 4

Problem 4 deserves extra time. Problem 4 of the Part 1 notebook is all about image augmentation with transformation matrices, and the written version lets you see exactly how a matrix moves pixels on a 3×3 example first.

Calculus (problems 5–6)

  • Problem 5, Partial Derivatives (8 pts): first and second partial derivatives of four two-variable functions. Part (b) asks you to find the critical points and decide which gives a minimum, which needs the second-order condition, i.e. the Hessian.
  • Problem 6, Recursive expression and derivatives (6 pts): given zₙ = wₙ₋₁zₙ₋₁ + bₙ₋₁, compute dzₖ/dzₖ₋₁, dzₙ/dz₁, and dzₙ/db₁.

Problem 6 looks like algebra practice, but it is backpropagation in miniature: along a chain of computations, the derivative of the final output with respect to an early variable is the product of the local derivatives. The same structure returns as full backprop in Lec 17–18 and HW3.

Probability (problems 7–10)

#ProblemWhat it tests
7Conditioned uniform difference (3 pts)X, Y i.i.d. Unif(−1, 1); find P(|X − Y| ≤ 0.5 | XY > 0). The hint says to draw the 2D square and look at areas
8Nearest-neighbor arc length (4 pts)Drop 20 uniform points on the unit circle; find P(D > t) for the arc distance from X₁ to its nearest neighbor, and the expected arc length in degrees
9Cancer screening (3 pts)Sensitivity 90%, false-positive rate 3%, prevalence 0.1%; find the posterior probability of disease given a positive test
10Follower Counts (3 pts)420 people seated randomly in a circle with distinct follower counts; find the expected number who have more followers than both neighbors

Each has one key move. Problem 7 turns a conditional probability into a ratio of areas. Problem 8 translates "D > t" into "none of the other 19 points may fall inside a certain arc." Problem 9 is Bayes' theorem, covered in the probability review in Lec 4. Problem 10 is linearity of expectation with indicator variables.

Pause after you finish problem 9. When prevalence is very low, a positive result from a decent test can mean something quite different from what intuition says. This base-rate issue comes back in another form in Lec 11–12 on evaluating classifiers.

Coding Part 1 (HW1.1): from DataFrames to tensor puzzles

The dataset is Fashion-MNIST: 28×28 grayscale images of clothing in 10 classes, loaded with torchvision. The notebook lists five learning objectives: data manipulation with numpy and pandas, visualization with Plotly, organizing and analyzing datasets, understanding data exploration and preprocessing, and practice with torch tensor operations.

The notebook's own grading table totals 39 points. The structure:

SectionPartsContent
Problem 00aWhy hold the data in a DataFrame
Problem 1: pandas and Plotly1a–1dCheck class balance, groupby(), plot the label distribution, show examples per class
Problem 2: data structure via clustering2a–2cRun K-means directly on pixels, evaluate the clusters, visualize them
Problem 3: training a classifier3a–3eTrain an MLP, write predictions back to the DataFrame, per-class accuracy, confusion matrix, prediction confidence
Problem 4: augmentation with transformation matrices4a–4jHorizontal flip, shift, blur, rotation, the gaps rotation leaves and bilinear interpolation, composing transforms, then augmenting test images and evaluating the classifier
Problem 5: Tensor Puzzles5a–5eUsing only broadcasting, indexing, arithmetic, and comparisons, write sum, outer product, diag, triu, and vstack in one line (< 80 characters)

Problem 2's K-means connects straight back to Lec 4. Problem 4 is the scaled-up version of written problem 4: still moving pixels with a matrix, but now rotated pixels no longer land on the grid, which is why bilinear interpolation comes in.

Problem 5's rules are strict: no torch.sum, view, reshape, or unsqueeze. It forces you to really understand broadcasting: two dimensions line up only when they are equal or one of them is 1. You will lean on this constantly when writing autograd and Transformers later.

To submit, download the notebook as .ipynb and zip it with the trained classifier.joblib.

Coding Part 2 (HW1.2): regression, a secret test set, and distribution shift

Part 2 turns to modeling and debugging. Its objectives are to build and evaluate classifiers, debug and analyze model performance, and explore ways to improve accuracy. Its grading table also totals 39 points.

It starts by redoing the Part 1 MLP pipeline: train_test_split, StandardScaler, MLPClassifier, and train/test accuracy. Then come four sections:

  1. Problem 5: Regression analysis. The task switches from classification to regression: predict each item's price from its image features. You join a price table onto the DataFrames, train a linear regression, compute MAE, MSE, and R², and finally compare the relative price error on misclassified versus correctly classified images. This ties directly into Lec 7–10.
  2. Problem 6: A new test set. Load a "super secret" test set, predict with the original model, look at per-class accuracy and the confusion matrix, and inspect where the errors cluster. Accuracy drops sharply, because the test images have been rotated.
  3. Problem 7: Solution 1, a rotation-invariant classifier. Augment the training data with random rotations and retrain, so the training distribution moves toward the test distribution.
  4. Problem 8: Solution 2, undo the rotations. Train an MLPRegressor to predict each image's rotation angle, rotate the test images back to 0°, and feed them to the original classifier. It is an example of a regression model used as preprocessing.

The last part, Problem 9, is Test Time Augmentation: using only the original MLPClassifier, with no new models trained, apply several transforms to each test image and aggregate the predictions to reach at least 44% accuracy. The notebook warns that the autograder runs your function repeatedly, so a slow one will time out.

To submit, zip the notebook with four .joblib model files.

The whole of Part 2 is about something that happens constantly in practice: the training and test distributions differ. The three solutions stand for three ideas: change the training data, change the test data, or look more than once at inference time.

MANUAL questions: the seven that go in the write-up

The final "Coding" section of the written PDF lists seven headings. Paste the figures or written answers for the notebook questions marked MANUAL under them:

SourceQuestionTitle
HW1.11cVisualizing Label Distribution
HW1.12cVisualizing Clusters
HW1.14gMatrix Multiply Questions
HW1.14jAnalysis of Augmentation Techniques
HW1.25cMAE, MSE, and R-Squared
HW1.26dAnalysis of Class Accuracies and Confusion Matrix
HW1.28dComparing Data Augmentation and Rotation Correction

The official files disagree in one place. The grading table at the top of HW1.2 also marks 6b, 7b, and 8a as "Manual," while marking 5c as not manual, which does not match the write-up template. As a self-learner, just answer all of them in your notes. If you are enrolled, follow the course announcements.

Doing it from outside Berkeley

  1. Write the written part in the LaTeX template. hw1_student.tex already lays out every problem, so filling it in gives you a PDF in the same format enrolled students submit.
  2. You do not have to run the notebooks on Modal. They are standard Jupyter notebooks. The Modal setup steps include creating a volume named cs189 mounted at /mnt/cs189, turning off AI code completion (the notebook says Modal's Tab completion gives incorrect outputs), and stopping the kernel when idle to save credits. You can also download the .ipynb and run it locally, but you will have to sort out data paths and package versions yourself.
  3. No hidden tests means writing your own checks. For the written part, verify with NumPy: approximate the matrix-power limits with a large n, compare against np.linalg.svd, and write a Monte Carlo simulation for each probability problem. For the coding part, sanity-check the numbers: on a class-balanced dataset, per-class accuracies should not be wildly different.
  4. Follow the Tensor Puzzles rules. Forbidden functions will still produce the right output, but then you have not practiced anything.

What to do tonight

  1. Open the folder with hw1.pdf and do M₁ from problem 2, then check it with np.linalg.matrix_power(M, 1000).
  2. Before working problem 9, write down your gut guess; compare after you compute.
  3. Open fashion_pt_1 and finish Problem 1 and Problem 5a, one for pandas and one for broadcasting.

Fall 2026 and further reading

Fall 2026 does not link its homework publicly on the course page, so HW1 cannot be compared yet.

On this site:

Series navigation

References