Skip to content

CS224U Analysis Methods I: Probing Shows You Representations, Feature Attribution Gives You Causal Guarantees

Sep 29, 20261 min
TL;DRThe Analysis methods unit of CS224U (Spring 2023) starts by grading three families of methods on a three-column scorecard. Probing is strong at characterizing representations but can't support causal claims. Integrated gradients only gives you a scalar about each representation, but it satisfies the sensitivity axiom, so it does come with a causal guarantee. This post covers slides 1–40, videos 33–35, and feature_attribution.ipynb, including where the notebook breaks in today's environment.

🌏 中文版

This post is based on the Spring 2023 offering of CS224U. It is part 11 of the Stanford CS224U guide series. It covers the first half of the Analysis methods unit: the overview, probing, and feature attribution. The schedule puts this unit on May 8, 10, and 15, 2023. I used three official sources: slides 1–40 of the Analysis methods in NLP deck (64 slides in total), videos 33–35 of the public playlist, and feature_attribution.ipynb in the course repo.

Access follows the course map definitions: A3 (historical offering). Slides, recordings, and notebooks are all public. What you can't get is Quiz 4 on Canvas and the classroom recordings.

Going one layer beneath behavioral evaluation

The previous post was about black-box testing: does the model look right from the outside? On slide 3, Potts splits evaluation into two kinds. Behavioral covers standard IID, exploratory, hypothesis-driven, challenge, adversarial, and security-oriented tests. Structural covers probing, feature attribution, and interventions. This unit is about the second kind.

He uses an even/odd detector to show why you need to cross over. Model 1 gets four, twenty one, thirty two, thirty six, and sixty three right. Look inside and it is a lookup table for exactly those five strings, with everything else defaulting to odd. So twenty two comes out wrong. Model 2 is smarter: it reads the last token, maps one through nine to even or odd, and still defaults to odd otherwise. This time sixteen breaks it.

In video 33 he says:

"no matter how many inputs we offer this model we will never get a guarantee for every integer string that it will behave as intended. For that kind of guarantee we need to look inside this black box."

He then names three positive guarantees everyone wants: free of pernicious social bias, safe in a given context, approved for a given use. Behavioral testing can show that a model has a problem. It can't show that it doesn't.

The three-column scorecard is the unit's spine

The whole unit fills in one table. The three goals are characterizing representations, making causal inferences, and having a path to improved models. The table below comes from Potts's spoken ratings in videos 33–37. The ratings on the slides are images, so they don't survive text extraction.

MethodCharacterize representationsCausal inferenceImprove models
ProbingStrongNoUnclear (whether multi-task training works is open)
Feature attributionWeak: a single importance scoreYes (with IG)No direct path
Intervention-basedYesYesYes (IIT, see the next post)

Potts is open about his stake. In the video he says the third family is the one he has been most involved in developing and the one he favors. Keep that in mind when reading the table.

Probing: reading a big model's hidden layers with a small model

How it works

Slide 16 breaks probing into four steps:

  1. State a hypothesis about the target model's internal structure, such as "this layer encodes part of speech."
  2. Pick a supervised task that serves as a proxy for that structure, such as a POS tagging dataset.
  3. Pick the place in the model where you think the structure lives.
  4. Train a supervised probe on that site.

In practice, BERT is just a machine that produces vectors. For each sentence, you pull the vector at the chosen site, pair it with a task label, build up an (x, y) dataset, and fit a small linear model on it. The schedule lists Tenney et al. as the probing reading. They probed each BERT layer and found part of speech emerging in the middle layers, dependency parses a bit later, and coreference later still.

A small discrepancy: the schedule says "Tenney et al. 2018," but the link goes to the ACL 2019 paper BERT Rediscovers the Classical NLP Pipeline, and the slide references list Tenney, Das, and Pavlick 2019. I go with the link and the slides.

Problem one: are you reading the target model or training a new one?

Slide 18 is blunt about it. A probe is itself a supervised model whose inputs happen to be the target model's frozen representations. That is hard to tell apart from training a classifier with a particular featurization. A more powerful probe finds more information, but some of that information may just be stored in the probe's own parameters.

Hewitt and Liang 2019 proposed the control task as a fix: same input/output format as the real task, but with labels assigned randomly and then held fixed, for example a random fixed POS tag for each word. Selectivity is the probe's score on the real task minus its score on the control task. Video 34 cites their result: a tiny probe with just two hidden units has the highest selectivity, and probes with many parameters have enough capacity to memorize the data, so their selectivity drops.

What to do: when you probe, report the control-task score and selectivity alongside probe accuracy.

Problem two: probes can't support causal claims

This is the point Potts cares about more. Slides 21–22 use an addition network. It takes three numbers and always outputs their sum correctly. Your hypothesis is that the first two numbers get added into an intermediate variable S1, the third is copied into w, and the output is S1 + w.

You probe L1 and it perfectly encodes the third input z. You probe L2 and it perfectly encodes x + y. The hypothesis looks confirmed, just in a different order. But lay out the weights and the output layer's weight on L2 is 0. L2 has no effect at all on the output. It really does store x + y, and nothing uses it.

A probe saying "this information is here" doesn't mean "the model relies on this information to decide."

A path to better models?

Slide 23 suggests one route: multi-task training, where you train on addition and also require one representation to encode z and another to encode x + y. Potts thinks it's an open question whether that actually induces modularity, and even if it does, you still don't get causal guarantees.

The last probing slide lists work on unsupervised probes: SVCCA, inspecting attention weights, the linear structural probe of Hewitt and Manning 2019, and others. These usually have no parameters of their own, so the "probe is too powerful" problem goes away. They still can't support causal claims.

Feature attribution: what is responsible for this prediction?

Two axioms

Slide 27 lists the methods supported by captum.ai: integrated gradients, gradients, saliency maps, DeepLift, deconvolution, LIME, feature ablation, feature permutation, and more. The unit goes deep only on integrated gradients (IG, Sundararajan et al. 2017). LIME is on the reading list too, but the slides only list it among captum's methods and don't discuss it.

Part of why Potts likes the IG paper is that it uses axioms to pin down what a good attribution should satisfy. He covers two of them:

  • Sensitivity: if two inputs differ only in dimension i and get different predictions, dimension i must get non-zero attribution.
  • Implementation invariance: if two models have identical input/output behavior, their attributions must be identical. Implementation details shouldn't matter.

The counterexample to inputs × gradients

The intuitive baseline is inputs × gradients: take the gradient with respect to a feature and multiply by the feature's value. It generalizes to any neuron in the network.

Slide 31 borrows the IG paper's counterexample. The model is M(x) = 1 − ReLU(1 − x), so M(0) = 0 and M(2) = 1. The input has one dimension and the outputs differ, so sensitivity requires non-zero attribution for that dimension. But inputs × gradients gives 0 at x = 0 (gradient 1 times 0) and 0 at x = 2 (gradient 0 times 2). The axiom is violated.

Slide 30 raises a separate conceptual problem. For a classifier, should attributions be computed with respect to the predicted label or the true label? When the model is accurate, the two barely differ. But you are often analyzing a bad model, and then they diverge. Potts deliberately trains a shallow classifier for a single iteration, and the two options give completely different mean attributions. His conclusion in video 35 is that there is no a priori reason to prefer either one. Be explicit about your assumptions and methods.

How integrated gradients works

The intuition behind IG is to look at counterfactual versions of the input. Pick a baseline (often the all-zeros vector), interpolate a series of points between it and the actual input, compute the gradient at each point, and aggregate. Slide 33 splits this into five steps:

  1. Generate the step vector α = [1, …, m]
  2. Interpolate between the baseline x′ and the actual input x
  3. Compute gradients at each interpolated point
  4. Approximate the integral by averaging
  5. Multiply by (x − x′) to scale back to the original input
Formula (slide 33)

IGi(M, x, x′) = (xi − x′i) · (1/m) · Σk=1..m ∂M(x′ + (k/m)·(x − x′)) / ∂xi

On the same counterexample, IG gives an attribution of about 1 for x = 2 with baseline 0, so the counterexample goes away. Video 35 notes that IG provably satisfies sensitivity.

The practical appeal of IG is that you can attribute with respect to any layer and any neuron in the model. You get some of probing's flexibility plus a causal guarantee. The full example on slides 35–39 uses Hugging Face's cardiffnlp/twitter-roberta-base-sentiment with captum's LayerIntegratedGradients on the embedding layer. The baseline is a same-length sequence of pad tokens that keeps only CLS and SEP. Captum's visualizer then colors the tokens.

The small challenge set uses sentences like "They said it would be great, and they were right." and "…they were wrong." The reporting verb said and right/wrong all get clear attributions. Potts takes this as reassurance that the model is using systematic cues.

His verdict on feature attribution: only an "OK" characterization of representations, since you get a scalar importance score; a causal guarantee, yes; and no direct path from IG to improving models.

Hands-on: running feature_attribution.ipynb today

The notebook goes in this order: two InputXGradients implementations (raw PyTorch and captum), the sensitivity counterexample (the notebook calls this section "selectivity examples"), a shallow classifier on make_classification synthetic data, error analysis for a bag-of-words SST classifier, and the RoBERTa example.

Things you can confirm from the repo:

  • The version string says "CS224u, Stanford, Spring 2022", a year older than the course site.
  • captum is not in requirements.txt. You have to pip install captum yourself, and the notebook itself says it isn't a required install.
  • The SST section needs local data. It reads data/sentiment and uses NLTK's stopwords list, so you have to download both first.
  • get_feature_names() breaks. The SST section calls DictVectorizer.get_feature_names(). requirements.txt only asks for scikit-learn>=1.0.2, and I confirmed on my local scikit-learn 1.9.0 that the method no longer exists. Use get_feature_names_out() instead.
  • ig_reference_implementation only interpolates correctly when the baseline is 0. It computes xx = (base + (k/m)) * (x - base), while standard interpolation is base + (k/m) * (x - base). Every baseline in the notebook happens to be 0, so the outputs are right, but a different baseline would give wrong results.

What to do: if you want a first pass tonight, run only the first two sections (InputXGradients and the sensitivity counterexample). They need only PyTorch and captum and no data downloads, and you'll see inputs × gradients return two zeros while IG returns roughly 1.

What this post can and can't confirm

Confirmed: the schedule, the slide content, the three videos, and the current state of the notebook on GitHub. Not confirmed: what Quiz 4 asks (Canvas requires a login) and any extra classroom discussion in 2023 (the recordings are on Panopto).

Further reading: the site's CS224N series has an interpretability guide covering Been Kim's agentic interpretability, which complements the probing/IG line here.

Series navigation: previous, Compositionality: COGS, ReCOGS, and HW3 | next, Analysis Methods II: causal abstraction, IIT, and DAS

References