CMU 07-280 Lecture 19: Turning Next-token Prediction into Geometry
Lecture 19 builds a minimal next-token model from two embedding matrices, dot-product similarity, softmax, and cross-entropy. Shared vector parameters replace the isolated count cells of an N-gram table.