Skip to content
All tags

#hardware

20 posts

CMU 11-868 L12–L13 TPU, JAX, and Pallas: One Attention Kernel on TPU, from XLA Fusions to Splash Attention

Across two lectures and more than 200 slides, Google's Srinath Mandalapu traces one attention computation from Python down to TPU VLIW instructions. L12 covers the JAX ecosystem, the memory and compute units of TPU Ironwood, and how XLA compiles attention into three fused kernels. L13 covers what XLA cannot do: using Pallas to control movement between HBM and VMEM yourself, writing FlashAttention, then adding block sparsity to get Splash Attention. The ideas match the GPU version. The difference is that on TPU the compiler does most of the scheduling, and Pallas is how you take loops and block sizes back into your own hands.

CS149 L12: From One Chip to a Whole Datacenter — Dataflow Hardware, Kernel Fusion, Parallelism Strategies, and the Memory Bottleneck

L12 is about moving data. The first part uses the SambaNova SN40L to explain dataflow architecture and metapipelining: running Llama 3.1 8B, the slides say the RDU needs about 3 kernel calls per token versus about 800 on a GPU, because it can fuse an entire decoder into one kernel. The middle part scales up to the datacenter: which collective each of TP, PP, EP, and DP requires, and why overlapping compute with communication decides how well you scale. The last part returns to energy and DRAM: moving a byte costs far more than computing on it, and memory controllers, burst mode, and HBM all attack the same problem. There is no public video for this lecture; this post relies on the slides alone.

CS149 L14 Cache Coherence: MSI, MESI, and False Sharing

When every core has its own cache, one address can have several copies, and different cores can see different values. Locks can't fix this; the hardware created it by replicating data. CS149 L14 defines what coherent means, then takes apart the snooping MSI protocol: before writing, broadcast BusRdX so everyone else invalidates. MESI adds an E state that saves the second transaction in read-then-write, and directories replace broadcast with point-to-point messages. The practical consequence for programmers is false sharing: two threads write different variables, but because they share a cache line, the line bounces between cores. In the lecture's demo it made the program three times slower.

CS149 L10: Why General-Purpose Processors Waste Energy — Hardware Specialization, Tensor Cores, TPU Systolic Arrays, and Dataflow Architectures

L10 starts from one equation: when power is capped, performance can only improve by spending fewer joules per operation, and a general-purpose processor spends most of its energy fetching, decoding, and moving data rather than computing. The slides' rule of thumb is that GPUs give about 10x better perf/watt than CPUs and fixed-function ASICs can reach 100–1000x. The lecture then judges the H100's Tensor Cores and TMA, Google's TPU systolic array, and reconfigurable dataflow architectures against the same checklist: tiled tensors, asynchronous compute and memory, and compute units talking directly to each other.

CS149 L3: Fast Processors, Slow Data — Latency vs. Bandwidth, and How ISPC Separates Abstraction from Implementation

The first half of L3 uses a highway and a laundry room to pull latency and bandwidth apart, then does the math: element-wise vector multiply runs at under 1% efficiency on a V100 because memory cannot feed the ALUs fast enough. The second half is about abstraction vs. implementation. ISPC lets you think in SPMD terms (a gang of program instances, each doing its share), while the compiler implements that with SIMD instructions. Mixing up the two layers is the most common source of confusion in the course.

CS149 L2: Multi-Core, SIMD, and Hardware Multithreading, and the Problem Each One Solves

The second lecture of CS149 Fall 2025 takes a loop that computes sin(x) and adds three ideas in turn: spend transistors on more cores (multi-core), let one instruction drive many ALUs (SIMD), and interleave several threads on one core to hide memory latency (hardware multithreading). The first two add compute; the third keeps that compute busy while waiting on memory. The conclusion is three requirements: enough parallel work, groups of work that run the same instructions, and more parallel work than ALUs so latency can be hidden.

CS149 PA1 + Written 1: Measuring Speedup on a Quad-Core CPU and Explaining Why It Isn't Linear

PA1 has little code and a lot of analysis. Its six programs cover work assignment across threads, SIMD masking, ISPC gangs and tasks, how input data shapes SIMD efficiency, a bandwidth-bound saxpy, and finding a K-Means hotspot with timers. Written 1 drills the same intuitions on paper: peak throughput, instruction dependencies, pipelining, latency hiding with multithreading, and SIMD divergence. Official grading uses Stanford's myth machines; you can run everything on your own hardware, but your numbers won't match the reference.

CS149 PA4 + Written 3: Moving Your Own Data on Trainium2 — NKI, SBUF/PSUM, and a Fused Conv+Maxpool

PA4 drops you onto a single NeuronCore of an AWS Trainium2 chip. No cache decides what stays on chip: you move data into SBUF (28 MiB) and PSUM (2 MiB) yourself with dma_copy, and the partition dimension tops out at 128. Part 1 teaches those limits and the cost of DMA through vector add and transpose. Part 2 asks you to rewrite convolution as a series of matmuls and fuse it with max pooling so nothing spills back to HBM. Written 3 drills the same idea with a line buffer, two back-to-back box blurs, softmax hardware, and metapipelining: keep intermediates on chip. The environment needs the course's private AMI and a paid capacity block, so for outside readers this assignment is effectively A2.

CS149 L11: Programming Specialized Hardware — ThunderKittens Tames H100 Asynchrony, Dataflow Replaces It with Metapipelines

L11 asks what programmers pay once hardware specializes for AI. On the H100, saturating Tensor Cores means 16×16 tiles, TMA moving data asynchronously, and producer and consumer warps running as a pipeline. That's hard to write, which is why DSLs like ThunderKittens exist. The other route is a dataflow architecture (SambaNova SN40L): describe the computation with parallel patterns such as map, reduce, and zip, and let the compiler handle tiling, metapipelining, and placement. The slides say this can fuse an entire Llama 3.1 8B decoder layer into one kernel.

CS149 L15 Memory Consistency: How Write Buffers Make r1 = r2 = 0 Possible

Coherence covers a single address. Memory consistency covers reads and writes to different addresses, and the order in which other threads see them take effect. CS149 L15 uses two threads and two variables to make the point: under sequential consistency, r1 = r2 = 0 is impossible, but the write buffer in every modern processor lets reads pass writes, so it becomes possible. TSO, PSO, and weak ordering relax more orderings in exchange for speed, and fences and synchronization primitives restore the orderings you need. The takeaway for application programmers is short: write data-race-free programs and use a synchronization library, and C11, C++11, and Java 5 guarantee you'll see sequential consistency.

CS149 L17–L18: Transactional Memory and Written 4, Handing "Make This Atomic" to the System

Coarse locks are easy to write but slow; fine-grained locks are fast but easy to get wrong. Transactional memory lets the programmer just declare atomic { } and leaves atomicity and isolation to the system. CS149 L17 covers the motivation (failure atomicity, composability) and the design space: data versioning is eager (undo log) or lazy (write buffer), and conflict detection is pessimistic or optimistic. L18 opens up STM runtime data structures and the McRT algorithm, then shows how HTM uses per-line R/W bits plus the coherence protocol to detect conflicts, ending with Intel Haswell's RTM. Written 4 ties MSI, LL/SC, locks and memory ordering, and fine-grained locking on a doubly linked list into four problems.

CS149 L1: Why Single Cores Stopped Getting Faster, and Why Fast Isn't Efficient

The first lecture of CS149 Fall 2025 defines speedup, then uses three classroom demos to show how communication and load imbalance eat into it. Next it explains why single-core performance stalled: superscalar execution runs out of instruction-level parallelism at about four instructions per clock, and clock frequency hits the power wall. So performance now has to come from more cores and specialized hardware. The last part turns to efficiency. A DRAM access takes about 60 times as long as an L1 cache hit, and moving 64 bits costs over a thousand times the energy of an integer op. Efficiency almost always comes down to accessing data efficiently.

MIT 6.5940 Lecture 4: Per-Layer Pruning Ratios, Fine-Tuning, and Hardware Support for Sparsity

MIT 6.5940 Lecture 4 finishes the pruning unit. Per-layer ratios come from sensitivity analysis, AMC (reinforcement learning), or NetAdapt (step-by-step with a lookup table). Fine-tuning uses 1/10 to 1/100 of the original learning rate, and iterative pruning pushes AlexNet from 5x to 9x. EIE, NVIDIA 2:4 sparsity, and TorchSparse/PointAcc show that sparsity only turns into speed with system support.

MIT 6.5940 Lecture 5: Number Formats, K-Means Quantization, and Linear Quantization

MIT 6.5940 Lecture 5 starts from one fact: an 8-bit integer add uses 30x less energy than a 32-bit float add. It reviews the bit layouts of INT, fixed point, FP32/FP16/BF16, FP8, and FP4, then covers two quantization methods. K-means quantization saves storage only, since computation stays in floating point. Linear quantization, r = S(q − Z), turns matrix multiplication, fully connected layers, and convolutions into integer arithmetic.

Taking Apart Two TTSB Crash Reports: Neither Was the Operator's Fault

Only four drone occurrences have entered Taiwan's official aviation safety statistics — because the threshold is 'substantial damage to a drone over 25 kg.' A hobbyist crash never enters the count. The two published investigation reports are the same model, the same manufacturer, and the same agency, and both probable causes were hardware failures: main rotor servo electrical failure in one, a fractured tail rotor pitch link in the other. In both, the flight control computer and the operator were explicitly cleared.

How to Read a Drone Spec Sheet: Which Lines Regulation Turned Into Boundaries

The three most important lines on a drone spec sheet are exactly the three lines spec sheets don't print. Taiwan's Drone Cybersecurity Testing Specification defines a product 'series' as units whose flight control, communications, and satellite positioning chip modules are all identical — regulation decides whether two drones are the same drone by those three modules, not by looks or endurance. And the weight field isn't a marketing number either: 250 g, 1 kg, 2 kg, 15 kg, and 25 kg are five separate legal thresholds.

Drone Industry Cycles: How the 2016 Bubble Burst, and What's Different This Time

The 2016 consumer drone bubble left specific wreckage: 3D Robotics stopped making hardware, GoPro recalled all 2,500 Karma units six weeks after launch and cut 15% of staff, Parrot cut 35% of its drone workforce, and Lily Robotics collapsed after taking $34M in pre-orders. In 2023 even Skydio — $570M raised — exited consumer, and three years later it is valued at $4.4 billion. This wave runs on a completely different engine, but three things are exactly the same.

The Drone Industry Map: Components, Regulatory Ceilings, and the Non-Chinese Supply Chain Rebuild

The global drone market is roughly US$69B in 2026 (IDTechEx). China holds about 80% of it (CSIS) and DJI over 70% of multi-rotor. The FCC put every foreign-made drone on its Covered List in December 2025; Taiwan's drone output jumped from NT$5.0B to NT$12.9B in one year, and Q1 2026 exports already beat all of 2025. This piece breaks down the five-layer supply chain, the four demand blocks, and the two ceilings holding back scale.

Taiwan's Drone Supply Chain: Where the 267 Companies Are, and Which Layer They're Stuck On

Of the 267 companies the Ministry of Economic Affairs counted, 164 are in northern Taiwan. But geography is not division of labor — Thunder Tiger's published bill of materials shows motors, batteries, frames, and propellers sourced locally, while flight control, comms/GPS, and camera modules go to US, European, and Japanese partners. Exports were only 23% of 2025's NT$12.9B output, and 88.1% of export value sits in the 2–7 kg weight band per Ministry of Finance statistics.

aiguide

2026 Personal AI Hardware Buying Guide: DGX Spark, Mac Studio, MSI AI Edge Compared

Comparing the NVIDIA DGX Spark, Apple Mac Studio M4 Ultra, ASUS Ascent GX10, MSI AI Edge, and more — helping you find the right local inference hardware.