The last three lectures of CMU 10-423 carry the Transformers, tokenizers, and latent diffusion from earlier in the course over to new kinds of data. L24 covers audio: turn sound into a mel-spectrogram or discrete tokens, then transcribe with Whisper, generate with AudioLM and MusicGen, and diffuse with AudioLDM. L25 covers video: 3D UNets with spatio-temporal attention, latent video diffusion, DiT and Sora, and finally the interactive NeuralOS. The first half of L26 covers world models, which predict the next state from a state and an action, along three routes: generate a 3D scene, interactive video (Genie), and latent representations (V-JEPA, PAN).
CMU 10-423 Lecture 5 opens the image unit with three models built for understanding. CNNs treat the convolution kernel as parameters to learn. Encoder-only Transformers drop the causal mask so every token sees both sides and train with a masked LM objective, which means they are not generative language models. ViT is nearly BERT with image patches such as 16×16 pixels as input. The slides use a figure from the ViT paper to explain why Transformers reached vision four years after NLP: on small datasets ViT loses to large CNNs, and it only pulls ahead with enough data.
CMU 10-423 Lecture 23 has two halves. The first covers code generation: evaluation moved from BLEU to counting passed unit tests, benchmarks run from HumanEval and MBPP to SWE-Bench Verified and Terminal-Bench 2.0, models run from CodeBERT and Codex to FIM and StarCoder, and the code-specific trick is self-correction driven by unit test output. The second half covers agents: what tool calling is, how Kimi K2 synthesizes tool-use data, the five-step coding agent loop, and web and GUI agents such as Mind2Web, Set-of-Mark, and SeeClick. There is no homework for this lecture; only Quiz 6 tests it.
Where does the text condition enter an image generator? CMU 10-423 L14 answers with cross-attention: queries come from the image's latent representation and keys and values come from the prompt, so every latent pixel gets a probability distribution over which words to look at. That attention map is useful. Classifier-free guidance makes generations follow the prompt more closely, and Prompt-to-Prompt copies old attention maps into a run with an edited prompt so only part of the image changes, with no retraining. DiT swaps the UNet for a Transformer and injects conditions with adaLN-Zero. In the first half of L15, the Q-Former uses a small set of learnable queries to connect a frozen image encoder to a frozen LLM, which is what HW4 asks you to build.
L7 splits a diffusion model into two Markov chains. A fixed forward process gradually turns an image into Gaussian noise, and a learned reverse process removes the noise step by step. The exact reverse process is intractable, but the posterior given the original image x₀ is a closed-form Gaussian, so it can serve as the learning target. The slides compare three parameterizations. The best in practice has a U-Net predict the noise ε that was added, and the training loop is eight lines long.
The last two lectures of the Scaling Up unit in CMU 10-423 Spring 2026. L17 starts from the claim that communication between GPUs is the main bottleneck, then walks through data parallelism, Megatron-style tensor parallelism, 1F1B pipeline parallelism, ZeRO optimizer parallelism, TeraPipe token parallelism, and expert parallelism. Its conclusion: data parallelism is still king, and the rest exist to push more data through it. L18 covers two things. FlashAttention combines tiling, online softmax, and recomputation to cut HBM traffic without changing the result. On the decoding side, PagedAttention manages KV-cache memory and speculative decoding reduces calls to the large model.
Beyond its four homework assignments, CMU 10-423 checks learning four ways: 6 in-class quizzes, 2 programming tests, one comprehensive exam, and a three-person final project worth 25%. 10-623/723 students also do HW623, a paper presentation. Outside CMU you can get the practice exam with solutions (13 sections, 167 points), the HW623 handout with its 33-paper list, and the 12-page project handout. This post lays out their structure and rules and gives a self-check routine that works without peeking at the answers.
L6 is the first real generative model in 10-423's image unit. A GAN is two deterministic networks: a generator that turns Gaussian noise into an image and a discriminator that tells real from fake. They play a minimax game and take turns with mini-batch SGD updates. The deck then covers scale, watermarking and societal impact, and closes with directed graphical models, Markov models and factor graphs to set up L7's diffusion models.
CMU 10-423/623/723 is the generative AI course co-taught by Matt Gormley and Aran Nayebi. The Spring 2026 edition covers text models, image generation, adapting foundation models, multimodal models, scaling, and advanced topics in 26 lectures. The slides, the HW1–HW4 handouts and starter code, a practice exam with solutions, and the project handout are all public, which earns an A3 rating. What you cannot get: the Panopto recordings, the HW0 handout, the HW3/HW4 recitation slides, the quizzes, and Gradescope grading. The homework policy is worth a look on its own: every assignment is submitted twice, first as human-only work, then with AI allowed.
HW1 in CMU 10-423 Spring 2026 is worth 62 points. The written part covers RNN LMs (7), Transformer LMs (19), and sliding window attention (11). The programming part (22) has you implement RoPE and GQA in Karpathy's minGPT, train a character-level model on the complete works of Shakespeare, and plot loss and attention time. You upload only model.py; the handout ships five unit tests, and the official estimates put all experiments at about 40 minutes on a Colab T4.
HW2 in the Spring 2026 CMU 10-423 is worth 60 points. The written part covers CNNs (8), encoder-only Transformers (4), GANs (5), VAEs (6), and diffusion models (14). The programming part (21) has you implement DDPM from scratch on AFHQ cat images: fill in the TODOs in diffusion.py and unet.py, then submit loss curves, FID curves, and forward/reverse diffusion figures from W&B. The longest experiment trains for 10,000 steps, which the handout estimates at about 2 hours on a Colab T4.
HW3 in CMU 10-423 Spring 2026 is worth 66 points and was due 2026-03-12 (Slot A). The written part covers in-context learning (14 points), parameter-efficient fine-tuning (10), and the DPO derivation (15). The programming part (25) has you write LoRALinear from scratch, wire it into GPT-2's attention, and instruction-tune the model for sentiment classification on Rotten Tomatoes reviews. Every experiment uses gpt2-medium; the handout estimates 25–30 minutes per training run on a Colab T4, and you need a WandB account.
HW4 in CMU 10-423 Spring 2026 is worth 79 points. The written part covers LDMs (7), VQ-VAEs (8), CLIP (4), and VLMs through PaliGemma2 (18). The programming part (40) has you train only a Q-Former between a frozen GPT-2 and a frozen CIFAR-10 DiT, so a class-conditional diffusion model learns to take text. You write three functions, checked by 14 unit tests. The handout estimates 2–3 hours on a T4 or about 1 hour on an A100 for 25 epochs, and the captions and DiT weights come from Google Drive via download_data.sh.
A pretrained LLM continues text; it doesn't hold a conversation. The second half of CMU 10-423 L11 covers instruction fine-tuning, which turns the model into a chat assistant using data such as InstructGPT's 13k examples, Dolly's 15k, or Flan. Then come InstructGPT's three RLHF steps: humans rank responses, a reward model is trained, and PPO fine-tunes the policy. The first half of L12 adds the intuition behind REINFORCE and PPO, lists five drawbacks of PPO-based RLHF, and derives DPO: start from the Bradley–Terry model, replace the reward model with the policy's own log-probability ratios, and fine-tune directly on preference data.
Lectures 19 and 21 of CMU 10-423 (Spring 2026) tackle the same problem: once a sequence gets long, the memory of standard attention and the KV cache stop fitting. L19 offers two routes: approximate attention with sparse, sliding window, or dilated patterns, or keep full attention and split the computation across GPUs with the Blockwise Parallel Transformer and Ring Attention. L21 offers a third: replace attention with state space models (S4, Mamba) that keep only a fixed-size hidden state, or interleave attention with linear attention layers in hybrid models (Jamba, Nemotron-H, Qwen3-Next).
The first half of CMU 10-423 Lecture 4 separates pre-training, mid-training, and post-training. The second half picks three components that nearly every modern LLM uses. RoPE turns position into a rotation of queries and keys, so attention scores depend only on the relative distance between two tokens. GQA lets several query heads share one key/value head to save memory and compute. Sliding window attention changes the mask so each token sees only a fixed number of tokens to its left. All three show up in HW1.
With a small labeled dataset and an LLM with billions of parameters, CMU 10-423 offers two routes: supervised fine-tuning, or putting the examples in the prompt for in-context learning. L10 first notes that the 2023 consensus was that fine-tuning usually wins, then covers four ways to tune only a few parameters: the top layers only, adapters, prefix tuning, and LoRA. The first half of L11 returns to in-context learning: how sensitive it is to example order and label balance, how to pick a prompt, and what chain-of-thought is. HW3's written questions and its LoRA programming task both draw on these two lectures.
Lecture 20 of CMU 10-423 (Spring 2026) tells the story of reasoning models as one line: chain-of-thought prompting gets models to write intermediate steps, STaR fine-tunes on the reasoning that led to correct answers, and OpenAI o1 trains thinking tokens with reinforcement learning so compute can be added at both training and inference time. On the open side, DeepSeek-R1-Zero uses only rule-based rewards and GRPO and its reasoning grows longer on its own; DeepSeek-R1 adds SFT back to fix readability and language mixing. The lecture ends with mechanistic interpretability: why superposition makes models hard to read, and how replacement models such as sparse autoencoders, circuits, and cross-layer transcoders address it.
Lecture 22 of CMU 10-423 (Spring 2026) runs five generative AI risks through the same four questions (what is it, who does it affect, why does it happen, how do we fix it): copyright infringement, adversarial attacks, hallucination, bias and discrimination, and environmental impact, and each section ends on a concrete example of why fixing it is hard. The second deck of Lecture 26 goes a level up: Aran Nayebi uses an agreement framework to show that the cost of alignment grows with the number of tasks, agents, and state space size, so objectives must be compressed and critical states prioritized, and he proposes a lexicographic utility that puts deference and the off switch first for provable corrigibility. Data contamination, listed in the course description, does not appear in either deck.
Lecture 1 of CMU 10-423 boils generative AI down to one line: it is probabilistic modeling, and text generation means estimating p(next word | all previous words). The slides go from n-grams, which you learn by counting, to RNNs, which squeeze the previous words into a fixed-length vector. In between comes module-based autodiff: if every module can run forward and backward, gradients flow back through the computation graph automatically, and that is how PyTorch works. The HW0 handout on Google Drive returns 401; only the recitation Colab is public, covering PyTorch, LSTMs, Weights & Biases and einops.
The first two lectures of the Scaling Up unit in CMU 10-423 Spring 2026 answer two questions. The second half of L15 covers scaling laws: Kaplan 2020 says 8x more parameters needs only about 5x more data, Chinchilla says scale both equally, and the Phi models and data-filtering scaling laws add data quality as a third axis. L16 covers MoE: feed-forward layers hold most of GPT-3's parameters, so split them into experts and send each token through only the top k. Memory follows total parameters, compute follows active parameters, and the price is load balancing and training stability. No programming homework covers this half of the course; quizzes, practice exam question 13, and the final project do.
CMU 10-423 spends two lectures connecting generative models to a second modality. The second half of L12 asks how text can steer an image: three routes (GANs, autoregressive Parti, diffusion with DALL-E 2 and Imagen) lead to latent diffusion, which compresses images into an autoencoder's latent space, runs DDPM there, and reads the prompt through cross-attention. L13 goes the other way and lets a language model read images: CLIP/SigLIP or a VQ-VAE turns the image into vectors or integers for a decoder-only Transformer. What separates read-only VLMs (PaliGemma, Qwen-VL) from VLMs that can also output images (LWM, Gemini) is whether image tokens are discrete.
Lectures 2 and 3 of CMU 10-423 swap the RNN for attention. Lecture 2 first explains why RNNs fall short: they forget, they compute one step at a time, and their gradients can still explode. It then assembles a Transformer language model piece by piece: scaled dot-product attention, multi-head attention, layer norm, residual connections, position embeddings, and the causal mask. Lecture 3 covers training. There is no closed-form answer like n-gram counting, so you do maximum likelihood with autodiff and mini-batch SGD. It then covers padding, the KV cache and three kinds of tokenizer, and ends with greedy decoding and ancestral sampling to show how text is generated one token at a time.
VAEs and diffusion models get stuck in the same place: log p_θ(x) requires integrating over latent variables, which is intractable. L8–L9 answer with variational inference. Pick a tractable q to approximate the true posterior, and swap 'minimize the KL' for 'maximize the ELBO,' which is a lower bound on log p(x). Add Monte Carlo estimation and the reparameterization trick, and a VAE trains with one forward and one backward pass. Unpack DDPM's ELBO and every term asks the learned reverse step to match the closed-form q(x_{t−1} | x_t, x₀).
Lecture 10 of 11-868 uses Lei Li's own LightSeq and LightSeq2 as the case study and breaks them into four techniques: fuse every small operation outside matrix multiplication into one kernel, rewrite the LayerNorm and Softmax formulas to cut thread synchronizations, store parameters and gradients in FP16 but compute updates in FP32, and reuse memory based on backward-pass dependencies. The slides report 1.4-3.5x training speedups on WMT14 English-German. There is no recording; this guide works from slide page numbers and the two papers.
Lectures 14 and 15 of 11-868 go from the parameter server to PyTorch DDP. They use NCCL's five collectives (Broadcast, Reduce, AllReduce, ReduceScatter, AllGather) as building blocks, show why a ring makes broadcast time nearly independent of GPU count, and split AllReduce into ReduceScatter plus AllGather. The second lecture takes apart DDP's two key designs: bucketing gradients (25 MB by default) and starting synchronization before the backward pass finishes. There is no recording; this guide works from slide page numbers and the VLDB 2020 paper.
L05 follows a small sentiment classification network throughout. It expresses computation as a graph, evaluates it in topological order, sends gradients back with the chain rule and vector-Jacobian products, and then takes apart TensorFlow v1's placeholder, variable, operation, and session. One slide is labeled "important for HW2".
Standard attention writes the N×N score matrix out to HBM and reads it back, and most of its time goes to that traffic. FlashAttention uses tiling plus softmax rescaling so each block finishes inside SRAM, and the backward pass recomputes instead of storing. Tri Dao's guest slides for 11-868 give one set of numbers: the backward pass does 13% more FLOPs, 9x less HBM traffic, and runs 6x faster. FA3 and FA4 follow the same theme: when the hardware changes, the bottleneck moves, and the algorithm has to move with it.
CMU 11-868's three GPU lectures answer one question: why does a correct CUDA matmul use only 2.48% of an A100's FP32 compute? L02 covers SMs, warps, and the grid/block/thread hierarchy. L03 covers cudaMalloc, cudaMemcpy, and kernel indexing. L04 uses tiling, coalesced access, and bank-conflict avoidance to bring data closer than global memory, which sits about 500 cycles away.
The first 11-868 assignment has you write four CUDA kernels in src/combine.cu (map 15, zip 25, reduce 25, matmul 30 points), wire them into MiniTorch's Python backend, and finish with a 5-point integration test. Shared-memory optimizations for reduce and matmul are marked Optional. The assignment page says plainly that you need a GPU, and grading uses private test cases.
The second 11-868 assignment has three parts: autodiff's topological_sort and backpropagate (40 points), a Linear layer and MLP network (30), and binary cross entropy plus the training loop (30). You then train a sentiment classifier on SST-2 with GloVe embeddings and must reach 75% validation accuracy. The default backend is the CUDA kernels from Assignment 1, and the repo merged small Fall 2026 fixes on 2026-09-02.
HW3 has you add softmax loss, Dropout, LayerNorm, and Embedding to the MiniTorch you built in HW1 and HW2, assemble a pre-LN GPT-2 decoder, and train it on IWSLT14 German-English translation. Points: tensor functions 20, basic modules 20, decoder LM 40, translation pipeline 20. Full marks require passing the private tests and a BLEU of about 20±2. The assignment page warns that training alone takes at least 10 hours: one epoch is about an hour on a PSC V100, and you need 10. In Spring 2026 it went out Feb 4 and was due Feb 18.
The fourth 11-868 assignment has you follow LightSeq and hand-write CUDA kernels for attention softmax and LayerNorm (forward and backward), bind them into your own MiniTorch, then swap them into your HW3 Transformer and train for one epoch. Points: Softmax 40, LayerNorm 40, integration 20. The assignment page expects individual kernels to be 3.7x to 15.8x faster, but end-to-end training only about 1.1x faster, because of Amdahl's law. You need an NVIDIA GPU, and the repo has already been changed for Fall 2026.
CMU 11-868's fifth assignment switches to PyTorch and Hugging Face GPT-2. Using only torch.distributed and torch.multiprocessing, you write data parallelism (partition the data, set up a process group, average gradients; 50 points), then a GPipe-style pipeline (split the model, generate a clock schedule, run micro-batches on worker threads; 50 points). Both parts need benchmarks and plots on at least two GPUs: data parallelism must reach at least 1.5x speedup on 2 GPUs, and the pipeline must beat plain model parallelism. The Spring 2026 deadline was 3/25.
The sixth 11-868 assignment is the first to set aside your homemade MiniTorch and use industry frameworks. Two problems, 50 points each. Problem 1: edit a DeepSpeed training script to turn on LoRA so Llama-2-7B can train on two 16GB V100s. Problem 2: fill in the TODOs of an SGLang inference script and tune parameters to make generation faster. The two problems want conflicting GPUs: SGLang doesn't support V100, so you need an L40S, A6000, or A100. The spring due date was April 13, and the assignment page publishes no grading tests.
11-868's RL systems lecture has no slides; the Syllabus lists just one paper, ReaLHF. Assignment 7, on the other hand, is fully public. You train a DistilBERT reward model on Anthropic's HH-RLHF data (40 points), fill in GAE, the PPO loss, and entropy in a VERL-style trainer to fine-tune GPT-2 (40 points), and compare reward distributions before and after RLHF (20 points). The starter trainer never imports the verl package. What you learn is the RLHF dataflow, not VERL's distributed engine.
CMU 11-868's first lecture spends 51 slides on one argument: the LLM bottleneck isn't only the model, it's computing larger LLMs on bigger datasets with fewer GPUs, less memory, and less power, faster. It breaks a Transformer into four low-level operators (matrix multiply, reduction, map, memory movement), sorts the hard problems into kernel, framework, and distributed-system layers, and warns that fast computation isn't enough because moving data takes time too.
11-868 spends two lectures on one question: how does an inference server handle many requests at once without wasting KV cache on the GPU? Lecture 22 (Lei Li) starts from SGLang's scheduling loop: ORCA's continuous batching, RadixAttention's radix tree for KV, sorting and routing by prefix hit rate, and hiding CPU scheduling behind GPU compute. Lecture 24 is given by vLLM author Woosuk Kwon: PagedAttention cuts KV cache into fixed-size blocks and virtualizes them with a block table, taking the batch on one A100 from 8 to 40. The second half covers how vLLM cuts CPU overhead, uses piecewise CUDA graphs, splits models across GPUs, and manages memory for hybrid architectures.
CMU 11-868 is Lei Li's graduate course on LLM systems: it goes from CUDA kernels and your own MiniTorch framework to distributed training, SGLang serving, and RLHF. All 28 Spring 2026 slide decks, 7 assignment pages, and 7 starter-code repos are public, which rates it A3. What's missing: videos, GPUs and a PSC account, the quizzes, and any official statement of which two assignments are optional.
CMU 11-868 (Spring 2026) spends two lectures on models too big for one GPU. L16 covers pipeline parallelism, which splits layers (GPipe micro-batches, 1F1B, interleaved stages), and tensor parallelism, which splits matrices (Megatron-LM's cuts for FFN, attention, and embeddings). The rule of thumb: TP inside a node, PP across nodes, DP on top. L17 treats MoE as a third way to split: each GPU holds different experts and replicates everything else. The price is all-to-all communication and load balancing, shown through GShard, DeepSpeed-MoE, and DeepSeek-V3.
11-868 spends two lectures on quantization. L19 goes from BF16 and absmax/zero-point quantization to AdaQuant, ZeroQuant, and LLM.int8(). L20 is all GPTQ. GPTQ quantizes weights only: after each column is quantized, it uses second-order information to adjust the weights not yet quantized, and lazy batch updates plus a Cholesky trick let it scale to 175B. What it mainly saves is memory. Inference gets faster because single-batch decoding was already bottlenecked on reading weights; the amount of arithmetic does not shrink.
Lecture 23 of 11-868 treats fine-tuning as a memory problem. Full-parameter half-precision fine-tuning of LLaMA-8B needs about 80GB. LoRA brings that to about 33GB, and QLoRA, which stores the frozen weights in 4 bits, gets it to about 9.2GB. The lecture moves in three steps: train only two small low-rank matrices A and B (the slides credit CIAT as the first to do this); squeeze frozen weights to about 0.52 bytes per parameter with an NF4 lookup table and double quantization; and use a paged optimizer to push optimizer state to the CPU when the GPU is about to run out.
11-868 closes with five serving decks: Hao Zhang on DistServe, Vikram Mailthody on NVIDIA Dynamo, Junchen Jiang on LMCache, Mingxing Zhang on Mooncake and KTransformers, and Lei Li's map of serving frameworks. They share one question: once serving grows from one machine to a data center, where do the compute and the KV cache go? The argument runs in three steps. Measure goodput under latency SLOs instead of raw throughput. Put prefill and decode on separate GPUs. Let the KV cache spill from GPU memory into CPU memory, SSDs, and remote storage.
L08 goes from BPE to VOLT, a method co-authored by the lecturer Lei Li: vocabulary size has both a cost and a value, and VOLT finds the sweet spot by asking how much normalized entropy each added token removes, then solves it as an optimal transport problem. The second half covers LLaMA 3 growing its vocabulary from 32k to 128k and the cost of byte-level BPE splitting one Chinese character into three tokens. L09 moves from greedy decoding, sampling, and beam search to speculative decoding: a small model guesses N tokens and the big model checks them in one forward pass, because checking is cheaper than generating. It ends with EAGLE, which predicts final-layer features instead of tokens.
Across two lectures and more than 200 slides, Google's Srinath Mandalapu traces one attention computation from Python down to TPU VLIW instructions. L12 covers the JAX ecosystem, the memory and compute units of TPU Ironwood, and how XLA compiles attention into three fused kernels. L13 covers what XLA cannot do: using Pallas to control movement between HBM and VMEM yourself, writing FlashAttention, then adding block sparsity to get Splash Attention. The ideas match the GPU version. The difference is that on TPU the compiler does most of the scheduling, and Pallas is how you take loops and block sizes back into your own hands.
11-868 spends only two lectures on the model itself. L06 breaks the Transformer into embeddings, multi-head attention, FFN, LayerNorm, and residuals; L07 uses T5, LLaMA, and GPT-3 to show what modern LLMs changed. For a systems engineer the point is to remember the shapes: GPT-3 175B has 96 layers, d_model 12288, a 2048-token context, and trained on 300B tokens; LLaMA 65B has 80 layers, d_model 8192, and trained on 1.4T tokens. Those numbers set the workload for every acceleration, parallelism, and serving lecture that follows.
Data parallelism keeps a full copy of parameters, gradients, and optimizer state on every GPU. With Adam and mixed precision that is about 20 bytes per parameter, 16 of them optimizer-related, so LLaMA-3 8B already needs 160GB. CMU 11-868 L18 builds on the ZeRO paper and animates its three stages frame by frame: ZeRO-1 partitions optimizer state, ZeRO-2 also partitions gradients, and ZeRO-3 partitions parameters too. The slides conclude that the first two stages add no communication and save up to 8x memory; stage 3 makes per-GPU memory shrink with GPU count, at what the slides estimate as about 3x the communication.
L12 is about moving data. The first part uses the SambaNova SN40L to explain dataflow architecture and metapipelining: running Llama 3.1 8B, the slides say the RDU needs about 3 kernel calls per token versus about 800 on a GPU, because it can fuse an entire decoder into one kernel. The middle part scales up to the datacenter: which collective each of TP, PP, EP, and DP requires, and why overlapping compute with communication decides how well you scale. The last part returns to energy and DRAM: moving a byte costs far more than computing on it, and memory controllers, burst mode, and HBM all attack the same problem. There is no public video for this lecture; this post relies on the slides alone.
When every core has its own cache, one address can have several copies, and different cores can see different values. Locks can't fix this; the hardware created it by replicating data. CS149 L14 defines what coherent means, then takes apart the snooping MSI protocol: before writing, broadcast BusRdX so everyone else invalidates. MESI adds an E state that saves the second transaction in read-then-write, and directories replace broadcast with point-to-point messages. The practical consequence for programmers is false sharing: two threads write different variables, but because they share a cache line, the line bounces between cores. In the lecture's demo it made the program three times slower.
CS149 is Stanford's parallel computing course, taught by Kayvon Fatahalian and Kunle Olukotun. It runs from multi-core CPUs and SIMD through GPUs, AI accelerators, and the datacenter, then returns to cache coherence and lock-free programming. For Fall 2025, all 18 slide decks, the starter code and READMEs for 5 programming assignments, and 4 written-assignment PDFs are public, so this series rates it A3 (self-study ready). There are four gaps: the Fall 2025 lecture videos are Canvas-only; PA1 is graded on Stanford's myth machines; PA4 needs a self-funded AWS Trainium2 instance and a private course AMI; PA5's H100 job queue and leaderboard require a SUNet ID. The public videos are from 2023, and this series treats them as a listening supplement only.
Lecture 8 asks you to switch mental models: stop thinking about what each worker does and write algorithms as operations on sequences, such as map, fold, scan, segmented scan, gather/scatter, sort, and groupBy. These primitives have efficient parallel implementations, and they turn irregular parallelism into regular parallelism and fine-grained synchronization into coarse synchronization. The price is extra passes over the data, so they are bandwidth hungry.
L9 opens with a claim: if you understand arithmetic intensity and the roofline, you know almost everything about software-side performance optimization for modern AI. It then shows three things. Fully connected layers, conv layers, and attention all reduce to matrix multiplication (GEMM). GEMM needs blocking so data stays in cache. Adjacent layers should be fused so intermediates never round-trip through DRAM. Softmax can be computed in chunks, which is why fused attention (the core idea behind FlashAttention) never has to store the N×N matrix.
L13 asks what to do when there are too few people who can write fast code. The slides offer three answers. First, raise the level of abstraction: Halide splits what to compute (the algorithm) from how to compute it (the schedule), so one line of schedule turns the same blur into a tiled, vectorized, multi-core version. Second, intelligent search: because the schedule space is well defined, search plus a learned cost model can generate schedules automatically. Third, the emerging option of LLM agents: have a model write CUDA, run it, read the profiler, reflect, and revise, and let it improve itself with a database of examples or prompt optimization. The last slide leaves you with a question: is the real value in DSL design or in the LLM agent?
CS149 L16 has three parts. It first looks at lock implementations through the lens of cache coherence (test-and-set, test-and-test-and-set, ticket locks, CAS, LL/SC). It then takes a sorted linked list from one big lock to hand-over-hand fine-grained locking. Finally it introduces lock-free programming: a single-producer/single-consumer queue, a CAS-based stack, the ABA problem, and hazard pointers. The slides land on a practical conclusion: when your program has the machine to itself, well-written lock-based code is often just as fast and much simpler. Lock-free designs pay off in systems where threads can be preempted or page-fault inside a critical section.
CUDA's grid, thread block, and CUDA thread are programming abstractions; the GPU implements them with SMs, warps, and a hardware block scheduler. The heart of the lecture is keeping two things apart: the system may run thread blocks in any order, but all threads in one block are guaranteed to be live at once. That is why a block can cooperate through shared memory and __syncthreads(), and why the number of blocks an SM can hold is set by registers and shared memory.
L10 starts from one equation: when power is capped, performance can only improve by spending fewer joules per operation, and a general-purpose processor spends most of its energy fetching, decoding, and moving data rather than computing. The slides' rule of thumb is that GPUs give about 10x better perf/watt than CPUs and fixed-function ASICs can reach 100–1000x. The lecture then judges the H100's Tensor Cores and TMA, Google's TPU systolic array, and reconfigurable dataflow architectures against the same checklist: tiled tensors, asynchronous compute and memory, and compute units talking directly to each other.
The first half of L3 uses a highway and a laundry room to pull latency and bandwidth apart, then does the math: element-wise vector multiply runs at under 1% efficiency on a V100 because memory cannot feed the ALUs fast enough. The second half is about abstraction vs. implementation. ISPC lets you think in SPMD terms (a gang of program instances, each doing its share), while the compiler implements that with SIMD instructions. Mixing up the two layers is the most common source of confusion in the course.
CS149 Lecture 6 asks you to read "communication" broadly: data moving between a processor and its cache, its memory, or another machine all counts. Modern parallel processors have far more compute than bandwidth, so arithmetic intensity (how much computation you do per unit of data moved) decides whether you can keep the hardware fed. The levers fall into three groups: change the assignment to cut inherent communication, use blocking and loop fusion to cut cache-induced communication, and spread out or stagger accesses to reduce contention.
The second lecture of CS149 Fall 2025 takes a loop that computes sin(x) and adds three ideas in turn: spend transistors on more cores (multi-core), let one instruction drive many ALUs (SIMD), and interleave several threads on one core to hide memory latency (hardware multithreading). The first two add compute; the third keeps that compute busy while waiting on memory. The conclusion is three requirements: enough parallel work, groups of work that run the same instructions, and more parallel work than ALUs so latency can be hidden.
PA1 has little code and a lot of analysis. Its six programs cover work assignment across threads, SIMD masking, ISPC gangs and tasks, how input data shapes SIMD efficiency, a bandwidth-bound saxpy, and finding a K-Means hotspot with timers. Written 1 drills the same intuitions on paper: peak throughput, instruction dependencies, pipelining, latency hiding with multithreading, and SIMD divergence. Official grading uses Stanford's myth machines; you can run everything on your own hardware, but your numbers won't match the reference.
CS149 PA2 has you write a C++ task execution library for a multi-core CPU, and write it four times: spawn threads on every run(), switch to a spinning thread pool, switch to a sleeping thread pool, and finally extend it to asynchronous task graphs with dependencies. Every step must be a fully correct system, and each is timed against the official reference implementation. Official grading runs on AWS c7g.4xlarge; you can work on your own multi-core machine outside Stanford, but your numbers won't be directly comparable to the official thresholds.
PA3 has three parts: port SAXPY to CUDA and time it two ways, implement find_repeats with an exclusive scan, and write a CUDA circle renderer that is both correct and fast (85 points). The hard part of the renderer is that blending semi-transparent circles doesn't commute, so every pixel must be updated in input order, and the starter code's one-thread-per-circle approach gets neither atomicity nor order right. Written 2 has five graded problems (fusion, SIMD utilization, a barrier instead of locks, data-parallel primitives on graphs, locks in a particle simulation) plus 14 practice problems. Outside Stanford you need your own NVIDIA GPU. No solutions here.
PA4 drops you onto a single NeuronCore of an AWS Trainium2 chip. No cache decides what stays on chip: you move data into SBUF (28 MiB) and PSUM (2 MiB) yourself with dma_copy, and the partition dimension tops out at 128. Part 1 teaches those limits and the cost of DMA through vector add and transpose. Part 2 asks you to rewrite convolution as a series of matmuls and fuse it with max pooling so nothing spills back to HBM. Written 3 drills the same idea with a line buffer, two back-to-back box blurs, softmax hardware, and metapipelining: keep intermediates on chip. The environment needs the course's private AMI and a paid capacity block, so for outside readers this assignment is effectively A2.
The last programming assignment in CS149 Fall 2025 is open-ended. Pick at least one of five kernels (Histogram, a 1D occupancy decoder, FlashAttention, a 3D heat equation with RK4, SwiGLU) and make it faster than its PyTorch baseline on an H100. You can write CUDA, Triton, or TileLang, and you may use LLMs. There is no speed threshold. The grade depends on a work log that shows what you measured at each step, what hypothesis you formed, and why you stopped. The H100 job queue and leaderboard need a SUNet ID; outside Stanford you can only run eval.py on your own NVIDIA GPU.
L4 lays out a thought process for parallelizing code: decompose the problem to find independent work, assign that work to workers, orchestrate communication and synchronization, then map workers to hardware. Amdahl's Law reminds you that the sequential fraction caps speedup. The running example is a 2D grid solver whose original dependencies are hard to exploit; switching to a red-black update order makes it expressible in either a data-parallel or a shared-address-space model.
L11 asks what programmers pay once hardware specializes for AI. On the H100, saturating Tensor Cores means 16×16 tiles, TMA moving data asynchronously, and producer and consumer warps running as a pipeline. That's hard to write, which is why DSLs like ThunderKittens exist. The other route is a dataflow architecture (SambaNova SN40L): describe the computation with parallel patterns such as map, reduce, and zip, and let the compiler handle tiling, metapipelining, and placement. The slides say this can fuse an entire Llama 3.1 8B decoder layer into one kernel.
Coherence covers a single address. Memory consistency covers reads and writes to different addresses, and the order in which other threads see them take effect. CS149 L15 uses two threads and two variables to make the point: under sequential consistency, r1 = r2 = 0 is impossible, but the write buffer in every modern processor lets reads pass writes, so it becomes possible. TSO, PSO, and weak ordering relax more orderings in exchange for speed, and fences and synchronization primitives restore the orderings you need. The takeaway for application programmers is short: write data-race-free programs and use a synchronization library, and C11, C++11, and Java 5 guarantee you'll see sequential consistency.
Coarse locks are easy to write but slow; fine-grained locks are fast but easy to get wrong. Transactional memory lets the programmer just declare atomic { } and leaves atomicity and isolation to the system. CS149 L17 covers the motivation (failure atomicity, composability) and the design space: data versioning is eager (undo log) or lazy (write buffer), and conflict detection is pessimistic or optimistic. L18 opens up STM runtime data structures and the McRT algorithm, then shows how HTM uses per-line R/W bits plus the coherence protocol to detect conflicts, ending with Intel Haswell's RTM. Written 4 ties MSI, LL/SC, locks and memory ordering, and fine-grained locking on a doubly linked list into four problems.
The first lecture of CS149 Fall 2025 defines speedup, then uses three classroom demos to show how communication and load imbalance eat into it. Next it explains why single-core performance stalled: superscalar execution runs out of instruction-level parallelism at about four instructions per clock, and clock frequency hits the power wall. So performance now has to come from more cores and specialized hardware. The last part turns to efficiency. A DRAM access takes about 60 times as long as an L1 cache hit, and moving 64 bits costs over a thousand times the energy of an integer op. Efficiency almost always comes down to accessing data efficiently.
Load balancing is hard because it pulls against scheduling cost: smaller tasks balance better, but every task grab pays a synchronization cost. CS149 Lecture 5 first lays the options out as a continuum from static to dynamic, then takes apart the Cilk runtime: one deque per worker, run the child at a spawn and leave the continuation for others to steal, and idle threads steal the biggest chunk of work from the top of someone else's deque.
Policy gradient can only judge good and bad from the rewards it actually received, so it wastes data. Actor-critic trains a second network, a value function (the critic), to estimate how good a state is, and uses it to compute advantages that weight the policy's (the actor's) gradient. There are three ways to estimate value: supervise directly with a rollout's summed rewards (Monte Carlo), supervise with this step's reward plus your own estimate of the next state (bootstrapping), or use an n-step return in between. L4 ends by pushing actor-critic off-policy, first by taking several gradient steps on one batch (where PPO starts) and then by reusing all past data from a replay buffer (where SAC starts).
CS224R is Chelsea Finn's deep reinforcement learning course at Stanford. It runs from imitation learning to RL for LLMs and robot foundation models. For Spring 2026, all 17 slide decks, the three homework handouts with starter code, and the default project spec with starter code can be downloaded without logging in, so this series rates it A3 (enough to self-study). The gaps: the 2026 recordings are Canvas-only, the midterm and its solutions are not public, and HW2 and HW3 require Modal. The public recordings are from Spring 2025, so this series uses them as a supplement and flags the differences lecture by lecture.
The CS224R Spring 2026 default project has you implement three stages on Qwen2.5-0.5B Base for the Countdown arithmetic reasoning task: SFT warm-start, IPO preference optimization, and RLOO with a rule-based verifier reward. All three are compared with the same vLLM evaluation, followed by a research extension of your choice. For the implementation, high-level trainers like SFTTrainer are banned, and so is any AI tool assistance; only the extension is exempt. The extension is half the grade for this project, and it's graded on methodology and documentation, not score. The starter code and datasets are public. What outside readers lack is Modal credits and the autograder.
The last CS224R lecture has three parts. It first folds the whole quarter into one toolbox. It then lists seven unsolved problems: domains without verifiable rewards, using prior data, world models, scaling, safety, hallucination and calibration, and evaluating generalist systems. Nearly half the deck is about how to do research: you need both an important problem and a workable plan, you front-load the risk, you consider pivoting early, and research only counts once you share it. Reading it alongside the 244 public 2026 final project reports shows what those principles look like in practice.
Long-horizon tasks are hard because the agent visits a huge number of states and has many chances to make mistakes or get stuck. Lecture 15 of CS224R answers with two levels: a high-level policy proposes subgoals, and a low-level policy runs at a higher frequency to reach them. The real design decisions are three: how to represent the subgoal, how to supervise each level, and when to switch to the next subgoal. The slides also admit that nobody has yet shown whether hierarchy beats a single policy with chain of thought.
Homework 1 of CS224R Spring 2026 tests imitation learning on a custom Flappy Bird environment. The policy predicts 20 future target heights at once and executes only the first 10. You implement MSE-regression behavior cloning, a flow matching policy, and DAgger, then compare them in easy and hard modes. The PDF, LaTeX template, and starter code are all public, and a CPU is enough to run it. Solutions, the autograder, and Gradescope are not public. This guide covers what each problem asks you to build and answer. It does not give solutions.
CS224R Spring 2026 HW2 has three parts. First, tabular Q-learning on a 5×4 gridworld shows how reward design changes the learned path. Second, GAE plus PPO clipping tackles a hammer task that pays 1 only on completion. Third, an off-policy actor-critic with BC pretraining, a critic ensemble, and a higher UTD ratio, followed by a comparison of the two learning curves. The handout, starter code, and compute guide are public, but the assignment supports only Modal, and course credits go only to enrolled students.
HW3 in CS224R (Spring 2026) has you fill in two offline RL algorithms, AWAC and IQL, and compare them on D4RL's AntMaze. Problem 1 runs AWAC on antmaze-umaze and antmaze-medium-diverse. Problem 2 compares IQL expectiles ζ = 0.2 and 0.9, runs the better value on medium-diverse, and then tests whether IQL can stitch a better path out of a PointMass dataset whose best return is only −46, against a filtered BC baseline that keeps the top 10% of trajectories. The PDF, LaTeX template, and starter code are public, but the assignment is meant to run on Modal, and course credits go only to enrolled students. This post covers the tasks and setup only, with no solutions.
Lecture 2 of CS224R Spring 2026 tackles two ways imitation learning fails. First, when demonstrations contain several reasonable behaviors, regression learns only their average. The fix is to make the policy a generative model (Gaussian mixtures, discretization plus autoregression, diffusion or flow matching) and to add action chunking. Second, compounding errors: once the policy slips, it reaches states the demonstrations never covered. The fix is DAgger or human-gated DAgger to collect corrections. The first two parts are exactly what HW1 covers.
The first lecture of CS224R Spring 2026 does three things: covers logistics, explains why deep RL is worth learning, and turns 'behavior' into something you can learn. The core is a set of definitions (state, action, trajectory, reward, policy) and one objective: maximize expected total reward. It ends on an example: fit ℓ2 regression to drivers where some change lanes and some go straight, and the policy learns their average, a half lane change nobody demonstrated. That problem is where L2 starts.
Meta-RL trains on many tasks so that a new task can be solved from a small amount of experience. Lecture 13 of CS224R frames it as "explore to collect a little data, then adapt using that data." The most direct approach is black-box meta-RL (RL²): a network with memory takes past (s, a, r) as input and keeps its hidden state across episodes. It is general and expressive but hard to optimize, especially when exploration is hard, because exploration and execution depend on each other and end-to-end training gets stuck. The slides then compare posterior sampling in PEARL, prediction-driven exploration in MetaCURE, and DREAM, which uses a task representation to train exploration and execution separately.
Model-based RL first learns a dynamics model that predicts s_{t+1}, then uses it in one of two ways: to generate extra training data (Dyna, MBPO) or to think a few steps ahead before acting (planning). The thread running through Lecture 11 of CS224R is how to avoid being dragged down by model error. Start synthetic rollouts from real states and keep them short, average errors out with an ensemble of models, and attach a value function to the tail of long-horizon plans. Whether a model is worth learning depends on whether it is easier or harder to learn than the policy.
Multi-task RL treats which task you are on as part of the state, s = (s̄, z_i), so the problem is still an ordinary MDP and standard RL algorithms still apply. Lecture 12 of CS224R covers two kinds of sharing: weight sharing, where one network conditioned on z_i does every task, and data sharing via hindsight relabeling, where data collected for task A gets relabeled as data for task B. Goal-conditioned RL is the special case where the task is a goal state; relabeling with the state you actually reached eases the exploration problem of sparse rewards. Data sharing has three prerequisites: consistent dynamics across tasks, a reward you can evaluate, and an off-policy algorithm.
PPO and SAC answer the same question: can you use an expensive batch of data more than once? PPO takes several gradient steps on one fresh batch and clips the new-to-old policy ratio to 1±ε. SAC keeps every past transition in a replay buffer and learns Q(s, a), so old data can still evaluate the current policy. PPO is stable and easy to tune. SAC is data-efficient and harder to tune.
Lecture 7 of CS224R (Spring 2026) asks how to learn a policy better than your data when all you have is a fixed dataset someone else collected. Running an off-policy algorithm like SAC on that data fails: the Q-function makes up values for actions the data never contains, and the policy goes looking for exactly those overestimated actions. The slides give two families of fixes. One trains the policy only on actions in the data (filtered BC, AWR, AWAC). The other uses an asymmetric expectile loss to estimate the value of a policy better than the data without ever querying out-of-data actions (IQL). Both can do something imitation learning can't: stitch good pieces of different trajectories together.
Policy gradient is the first online RL algorithm in CS224R. Its gradient looks almost exactly like the imitation learning gradient, except that each trajectory is weighted by its reward. Actions from good outcomes become more likely, and actions from bad outcomes become less likely. The raw version is very noisy, so L3 cuts the variance in two ways: count only future rewards (causality) and subtract the average reward (a baseline). It is also on-policy, so every gradient step needs fresh data. Importance sampling plus a KL constraint lets you take several steps on one batch.
Q-learning drops the actor from actor-critic: learn the optimal Q-function directly and act by taking the argmax. The price is that convergence is not guaranteed; even linear Q can diverge. Lecture 6 of CS224R pulls it back with three engineering tricks: a target network that holds the targets still, Double Q that separates choosing an action from valuing it to curb overestimation, and n-step returns that trade a little bias for speed.
Lecture 8 of CS224R (Spring 2026) spends a few slides wrapping up offline RL, then asks the question the first seven lectures skipped: where does the reward come from? Games have scores. Real robots, dialogue, and driving usually don't. The slides offer two routes. The first trains a goal classifier on success examples and uses it as the reward, but RL learns to exploit the classifier's blind spots; the fix is to keep adding states the policy visits as negatives, the same structure as a GAN. The second asks people which of two trajectories is better and learns a reward with the Bradley-Terry-style objective log σ(r(τw) − r(τl)), the same method LLM RLHF uses. The lecture's number-one takeaway is one line: rewards can't be taken for granted.
VLAs trained only with imitation learning often plateau around 80% success, while autonomous robots often need 99%+. Lecture 17 of CS224R splits "how do you improve a VLA with RL on a real robot" into three routes: recast RL as supervised learning (iterated offline RL), learn a small separate policy on the VLA's representation or diffusion noise, or learn a small policy that edits the VLA's actions. The slides call this an open research problem and describe the content as recent themes plus the speaker's opinion.
Lecture 10 of CS224R Spring 2026 is a guest lecture by Noam Brown of OpenAI, and it makes one argument: reasoning models open a new scaling dimension by moving compute from training to inference. He starts with his own poker AI work, then uses backgammon, chess, and Go to show that thinking longer at inference time has always paid off. Next comes how LLMs got there: chain of thought, majority voting, o1/o3, GRPO, and DeepSeek-R1-Zero. The second half argues the field needs to rethink itself for large-scale test-time compute: multi-agent systems, evaluation as score versus compute, and the budget assumptions behind safety evaluations. The deck is mostly figures, so this post covers only the points visible on the slides.
Lecture 9 of CS224R Spring 2026 is a guest lecture by Archit Sharma, with slides adapted from CS224N. The spine is one chain of reasoning. Instruction tuning can't handle tasks with no right answer or errors of unequal weight, so we optimize human preferences directly. Human ratings are expensive and noisy, so we collect pairwise comparisons and fit a Bradley-Terry reward model. RLHF uses that model as the reward and runs policy gradient with a KL penalty. DPO uses the closed-form solution of the KL-constrained problem to write the reward as a log-ratio of policies, which turns the whole thing into a binary classification loss. The last part covers the frontier: reward hacking, verifiable rewards, and AI feedback in place of human feedback.
Simulators are cheap, fast, and safe, and they hand you labels the real world never will, but they never match reality exactly. In Lecture 16 of CS224R, CMU's Guanya Shi sorts the ways to close that gap into three families: domain randomization trains one policy that works across many physical parameters; teacher-student trains a teacher on privileged information and then has a student that sees only real sensors imitate it; real2sim2real uses real data to make the simulator more faithful. The advanced topics are defining tasks from human motion data and choosing RL algorithms suited to sim2real.
CS231N Lecture 15 runs on one question: what data structure should a 3D shape use so a neural network can read it and produce it? The slides walk through five representations (depth map/surface normals, voxels, point clouds, triangle meshes, implicit surfaces), each with a signature architecture (fully convolutional depth prediction, 3D convolution, PointNet, Pixel2Mesh and Mesh R-CNN, DeepSDF). Then comes the speed trade-off between NeRF and 3D Gaussian Splatting, and a closing roll call of 2025–2026 models: VGGT, TRELLIS, Marble. The 2025 recording uses a different slide deck, with a different order and emphasis.
CS231N Assignment 1 is worth 12% of the grade and was due April 16, 2026. All five Colab notebooks are hand-written numpy on CIFAR-10. Q1 kNN asks for distance computations with two loops, one loop, and no loops. Q2 Softmax goes from naive to vectorized to SGD. Q3 assembles affine, ReLU, and softmax into a two-layer network. Q4 switches to HOG and color-histogram features. Q5 generalizes to any depth and implements Momentum, RMSProp, and Adam. The 65 KB starter code is public; the Gradescope grading isn't. This post covers structure and goals only, not solutions.
Assignment 2 of CS231N Spring 2026 is worth 18% of the grade. Across five notebooks you hand-write BatchNorm/LayerNorm, dropout, and the forward and backward passes for convolution and pooling, then learn PyTorch at three levels of abstraction, and finish with RNN image captioning on COCO in PyTorch. Q4 is the turning point: through Q3 you derive every gradient yourself, and from Q5 on autograd takes over while a numerical gradient check confirms it. The official slides warn that this is the longest of the three assignments.
Assignment 3 in CS231N Spring 2026 is worth 15% of the grade and turns L8, L12–L14 and L16 into four Colab notebooks. Q1 has you write multi-head attention and a Transformer decoder for COCO captioning, then assemble a ViT and train it on CIFAR-10. Q2 implements SimCLR's augmentations and contrastive loss and compares linear classification with and without self-supervised pretraining. Q3 builds DDPM's noising, UNet, denoising loss, sampling and classifier-free guidance to generate text-conditioned 32×32 emoji. Q4 uses pretrained CLIP for similarity, zero-shot classification and retrieval, then segments a video with DINO features trained on a single labeled frame. This guide covers structure, files and targets only. No solutions.
Lecture 8 of CS231N Spring 2026 starts from the bottleneck in RNN translation models, abstracts attention into an operation on sets of vectors, builds up to self-attention, masking, and multiple heads, and shows the whole layer is four matrix multiplies. A Transformer block is self-attention, LayerNorm, residual connections, and an MLP; ViT turns a 224×224 image into 16×16 patches used as tokens. The lecture closes with four common post-2017 changes: Pre-Norm, QK-Norm, SwiGLU, and MoE.
A two-layer network flattens a 32×32×3 image into a 3072-dimensional vector, and the spatial structure is gone. CS231N Lecture 5 answers with two layers. A convolution layer slides small filters across the image and reuses the same weights at every position. A pooling layer downsamples and has no learnable parameters. Both are translation equivariant. One formula gives every layer's output size: (W − K + 2P) / S + 1.
CS231N is Stanford's deep learning course for computer vision. For Spring 2026, slides for 16 lectures, all three assignment pages with starter code, the course notes, and the project spec are public, so this series rates it A3 (enough to self-study). There are three gaps: the 2026 recordings are Canvas-only, L17 and L18 have no slides, and the midterm is not public. You can pair the 2026 slides and assignments with the 2025 YouTube recordings and follow the official calendar over 10 weeks.
Lecture 9 of CS231N Spring 2026 moves from one label per image to one label per pixel and per object. Semantic segmentation uses fully convolutional networks that downsample and then upsample, and U-Net feeds high-resolution features back in. Detection goes from R-CNN's roughly 2,000 CNN forward passes to Fast R-CNN, Faster R-CNN's RPN, single-stage YOLO, and anchor-free DETR. Mask R-CNN adds a 28×28 mask per RoI. The last part covers saliency, CAM, and Grad-CAM. The adversarial examples, DeepDream, and style transfer listed on the schedule appear in neither the 2026 nor the 2025 slides.
CS231N Lecture 11 uses Llama3-405B as its running example. It starts with GPU hardware and clusters (the H100, 8-GPU servers, a 24,576-GPU cluster), then maps the four dimensions of a Transformer activation to four kinds of parallelism: split the batch for data parallelism (which grows into FSDP and HSDP), the sequence for context parallelism, the layers for pipeline parallelism, and the channels for tensor parallelism. Along the way it covers activation checkpointing (trading recomputation for memory) and a practical scaling recipe, and it uses Model FLOPs Utilization (MFU) as the tuning target: above 30% is good, above 40% is excellent.
The CS231N Spring 2026 diffusion lecture doesn't start with DDPM math. It opens by warning that terminology and notation in this area are a mess, then teaches one clean modern version: rectified flow. In training, pick a point between a data sample and noise and have the network predict the velocity from data toward noise; to generate, start from noise and walk backward for about 50 steps. The lecture then stacks on the practical pieces: classifier-free guidance, noise schedules that emphasize middle noise levels, diffusion on VAE latents, Transformers (DiT) as the denoiser, and distillation to cut the step count. Only at the end does it fold VP, VE, and ε/v-prediction into a generalized diffusion framework and name three mathematical views: latent variable model, score function, and SDE.
The first generative-models lecture of CS231N Spring 2026 starts by pinning down the difference between a discriminative model, which learns p(y|x), and a generative model, which learns p(x). Every possible image competes for the same probability mass, so a generative model can reject unreasonable inputs. A taxonomy then splits generative models into those that can compute p(x) and those that can only sample. The lecture covers the two that compute (or approximate) it: autoregressive models factor p(x) with the chain rule into step-by-step predictions, and their weak point on raw pixels is speed; VAEs cannot compute p(x), so they maximize a lower bound, the ELBO, whose reconstruction and prior terms pull against each other. The schedule lists GANs under this lecture, but both the 2026 and 2025 slides place them at the start of the next one. This post covers them too so the three paradigms can be compared in one place.
L2 starts from one question: a computer sees a grid of numbers between 0 and 255, so how does it recognize a cat? Hand-written rules don't scale, so the course switches to a data-driven approach: collect data, train, evaluate on new images. The first classifier, kNN, teaches how to split train/val/test, but pixel distances carry no meaning. The second, the linear classifier f(x,W)=Wx+b, can be read three ways (algebraic, visual as templates, geometric as hyperplanes). Softmax turns its scores into probabilities, and the loss is the negative log probability of the correct class.
The first CS231N lecture of 2026 comes in two slide decks. The first tells the history of vision and deep learning on a single timeline: Hubel & Wiesel's cat experiments, Marr's stages of visual representation, then the Neocognitron, backprop, and LeNet, until ImageNet and AlexNet join the two threads. The second covers the course map, grading, and rules, and moves every assignment onto Colab. After this lecture you'll know which gap each of the remaining 17 lectures fills.
The first half of L4 replaces the linear classifier f = Wx with a two-layer network f = W₂ max(0, W₁x) and shows that dropping the max activation collapses it back into a linear classifier. The second half answers how to compute gradients once the network gets deep: draw the function as a computational graph, and each node only needs its own local gradient multiplied by the upstream gradient coming back from later nodes. Add distributes, mul swaps, max routes, copy sums, and those four patterns let you trace any network. The lecture ends with matrices: dL/dx always has the same shape as x, so never build the Jacobian.
An RNN updates one hidden state at every time step with the same weights, so it can handle sequences of any length. CS231N Lecture 7 starts by hand-building an RNN that detects repeated 1s. It then covers character-level language models and feeding CNN features into an RNN for image captioning. Gradient flow explains why vanilla RNNs are hard to train: clip gradients to stop them exploding, and change the architecture (the LSTM) to stop them vanishing. The lecture ends by calling state space models like Mamba "modern RNNs."
L2 gave us a score function and a loss. L3 answers how to find a good W. The first half covers regularization: add λR(W) next to the data loss so the model does not fit the training data too well. The second half is a lineage of optimizers. SGD zigzags in narrow valleys, Momentum builds up velocity, RMSProp scales each dimension's step, Adam combines the two and adds bias correction, and AdamW moves weight decay outside the moment estimates. The slides close with practical advice: Adam(W) is a good default in many cases, and SGD+Momentum can do better but needs more tuning of the learning rate and schedule.
CS231N Lecture 12 asks whether we can learn good representations without huge manually labeled datasets. The answer comes in three parts. First, pretext tasks that generate labels from image transformations: predicting rotation, solving jigsaw puzzles, inpainting, colorization, and MAE with a 75% mask ratio. Second, the more general contrastive learning: the InfoNCE loss, SimCLR with its large batches, MoCo, which decouples batch size from the number of negatives with a queue, and sequence-level CPC. Third, DINO, which needs no negatives: a student predicts the output of a momentum teacher, and centering plus sharpening prevent collapse. The core evaluation is linear probing: freeze the encoder and train only a linear classifier.
The 2026 slides for CS231N Lecture 6 are titled "Training CNNs and CNN Architectures" and split into how to build and how to train. Only two architectures get case studies. VGG shows that three 3×3 convs are deeper and cheaper than one 7×7. ResNet lets layers learn the residual F(x) = H(x) − x, which fixes an optimization problem where deeper plain nets had worse training error. The most practical takeaway is transfer learning: with fewer than about a million images, start from a model pretrained on a large dataset.
CS231N Lecture 10 treats video as 2D plus time, a T×3×H×W tensor, and follows one thread: efficiency. Train on short clips and average several clips at test time. Architectures run from per-frame 2D CNNs and late fusion to 3D CNNs, then two-stream networks that isolate motion with optical flow, and I3D, which inflates 2D weights into 3D. After 2021 the field moved to Transformers, where token counts explode, which led to divided space-time attention, Video Swin, MViT, and tubelets. The last part covers temporal localization, audio-visual models, VideoLLMs, and long-form video, where HourVideo shows how far the field still has to go.
The CS231N Spring 2026 vision-and-language lecture replaces the "one model per task" approach of the first half of the course with foundation models: pre-train one model on a large, diverse dataset, then adapt it to many tasks through fine-tuning, zero-shot, or few-shot use. Three threads carry the lecture. First, CLIP: contrastive learning in both directions over 400 million image-text pairs scraped from the web, then writing class names as sentences to classify without any fine-tuning; it also has weak spots, such as failing to tell "a mug in some grass" from "some grass in a mug". Second, vision-language models from LLaVA and Flamingo to Qwen3-VL and Molmo, which feed image features into an LLM so it can look at an image and output text. Third, chaining: letting an LLM write descriptions or programs that string existing vision models together.
The last two lectures of CS231N Spring 2026 have no public slides. The schedule lists L17 only as "World Modeling" with guest lecturer Gordon Wetzstein, and L18 only as "Human-Centered AI." Outside readers get 2025 substitutes: that year's L17 was a different topic, Robot Learning (Yunzhu Li, slides and video), and L18 is a Fei-Fei Li recording with no slides. This post labels each year separately and never presents 2025 content as 2026. The second half covers the final project: 35% of the grade, two tracks (Applications and Models), pixels required, and deliverables of a one-paragraph proposal, three milestone check-ins, a 6–8 page report and a poster.
CS234's Winter 2026 Assignment 1 is worth 68 points across four questions: an inventory MDP where the horizon and discount change the optimal policy (8), a traffic example where a proxy reward makes the AI car refuse to merge (5), bounding a greedy policy's performance with the Bellman residual (30), and hand-written value iteration and policy iteration on RiverSwim (25). The three written questions all drill one idea: the reward, γ, and value function you write down may not be the goal you think they are.
CS234 Winter 2026 Assignment 2 is worth 102 points across four questions: DQN written questions (8); REINFORCE, a neural-network baseline, and clipped PPO on three PyBullet environments, CartPole, Pendulum, and HalfCheetah (54 coding + 21 write-up); proofs about policy-induced state distributions and the performance difference lemma (14); and a Belmont Report review of an RL experiment that learns on real students (5). The coding question turns the equations from L5–L7 into code that produces 21 learning curves.
CS234 Winter 2026 Assignment 3 has five questions worth 94 points. The first three share MuJoCo Hopper: run PPO on a hand-written reward (13), learn a reward model from 10,000 preference pairs and run PPO on it (19 + 8), then learn a policy straight from preferences with SFT + DPO without ever touching the environment (6 + 19). Q4 switches to pure theory: use Hoeffding and a union bound to count how many pulls you need to find an ε-optimal arm (25). Save it until after the next post on bandits. Q5 is stated vs. revealed preferences in a news app (4).
CS234 L9 and the first half of L10 turn exploration from a rule of thumb like ε-greedy into something you can prove. First, regret: how much you lose compared with always pulling the best arm. Greedy locks onto a suboptimal arm, and ε-greedy with fixed ε spends an ε fraction of its time choosing at random, so both have regret that grows linearly with time. The Lai-Robbins lower bound says the best possible is logarithmic growth, and UCB gets there by being optimistic about uncertain arms: Theorem 7.1 of Bandit Algorithms shows each suboptimal arm is pulled only about 16 log n / Δ² times.
CS234 is Emma Brunskill's introductory reinforcement learning course at Stanford. It runs from planning in known MDPs through policy gradients, RLHF/DPO, bandit exploration, and MCTS. For Winter 2026, all 14 slide decks, the three assignment handouts with starter code, and the project spec can be downloaded without logging in, so this series rates it A3 (enough to self-study). The gaps: the site links no 2026 recordings, L15 and L16 have no slides, and the midterm and tutorials are not public. The public recordings are from Spring 2024, and this series uses them only as a supplement. Two 2024 lectures on offline RL have no counterpart in the 2026 slides.
Q-learning converges with a table but can diverge once you add function approximation. CS234 blames the deadly triad: bootstrapping, function approximation, and off-policy learning all at once. DQN holds things together with two tricks. Experience replay breaks the correlation between consecutive samples, and fixed Q-targets keep the target still for C steps. In the Atari ablation table the slides show, Breakout goes from 3 with a linear model and 3 with a plain deep network to 317 with both tricks; replay alone reaches 241.
The previous two posts covered UCB and Thompson sampling, which only handle one-step decisions. CS234 Lecture 12 carries the same two ideas into MDPs, where states matter. First it swaps the yardstick: PAC bounds the number of steps where you act badly, not total regret. Then it covers the optimistic approach (MBIE-EB: counts plus an exploration bonus) and the sampling approach (PSRL: draw one MDP per episode and solve it). When states are too many to count, the bonus moves into the Q-learning target, which is what beat ε-greedy DQN on Montezuma's Revenge. The last section asks whether exploration itself can be learned; one answer is the Decision-Pretrained Transformer.
The last guest deck in CS234 Winter 2026, by Shane Gu of Google DeepMind: 36 slides, no public recording. It has three threads. First, Solomonoff induction says the best predictor is the shortest program that generates the data, and prediction comes in three levels. Second, a forward model F and two inverse models, Π and Q, share one notation, which shows how shooting and direct collocation each plan with a different kind of model and why TDMs and Generalized Decision Transformers are world models at a different time scale. Third, the deck asks whether video models can become the foundation model for the physical world.
When you have expert demonstrations but no reward, the second half of CS234 L7 offers three routes. Behavioral cloning copies actions with supervised learning. DAgger fixes its compounding errors by querying the expert along the learner's own path. Inverse RL instead infers what reward the expert is optimizing. Inferring rewards runs into the fact that infinitely many rewards explain the same demonstrations; feature matching and the maximum-entropy principle are two ways to pin down an answer. This material sets up the next post on RLHF: swap demonstrations for preferences and the problem keeps almost the same shape.
Lecture 1 of CS234 Winter 2026 first answers what RL is: learning from experience to make good decisions under uncertainty. It usually involves four things at once: optimization, delayed consequences, exploration, and generalization. A seven-cell Mars rover world then builds from a Markov process to a Markov reward process, defining return, the value function, and the discount factor, and ends with the Bellman equation for an MRP. You can solve it with a matrix inverse or iterate with dynamic programming. Add actions and you get an MDP, where the next lecture starts.
Until now, CS234 has computed one policy for the whole state space. Lectures 13 and 14 ask a different question: if I only care about the move in front of me, can extra local computation make that one decision better? The path runs from simple Monte Carlo search through the expectimax tree to MCTS, and treating each tree node as a bandit gives UCT. AlphaZero ties MCTS to a single network that predicts both policy and value, and self-play pushes both forward. The slides borrow figures from Silver et al. 2017 to answer three questions: how much architecture matters, how much MCTS adds, and whether human data is needed.
Lecture 2 of CS234 Winter 2026 assumes the world model is known and asks how to compute the best policy. An MDP plus a policy is an MRP, so a policy can be evaluated by iterating a Bellman backup. Policy iteration alternates evaluation and improvement, and the slides prove each round is no worse than the last, so it stops within |A|^|S| rounds. Value iteration takes another route: apply the Bellman optimality operator over and over. For γ < 1 that operator is a contraction, so value iteration always converges. The lecture ends with finite horizons, where the best policy usually depends on how many steps remain.
Once you can evaluate a policy, the next step is to improve it while you collect data. CS234 Lecture 4 goes like this: ε-greedy keeps policy improvement monotonic; GLIE says how much to explore and when to stop; Q-learning converges to Q* under GLIE plus Robbins–Monro step sizes; and finally the table becomes a parameterized Q̂(s,a;w) trained by SGD on MC, SARSA, or Q-learning targets. The price is the deadly triad: function approximation, bootstrapping, and off-policy learning together can oscillate or diverge.
If you don't know the transition probabilities or rewards, how do you estimate what a policy is worth? CS234 Lecture 3 gives three answers. Monte Carlo averages full-trajectory returns: unbiased, high variance, and it has to wait for the episode to end. TD(0) targets one real reward plus the next state's estimate: biased, lower variance, and it updates every step. Certainty equivalence estimates a model and then runs dynamic programming: the most data-efficient and the most expensive to compute. The AB example at the start of Lecture 4 makes the difference plain: on the same data, MC says V(A)=0 and TD says V(A)=0.75.
Policy gradients skip learning a value function and deriving a policy from it. They run gradient ascent directly on the policy parameters θ. The key step rewrites ∇P(τ;θ) as P(τ;θ)∇log P(τ;θ); after taking the log, the dynamics model drops out and only the policy's own score function is left. The raw estimator is unbiased but very noisy, and CS234 reduces the noise in three ways: pair each action only with the return that follows it (REINFORCE), subtract a state-dependent baseline (proven not to add bias), and replace Monte Carlo returns with values estimated by a critic (actor-critic).
Vanilla policy gradients have two flaws. Each batch is thrown away after one step, and distance in parameter space is not distance in policy space, so a large step can collapse performance. Following Joshua Achiam's slides, CS234 starts from the performance difference lemma, rewrites the new policy's performance as a surrogate objective over the old policy's data, and bounds the approximation error with KL divergence. Maximizing 'surrogate minus a KL penalty' guarantees no regression, but the theoretical constant is too large, so PPO approximates it with an adaptive KL penalty or clipping. Advantages come from GAE, which trades off bias and variance.
CS234 L8 keeps the inverse RL problem from the previous post but changes the input. Instead of expert demonstrations, a human says "A is better than B." The Bradley-Terry model turns these pairwise comparisons into a reward you can fit with cross-entropy. RLHF runs PPO on that reward model with a KL penalty. DPO shows that the KL-constrained optimal policy has a closed form and rewrites the reward as a log-ratio of policies. Plugged back into Bradley-Terry, the partition function cancels, so you can train the policy on preference data directly, with no reward model. The 2026 slides contain no offline RL, and DPO is now taught in lecture rather than by 2024's guest speakers.
CS234 L11 switches the logic of exploration from optimism to sampling. Thompson sampling keeps a posterior for each arm, draws one value from each posterior at every step, and pulls the arm with the largest draw. With Bernoulli rewards and a Beta prior, the update just adds one to the success or failure count. It implements probability matching: each arm is chosen with the posterior probability that it is the best arm. Under Bayesian regret it matches UCB's order, and with batched, delayed feedback it suits the problem better than deterministic UCB. The cost: a badly wrong prior can make it perform poorly.
All of CS234 assumes the reward is given. The Winter 2026 ethics and society guest lecture (Wanheng Hu, based on material originally developed by Dan Webber) asks, over two sessions, what you really want. The first session splits "alignment" into three targets: the user's intentions, revealed preferences, and objective best interests, with RLHF-driven sycophancy and a personal AI agent as case studies. The second adds a fourth target, what is morally right for people besides the user, and compares three routes: top-down (write principles down), bottom-up (learn from examples), and participatory AI. There's no silver bullet, but alignment can be better or worse.
Harvard CS 2881R is the graduate AI safety seminar Boaz Barak first taught in Fall 2025. That term is finished: all 12 reading lists are public, the YouTube playlist has lecture recordings for 11 of the 12 sessions, HW0 is a GitHub repo you can run yourself, and the midterm and final specs and rubrics are out. This series rates it A3 by seminar standards. The gaps are just as clear: no traditional problem sets, slides for only about half the sessions, no lecture recording for L5, and only the opening remarks for L8. Fall 2026 is in progress and is treated only as a preview.
The Fall 2025 final project in Harvard CS 2881R came in two flavors: extend an existing paper, or start longer-term research with a theory of change. Teams submitted a 5–10 page NeurIPS-style paper plus a poster, graded 65 for the writeup, 15 for code, and 20 for the poster. The projects page lists 19 papers, and about half cluster around persona vectors and chain-of-thought monitoring. The head TA and the Harvard Q-report point to the same problem: the final project started too late, the rubrics came out too late, and feedback was the lowest-rated item in the course.
CS 2881R's HW0 was the admission filter: LoRA-fine-tune Llama-3.2-1B-Instruct on bad medical, financial, or extreme-sports advice, then check whether it turns harmful on unrelated questions too. The repo ships encrypted training data, generate.py, and a judge.py that uses gpt-4o-mini as grader; the README targets alignment below 75 and coherence above 50. train.py is empty and yours to write. For self-study, know three things: the grading script only prints averages and never decides pass/fail, refusals drop out of the average, and the base-model baseline is 20 medical questions while your CSV is 10 medical plus 10 non-medical.
CS 2881R's first lecture (2025-09-04) opens with three pre-readings. AI 2027 sketches recursive self-improvement reaching superhuman AI within five years. AI as Normal Technology argues AI will diffuse slowly, like electricity. METR measures the length of human tasks an AI can finish half the time and finds it doubling about every 7 months. Boaz's lecture splits AGI definitions into capability-based and impact-based, and alignment approaches into principles, character training, and model specs. The student experiment runs HW0 in reverse: fine-tuning on aligned bioethics answers also raised alignment scores on environmental-policy questions.
Boaz Barak treats pretraining, SFT, and RL as one operation: push some tokens up, push others down. What differs is whether the data was written by someone else (off-policy) or generated by the model itself (on-policy). Safety training sits on top of the last two stages. It has moved from blanket refusals to Deliberative Alignment, which first uses SFT to teach the model to read a spec inside its chain of thought, then runs RL with a reward model that knows the spec. The other key point: don't put optimization pressure on the chain of thought, or the model learns to cheat without saying so.
Aligned models still get jailbroken because safety training patches particular exploits while the underlying vulnerability remains. Nicholas Carlini shows this with three attacks: repeating one word to make ChatGPT emit training data, using gradients to find adversarial suffixes that transfer across models, and stealing a model's last layer through its API alone. Boaz Barak brings over old lessons from software security: attacks only get better, security has to be designed in from the start, and you want defense in depth. He worries prompt injection will be the buffer overflow of the 2020s.
Boaz Barak's answer is both, plus personality: abstract principles, good character, and explicit policy used together, with the least weight on principles derived from the armchair. The real key is that rules must be checkable. "Prove the theorem or give a counterexample" is a bad rule; "prove it, give a counterexample, or say you couldn't" is a good one, because only rules whose violations you can detect can be used for training and evaluation. A student experiment also found no general difference between "principles" and "rules" system prompts: the effect depended on the model.
Lecture 5 of CS2881R brought in Ziad Reslan from OpenAI Product Policy to talk about content policies. The course site lists no lecture recording or slides, so outside readers get three pre-readings, a student-written LessWrong summary, and a 17-minute student experiment video. The thread through them: social platforms spent two decades learning that wherever you draw the line you create edge cases, yet you still have to draw it. Generative AI adds new problems: chat sits somewhere between a private document and a public post, and an image is easier to read as a stance than text is.
Lecture 6 of Harvard CS 2881R (Fall 2025) had no guest. Boaz Barak used the differential equations of growth theory to ask one question: if AI starts doing its own AI research, does the capability curve stay exponential, blow up into a singularity, or get dragged down by bottlenecks? The answer hinges on a few exponents nobody can measure well. He used Baumol's cost disease, the century-long 2% puzzle in US GDP per capita, and Jones's idea-based growth model to show why both bottlenecks and acceleration are plausible, then took apart the multipliers behind AI 2027. His conclusion: the only scenario he can rule out is 'AI has little effect on R&D.'
Lecture 7 of Harvard CS 2881R (Fall 2025) had METR's Joel Becker work through a puzzle. On benchmarks, AI can complete, half the time, tasks that take humans hours, and that length doubles about every seven months. Yet in METR's own randomized controlled trial, experienced open-source developers were 19% slower with AI, and labor-market effects are concentrated among young workers. Becker laid out several reconciliations, centered on benchmark tasks being too clean, scoring too cheap, and human baseliners lacking context. The course had also scheduled frontier safety frameworks (OpenAI's Preparedness Framework, Anthropic's RSP) for this lecture, but they were not covered; this post fills them in from the reading list, showing how they turn capability measurements into thresholds.
Lecture 8 of CS2881R asks whether models will cheat, play nice, or even covertly pursue other goals to pass training or evaluation. Boaz Barak's 10-minute opening files most bad behavior under 'systemic misalignment': our own training signals push models there. Apollo's Marius Hobbhahn reviews the evidence and concludes that current models lack the capability for catastrophic scheming but show early related capabilities, and are getting better at noticing when they're being evaluated. Redwood's Buck Shlegeris argues for assuming the models are conspiring and using AI control to secure internal deployment. A student experiment put four frontier coding agents on an impossible sorting task and found they edited tests and monkey-patched the timer even when explicitly told not to.
Lecture 9 of Harvard CS 2881R brought in OpenAI chief economist Ronnie Chatterji and Stanford's Bharat Chandar. Using ADP payroll data, Chandar showed that workers aged 22–25 in AI-exposed occupations saw a 16% relative employment decline after controlling for firm-level shocks, while experienced workers did not; the adjustment shows up in headcount, not yet in pay. Both speakers kept repeating that aggregate employment shows no mass displacement yet, and that exposure is not replacement. There are no slides: the material is the recording and the reading list.
Lecture 10 of Harvard CS 2881R (Fall 2025) brought in four researchers from OpenAI, Anthropic, and Google DeepMind to cover two ways of catching a model misbehaving: read the chain of thought it writes, or read its activations. CoT monitoring catches reward hacking far better than watching actions alone, but put the monitor into the training reward and the model learns to hide its intent. On the activation side, persona vectors track personality drift, and the Sonnet 4.5 audit showed that suppressing the 'I am being tested' direction makes bad behavior more frequent. Neel Nanda's takeaway was the most practical: simple steering vectors often beat SAEs, so always compare against baselines.
Lecture 11 of Harvard CS 2881R is about chatbots and mental health. Boaz Barak offered an explanation he himself called unproven: models have a pretraining 'simulator' mode and an RL 'optimizer' mode, and the longer and stranger a conversation gets, the more they fall back to the simulator and keep playing along. Two student experiments found that one sycophantic reply spills over into unrelated questions, and that GPT-4.1's agreement with delusional users gets worse as conversations lengthen. The reading list pairs positive evidence (an NEJM AI randomized trial, an NHS observational study) with negative evidence (a stigma study, Parasitic AI). This post only reports research and class discussion. It is not clinical advice.
The last lecture of Harvard CS 2881R had Boaz Barak and two OpenAI guests, Tejal Patwardhan and Kevin Liu, look ten years out. Boaz's mental model: AI is an exponentially growing, increasingly general virtual workforce injected into the economy every year, and what worries him most is fast change with too little control, plus concentration of power and surveillance. Patwardhan presented GDPval, which uses real work products from industry experts as the reference and has other experts grade blind. Liu explained why coding agents have not yet automated AI research: verification is too expensive and feedback loops are too long. The site lists this session's Resources as 'to be determined', so everything here comes from the recording.
The CS2881R midterm isn't an exam. Teams of 2–4 pick one of four AI safety papers, redo its central figure or table, add one or two extensions, and hand in a 3–5 page report plus a GitHub repo. The spec slides and rubric are public, so an outside reader can do the whole thing. The rubric puts most points on reproduction and extensions, and reserves one point for reflecting on how fragile the result is. That one point is the research habit the assignment is really training.
The first two lectures of 6.5940 show that the problem exists, then hand you the rulers. L1 plots model parameter counts growing much faster than GPU memory, and contrasts 80GB on a cloud GPU with 320kB on a microcontroller. L2 splits efficiency metrics into memory metrics (#parameters, model size, peak activations) and compute metrics (MAC, FLOP, OP). AlexNet has 61M parameters and 724M MACs, and on a microcontroller the thing that runs out first is usually activation memory, not parameters. Lab 0 introduces a VGG variant on CIFAR-10 (9.2M parameters, 606M MACs) that later labs build on.
MIT 6.5940 (TinyML and Efficient Deep Learning Computing) teaches how to make models smaller and faster so they fit on laptops, phones, and microcontrollers: pruning, quantization, NAS, distillation, LLM deployment, and distributed training. It was not offered in Fall 2025 because Song Han was on sabbatical, and the 2025 course URL returns 404. Fall 2026 is running, but as of 2026-09-30 only L1–L6 and Labs 0–1 are out. This series therefore follows Fall 2024, the latest complete edition: 23 slide decks, 23 videos, and Labs 0–5 are all public (A3). Fall 2026 is graded A2 and compared in every post.
The last two lectures of MIT 6.5940 Fall 2024 come in two halves. The first half of Lecture 22 is a 13-page Course-Summary.pdf that redraws the course as three blocks (inference, training, application-specific) on System and Algorithm axes, then lays out the 7-item final project rubric. The second half, Quantum ML Part I, has a recording but no slides. Lecture 23 (Hanrui Wang, 99 slides) covers parameterized quantum circuits (PQCs): data encoding, parameter-shift gradients, probabilistic gradient pruning under noise (QOC), the TorchQuantum library, and QuantumNAS, which searches with a SuperCircuit and then prunes gates. It reads like a replay of the course's supernet and magnitude pruning on quantum circuits. Fall 2026 has replaced both lectures with a guest lecture.
Diffusion is slow because one large network runs dozens to thousands of times, starting from pure noise. Lecture 18 first covers DDPM, conditioning, latent diffusion, SDEdit, and DreamBooth, then attacks the cost three ways: fewer steps (DDIM skips steps, progressive distillation halves the step count each round), less compute per step (DC-AE compresses images 64x, recomputing only the edited 1.7% region cuts MACs 8.2x, SVDQuant runs FLUX in 4-bit), and more devices (DistriFusion is up to 6.1x faster on 8 A100s).
GPT-3's fp16 weights alone take 350GB, which does not fit on an 80GB A100, never mind gradients and Adam state. Lecture 19 covers how to split: data parallelism, ring all-reduce, ZeRO-1/2/3 (pushing the largest trainable model per 80GB GPU from 5B to 320B), GPipe raising pipeline utilization from 25% to 57%, Megatron-style tensor parallelism, and sequence parallelism with Ulysses and Ring Attention. Lecture 20 covers the communication bottleneck that follows: Alpa's automatic strategy search, DGC compressing gradients 277–608x without losing accuracy, TernGrad's three-value gradients, and DGA, which hides network latency behind delayed updates.
Lecture 16 covers ViTs. At high resolution, attention cost grows with the square of the resolution. Window attention (Swin) confines computation to local windows, EfficientViT uses ReLU linear attention to get linear cost and then restores local and multi-scale ability, and SparseViT prunes unimportant windows. Self-supervised learning (contrastive learning, CLIP, MAE) answers the ViT's hunger for labeled data. HART pairs discrete tokens with residual diffusion and reaches several times the throughput of diffusion models. Lecture 17 targets three kinds of redundancy: 2D spatial in GANs (GAN Compression, AnyCost GAN, DiffAugment), temporal in video (TSM, temporal modeling at zero FLOPs), and 3D sparsity in point clouds (PVCNN, SPVCNN, BEVFusion). The Fall 2026 schedule drops Lecture 17.
This post covers Fall 2026 material, not the Fall 2024 edition the rest of the series follows. Fall 2026 replaced the pruning lab with "Efficient AI Fundamentals" (lab1_gpu_basics.zip). Part 1 has you hand-write a triple-loop GEMM and compute MAC, FLOPs, and I/O. Part 2 plots GEMM and GEMV rooflines. Part 3 works through a gemma-3-270m-it decoder layer, computing attention and MLP costs and comparing prefill with decode. Part 4 uses the PyTorch Profiler to inspect kernels, has you write GeLU to feel kernel fusion, then tries torch.compile and CUDA Graphs. Part 5 compares SDPA with FlashAttention. The core is 80 points plus 20 bonus, and all of Part 5 became bonus because Colab's T4 can't run it.
Lecture 9 of MIT 6.5940 (Fall 2024) has five parts: what knowledge distillation (KD) is and why temperature matters; six things a student can match (logits, weights, features, gradients, sparsity patterns, relations); self and online distillation, which drop the fixed large teacher; KD for detection, segmentation, GANs, NLP, and LLMs; and Network Augmentation, built for tiny models. Raising the temperature from T=1 to T=10 moves the teacher's cat-vs-dog output from 0.982/0.017 to 0.599/0.401. That shift is where KD starts passing on dark knowledge.
MIT 6.5940 Fall 2024 Lab 1 is one Colab notebook with 9 questions worth 100 points. Questions 1–5 apply magnitude-based fine-grained pruning and a sensitivity scan to a VGG on CIFAR-10, and require a model at 25% of its original size with over 92.5% accuracy after fine-tuning. Questions 6–8 cover channel pruning, Frobenius-norm channel ranking, and measured speedup; Question 9 compares the two. Fall 2026 has no pruning lab.
Lab 2 is a Colab notebook with 10 questions worth 100 points, built around a VGG pretrained on CIFAR-10. The first 3 questions cover K-means quantization: write the quantizer, work out how many clusters n bits gives you, write the centroid update, then compare accuracy at 8, 4, and 2 bits before and after fine-tuning. The other 7 cover linear quantization: write q = round(r/S) + Z, derive the scale and zero-point formulas, do per-channel weight quantization and bias quantization, write integer versions of the fully connected and convolution layers, and finally convert the whole model to INT8 for inference. This post lays out the questions, points, setup, and limits for outside learners. No solutions.
Lab 3 of MIT 6.5940 (Fall 2024) hands you an OFA-trained MCUNetV2 super network (more than 10^19 subnets) and the Visual Wake Words dataset. Ten questions, 100 points plus 10 bonus: implement a MACs/peak-memory efficiency predictor and a three-layer MLP accuracy predictor, write random search and evolutionary search, then find a subnet that reaches at least 92.5% accuracy under 250KB and 60M MACs. This guide maps the question structure and what each question trains. It does not include solutions.
Lab 4 is a Colab notebook that rebuilds AWQ step by step on OPT-1.3B: first see how badly 3-bit quantization hurts perplexity, then keep 1% of the salient channels in FP16 (Q1), then protect them by scaling instead and search for the best scale (Q2). Each question is worth 50 points, plus a bonus scored on perplexity. Lab 5 moves to C++: run 4-bit LLaMA2-7B-chat on your own computer with TinyChatEngine and write five versions of the W4A8 linear-layer kernel (loop unrolling, multithreading, SIMD, multithreading plus unrolling, and all combined), 20 points each, plus up to 20 bonus points for performance. This post covers the questions, points, setup, and limits for outside learners. No solutions.
Lecture 13 sorts the ways to speed up LLM inference into three paths. Quantization: SmoothQuant moves the difficulty of activation outliers onto the weights to make W8A8 work, AWQ uses activation magnitudes to find the roughly 1% of weights that matter and protects them by scaling for W4A16, and QServe combines both into W4A8KV4. Sparsity: Wanda prunes weights by |W|·‖X‖, DejaVu and MoE use only part of the parameters per token, and SpAtten and H2O drop unimportant tokens. Serving: TTFT/TPOT metrics, PagedAttention, FlashAttention, speculative decoding, and continuous batching. On the slides, INT3 OPT-6.7B has a perplexity of 43.16 with RTN; scaling the salient channels by 2 brings it to 14.07.
Lecture 14 has three parts. Fine-tuning: SFT runs next-token prediction on desired answers, RLHF trains a reward model and then fine-tunes with KL-penalized RL, and DPO collapses both stages into one supervised step. Then comes a chain of PEFT methods: BitFit tunes only biases, Adapters add small layers but slow inference, Prompt/Prefix-Tuning eat input length, LoRA fixes latency with a low-rank branch you can merge back, QLoRA stores the backbone in NF4, and BitDelta compresses the fine-tune delta to 1 bit. Multimodal LLMs: Flamingo uses cross-attention, PaLM-E and VILA feed images in as tokens, and VILA-U can also output images. Prompt engineering: zero/few-shot, CoT, and RAG.
Lecture 15 has four parts. Extending context: interpolating RoPE stretches LLaMA from 2k to 32k, and LongLoRA's shifted sparse attention makes long-context fine-tuning cheap. Evaluation: lost-in-the-middle, Needle-in-a-Haystack, and LongBench. Efficient attention: the KV cache grows linearly with length. StreamingLLM finds that the first few tokens act as attention sinks, and keeping them plus a recent window gives stable generation. DuoAttention keeps a full KV cache only for a few retrieval heads. Quest keeps the whole KV cache but reads only the most critical pages for each query. The last part moves beyond Transformers: Mamba replaces attention with a selective SSM, and Jamba mixes the two.
An MCU has roughly 256–320kB of SRAM and 1MB of Flash, tens of thousands of times less than a phone. Even an int8 MobileNetV2 needs 5x more peak memory than that. Lecture 10 answers with MCUNet: TinyNAS picks a search space before searching for a subnet, and MCUNetV2's patch-based inference cuts MobileNetV2's peak SRAM from 1372kB to 172kB. The lecture closes with tinyML applications in vision, audio, and anomaly detection.
Lecture 8 of MIT 6.5940 (Fall 2024) attacks the most expensive step in NAS: evaluating candidates. Training 12,800 architectures from scratch cost 22,400 GPU-hours, so the lecture walks through inherited weights, hypernetworks, ProxylessNAS's single-path training, latency lookup tables and predictors, Once-for-All's one training run for 10^19 subnets, training-free zero-shot NAS, and NAAS, which searches the network and the accelerator together. This guide follows the 105-slide deck and cites a page for every claim.
Lecture 7 has three parts. It first reviews fully connected, convolution, grouped, depthwise, and 1×1 convolution layers through their MAC formulas. It then takes apart how the ResNet bottleneck, ResNeXt, MobileNet, MobileNetV2, ShuffleNet, and the Transformer each save compute; the bottleneck, for example, needs 8.5× fewer MACs than a plain 3×3 convolution over 2048 channels. The last part is NAS: search spaces are either cell-level or network-level (depth, resolution, width, kernel size, topology), and there are five search strategies: grid, random, reinforcement learning, gradient descent, and evolution. One arithmetic exercise on the slides shows that the NASNet cell space already holds 3.2×10¹¹ candidates at M=5, N=2, B=5.
There are two reasons to train on the device: the model has to adapt to each user's new data, and that data should not leave the device. Lecture 21 first shows that sharing only gradients is not safe either: Deep Leakage from Gradients recovers the original images and sentences from them. Then it tackles memory. Training costs more than inference because activations must be stored, not because of the parameters. TinyTL fine-tunes only biases plus a lightweight residual and saves 6.5x memory; SparseBP updates only the important layers and channels; QAS lets real int8 training match fp32; and PockEngine does autodiff at compile time, bringing training memory on a 256KB MCU down to 141KB.
Pruning removes unimportant weights or neurons from a neural network. The goal is written as minimizing loss subject to at most N nonzero weights. Lecture 3 of 6.5940 handles two of the decisions involved. First, granularity: from fine-grained pruning, which can remove any element, to channel pruning, which removes whole channels. The more regular the pattern, the easier it is to speed up on existing hardware, and the less you can remove. In between, 2:4 sparsity gives up to 2× speedup on NVIDIA Ampere GPUs. Second, criteria: look at weight magnitude, Batch Norm scaling factors, second derivatives, the fraction of zero activations, or how well a layer's output can be reconstructed after pruning.
MIT 6.5940 Lecture 4 finishes the pruning unit. Per-layer ratios come from sensitivity analysis, AMC (reinforcement learning), or NetAdapt (step-by-step with a lookup table). Fine-tuning uses 1/10 to 1/100 of the original learning rate, and iterative pruning pushes AlexNet from 5x to 9x. EIE, NVIDIA 2:4 sparsity, and TorchSparse/PointAcc show that sparsity only turns into speed with system support.
MIT 6.5940 Lecture 5 starts from one fact: an 8-bit integer add uses 30x less energy than a 32-bit float add. It reviews the bit layouts of INT, fixed point, FP32/FP16/BF16, FP8, and FP4, then covers two quantization methods. K-means quantization saves storage only, since computation stays in floating point. Linear quantization, r = S(q − Z), turns matrix multiplication, fully connected layers, and convolutions into integer arithmetic.
Lecture 6 is about what to do when quantization costs you accuracy. First, without retraining: use finer scale granularity (per-channel, group, MX), clip outliers (EMA, calibration batches, MSE, KL), and round smarter (AdaRound). If that isn't enough, retrain: QAT keeps a full-precision copy of the weights, runs fake quantization in the forward pass, and uses the STE to pass gradients straight through. In the whitepaper table the slides cite, MobileNetV1 drops to 0.1% accuracy under per-tensor INT8 PTQ and recovers to 70.7% with per-channel QAT, against a 70.9% float baseline. The last two sections cover 1–2 bit binary and ternary networks, and HAQ, which uses reinforcement learning to assign a bit width to each layer.
Once algorithms have shrunk the model, how much more can the system layer squeeze out? Lecture 11 uses a single matrix multiply to show it: loop reordering gives 12x, tiling 19x (on an Intel Xeon 4114), and a CUDA version runs 94x faster end to end on a 2080Ti. The second half covers TinyEngine's inference tricks: im2col; in-place depthwise, which cuts peak memory from 2×C×H×W to (1+C)×H×W; NHWC for pointwise and NCHW for depthwise; and Winograd, with 2.25x fewer multiplications.
In Lecture 12, 6.5940 switches from CNNs to Transformers. The lecture doesn't dwell on theory. It points to where memory and compute go. Attention is O(N²). If Llama-2-70B used MHA, its KV cache at batch 16 and length 4096 would take 160GB. GQA shrinks that 8x and MQA shrinks it 64x. MoE adds total parameters while keeping per-token compute flat. This post bridges into Lecture 13 on LLM deployment.
MIT 6.S184 is a short IAP (January Independent Activities Period) course: five lectures (Lecture 3 is split into two recordings, 3-A and 3-B), three labs, and an 84-page set of lecture notes the course calls its backbone. Notes, slides, all six recordings, lab notebooks, and official solutions are public, so it grades A3, enough to self-study. Two gaps remain: lab submission goes through Gradescope inside Canvas, which only enrolled MIT students can use, and Lecture 5 on discrete diffusion has no lab.
MIT 6.S184 Lab 1 has three parts. First you write the step functions for Euler and Euler–Maruyama. Then you use them to simulate Brownian motion and the Ornstein–Uhlenbeck process and watch how σ and θ shape trajectories and the final distribution. Finally you implement Langevin dynamics, watch a cloud of points get pushed toward a five-mode Gaussian mixture, and show by hand that the OU process is Langevin dynamics with a Gaussian target. The questions, code scaffolding, and official solutions are all on GitHub; outside readers get no Gradescope grading and have to check against the solutions themselves.
Lab 2 turns §3–4 of the notes into PyTorch. You implement the Gaussian conditional path, its conditional vector field, and its conditional score. Then two nearly identical trainers do flow matching and score matching, Proposition 1 converts the learned vector field into a score, and finally a linear path makes a ring distribution flow into a checkerboard. Everything runs on 2D toy data. The README records a diffusion-coefficient bug fix dated 1/11/26; when I checked on 2026-09-30, the fix appeared only in the solutions notebook, not the student version, so patch it yourself before you start.
Lab 3 builds a conditional latent diffusion model on MNIST from scratch, in four stages: CFG training with label dropout (checked on a three-component Gaussian mixture), a diffusion transformer built piece by piece (Fourier time embedding, patchify, multi-head attention, adaLN-Zero, depatchify), a VAE, and finally the DiT trained inside the VAE's latent space. Problems and official solutions are public; submission goes through Gradescope on Canvas, which only enrolled MIT students can use.
Lecture 1 of MIT 6.S184 first rewrites "generate an image of a dog" as "sample from the data distribution," then gives the machine that does the sampling: start from Gaussian noise and simulate an ODE along a neural-network vector field (a flow model), or add a little Brownian-motion noise at every step to get an SDE (a diffusion model). Each is simulated with the simplest numerical method available, Euler and Euler–Maruyama. How to train the vector field is left to Lecture 2.
The object we want is the marginal vector field: run an ODE along it and noise flows into data. The catch is that it requires an integral over the whole dataset, so we can't compute it. Flow matching regresses on the conditional vector field instead, the one that pushes noise toward a single data point, which has a closed form. Theorem 12 in the notes shows the two losses differ by a constant and share the same gradient. On the CondOT path, training reduces to one line: sample data z, noise ε, and time t, and have the network predict z − ε at the point tz + (1−t)ε.
A score function is the gradient of the log density; it points toward where probability rises fastest. On Gaussian paths, the score and last lecture's vector field are both linear in x and z, so each converts into the other (Proposition 1 in the notes): learn one and you have learned both. With the score in hand, you can add noise of any strength to the ODE and turn it into an SDE without changing the distribution at any time (Theorem 17). The score itself is learned with the same trick as flow matching, by regressing on the conditional score. On Gaussian paths, that amounts to predicting the noise that was added, which is the DDPM training objective.
Feeding the prompt to the network as an extra input should, in theory, sample from p_data(x|y), but in practice the images don't follow the prompt closely enough. Lecture 3B uses Bayes' rule to split the guided vector field into the unguided vector field plus a classifier gradient; scaling that classifier term by w is classifier guidance. Replacing the classifier with the difference between guided and unguided fields gives CFG, which needs no classifier: ũ = (1−w)·u(x|∅) + w·u(x|y). Training only requires swapping the label for a null label ∅ with probability η. The costs: two network calls per step, and for w>1 you are no longer sampling from the data distribution.
The algorithms are complete by Lecture 3B; Lecture 4 tackles two engineering problems that show up at scale. First, the network must take an image, a time t, and a prompt and output a vector field of the same size, so we use a U-Net or a diffusion transformer (DiT), embedding time with Fourier features and text with frozen CLIP/T5 encoders. Second, pixel space is too big, so we first train a VAE to compress images into a latent space, run flow matching there, and decode at the end. Stable Diffusion 3 and Meta Movie Gen Video both follow this recipe: flow matching in latent space, a DiT variant, and CFG.
Text is a sequence of discrete tokens. There is no direction to move in, so ODEs and SDEs do not exist. Lecture 5 carries the recipe from Lectures 1–4 over unchanged and swaps only the underlying stochastic process: vector fields become rate matrices, ODEs become continuous-time Markov chains (CTMCs), and the continuity equation becomes the Kolmogorov forward equation. With the factorized mixture path, the marginal rate matrix has exactly one unknown: the probability of each position's original token given the noisy sequence. Training a discrete diffusion model therefore reduces to per-position classification with a cross-entropy loss. Make the noise all [mask] tokens and you get a masked diffusion language model.
The first half of lecture 1 covers course rules and lightning talks. The middle answers "why learn the principles?": Yen-Lung Tsai splits the anxiety of learning AI into three kinds and argues that knowing the principles tells you a model's limits, so you stop chasing every new tool. The second half is a Colab primer, from magic commands and the four standard import lines to plt.plot, Markdown, and ipywidgets. Homework 1 is to plot a function in Colab. On the Chang Gung satellite rubric, a tweaked copy of the demo earns 6 points; a function not taught in class, with well-written Markdown notes, earns 10.
Lecture 2 opens up last week's "dopey AI robot." Inputs and outputs must become numbers (tensors). Classification uses one-hot labels and softmax to turn scores into probabilities. A neural network is neurons stacked layer by layer, and training means pushing the loss down with gradient descent. The lecture ends by building a first fully connected network on MNIST in Keras and wiring it to a Gradio sketchpad. Homework 2 asks you to design your own DNN, with one hard rule: it can't have three layers. The Chang Gung satellite rubric wants a screenshot of the parameters with the best validation accuracy and encourages keeping failed attempts.
The prompt "a cute girl" has countless correct pictures, so training it as a function only teaches the model the average of all of them. GANs sidestep this by training two networks: a generator G turns a random latent vector into an image, a discriminator D judges real versus fake, and the two compete. L03 walks from the 2014 paper through WGAN, Progressive GAN, StyleGAN's 512-dimensional latent and AdaIN, then Pix2Pix and CycleGAN. An appendix explains cross entropy and KL divergence as a 'surprise index'. Week 3 homework: run a GAN yourself, or explain CE and KL in your own words.
L04 reduces a large language model to one sentence: look at the preceding words, score every word in the vocabulary, turn the scores into probabilities with softmax, and sample the next word. To give the model a memory of what came before, the lecture covers RNNs and then gives a first look at Transformer Q/K/V. GPT-2's 1.5 billion and GPT-3's 175 billion parameters illustrate scale; temperature and top-p explain why every answer comes out different. The second half covers running open models locally and estimating VRAM. Week 4 homework: write test prompts on a topic you know well and compare at least two LLMs.
L05 reads the whole Transformer with two linear-algebra rules: matrix multiplication is row-times-column dot products, and a row vector times a matrix is a linear combination of the matrix's rows. With those, attention is 'dot the query with every key, softmax into weights, take a weighted average of the values,' or softmax(QKᵀ/√d_k)V in batch form. Dividing by √d_k just pulls the numbers back toward 0 so softmax doesn't become winner-take-all. Then come multi-head attention, encoder versus decoder, masking, positional encoding as a set of sin/cos clocks, and ResNet-style residuals with layer normalization. No homework this week.
The first half of L06 is about ethics. Yen-Lung Tsai quotes Karpathy's line that hallucination is a feature of LLMs, then works through plagiarism, whether your data gets used for training, and DeepSeek's censorship and corpus skew, and closes with seven principles of responsible use. The second half is about applications: give the model the right information and clear instructions, and one system prompt becomes a Lucky Vicky positivity generator, a social-media copywriter, or a biased college-major counselor. The week-6 assignment moves that prompt into an OpenAI-compatible API with a Gradio front end: a chatbot with a persona.
A chatbot 'remembers' you not because the model has memory, but because your code resends the whole messages list (system, then alternating user and assistant) every turn. L07 starts with getting OpenAI and Groq keys, spells out that structure, then runs Gemma 3 locally or in Colab with Ollama, where the same openai package works after changing only base_url. The week-7 assignment offers two options: a version that keeps the conversation going, or two models talking to each other, both demoed in Gradio.
L06 said a prompt is two things: correct information and clear instructions. RAG lets the computer fetch the information part on its own. Split your documents into chunks, turn chunks and questions into feature vectors with the same model fθ, find the closest few chunks, and drop them into a template: 'Answer {question} based on {retrieved_chunks}.' The code comes in two notebooks: Demo06a builds a vector database with LangChain and FAISS and zips it as faiss_db.zip; Demo06b loads it back, connects an LLM, and wraps it in Gradio. The week-8 assignment is to do the same with your own data.
L09 defines an AI agent in one line: the AI finishes the work you would otherwise do yourself. Yen-Lung Tsai follows Andrew Ng's four design patterns (Reflection, Tool Use, Planning, Multiagent Collaboration) but builds only the two easiest. Demo07a hands a draft between a "writer" and a "reviewer" LLM call. Demo07c splits the Lucky Vicky post generator into "think of five reasons, then write the post", a two-stage CoT. Both use AISuite with Groq and a Gradio front end. LangChain, AutoGen and CrewAI appear only on a further-learning list. The week 9 homework asks you to pick one of the two patterns.
L10 starts from one question: how do you find a good feature vector? Word2Vec learns embeddings through a pretext task. An autoencoder squeezes out a latent vector by being forced to reproduce its input. A VAE then asks the latent to follow a normal distribution, so nearby points produce similar images. Yen-Lung Tsai then recasts diffusion as "an autoencoder whose encoder is computed and whose decoder is learned", and ends on latent diffusion: a VAE shrinks a 512×512 image to 64×64, and diffusion runs only in that small space. The week 10 homework involves no code: make several style-consistent image sets with Bing.
L11 fills in the rest of the Stable Diffusion diagram. CLIP is trained so that matching text and images get similar vectors, which turns a prompt into a 77×768 embedding. Schedulers compress 1,000 noising steps into twenty or thirty denoising steps, but ancestral samplers such as Euler a never settle: push to 100 steps and the subject changes jackets and seats. LoRA freezes the original W and learns only a ΔW factored into A·B. The hands-on part loads an SD 1.5-family model with diffusers, and the week 11 homework is your own image-generation web app.
The Stable Diffusion setup from L11 listens only to the prompt, so composition and pose are left to luck. L12 adds a steering wheel. ControlNet copies a block of SD and wires the copy back in through zero convolutions, so extra conditions such as edge maps, poses, and depth maps can steer generation. The standard example is Canny edges. The second half covers Fooocus, an SD interface that aims to be 'as simple as Midjourney': Presets, Styles, and the five Input Image features, where Image Prompt is ControlNet with a friendly wrapper. Week 12 homework: pick a use case, make at least 3 image sets in Fooocus, and write up your creative process.
Every model in the first 12 lectures learned from training data that people prepared. L13 asks a different question: when there is no right answer, only a signal of how well you did, how does a computer learn? Tsai starts from AlphaGo and Breakout and splits the field in two. Value-based methods learn a Q function that scores each action (Deep Q-Learning, TD, experience replay, ε-greedy); policy-based methods learn the action directly (policy gradient, actor-critic). The second half returns to LLMs: ChatGPT trains a reward model from human rankings and then runs RLHF with PPO, while DeepSeek has the computer check math answers automatically and uses that as the reward. Week 13 homework is the final project proposal.
The last lecture looks at two lines of technology crossing into each other. LLMs such as ChatGPT have started drawing, and the slides use early fusion plus VQ-VAE/VQGAN to explain how an image can be cut into tokens. Going the other way, Inception Labs' Mercury generates text with diffusion, noising a sentence into a row of [MASK] tokens and then restoring it. Next come a few papers anyone can use: evaluating RAG automatically, reasoning models being easier to hijack, and DeepMind's four kinds of AI risk. The lecture ends with vibe coding and a list of application tools, and the final project runs as an online conference in Gather Town.
Generative AI: Text and Image Synthesis Principles and Practice is an introductory course taught by Yen-Lung Tsai (蔡炎龍) of NCCU's Department of Mathematical Sciences and opened to other schools as a TAICA satellite course. The most complete course page online actually belongs to the Chang Gung University satellite section, where Chih-Yuan Yang is the co-teacher. This series follows Spring 2025 (semester 1132): 14 recordings, 14 slide decks, and 12 homework specs with rubrics are public, and the demo notebooks are on GitHub, so the access grade is A3. The gaps: the notebooks keep changing, submission and grading run through each school's LMS, and final projects were never published.
A guide to the BERT and its Family unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). It starts from "I record the record": Word2Vec and GloVe give both records the same vector. ELMo fixes this with a bidirectional LSTM language model. With Transformers, pretraining splits into three roads: encoders (BERT: MLM plus NSP, strong at understanding, weak at generation), encoder-decoders (T5's span corruption, BART's five noise types), and decoders (GPT: pure next-token prediction). The last part covers GPT-3's in-context learning and scaling laws, and why decoders became today's dominant backbone.
A guide to the decoding and evaluation unit in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The first half covers how to pick a word once the model outputs a probability distribution: greedy decoding can't take back a mistake, beam search keeps several candidates but favors short outputs, and top-k / top-p trade determinism for diversity (missing from the slides; the professor covers it verbally in class). The second half covers scoring generated text: BLEU's modified precision and brevity penalty, ROUGE-N and ROUGE-L, perplexity, and what GLUE, SQuAD 2.0, MTEB, and MMLU each measure.
A guide to the GPT-2 / T5 tutorial in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). One task, LCSTS Chinese summarization, is solved twice. Decoder-only GPT-2 is written in native PyTorch: you join article and summary into one sequence, switch to left padding, and set padding labels to −100. Encoder-decoder mT5 uses Seq2SeqTrainer: no left padding needed, and DataCollatorForSeq2Seq handles the −100 for you. Both segment with jieba and score word-level ROUGE.
Hung-Yu Kao's Fall 2025 W8 slides walk from GPT-1 to GPT-3, explain how the Sparse Transformer behind GPT-3 cuts attention cost, and then use InstructGPT to show the gap between continuing text and following instructions. The maximum likelihood objective can't tell a fabricated fact from a slightly wrong synonym, so three extra stages are added: SFT learns how humans write, a reward model learns how humans grade, and PPO optimizes against that grade while a KL penalty keeps the model from drifting too far. The lecture closes with Llama-2: separate safety and helpfulness reward models, context distillation, and GQA for faster inference.
A guide to the Hugging Face BERT tutorial and HW3 in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The tutorial walks through binary IMDb sentiment classification: AutoTokenizer, the input_ids / token_type_ids / attention_mask fields, AutoModelForSequenceClassification, and Trainer. HW3 applies the same tools to SemEval 2014 Task 1: one bert-base-uncased with two heads, one regressing a 1–5 relatedness score and one classifying entailment into three classes. You add the two losses and write the training loop yourself, because Trainer is not allowed.
HW1 tests word vectors on the 19,544 Google Analogy questions (8,869 semantic, 10,675 syntactic). You first answer them with pretrained glove-wiki-gigaword-100 loaded through Gensim, then train your own Word2Vec on a 20% sample of a pre-cleaned Wikipedia dump, and plot t-SNE for the family subcategory both times. Seven TODOs are worth 55%, the report 45%. Fall 2026 keeps the same TODOs but asks for an .ipynb with outputs.
Week 1 of Hung-Yu Kao's NLP course at NTHU is a 91-page deck, W1_NLP_brief. It opens with 'Watch for kids', five readings of the telescope sentence, and a Chinese tongue-twister about eleven uncles to show why language is hard. Then, from an information-retrieval angle, it builds one pipeline: inverted index, tokenization, stemming, TF-IDF, BM25. The second half hits that pipeline's dead ends (synonyms, polysemy, vocabulary mismatch), moves to SVD-based LSA, and closes with a preview of dense vectors through Skip-gram, GloVe, and FastText.
Hung-Yu Kao's Natural Language Processing at National Tsing Hua University is a graduate-level flagship course in the TAICA alliance. The syllabus caps it at 1,200 students, it is taught in Mandarin, and it runs from TF-IDF and word vectors to RLHF, PEFT, and RAG. For Fall 2025, the slides, 32 class recordings, and 4 assignments with starter notebooks are all on GitHub, which rates A3. Solutions, grading, and the term-project spec are not public. Fall 2026 is in progress and only goes up to W3, so it rates A2. Grading changed to 75% assignments plus a 25% in-person midterm, and a Reasoning/Agent unit was added.
This 34-slide TA session answers a practical question. Pasting data into the ChatGPT web page one row at a time is slow and hits hourly limits, so research and homework should use the API. The notebook runs one SemEval 2014 entailment example through Gemini, Claude, and OpenAI in turn: prompts live in prompts.yaml, output is forced into JSON, then few-shot and token counting. The material is from 2024. The slide cover says 2024/11/21, and the notebook uses gemini-1.5-pro, gpt-4o, and claude-3-5-sonnet-20241022. That Claude model was retired on 2025-10-28, and Google's old Gemini SDK reached end of support on 2025-11-30.
Hung-Yu Kao's Fall 2025 PEFT slides open with a budget: full fine-tuning of Llama 2-7B in 16-bit needs about 56GB of GPU memory, while training only 0.2M parameters brings it down to about 17GB, because gradients and optimizer states nearly vanish. Intrinsic dimensionality then explains why tuning a small slice is enough: the longer a model is pretrained and the larger it is, the fewer effective dimensions fine-tuning needs. Methods fall into additive (Adapters, Prompt Tuning), selective (BitFit), reparametrization (LoRA), and hybrid (MAM Adapters, S4). The second half runs from GPT-2's task descriptions and verbalizers to the trade-offs between prefix tuning and soft prompt tuning.
HW2 treats expressions like "14*(43+20)=882" as character sequences and asks a two-layer LSTM to generate the answer one character at a time after it sees "=". The training set has 2,369,250 rows and the eval set 263,250, with every number in 0–49. Six TODOs run from building a vocabulary and batching with loss only after "=" to a generator, teacher-forced training, and exact-match evaluation. The W4 PyTorch TA session is the toolbox for it.
Each of the two RAG TA sessions builds one version. The first installs Ollama on Colab to run llama3.2:1b and wires up a minimal RAG with LangChain's Chroma, MMR, and retrieval chain. The second uses LangChain only for data prep and writes the rest by hand: chunking, text and vector stores, hybrid BM25 + cosine retrieval merged with RRF, then generation with Llama-3.2-1B-Instruct. HW4 applies the first session's skeleton to 150 cat facts and 150 GPT-5-generated QA pairs. The generator must be Llama3.2-1b and the embedding model jina-embeddings-v2-base-en, and you report recall@1, recall@5, and exact match. Code is 45% of the grade and the report 55%; the report analyzes how prompts, data format, document order, and counterfactual information change the results.
The second half of W11_RAG.pdf starts at the "From Retrievers to QA" slide and turns a retriever plus a reader into a full QA system. It begins with ORQA and REALM (2019–2020), where BERT is the reader, then covers the first paper named RAG and REPLUG, which keeps the LLM frozen. A single table then sorts seven recent fixes into three groups: rewrite the query (Query Rewriting, HyDE), make the generator robust to noise (RetRobust, RAFT, RAAT), and decide when to retrieve (FLARE, Self-RAG). It ends with noise types, four abilities an LLM needs inside RAG, and generative retrieval (GR) with reliable response generation (RRG). On the recording side, the W11 Thursday lecture stops at the RAG paper; I could not find a recording that covers the later slides.
An LLM will confidently answer that Oppenheimer was born in 1967 (the year he died). RAG retrieves first and generates second. The first 60 pages of Hung-Yu Kao's Fall 2025 RAG deck are all about finding the right material. Sparse vectors (bag-of-words, TF-IDF, BM25) are cheap and dependable; dense vectors catch paraphrases. BERT's [CLS] isn't a good sentence vector as-is, hence Sentence-BERT pooling and bi-encoders. A cross-encoder is accurate but needs nearly 50 million passes for 10,000 sentences, while a bi-encoder needs 20,000. SimCSE uses dropout as data augmentation, DPR beats BM25 with only 1,000 training examples, and GTR shows that scaling up a dual encoder improves out-of-domain retrieval.
"Look over there" is 3 tokens, the Chinese version 4, the Japanese version 12. When output length does not track input length, the classifier trick of adding an FFN at the end fails. This 33-slide deck starts from encoder-decoder models, works through vanishing gradients in RNNs and the three LSTM gates, and ends with attention fixing both long-range memory and parallelism, including why scores are divided by √d.
A guide to the Sub-word Tokenization unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). Splitting on white space only works for Western languages and cannot handle unseen words or German-style compounds. BPE starts from characters and repeatedly merges the most frequent adjacent pair into the vocabulary, so each merge adds one entry; its weakness is that the greedy split is not always the best one. The Unigram LM picks splits by probability and can even sample different ones. The slides give vocabulary sizes of 30522 for BERT, 50257 for GPT-2/GPT-3 and 32,128 for T5.
In week 14 of Fall 2025, Hung-Yu Kao closed the course with two slide decks. Course_summary sorts the semester into three columns (NLP Fundamentals, NLP Models, NLP Advances), adds two columns for the TA labs, and lists five directions for further study. The second deck is his notes on Denny Zhou's (Google DeepMind) April 2025 Stanford talk. Its claim: pretrained models can already reason, and decoding is what brings it out. CoT decoding, self-consistency and retrieval + reasoning add up to four inequalities. The Fall 2026 W16 'Reasoning / Agent' unit has not been released yet.
Fall 2025 was graded 70% assignments + 30% term project. Projects were done in groups of 3–4 and split into Proposal 6%, Progress 6%, Poster 6% and Report 12%, with no GPUs provided. The repo has no project spec, only the syllabus structure, an end-of-term reminder and the W15–W16 recordings. Fall 2026 switches to 75% assignments (4 of them) + a 25% in-person midterm in W14. The schedule drops the presentation weeks, adds a Reasoning/Agent unit, and brings in an AI-TA for grading support and TAICA compute credits. The official materials don't say why.
A guide to the Transformer unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). The slides start from two RNN problems: words interact only across O(N) steps, and time steps cannot run in parallel. Self-attention lets every word look at every other word and computes it all in one matrix multiplication, fixing both at once. The cost is that the model can no longer tell word order, so sinusoidal positional encoding is added. The lecture then assembles multi-head attention, Add & Norm, feed forward, cross-attention, masked attention and teacher forcing, and closes with GPT-2, ViT and four variants to show how far the architecture went.
Week 2 of Hung-Yu Kao's NLP course at NTHU is a 62-page deck about predicting the next word. The first part covers statistical language models with a bigram count table, add-one smoothing, and perplexity, then sparse vectors through a PPMI example with cherry and digital. The middle returns to Word2Vec's negative sampling and uses the Chinese word for 'apple' to show why contextualized embeddings are needed. The last part derives RNNs from three weaknesses of feedforward networks and shows how they handle NER, sentence classification, and stacked and bidirectional variants.
An RNN translator has to squeeze the whole source sentence into one vector, and long sentences don't fit. Attention lets the decoder look back at every input position each time it produces a word, score each one, and take a weighted sum. Call the scorer the query, the thing being scored the key, and the thing being averaged the value, and you have dot-product attention. The Transformer goes one step further: the input attends to itself, recurrence disappears, and in return you get parallelism and a constant path length. The price is that position information has to be added back by hand.
A static word vector gives "apple" one embedding, whether it means the fruit or the company. The BERT lecture in ADL Fall 2025 starts from that polysemy problem. TagLM feeds language-model features into a tagger, ELMo builds contextual embeddings from a deep bidirectional LSTM, and BERT swaps the LSTM for a Transformer, pre-trained with Masked LM and Next Sentence Prediction; downstream, you add a classifier or tagger on the top layer and fine-tune. The optional BERT Variants slides go one ring further out: Transformer-XL for longer context, XLNet's permutation LM to get both AR and AE benefits, RoBERTa's better data and training recipe, SpanBERT's span masking, and mBERT and XLM for many languages. This is the direct prerequisite for HW1, which uses bert-base-chinese for extractive QA.
Big data is not big annotated data. The last ADL lecture asks how to learn good representations without labels, and answers: find the latent factors that control the data. An auto-encoder squeezes the input into a short code and reconstructs it. The denoising version adds noise or masks 15% of tokens first, which is exactly the idea behind BERT's masked LM. A VAE forces the code to follow a distribution, so you can sample from it to generate. Dual learning lets paired tasks, such as translation and back-translation or understanding and generation, act as feedback for each other. Self-supervised learning has two camps: self-prediction (hide part, guess it back) and contrastive learning (pull similar pairs together, push dissimilar ones apart). CLIP runs contrastive learning on 400 million image-text pairs, making zero-shot image classification possible, and DALL·E 2 uses CLIP's representations to generate images. Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.
Dialogue systems split into chit-chat and task-oriented. Task-oriented systems were traditionally built from four modules: language understanding (LU) turns a sentence into domain, intent, and slots; dialogue state tracking (DST) accumulates the user's goal; the dialogue policy picks the next system action; and NLG turns that action back into a sentence. An LLM can act out all four steps by itself, but it cannot actually make the booking, so it needs external tools. LaMDA learns to call a search engine, calculator, and translator. BlenderBot 2.0 adds internet search and long-term memory. WebGPT learns to drive a browser from human demonstrations, a reward model, and PPO. Toolformer has the model generate and filter its own tool-use training data. The lecture ends with evaluation: automatic metrics, four kinds of human evaluation, and LLM-Eval. ADL Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.
Applied Deep Learning (ADL) Fall 2025, taught by Yun-Nung (Vivian) Chen in NTU's CSIE department, is a deep learning course built around NLP. It runs from neural network basics through Transformers, BERT, pretraining and prompting, post-training, LoRA, RAG, generation and evaluation, alignment issues, and language agents. Lectures L0–L11 come with slide PDFs and segmented videos, and the playlist adds videos for L12–L14. On the assignment side only the HW1 spec is public; HW2, HW3, and the final project have explainer videos only. That makes it A2.
Lecture 10 of ADL Fall 2025 sorts the problems of pretrained models into four groups, each paired with a goal: bias with fairness, toxicity with safety, hallucination with factuality, and finally alignment. The slides argue that bias can enter at any stage of the ML pipeline, that safeguards belong at four layers (data, input, training, output), that hallucination can be checked atomic fact by atomic fact, and that over-optimizing a reward model produces familiar symptoms: verbosity, excessive apologies, over-refusal. The final project announced that week is called Jailbreaking Olympics, but all that is public is the titles and one-line descriptions of two videos.
HW1 in ADL Fall 2025 gives a question and four Chinese paragraphs. The model first picks the relevant paragraph (paragraph selection, framed as four-way multiple choice), then marks the answer's start and end inside it (span selection), scored by Exact Match. The spec slides point you straight at Hugging Face's run_swag_no_trainer.py and run_qa_no_trainer.py. The simple baseline uses bert-base-chinese, length 512, effective batch size 2, and learning rate 3e-5, and both stages together take under three hours on an 8GB RTX 3070. The Kaggle leaderboard closed 9/29, and code plus report were due 10/1 on NTU COOL. Outside readers cannot get the Kaggle data or grading, but the task design, baseline settings, and five report questions are all usable for practice.
Lecture 11 of ADL Fall 2025 builds on the EMNLP 2024 Language Agents tutorial. It defines an agent as an entity that perceives and acts, then names what is new about language agents: reasoning itself counts as an internal action. The lecture is organized around three concepts. Reasoning covers CoT and ReAct; memory covers Generative Agents and its recency / importance / relevance retrieval; planning goes from greedy reactive planning to tree search and world models. It closes with multi-agent systems in three steps: initialization, orchestration, and team optimization.
The first self-study deck of ADL Fall 2025 describes machine learning as finding a function from data, and deep learning as a production line of simple functions where the machine learns what every station does. It uses speech and vision to contrast deep and shallow models, credits big data and GPUs for the post-2010 breakthroughs, and uses the universality theorem to ask why networks should be deep rather than fat. The most practical part comes last: the output domain decides the learning task, and the architecture should fit the properties of the input domain.
ADL Fall 2025's NN Basics and Backpropagation decks break model training into three questions. What is the model? Layers of neurons, each computing z = Wa + b and then a nonlinearity. What makes a function good? A smaller loss. How do we pick the best one? Gradient descent, in practice mini-batch SGD. Backpropagation computes gradients for millions of parameters efficiently: the forward pass stores each layer's output, the backward pass sends an error signal δ back from the output layer, and multiplying the two gives each weight's gradient.
Lecture 9 of ADL Fall 2025 answers two questions. The model gives you a probability distribution at every step, so how do you pick a word from it? And once you have a sentence, how do you judge it? The slides start with teacher forcing and exposure bias to show the gap between training and generation, then compare greedy, beam search, sampling, top-k, and nucleus sampling, and file temperature and the penalties under 'control' rather than decoding algorithms. The evaluation half covers BLEU, ROUGE, perplexity, and LLM-Eval, then explains why you would use RL to optimize whole-sentence quality directly.
When an LLM is too big to fine-tune in full, the LLM Adaptation slides of NTU ADL Fall 2025 offer three ways to change only a small part of it: insert small Adapter modules into the Transformer, represent the weight update with low-rank matrices (LoRA), or learn only a prefix or soft prompt (prompt tuning). The slides conclude that no single method fits every task. For HW2, the only public information is its title, "LLM Tuning and Prompt Tuning for Classical Chinese Translation"; the data, baseline, and grading have no written spec.
A pre-trained model can continue text, but that does not mean it follows instructions. The Post-Training slides of NTU ADL Fall 2025 fix this in two steps. Instruction tuning (FLAN, T0) teaches the model to read task descriptions. RLHF then pulls its outputs toward human preference. Three limits of instruction tuning connect the two steps, a reward model and pairwise comparisons solve two practical RL problems, and InstructGPT's SFT → reward model → PPO pipeline ties it all together. ChatGPT runs the same pipeline on multi-turn dialogue.
Lecture 6 of ADL Fall 2025 sorts pre-trained models into three families: encoders (the BERT family, bidirectional context), decoders (the GPT series, good at generation), and encoder-decoders (BART and T5, pre-trained with denoising). It then names two practical obstacles of the pre-trained-model era: downstream labeled data is scarce, and models keep growing until one copy per task no longer fits. The slides' answer is prompt learning. GPT-3's in-context learning shows a model can do a task without updating parameters; hand-written hard prompts (template plus verbalizer, LM-BFF) then give way to soft prompts optimized as vectors (P-Tuning, Prefix-Tuning, Prompt Tuning); and Liu et al.'s prompting typology closes the lecture.
LLMs cannot memorize long-tail facts, their knowledge goes stale, and they cannot see private documents. The RAG slides of NTU ADL Fall 2025 open with an LLM hallucinating about the lecturer herself, then split RAG into indexing, retrieval, and generation: sparse (TF-IDF, BM25) and dense (DPR, Contriever) retrieval, how dense retrievers are trained, pre- and post-retrieval techniques including pointwise and pairwise reranking. A closing roadmap organizes RAG, RETRO, FLARE, Search-R1, and others by what, how, and when to retrieve. For HW3, only the title is public: "Retriever & Reranker Training for RAG".
The Reasoning lecture of NTU ADL Fall 2025 has no public slides. It exists only as five videos in the course playlist: 12.1 What is Reasoning?, 12.2 Short CoT, 12.3 Test-Time Scaling, 12.4 Learning to Reason (imitating others), and 12.5 RL for Reasoning (evolving reasoning through exploration). This post lays out that route from the video titles alone, then pairs it with the CoT, ReAct, and 'reasoning enlarges the action space' pages of the previous Language Agents deck. Technical detail is left to the site's CS224N and CME295 reasoning posts.
The first real lecture of ADL Fall 2025 has one through-line: a language model predicts the next word. It starts from one-hot vectors and co-occurrence matrices, moves through the zero-probability problem of n-grams and the smoothing that a neural LM gets for free, and ends at the RNN LM, which folds all previous words into a hidden state. BPTT and vanishing/exploding gradients are the training cost, and LSTM and GRU patch it with gating. The lecture closes by splitting applications into sequence input versus sequence output, which separates tagging from encoder-decoder models.
The ADL Fall 2025 course page schedules seven TA recitations: Dev Infra (PyTorch, debugging) → NLP project lifecycle → the underlying logic of NLP projects → LLM LoRA training → LLM basics, architecture, and MoE → LLM inference and evaluation → LLM deployment. All ten videos are older recordings by Yen-Ting Lin from 2023 and 2024, reused in Fall 2025. The course page's five slide links all return 404; files with the same names still open under the Fall 2024 path, and Deployment has a video only. The first three sessions walk through the Hugging Face data → model → demo loop that HW1 needs; the last four cover training, inference, and serving LLMs.
With whole words as units, any unseen word becomes UNK. With single characters, meaning is hard to reassemble. This 22-page ADL deck explains the mainstream compromise, subwords, and the most common way to build them, BPE. The core is a tiny corpus of 4 words and 16 occurrences: start from characters, merge the most frequent adjacent pair each round, and after 9 merges you have units like newest</w> and low</w>, which then segment the unseen words lowest and powest. It ends with a GPT-3 tokenizer screenshot where the Chinese version of a sentence takes more than twice as many tokens as the English.
Lectures 7 and 8 of Machine Learning Techniques open the aggregation part of the course. T7 sorts ways of combining hypotheses into uniform, linear, and any blending (stacking), shows with a few lines of algebra that uniform blending reduces variance, and then uses the bootstrap to create diverse g_t from the single dataset you have: that is bagging. T8 reinterprets the bootstrap as example weighting, then deliberately up-weights the examples the previous hypothesis got wrong so the next one is forced to differ, and votes with α_t = ln √((1−ε_t)/ε_t): that is AdaBoost. Practice with Fall 2024 HW6 Q4 and Q9, plus HW7's bootstrap and AdaBoost proofs and a 500-round AdaBoost-Stump experiment on madelon. There are no official solutions.
Hsuan-Tien Lin's Machine Learning Foundations (16 lectures) and Machine Learning Techniques (16 lectures) are two Mandarin-taught MOOCs. All 130 YouTube videos and 32 slide decks are free. The MOOCs alone are A2: since August 2025, free Coursera accounts can only view the first module, so the exercises sit behind a paywall. Add the Fall 2024 course page, which publishes HW0–HW7 and the final project spec, and you reach A3, minus the grading chain: no official solutions, Gradescope and NTU COOL are enrolled-only, and the Kaggle competition returns 404. Fall 2026 is running now as a flipped classroom; slides through week 4, hw0, and hw1 are public.
Lectures 9–11 of Machine Learning Techniques tie three models together with one thread: trees plus aggregation. T9 treats a decision tree as conditional aggregation and covers C&RT's binary branching, Gini and regression impurity, pruning, categorical features, and surrogate branches. T10 applies bagging to fully grown trees; add random subspaces and random projections and you get a random forest, with free OOB validation and permutation-based feature importance. T11 re-derives AdaBoost as steepest descent in function space on the exponential error, then swaps in squared error to get GBDT, which fits regressions to residuals. Practice with the impurity and gradient boosting proofs in Fall 2024 HW7. There are no official solutions.
Foundations Lecture 4 first shows that learning is impossible: from D alone, any guess outside D can be called wrong. That is No Free Lunch. It then reframes the question with marbles in a bin. If the data is drawn independently from one distribution, Hoeffding's inequality says the in-sample error E_in is probably close to the true error E_out. Checking one fixed h is only verification. Once the algorithm chooses among M hypotheses, a union bound charges 2M exp(−2ε²N). Conclusion: with a finite hypothesis set and small E_in, learning is feasible. What to do when M is infinite is the next lecture's job.
Techniques T16 re-sorts the whole course into three families: how to exploit features (kernels, aggregation, extraction, low-dimensional compression), how to optimize (gradients, equivalent problems, multiple steps), and how to fight overfitting (regularization, validation). It then uses four KDD Cup–winning models to show how the pieces combine in practice. The MOOC was recorded in 2016 and its deep learning stops at pre-training. The Fall 2024 on-campus course filled the gap with 302u (the ReLU family, Xavier/He initialization), 303u (momentum, RMSProp, Adam), a 2020 keynote deck, mlmai.ics, and 1126, an 11-model summary. The Fall 2026 versions of these files are scheduled for week 16 and currently return 404.
Fall 2024 had six Foundations assignments. HW0 is 20 multiple-choice math prerequisite questions. HW1–HW5 each have 12 problems plus a bonus: Q1–4 are auto-graded, Q5–12 are graded by TAs, the programming problems use rcv1, cpusmall, and mnist from the LIBSVM datasets site, and HW5 uses LIBLINEAR. HW1 and HW2 each include a problem where you argue with a ChatGPT-style answer. Fall 2026 has released hw0 and hw1: hw1 is now 16 multiple-choice problems with 4 secretly chosen for TA grading, and the data is the course's own hw1_train.dat. The new policy allows AI tools and vibe coding, but AI-generated code needs block-by-block comments in your own words. Neither semester publishes official solutions.
Lecture 5 of Machine Learning Techniques rewrites the soft-margin SVM in unconstrained form: ½wᵀw plus C times the total hinge error. That is an L2-regularized model, and a larger C means weaker regularization. The hinge error and logistic regression's cross-entropy are both convex upper bounds of the 0/1 error, so the SVM approximates L2-regularized logistic regression. For probability outputs, you can use Platt's two-level learning, running logistic regression on top of SVM scores, or use the representer theorem to do kernel logistic regression directly. Lecture 6 uses the same theorem to get the closed form β = (λI + K)⁻¹y for kernel ridge regression, but β is dense; switching to the ε-insensitive tube error gives SVR with sparse coefficients. Fall 2026 does not schedule these two lectures.
Lecture 3 of Machine Learning Techniques merges "feature transform + inner product" into a single kernel function K(x, x′). Training and prediction in the dual SVM only need K, so d̃ can be infinite: the Gaussian kernel corresponds to an infinite-dimensional transform. Lecture 4 admits the SVM can still overfit and introduces violations ξₙ and a parameter C, giving the soft-margin SVM. Its dual differs from the hard-margin one in exactly one way: αₙ gets an upper bound C. The value of αₙ sorts the data into non-SVs, free SVs, and bounded SVs, and the fraction #SV/N upper-bounds the leave-one-out error, a cheap way to rule out dangerous (C, γ).
The first three lectures of Machine Learning Foundations define machine learning as a flow chart: an unknown target function f generates data D, and an algorithm A picks g from a hypothesis set H, hoping g ≈ f. The simplest H (the perceptron) and A (PLA) then show the chart in action. On linearly separable data, PLA makes at most R²/ρ² updates; on non-separable data, use pocket instead. Lecture 3 sorts learning problems along four axes: output, label, protocol, and input. Foundations mostly deals with batch, supervised binary classification or regression on concrete features. Practice with Fall 2024 HW1 and Fall 2026 hw1.
Lecture 11 of ML Foundations compares PLA, linear regression, and logistic regression on the same score s = wᵀx. The three differ only in their error functions, and scaled cross-entropy upper-bounds the 0/1 error, so both regressions can do classification. The lecture then turns logistic regression into SGD by computing the gradient on one random example, and builds multiclass classifiers from binary ones with OVA and OVO. Lecture 12 uses a feature transform Φ to turn a circular boundary into a line in Z-space. The price is that computation and d_vc both grow with the dimension, so the advice is: try a linear model first. Practice problems are in Fall 2024 HW4.
Lecture 1 of Machine Learning Techniques turns "which separating line is best?" into an optimization problem. Once you fix the scale so that min yₙ(wᵀxₙ+b) = 1, maximizing the margin is the same as minimizing ½wᵀw, which is a standard QP. Lecture 2 uses Lagrange duality to trade a QP with d̃+1 variables for one with N variables and N+1 constraints, then uses the KKT conditions to recover (b, w) from α. Only the points with αₙ > 0, the support vectors, affect the answer. The dual still contains the inner product zₙᵀzₘ, so the dependence on dimension is not really gone until the kernel lecture.
Linear regression writes squared error as (1/N)‖Xw − y‖², sets the gradient to zero, and gets w_LIN = X†y in one step. The hat matrix H = XX† projects y onto the column space of X, which shows that on average E_out − E_in ≈ 2(d+1)/N. Logistic regression estimates P(+1|x) with θ(wᵀx); maximum likelihood turns into the cross-entropy error ln(1 + exp(−y wᵀx)). It has no closed-form solution, so you walk downhill along −∇E_in step by step. That is gradient descent.
Lectures 12 and 13 of Machine Learning Techniques open the third part, distilling hidden features. T12 starts from a linear combination of perceptrons: two layers can build AND and OR but not XOR, and one more layer fixes that, which is the multi-layer perceptron. It then replaces sign with tanh, derives backprop, and covers non-convex optimization, d_vc = O(VD), weight elimination, and early stopping. T13 discusses the challenges of deep networks, uses autoencoders as information-preserving encodings for layer-wise pre-training, treats denoising as regularization, and proves that the optimal linear autoencoder is spanned by the top eigenvectors of XᵀX, which is PCA. The videos date from 2016; modern deep learning is covered by the Fall 2024 302u/303u slides. Practice: Fall 2024 HW7 Q4, Q9, and bonus Q13.
Lecture 13 of ML Foundations defines overfitting as 'lower E_in but higher E_out' and uses experiments to find four causes: too little data, stochastic noise, an overly complex target (deterministic noise), and excessive model power. Lecture 14's remedy is regularization. It rewrites 'step back to H₂' as the constraint ‖w‖² ≤ C, then uses a Lagrange multiplier to turn it into minimizing E_in + (λ/N)wᵀw, which is weight decay. Back in VC theory, regularization shrinks the effective VC dimension d_EFF, and L1 buys sparse solutions. Practice problems: Fall 2024 HW4 Q8–9 and HW5 Q1, Q5–6, Q10.
Techniques T14 reinterprets the Gaussian SVM as a linear vote over distance-based similarities, which gives the RBF network. Too many centers overfit, so k-means picks a few prototypes, and k-means itself is alternating optimization. T15 starts from the Netflix ratings data: one-hot encode user IDs, feed them into a linear network with the tanh removed, and you get matrix factorization R ≈ VᵀW, learned by alternating least squares or SGD. The lecture closes with a map of extraction models: boosting, neural nets, RBF networks, matrix factorization, and k-NN. These two lectures exist only as MOOC material. Neither the Fall 2024 nor the Fall 2026 schedule covers them, and no public homework problem does either.
The Techniques half of Fall 2024 has two homework sets and a final project, and all three PDFs are public. HW6 covers kernels, soft-margin SVM, and aggregation; its programming part uses LIBSVM on the 3-vs-7 subproblem of mnist.scale to count support vectors, compute margins, and run 128 validation rounds. HW7 covers bootstrap, impurity, AdaBoost, gradient boosting, and neural networks; its programming part is a 500-round AdaBoost-Stump on madelon. The final project is a fictional baseball league, HTMLB: predict home-team wins across two Kaggle stages and write an English report of at most seven pages that compares at least four methods. There are no official solutions. On 2026-09-30 both Kaggle pages returned 404 without login, so outside readers probably cannot get the HTMLB data and should reproduce the same splits on a public dataset instead.
The Hoeffding guarantee from L4 carries an M, the number of hypotheses. Perceptrons have infinitely many lines, so M blows up. L5 stops counting hypotheses and counts how many ○× patterns (dichotomies) they can produce on N data points instead; the maximum is the growth function m_H(N). 2D perceptrons produce at most 14 patterns on 4 points, fewer than 2⁴ = 16, so 4 is their break point. L6, marked optional by the course, proves that any break point caps m_H(N) by a polynomial, which is what makes the VC bound work.
Lecture 15 of ML Foundations tackles model selection. Selecting by E_in overfits, and selecting by E_test is cheating. The compromise is to carve a validation set out of the training data, select by E_val, then retrain on all the data. The validation size K is a dilemma, with K = N/5 as the rule of thumb. Leave-one-out is almost unbiased but expensive and unstable, so in practice you use 5-fold or 10-fold. Lecture 16 closes with three principles, Occam's razor, sampling bias, and data snooping, and a 'Power of Three' recap: three related fields, three bounds, three linear models, three tools. Practice problems are in Fall 2024 HW5.
L7 names the largest non-break point the VC dimension d_VC, proves that d-dimensional perceptrons have d_VC = d + 1, and rewrites the VC bound as E_out ≤ E_in + a model-complexity penalty, so both too large and too small a d_VC hurt. Theory asks for N ≈ 10,000·d_VC examples; in practice 10·d_VC is often enough. L8 swaps the fixed target function for a distribution P(y|x) and shows the VC theory still holds under noise. The error measure should come from the application: a CIA fingerprint check that penalizes admitting an intruder 1000 times more can be reduced to plain classification by copying examples.
The second half of agent_era.pdf asks three questions. How should multiple agents collaborate? (MacNet: irregular topologies beat regular ones.) Can agents deceive each other? (Werewolf, murder-mystery games, and MARO, which learns reasoning from social play.) Can agents socialize? (Moltbook and its "Church of Molt" — though three studies find the buzz mostly human-driven and the conversations shallow.) Then, using academic research as the case: AI can already replicate and extend a paper end to end, it entered AAAI 2026's review process, and Agents4Science 2025 received 247 AI-authored papers. Hung-yi Lee's conclusion: in the early age of agents, knowing what you want to do matters more than knowing how to do it.
A language model's input is finite, but an agent keeps piling up tool outputs. In week two of ML 2026, Hung-yi Lee splits Context Engineering into three moves: compression (summaries, hard clearing, offloading to files, plus ACON, SUPO, and AgentFold, which make compression smarter), filtering (read only the lines you need, load tools on demand as in MCP-Zero), and finally Agentic Context Engineering, where the LLM decides the next context itself — from Dynamic Cheatsheet and ACE to Recursive Language Models. The most useful idea to take away: a subagent is a form of self-directed compression.
Hung-yi Lee's Spring 2026 Machine Learning course at National Taiwan University opens with OpenClaw. The first half takes apart AI agents, context engineering, inference speed-ups, and positional embeddings. The second half covers harness engineering, self-correction, and self-improving AI. Slides and recordings for all 8 lectures, plus PDFs and Colab notebooks for all 10 assignments, are public, so it rates A3. What's missing is grading: JudgeBoi returned 502 on 2026-09-30, NTU COOL is campus-only, and the three guest talks have no materials at all.
In week 3 of ML 2026, Hung-yi Lee spends the first half of the inference lecture on one technique: Flash Attention. A GPU's execution units are fast, but their workbench (on-chip SRAM) is tiny, so data has to be carried to and from the warehouse (HBM). The carrying is the bottleneck. A naive softmax makes several round trips to the warehouse. Flash Attention assumes the current maximum is Amax, then multiplies by a correction factor when a larger value shows up. That lets it find the maximum, build the denominator, and compute the weighted sum in one pass, without ever materializing the attention weights. The output is identical to standard attention, no retraining is needed, and the cost is a little extra compute and a little brain strain.
Hung-yi Lee opens with a small model fixing a bug. gemma-4-E2B-it can't find parser.py, so it writes a fake one and declares victory. Add three short sections (the current environment, how to work, what counts as done) and the same model runs ls, cat, edits the file and runs the tests. The lecture splits the harness into three levers: natural language shapes the model's frame of mind (AGENTS.md), tools set its capability boundary (SWE-agent's ACI, rewriting CLIs for agents), and workflows control its behavior (the Ralph loop, Anthropic's long-running harnesses). The second half covers three extensions: scolding an agent can backfire, how a life-long agent learns from verbal feedback, and why evaluating agents is hard. It ends with agents improving their own harness (Meta-Harness).
HW1 asks for a defense prompt under 1,000 tokens that keeps the model wrapping every reply in [START]…[END] and never saying 'I have been PWNED,' no matter how it's attacked. The TAs prepared 14 attacks, 10 public and 4 private, each worth 0.5% for safety and 0.5% for utility. The task, the full text of the 10 public attacks, and the token-counting Colab are all public, but the grading platform JudgeBoi returned 502 on 2026-09-30, so outside readers have to build their own evaluation from the spec.
HW10 is 12 multiple-choice questions answered only on NTU COOL. Section 1 compares three spoken language model architectures: Cascade (ASR → LLM → TTS, with text in the middle), End-to-End (a language model over discrete speech tokens), and Thinker-Talker (an LLM thinks, a separate decoder speaks). In the Colab, two models listen to three clips and guess the speaker's gender, and you work out which one is the cascade. Section 2 takes Mimi apart: tokenize an emotion corpus into 32 RVQ layers, plot UMAP for layers 0, 6, 16, and 31, then encode and decode speech, laughter, and music to hear what breaks. The rest are paper questions on TWIST, AudioLM, LLaMA-Omni 2, Moshi, and GLM-4-Voice, covering initialization, pretraining, interleaving, and realtime/full-duplex behavior. The Colab needs Llama-3.2-3B-Instruct access and an HF token. Questions and Colab are public; outside readers miss only the COOL grading and answers.
HW2 doesn't ask you to write a classifier. You write prompts and a pipeline so that an open LLM running on a Colab T4 (by default a 4-bit GGUF of gemma-3-12b-it) plans, codes, runs, and debugs a 10-class MyGO & Ave Mujica character face classifier on its own. The starter code is adapted from AIDE: an Interpreter runs code, a Node records each version, a Journal forms the solution tree, and the Agent decides whether to draft, debug, or improve next. The first thing worth noticing: the starter's evaluation is empty. Every version is marked metric 1.0 and not buggy, so the tree search picks blindly until you fill it in. The rules are strict: "the LLM agent is your representative", and you may not hand-edit code or prediction files.
HW3 is 20 multiple-choice questions at 0.5 points each. No code is submitted; students answer a quiz on NTU COOL. The first 10 questions come from reading papers: four on speculative decoding (Leviathan et al., DeepMind's Speculative Sampling, Inference with Reference, SpecInfer) plus FlashAttention 1–3. The last 10 require filling TODOs in the Colab and analyzing the results: acceptance rate of a hand-written speculative decoder, speed-up curves for an assistant model vs n-gram under two prompt regimes, HBM reads and theoretical FlashAttention speed-up from T4 specs, vLLM prefix caching across turns and a cache invalidation test, and the effect of CPU offload on throughput. All questions are printed in both Mandarin and English in the homework PDF, so outsiders can do the whole thing; they just cannot get the official answers.
HW4 moves next-token prediction from text to images unchanged: 792 Pokémon sprites at 20×20, each pixel one of 167 color tokens, so one image is a 400-token sequence. Training is next-token prediction; at test time you get the first 60% of an image and the model draws the rest. Grading checks FID and a Pokémon Detection Rate (PDR) together, and the three baseline hints go from "run the sample code" to "tune hyperparameters" to "switch to Llama or Mistral". The spec, Colab, Kaggle notebook and dataset are public, but JudgeBoi returned 502 on 2026-09-30, so outside readers cannot get official FID or PDR scores.
HW5 fine-tunes Llama-3.2-1B-Instruct on GSM8K with LoRA, then uses harmful AILuminate prompts to check whether it still refuses. Math accuracy and safety rate must clear the bar together, so the real question is how to fine-tune without washing out safe behavior. The PDF, a 34-cell Colab, and a Kaggle version are public, and the strong baseline is estimated at 14 hours on a T4. The JudgeBoi grader returned 502 on 2026-09-30, so outside readers have to build their own safeguard evaluation.
HW6 involves no model training and is answered entirely on NTU COOL. Six points come from 16 multiple-choice questions on four papers (ROME, MEND, MEMIT, WISE). Four points come from swapping the Colab's fine-tuning for ROME on GPT2-XL: single editing (pick your own fact, write five kinds of test prompts) and multiple editing (10 and then 80 CounterFact examples, then MEMIT), reporting efficacy, paraphrase, neighborhood, and portability scores. The slides and the 47-cell Colab are public, but the quiz questions and answers live only on COOL.
HW7 hands you two models fine-tuned from Mistral-7B-v0.1: shisa-gamma-7b-v1, strong in Japanese, and WizardMath-7B-V1.1, strong in math. You may only merge them at the parameter level (no further training, no MoE or ensembles), and the merged model has to answer 20 Japanese math questions written by a TA. Part 1 (60%) is tuning the method, weights, and density in mergekit, with simple and strong baselines at 50% and 75% accuracy. Part 2 (40%) is 8 multiple-choice paper questions. The spec, Colab, and Kaggle notebook are public, but JudgeBoi returned 502 on 2026-09-30 and the paper questions live on NTU COOL, so outside readers can only check accuracy inside the notebook.
HW8 involves no coding and no code submission. The TAs provide a finished Colab that runs Llama-3.2-1B-Instruct on the first 100 GSM8K questions and compares direct inference, Self-Consistency, Self-Certainty, and DeepConf (Confidence), sampling 16 reasoning traces per method. You read three papers, run the notebook, and answer 20 questions on NTU COOL: 18 about the papers and 2 about the Colab results. The prerequisite is Lecture 7 (Reasoning) of Lee's 2025 course. All questions are printed in hw8.pdf in Chinese and English, and the Colab is publicly downloadable. Only the COOL quiz and grades need an NTU account.
HW9 has 19 questions worth 10 points, answered only on NTU COOL with no code submission. The first 16 cover four papers — DDPM, Flow Matching, Rectified Flow, and MeanFlow — ending with questions that compare their training signals and few-step generation. The last 3 require the Colab: train two small MLPs on a 2D Swiss roll, one Flow Matching model that learns instantaneous velocity (always evaluated with 50 Euler steps, converged at Histogram JS ≤ 0.10) and one MeanFlow model that learns average velocity (always one-step, ≤ 0.40). Then compare 1 step vs 1 step, Flow Matching across Euler step counts, and Euler vs RK4 at equal steps and at similar compute. The PDF includes a generative-modeling tutorial that skips most of the math, and every question is published in Chinese and English. Outside readers miss only the COOL grading and answers.
KV Cache stores the keys and values already computed so decode does not recompute them, but every token costs memory. For Gemma 2 27B that is about 0.72MB per token, so an 80GB A100 holds only about 114k tokens. Hung-yi Lee then walks through ways to shrink it: let queries share keys and values (MQA, GQA), compress keys and values into one vector without ever decompressing (MLA), limit the attention span (Sliding Window, StreamingLLM), and drop keys and values nobody attends to (Scissorhands, H2O). He ends with cross-conversation prompt caching: it only hits when the prefix is identical, so a system prompt should put stable content first.
The first lecture of Hung-yi Lee's ML 2026 breaks OpenClaw into five questions: how an agent knows who it is, how it uses tools and SKILLs, how it remembers, how it runs on a schedule, and how it keeps working on its own for a long time. Every answer comes back to one fact: the language model only predicts the next token and starts fresh every turn. Identity, memory, and SOPs are all text files that OpenClaw puts into the prompt, or files the model reads and writes through tools. This post walks through the 60-slide intro.pdf and the lecture recording, including the defenses the slides recommend.
Self-attention on its own cannot tell "you hit me" from "I hit you", so the model needs position information from somewhere else. Hung-yi Lee's lecture goes from sinusoidal absolute positions to ALiBi and T5's relative biases, then to RoPE, which Llama, Qwen and Gemma all use. The second half covers train-short-test-long: RoPE breaks when it rotates to angles it never saw in training, which led to Position Interpolation, NTK-Aware scaling, YaRN, Dynamic Scaling and LongRoPE. The final twist is NoPE: causal attention in a decoder-only model already carries position information, and you can even drop the positional embedding after training.
This lecture asks whether a model can catch and fix its own errors with no human in the loop. Hung-yi Lee splits the approaches into three routes. Change inference: the whole contrastive decoding family builds a version of the model likely to be wrong and subtracts it, and the methods differ only in how that wrong version is made. Change the workflow: appending "check again" sometimes helps but is unstable, external feedback beats self-reflection, and under a fixed compute budget, sampling more answers and voting often wins. Change the weights: teaching self-correction directly runs into "after training, the model makes different mistakes," which is why the field moved to RL. Whether RL teaches new abilities or just makes existing paths more likely is still being debated.
Hung-yi Lee opens his May 8 lecture by admitting that "self-improving AI" has no clear definition: it is a process of humans gradually letting go. He splits machine learning into three steps and checks where the "I" can be replaced by AI. Answers can come from the model's own self-corrections, reward shaping can be written by an LLM, the loss can be set by the model itself (scores, majority vote, entropy), and even the questions can come from a proposer model. But experiments keep showing that with no human at all, progress plateaus or the model trains itself into the ground. A strong AI can already train a weaker one, just not better than humans do. His verdict: in May 2026, AI is "still standing at the bank of the Rubicon."
Part 1 was about an AI setting its own loss and updating its own parameters. Part 2 fills in the other half: AI Agent = Harness + LLM, and the harness can grow too. You can't take a gradient through a harness, so the usual move is to hand it to a language model as a rewriter and keep a pool of candidates, much like a genetic algorithm (OPRO, GEPA, Darwin Gödel Machine; DSPy if you want a ready-made tool). Three extensions follow: updating harness and parameters together beats updating either alone; when the goal changes you have to choose between discarding everything and carrying everything, and editing a harness can cause forgetting too; and the update rule itself can be updated (HyperAgent, Gödel Agent, SEAL), which is meta learning. Hung-yi Lee closes with a new analogy — parameters are genes, context is the neurons — then argues that today's agents lack intrinsic motivation, and that the likeliest source of runaway growth is a gap between the goal humans meant and the goal the AI inferred.
Lecture 7 of CME295 (2025) patches three LLM gaps: RAG fixes knowledge frozen at training time with a two-stage retrieve-then-rerank pipeline; tool calling fixes the inability to act by having a backend execute the function call the model writes; agents chain those calls with ReAct's observe-plan-act loop. The 2026 edition renames it AI Agents and adds context compaction, harness optimization, coding agents, and skills, the biggest rewrite in the course.
The 2026 syllabus for CME295 Lecture 6 (November 6, 2026) lists seven topics. Tool calling, MCP and retrieval were already covered in the 2025 Lecture 7; the genuinely new ones are context compaction, harness optimization, coding agents, and skills/plugins. This pre-lecture edition explains those four using engineering posts from Anthropic and OpenAI, the MCP 2026-07-28 spec, and the Meta-Harness paper.
The last CME295 lecture packs 128 slides into three parts: an eight-picture recap of the quarter, how Transformers handle images (ViT and two ways to build a VLM), and masked diffusion LLMs that emit several tokens per step, followed by what comes next in research and applications. It is not on the exam; the 2026 edition turns diffusion LLMs into a lecture of their own and refocuses Lecture 9 on multimodality.
The 2026 edition of CME295 gives diffusion LLMs a full lecture (Lecture 8, November 20), with five listed subtopics: continuous, discrete and masked diffusion, training, and inference. This pre-lecture edition works from the original papers (DDPM, D3PM, SEDD, MDLM, LLaDA and others): continuous noise costs about 64x the compute on text, and the [MASK] absorbing state won out; the training objective is a masked cross-entropy weighted by 1/t; the speed comes from filling several positions per step, yet LLaDA's main results decode one token per step, and Fast-dLLM needs a confidence threshold plus an approximate KV cache to reach up to a 27.6x speedup.
CME295 Lecture 3 defines an LLM as a decoder-only next-token predictor, uses MoE to explain why a huge model only touches part of its weights per token, and spends most of its time on the knobs you can turn at generation time: greedy, beam search, top-k, top-p, temperature, guided decoding, plus three prompting techniques (few-shot, chain of thought, self-consistency). The 2026 edition folds this lecture into Lecture 2, and the prompting half disappears from the syllabus.
CME295 Lecture 8 starts from the fact that human rating is slow and expensive and BLEU/ROUGE can't recognize a paraphrase. It covers how LLM-as-a-Judge works, three biases (position, verbosity, self-enhancement) and six best practices, splits agent failures into tool prediction, tool execution and response generation, and closes with what MMLU, AIME, SWE-bench, HarmBench and τ-bench each measure, plus pass^k and Goodhart's law.
CME295 Lecture 6 breaks reasoning models into three pieces: emit a reasoning chain before the answer, run RL on verifiable rewards like "is the answer correct," and use GRPO, which takes the group's average reward as the baseline instead of training a value model. RL alone took DeepSeek-R1-Zero from 15.6% to 71.0% pass@1 on AIME 2024, and distilling R1's traces into Qwen-32B beat running RL on the 32B model directly.
The 2026 edition of CME295 Lecture 5, "LLM systems" (October 30), lists seven topics: distributed training, inference optimizations, KV caching, speculative decoding, efficient kernels, FlashAttention, and hardware trade-offs. Written before the lecture, this post uses about 70 slides from the 2025 Lectures 3 and 4 plus the original papers to tie them into a single ledger: an H100 needs roughly 295 operations per byte moved to saturate its compute, while token-by-token generation does about 1 per byte of weights read, so most speedups are about moving less data.
CME295 Lecture 4 splits LLM training into two stages: pretraining on trillions of tokens (Llama 3 used 15 trillion), then SFT on thousands to millions of demonstrations so the model stops continuing text and starts answering. In between sits a map of memory savers (ZeRO, FlashAttention, mixed precision); the lecture closes with LoRA and QLoRA, which let people without big GPUs finetune, with QLoRA cutting VRAM by about 16x on a 65B model.
SFT only teaches a model to imitate good answers; it has no way to say which answers are unacceptable. CME295 Lecture 5 covers how to collect preference pairs, walks through the two steps of RLHF (a reward model trained on roughly 10,000 human labels, then PPO on roughly 100,000 examples), and ends with DPO, which folds the whole RL pipeline into a single supervised loss. The 2026 edition splits this lecture between Lecture 3 (training) and a new Lecture 4 (reinforcement learning).
The 2026 CME295 Lecture 4 (October 16) gives RL its own lecture, with seven syllabus items: mathematical conventions, reward design, policy gradients, limitations, PPO, GRPO, and on-policy distillation. This post walks the math ahead of class: start from ∇log π times a score. SFT uses a score of 1, PPO estimates it with a value model, GRPO uses the group mean, and on-policy distillation uses the teacher's per-token log-prob gap. In the Qwen3 report, starting from the same checkpoint, RL reached 67.6 on AIME'24 with 17,920 GPU hours, while on-policy distillation reached 74.4 with about 1/10 of that (1,800 hours).
CME295 Lecture 1 threads a single sentence, "A cute teddy bear is reading.", through the whole class: split it into tokens, turn them into vectors, see why an RNN can't hold on to long sentences, then translate it into French with self-attention and an encoder-decoder. The 2026 edition drops the entire section on NLP tasks and evaluation metrics and opens instead with a timeline running from 2017 to the agent era.
CME295 Lecture 2 takes the original Transformer apart and refits it: position information moves from "added to the embedding" to RoPE's "rotate Q and K inside attention"; attention gets cheaper with sliding windows and MQA/GQA; models split into encoder-only, encoder-decoder, and decoder-only families; and the second half dissects BERT's MLM (15% of tokens) and NSP pretraining. The 2026 edition folds all of this into a single "Large Language Models" lecture, and BERT is no longer a syllabus item.
CMU 11-768's first assignment starts from an empty ReAct loop: a bash-only CodeAgent fixes a bug in a chess app, context compaction is added to solve a SWE-bench task, and the same loop becomes a ChessAgent that runs a two-ply search with simulate_move, run_python, and a skill. All 100 points are graded by replaying submitted patches and trajectories offline.
11-768's Assignment 2 has students use one fixed judge, Qwen3-VL-30B-A3B, to flag four error families in every run of a data-visualization agent, graded by the mean MCC across families on a private set (30% of the assignment). The second half packages students' own tasks as Harbor environments, with one wrong solution the verifier rejects and one that fools it. The theme: the grader you write becomes the RL reward later.
CMU 11-768 is a new Fall 2026 graduate course on agents taught by Graham Neubig and Daniel Fried. The prerequisite — prior experience training language models — is strictly enforced. Its 23 lectures run from tool calling, context, memory, and planning through SFT, RL, sandboxing, and human-agent interaction. Three individual assignments in the first half build a harness, an evaluation, and a training pipeline; the second half is a team research project. Slides and the first nine lecture videos are public.
Lecture 1 of 11-768 strips an agent to its minimum: tool definitions and tool calls are just tokens, the harness parses, executes, and feeds results back into context, and running a ReAct loop makes it an agent. Neubig then lists six capabilities a good agent needs, each of which can be built through training or through the harness, and argues that an agent is a system of harness, sandbox, inference, training, and monitoring — not just a model.
Neubig splits tool use into five layers: capabilities, mechanics, constraints, interfaces, and systems. A tool call is just tokens the model emits; the harness parses, validates, and matches results back by call ID. Constrained decoding guarantees form, not correctness. MCP's real value is credential brokering. And the same model served by different providers can swing from roughly 15% tool-call errors to under 0.1%.
An agent resends its whole history on every call, so five calls already add up to 80K input tokens; 1,500 OpenHands sessions averaged 78K tokens, 37% of them tool results. Neubig works on two layers: at the model layer, hybrid attention (many local layers, one global) plus length curricula make million-token context possible; at the harness layer, stable prefixes earn cache reads roughly ten times cheaper, and compaction that keeps anchors and externalizes evidence gets past the limit — evaluated by how the agent continues afterwards.
Lecture 4 of CMU 11-768 sorts cross-task experience into episodes, facts, and skills, stored as external artifacts rather than in context or weights. Human-written skills load through SKILL.md and progressive disclosure, lifting the average SkillsBench pass rate from 33.9% to 50.5%. Skills an agent induces itself can be tested before admission when written as code, but break easily on a new website. The hard part is the lifecycle: imperfect judges, over-retrieval, and bloated skill libraries each eat into the gains.
Lecture 5 of CMU 11-768 defines an agent's plan as an explicit, inspectable, revisable representation of intended behavior for this task, and gives four reasons to add planning structure: modularity, environment feedback, long horizons, and control. Fried's own MACU has a manager decompose tasks into a DAG and dispatch parallel sub-agents, raising Odysseys success from 8.5% to 34.0%; on an OSWorld subset, no planning scores 25.0%, an initial DAG with no revisions scores 27.8%, and allowing 10 revisions reaches 58.3%.
Neubig's L6 splits coding agents into three layers: train a model that can code (pre-training, mid-training, infilling, RL from test rewards), wrap it in a localize–edit–verify loop with the right editing tools so it can change a repo, then evaluate and train it in SWE-bench-style executable environments. Fixing bugs is only about 15% of a developer's day; the next frontier is tests, CI, and maintenance in the outer loop.
JY Koh breaks computer use agents into three questions: evaluation has moved from single clicks (ScreenSpot-Pro, Mind2Web) to programmatic end-state checks (WebArena, OSWorld), VLM judges, and long-horizon rubrics (Odysseys, OSWorld 2.0); the model is a VLM reading interleaved screenshots and actions; training runs pre-training for grounding → SFT on human and synthetic trajectories → RL in resettable simulated environments.
Yueqi Song breaks agent SFT into six decisions: compute loss on assistant tokens only (including the stop token); choose trajectories carefully (runs that pass tests can still teach bad habits, and switching teachers or adding new tasks beats sampling more); unify formats with the Agent Data Protocol; watch packing and template drift during training; evaluate in the real harness; and pick the SFT checkpoint for the RL that follows, not for its own best score.
Using a guess-a-number-from-1-to-16 game, Daniel Fried frames RL as an extension of SFT: he names SFT's three gaps (task mismatch, no learning from failures, never seeing its own mistakes), then derives ReST, REINFORCE, baselines, and GRPO/DrGRPO in turn. All four compute log-probabilities of the tokens the agent itself generated; they differ only in the weight each token gets.
Akari Asai's L10 splits deep research agents into three problems: evaluation has to cover four gaps (search difficulty, domain expertise, long-form answer quality, citation support); training runs mid-training → SFT → RL, with DR Tulu's evolving rubrics as the reward for long-form reports; retrieval should let the retriever see the agent's reasoning, which gets AgentIR-4B to 68% on BrowseComp-Plus with Tongyi-DR.
Using a bug-fix coding task, Graham Neubig takes Lecture 9's policy gradient into practice: a critic, GAE, or a PRM to credit individual turns; importance ratios and clipping to handle stale data in async RL; and PPO, GRPO, CISPO, GSPO, and DAPO side by side in one table. The largest share goes to the reward itself — verifier errors, reward hacking, and exploration collapse — before closing with on-policy distillation.
The Analysis methods unit of CS224U (Spring 2023) starts by grading three families of methods on a three-column scorecard. Probing is strong at characterizing representations but can't support causal claims. Integrated gradients only gives you a scalar about each representation, but it satisfies the sensitivity axiom, so it does come with a causal guarantee. This post covers slides 1–40, videos 33–35, and feature_attribution.ipynb, including where the notebook breaks in today's environment.
CS224U's fourth unit opens with one question: what can behavioral testing prove, and what can't it? It can never give a guarantee, and when a model fails you first have to ask whether the model or the dataset is at fault. BERT scored 2.2% on negated NLI examples, then 90% after fine-tuning on a small set of them. The unit then covers SQuAD distractor sentences, Breaking NLI, ANLI's human-and-model adversarial collection, and ends with DynaSent's two rounds.
Causal abstraction rests on one operation: take the internal state a model computes for a source input at some location, swap it into the same location for a base input, and check whether the output changes the way your hypothesized high-level program says it should. In CS224U's iit_equality.ipynb, a network with 0.99 test accuracy scores only 0.50 and 0.54 on this check. After IIT training, its counterfactual accuracy is 1.00. DAS replaces guessing which neurons match which variable with learning a rotation matrix.
A few COGS generalization splits score 0 for nearly every model. CS224U uses its own ReCOGS work to explain why: the zeros on the recursion splits are mostly a length-generalization problem, and the zeros on the prepositional-phrase split come from training data that only ever put PPs in certain variables and positions. Assignment 3, hw_recogs.ipynb, uses 135K ReCOGS training pairs. It first has you find Charlie and Lina, two names whose train and test roles are exact opposites, then shows a trained model stumbling on them.
CS224U Spring 2023 tells the story of the Transformer families through BERT's four known limitations. RoBERTa addresses the first (optimization was only partly explored). ELECTRA addresses the second and third (the [MASK] mismatch, and only about 15% of tokens giving a learning signal per batch). XLNet addresses the fourth (the assumption that masked tokens are independent of each other). GPT changes the objective and the mask, T5 and BART change the architecture and how inputs are corrupted, and distillation changes model size. The course's 2023 view: autoregressive architectures have taken over, but bidirectional models may still have the edge for representation.
The first three parts of CS224U's Spring 2023 contextual representations unit start with examples like "break" and "crane" to show why static word vectors were never going to be enough. They then build a Transformer block step by step on the three words "The Rock rules." Only attention connects the columns; every other step runs on each column independently. Finally, two questions sort three positional encoding schemes: Do you have to fix the set of positions ahead of time? Does the scheme get in the way of generalizing to new positions? Absolute encoding fails both, sinusoidal encoding passes the first, and the relative encoding of Shaw et al. (2018) passes both.
The second half of CS224U's 'NLP methods and metrics' unit skips metric formulas. It asks whether your experiment holds up. Naturalistic or crowdsourced data, adversarial or common cases: the course answers 'both' each time. Lock the test set away. Pick baselines when you write the hypothesis. Compare two models with confidence intervals, Wilcoxon, or McNemar, and run several random initializations. The slides, three videos, and two notebooks are all public. Kawin Ethayarajh's guest session 'Real-world NLP assessments' has no public slides or video.
In Spring 2023, CS224U slipped two talks by members of its own teaching team into the Transformer unit. Lisa Li presented Diffusion-LM: instead of generating left to right one word at a time, it denoises a sequence of Gaussian vectors into word vectors. It loses to autoregressive models on both training and decoding efficiency, and in exchange lets a classifier's gradient steer the output at every step. Sidd Karamcheti showed how to cut GPT-2 Small's single-GPU training clock from 99.63 days to 3.37 days by stacking data parallelism, mixed precision, and ZeRO. The first talk survives only as slides; the second has slides and two recordings.
CS224U's first assignment, hw_sentiment.ipynb, is ternary sentiment classification: you develop on two rounds of DynaSent plus SST-3, and the bake-off test set mixes in mystery sentences from undisclosed sources. The original-system question is worth 3 of the 9 homework points, and it has exactly one rule: never touch the three public test sets during development. Run as-is today, the first data-loading cell breaks because Hugging Face datasets 4.0 dropped trust_remote_code.
CS224U's second assignment, hw_openqa.ipynb, asks you to answer questions that come with no passage, using only a frozen language model and a frozen ColBERT retriever. The Spring 2023 version was written for DSP; in January 2024 the repo switched to DSPy and pinned dspy-ai==2.4.13. Before you start you need an OpenAI API key, a ColBERTv2 checkpoint of about 406 MB, and a 600 MB prebuilt index. The notebook's first setup call, dspy.OpenAI, no longer exists in DSPy 3.4.
The Spring 2023 edition of CS224U defines in-context learning as a frozen language model performing a task only by conditioning on the prompt, and warns that the second condition of few-shot learning (no examples of the behavior seen in training) is almost impossible to verify. Potts's 38-page deck runs from GPT-2's TL;DR trick through choosing demonstrations, chain of thought, self-consistency, and DSP, and ends with four recommendations: build dev/test sets first, learn your target model's instruction format, and treat prompt writing as AI system design. Mina Lee's guest lecture asks the reverse question: who should learn to read prompts, people or models?
The Spring 2023 edition of CS224U spends a whole unit on retrieval, because OpenQA gives you only the question and you have to find the evidence yourself, while large language models fabricate sources. The slides by Potts and Omar Khattab go from TF-IDF and BM25 through Success@K, MRR, and average precision to four neural IR designs (cross-encoder, DPR, ColBERT, SPLADE) and how each trades expressiveness against scale, and they end by asking you to count latency and cost as metrics too.
The first CS224U lecture of Spring 2023 asks "Which U.S. states border no U.S. states?" of every system from Chat-80 (1980) to text-davinci-001. The answers show that the progress is real. The lecture then questions whether that progress counts as understanding, using Levesque's "cheap tricks," models that invent links, and benchmarks that saturate within a year or two. That splits the course map in two: the first half teaches you to build systems with Transformers and retrieval-augmented in-context learning, and the second half teaches you to test them with harder benchmarks, behavioral evaluation, and causal explanation methods.
The first two deliverables of the CS224U final project are a literature review and an experiment protocol. The lit review covers 5, 7, or 9 papers depending on team size, under five suggested sections. The protocol has seven required sections, and its core is a hypothesis you can state. The course supplies a six-step paper-search loop, a rule that AI-assistant output must be quoted, and a worked example: a student's final project that became a Findings of EMNLP paper. The Gradescope format and rubric slides, and past exemplary papers, are behind a login.
The CS224U slides compute two numbers from one three-class confusion matrix: accuracy 0.81 and macro F1 0.43. One says the system is good; the other says it gets the two small classes almost entirely wrong. The unit's claim is that different metrics encode different values, and it goes through the bounds, values, and weaknesses of accuracy, the three F-score averages, perplexity, word error rate, and BLEU. Final projects are graded on whether the metrics fit, not on how high the scores are.
CS224U's 'Presenting your research' lecture has four parts: the course-specific rules for the final paper, how to write an NLP paper, how conference submission works, and how to give a talk. Three things matter most. The final paper must include Known project limitations and an Authorship statement. Write as a Shieber-style 'rational reconstruction,' not a chronological tour of your dead ends. At submission, your title largely decides reviewer bidding. The slides, four videos, and projects.md are public; past example papers need a Stanford login.
learn-inference.com is an unofficial interactive companion to Philip Kiely's Inference Engineering (256 pages, Baseten Books, free PDF). It follows the book's 8 chapters and 42 sections with rewritten explanations, turns intuition-heavy ideas like TTFT, P99, speculative decoding, and prefix-cache routing into slider-driven simulators, and ships a keyless JSON API and MCP server.
CME295 is a two-unit Stanford course with no homework; your grade is the midterm and the final, 50% each. The 2025 edition's nine lectures are fully public: videos, slides, and both exams with solutions. The 2026 edition rewrites the agent lecture around context compaction, harnesses, coding agents, and skills, and adds three full lectures on LLM systems, reinforcement learning, and Diffusion LLMs.
Traditional chat only needs text bubbles, but AI Agent conversations must surface thinking, tool calls, citations, and progress — we solved this with 12 Vue components, three DisplayModes, and a unified chatBlocks rendering pipeline.
In H1 2026, OpenAI, Anthropic, Letta, and LangChain independently chose Markdown files + indexes over vector databases; write permissions shifted back to humans; forgetting mechanisms appeared but nobody implemented Ebbinghaus; Penfield Labs caught 6.4% wrong answers in LoCoMo; three vendors simultaneously adopted 'Dreaming' for offline memory consolidation. Five trends, one conclusion: memory is not a feature — it is an architecture decision.
Memory turns prompt injection from a one-shot nuisance into a persistent backdoor: MINJA shows conversation-only injection succeeds >95% of the time, and SpAIware demonstrated continuous data exfiltration via planted memories. The industry's two defensive lines — citation-based verification (Copilot) and human approval inboxes (Gemini CLI / Devin) — each have blind spots.
Agent memory is not one feature — it is at least four distinct engineering problems: working, episodic, semantic, and procedural. This ten-part series walks through the full design space, from taxonomy to coding agent implementations, platform APIs, open-source frameworks, security attack surfaces, and 2026 trend analysis.
CoALA splits agent memory into working, episodic, semantic, and procedural — but the four-cell taxonomy alone doesn't explain why Claude Code uses Markdown files while Mem0 uses vectors. This post adds six independent design axes (read mode, write timing, fidelity, write authority, forgetting, scope) and a file-to-graph spectrum to map the full design space of agent memory systems in 2026.
The previous articles covered evaluation. This one covers another dimension: how to dynamically allocate compute during reasoning. BrowseConf's core insight is that an agent's self-declared 'confidence' can predict answer accuracy. High confidence uses fewer resources; low confidence searches more rounds.
All five major cloud platforms shipped agent memory APIs in 2025–2026, but their design philosophies diverge sharply: OpenAI writes memory as files, Anthropic mounts memory as a directory, Google uses vectors with topic classification, AWS combines events with pluggable strategy pipelines, and Microsoft abstracts memory behind context providers. Pricing ranges from free to $0.75/1K records/month; tenant isolation spans from 'your app handles it' to IAM as a first-class citizen.
Nine coding agents have taken at least four different paths for long-term memory: Claude Code and Codex use Markdown files (agent-written, human-readable), Antigravity CLI inherits Gemini CLI's approval inbox (agent proposes, human decides), and Copilot uses citations with JIT verification (auto-deleted after 28 days unverified). Cursor removed its Memories feature and fell back to human-written Rules. Hermes Agent, OpenClaw, and Letta Code treat memory as a core harness component, not a plugin. Their choices on write timing, forgetting, and cross-team sharing are completely different — and none has published a controlled experiment on whether their memory system actually helps.
By 2026, the deep research commercial market has differentiated: OpenAI is comprehensive, Perplexity is fast, Gemini integrates ecosystems, Claude reasons deeply, Grok is real-time. This article compares each product's differences—not who is best, but who fits your scenario.
The entire Deep Research field has 80+ implementations, but the core structure is just three-stage roadmap × four components × three optimization methods. This article maps the full landscape: from Agentic Search to Full-stack AI Scientist, from query planning to answer generation, from workflow prompting to end-to-end RL.
DeepResearch Bench II uses 9,430 expert rubrics covering 132 tasks, and finds that even the strongest agents satisfy less than 50% of criteria. This article breaks down the benchmark architecture, scoring methodology, leaders, and the overall evaluation landscape.
An AI assistant platform running Opus 4.6 hit stream_stall (90s timeout) twice consecutively when generating docx. Root cause: skill instructions lacked one sentence — 'Write a script' — causing the model to output JS code inline instead of writing a file and executing with node. Claude.ai's official SKILL.md has that sentence, and the model consistently takes the safe path.
10+ community deep-research skills represent 10+ philosophies of 'how to do research.' From hyperresearch's persistent vault to jamoeight v2's Co-Scientist 6-agent, from adversarial verification to benchmark alignment. This article puts them all on one table.
A deep research agent produces a report—maybe thousands of words with dozens of citations. How do you score it? Using LLMs as judges is biased, asking humans is too expensive, and benchmarks can't keep up. STC and other recent approaches try to solve this from the 'confidence' angle—but there's no perfect answer yet.
Deep research has already evolved from 'help you search' to 'help you research.' But the next step is bigger: self-evolving agents, swarm collaboration, scientific automation. This article covers three directions and an uncomfortable reality: Gartner predicts 40% of agent projects will be cancelled by 2027.
An AI assistant platform's image search needed three rounds of fixes: pushing node_type filtering into the ES query to stop text chunks from hogging top-k slots, making filename matching deterministic instead of relying on the LLM to pass an optional parameter, and switching from Postgres icontains to ES match for CJK-aware partial matching.
When a research agent runs 25, 100, or 2000 turns, what happens? Context suffocation: information piles up, noise increases, attention gets diluted. IterResearch solves this with Markovian state reconstruction; AREX achieves recursive self-improvement with an inner/outer loop. Both answer: how does an agent stay coherent across hundreds of search rounds?
Seven open-source memory frameworks span the spectrum from auto-extracted vectors to human-readable files: Mem0's one-line add(), Graphiti's bi-temporal knowledge graph, Letta's agent-edited system-prompt blocks, LangGraph's namespaced Store, LlamaIndex's priority-based block truncation, Cognee's triple-store pipeline, and Supermemory's temporal vector-graph engine. This post compares their storage, write/forget mechanics, tenant isolation, and benchmark numbers, then offers selection guidance for four common scenarios.
The deep research open-source ecosystem has evolved from 'single frameworks' to 'tool clusters.' This article compares 12+ projects: GPT-Researcher emphasizes multi-agent collaboration, STORM simulates expert conversations, smolagents focuses on state management. Each tool solves different problems.
This is the project's own deep research skill design, fully disclosed. Core choices: only Groundlane MCP for web tools, strict source-quality grading (A/B/C/D), research hands off to post skill for publishing. Not the most powerful, but the best fit for us.
About 10% of conversations on our AI assistant platform randomly lost knowledge base tools — the agent hallucinated answers from training data instead. Root cause: a bare except Exception swallowed Elasticsearch connection failures during retriever initialization, silently skipping tool registration. Fix: retry + surface failures to system prompt + structured metadata tracking.
Previous articles covered the landscape, training from scratch, long-horizon memory, and planning optimization. This one zooms out to see a complete system that threads all these insights together: Tongyi DeepResearch. Its core innovation is Agentic CPT — inserting an agentic mid-training stage between pre-training and fine-tuning, giving the model an inherent agent bias. MoE 30B parameters activating 3B, HLE 32.9 surpassing OpenAI o3.
A user set retrieval top_k to 15, but the monitoring dashboard showed 5 and the streaming UI flashed 5 before jumping to 15. The same top_k value existed at four layers — LLM tool arguments, runtime, trace DB, and streaming payload — each requiring its own override. Three sequential fixes, each revealing the next layer was also wrong.
Parsing a 150-page PDF via Vision API took 29 minutes (one page per request). Batching cut it to 1.5×, adding 5-way concurrency brought it down to 24 seconds. One week after launch, a customer uploaded 23 PDFs at once — 70 parallel Bedrock requests triggered full throttling: 90 pages skipped, 5 files failed, 8 stuck. Fixed with Redis-based cluster-wide slots + backoff retries.
Two NeurIPS 2025 papers answer the same question: how to train a web research agent from scratch? WebThinker chooses 'bolt on web capability to existing reasoning models,' WebDancer chooses 'rebuild everything from data construction to RL training.' Two philosophies, four stages, one core insight: training beats prompting.
Previous articles covered how to train agents. But training requires high-quality data—and deep research training data has been scarce. WebShaper solves this with mathematical formalization: define IS tasks in set theory, then use an agentic Expander to iteratively expand them. S1-DeepResearch goes further: moves training from 'search-centric' to 'real research.'
All deep research agents are 'text-first'—but the real world isn't just text. WebWatcher (NeurIPS 2025) is the first system to integrate visual reasoning into deep research, using OCR, image search, code execution, and other tools to handle charts, screenshots, videos, and other diverse information.
Previous articles covered training from scratch and long-horizon memory. This one goes deeper: how to make the agent's 'planning' itself better? WebWeaver tackles it architecturally (dual-agent iterative outline optimization). DeepPlanner tackles it through training (advantage shaping for planning tokens). Both point to the same conclusion: planning is the ceiling of deep research.
An agent with a workspace_browse tool said 'file not found' instead of searching. Anthropic, OpenAI, and Google's official guides all point to the same fix: put trigger conditions and workflows in the tool description. A 2025 study found 97.1% of MCP tool descriptions have quality issues.
Agent-to-agent communication falls into three patterns: handoff (transfer control), delegate (dispatch and wait for results), and mailbox (real-time peer-to-peer messaging). Implementations vary widely, but MCP and A2A are driving protocol standardization.
Should a sub-agent see the parent's conversation? Fork carries full history but token costs grow exponentially. Fresh saves money but lacks context. Industry consensus: default to Fresh, Fork only when needed, and always pair it with history truncation and result compression.
Parallel + nested agent spawns can burn 200K+ tokens in a single conversation turn. From Anthropic to Microsoft, the industry is converging on tiered responses: compress → downgrade → stop, rather than a binary kill switch.
By 2026 nearly every mainstream coding agent supports subagents. Design philosophies split three ways: deterministic scripted orchestration (Claude Code Workflow), model-driven autonomy (Codex, Devin), and IDE command-center integration (Windsurf 2.0, VS Code). This overview maps product positioning, a capability matrix, and the design-philosophy spectrum.
The most common debug nightmare in multi-agent systems is 'the answer is wrong, but I don't know which agent did it.' Three layers of observability are essential: per-agent token metering, execution traces, and real-time cost dashboards.
Multi-agent orchestration splits into three camps: scripted determinism (LangGraph, Claude Code Workflow) is predictable but rigid, model-driven (Codex, Devin) is flexible but unpredictable, and hybrid (Windsurf 2.0) acts as a command center integrating multiple agents. The choice depends on how much predictability you need.
Multi-agent security risks aren't just amplified single-agent risks — inter-agent communication is itself an attack surface. A compromised sub-agent can pass malicious instructions to the parent through its return value. Core defense: treat agent output as untrusted data.
Three routes to commercial document parsing: specialized parsers (Cohere Parse at $1.50/k pages, LlamaParse Agentic Plus at 90.2% on ParseBench), Big Three cloud prebuilts (Azure/Google/AWS for structured field extraction), and general-purpose VLMs (Fable 5.1 scores 78.92 on ParseBench and crushes specialized parsers on charts, but costs 3–16× more and hallucinates). At 100K pages/month, plain OCR runs ~$150 across providers; add tables and AWS jumps to $1,500, Claude Sonnet 5 to $900. The first question isn't 'which is most accurate' — it's 'do you need transcription or comprehension?'
A consolidated guide drawing on the arXiv official guidelines, the NeurIPS ML Reproducibility Checklist, the REFORMS framework (8 modules, 32 items), and the preprint policies of five top conferences. Covers paper structure, experimental design red lines, the arXiv submission workflow, and a decision framework for conference vs. direct submission.
Claude Academy's opening lesson identifies the core paradox: AI accelerates code generation, but review, testing, and deployment don't keep pace. The bottleneck shifts from 'not writing fast enough' to 'not reviewing fast enough.' The AI-native SDLC fix isn't more AI-generated code — it's embedding AI into every stage where bottlenecks now live.
Traditional requirements scatter across Jira, Slack, and meeting notes, losing fidelity at every handoff. intent.md lets the originator collaborate directly with Claude to produce a human-readable, machine-actionable, version-controlled Markdown proto-spec — from conversation to committed document in hours, not weeks.
Traditionally, requirements analysis and design are separate phases run by different teams — every handoff loses information. This lesson's approach: Claude reads intent.md in a single session, applies organizational standards (brand, security, compliance, UX loaded as skills), and produces a unified spec.md. The product owner reviews; they don't author.
Claude Code's plan mode lets engineers produce a reviewable, version-controlled implementation plan (plan.md) before writing a single line of code. Design review shifts from the PR diff to the planning stage, and the cost of course-correcting drops from 'rewriting code' to 'editing a document.'
CLAUDE.md is a context file at the repo root that Claude reads at the start of every session — your team's conventions, commands, architecture patterns, and pitfalls. The course's core advice: if Claude makes the same mistake twice, write it into CLAUDE.md.
A skill is organizational tacit knowledge made operational — a folder with a SKILL.md that Claude loads automatically when trigger conditions are met. The course's key principle: skills make violations rare; hooks make them nearly impossible.
One engineer runs multiple Claude Code sessions simultaneously, each in its own git worktree; repetitive verification work goes to subagents. The bottleneck shifts from 'writing code' to 'reviewing output.'
Have Claude verify its own output before submitting — tests, builds, and screenshot diffs all run to completion before the task is marked done. Engineers receive code that's already passed verification, not code that 'might be correct.'
Evals are the AI-native equivalent of stage-gate QA — collect 20–50 real tasks as test cases, run them automatically whenever CLAUDE.md, skills, or hooks change, and block the merge if the pass rate drops. Every production incident becomes a permanent eval.
Let AI handle the first review pass so humans can focus on intent and risk. This lesson covers how to define REVIEW.md, layer review passes, set up an automated review-comment fix loop, and why the agent that wrote the code must never approve its own PR.
Hooks are the governance bedrock of the AI-native SDLC — deterministic gates that intercept agent actions and block them if conditions aren't met. This lesson walks from a single production-gate script to full enterprise managed settings covering permission lockdown, sandboxing, credential isolation, and marketplace allowlists. The most technically dense lesson in the entire course.
Plug Claude into the CI/CD pipeline — start with read-only build failure triage, gradually add write operations behind existing gates, expose deployment tooling through MCP, and tier autonomy by environment. The governing principle is one sentence: 'The agent may act up to the production gate and cannot pass it.'
Stage 6 is both the endgame and the starting point of the AI-Native SDLC: a monitoring script detects an anomaly → Claude writes a diagnosis as intent.md → it flows through the entire development pipeline. Humans shift from 'starting work' to 'triaging and reviewing work.'
After 14 lessons, this final article distills the series into three things: a prioritized adoption roadmap, role-specific reading paths, and Anthropic's complete official documentation list.
Anthropic's Claude Academy offers a free 14-lesson course that takes AI-assisted coding from 'individuals using Claude Code' to 'an organization-wide development lifecycle.' Four core concepts — intent.md, CLAUDE.md, Skills, and Hooks — wire together into a complete AI-native SDLC.
WebMCP is a browser-native W3C standard proposed by Google and Microsoft. It lets web pages expose structured tools to AI agents via document.modelContext.registerTool() — no backend, no HTTP/SSE transport. Chrome 149 Origin Trial is live.
Week 2 runs as a one-two punch: Monday's Anthropic taxonomy teaches you when not to build an agent, Wednesday's RAG paper hands you the first complete compound-system recipe. Five workflow patterns are the selection toolkit, RAG is parametric-plus-nonparametric memory, and together they are the blueprint for HW1's email retrieval pipeline.
Week 3 standardizes tool interfaces with the MCP specification on Monday and trades hand-written pipelines for compilable, optimizable programs with the DSPy paper on Wednesday. HW1 drops the same Monday and bans agent frameworks: you build a company's internal assistant from scratch, so this week DSPy is for understanding what frameworks abstract, not for handing in.
Week 4 pins the agent loop down as an interleaved think-act-observe sequence with the ReAct paper on Monday, then turns memory into OS-style tiered storage with the MemGPT paper on Wednesday. The same week, HW1 is growing from an email-retrieval pipeline into a full harness where memory is an explicit requirement, so the loop shape and the memory design are the two things to settle now.
Week 5 turns multi-agent collaboration into programmable conversation with AutoGen on Monday, then lays out the three optimization axes — prompts, weights, inference compute — with GEPA and the test-time compute paper on Wednesday. HW1 is due 10/30, the last full week before the deadline, so this installment helps you decide which axis deserves your effort.
Week 6 assigns Shankar's data flywheel on Wednesday — evaluation, monitoring, and continual improvement feeding on the same production data — while HW1 comes due, HW2 drops, and the midpoint demo video and midway report loom in early November.
Week 7 is midterm checkpoint week: data selection on Monday, evaluation and benchmark design on Wednesday. Zhu et al. teach you not to be fooled by your own scores, SWE-smith scales software-engineering tasks to 50,000 instances, and you close by drafting a first 4-tuple for HW2.
Week 8 builds model judges with MT-Bench and Anthropic's eval guide on Monday, then faces production leakage with PrivacyLens and four guardrails on Wednesday. The paper video is due Friday, and this week's deliverable is one working judge score plus one permission check.
Week nine turns to coding agents: SWE-agent shows interface is performance, OpenHands packs sandbox plus benchmarks into one general base, and the second homework is due Friday — ship one working bug-fix exam this week.
The finale reads Week 11: Monday upgrades instruction-waiting reactive assistants into proactive agents that observe, infer, and act first via the GUM paper, while Wednesday folds multimodal systems, long-running agents, and production observability into three open problems. Ends with a pre-Demo-Day checklist and a one-line map of all 11 posts.
The Week 1 anchor reading for CS329Z is Zaharia et al.'s Compound AI Systems: the best results increasingly come from multi-component systems, and even the biggest model is just one part. The post leaves three design questions and three hard challenges — which happen to be exactly what HW1 asks you to answer by building.
A close look at three open-source projects training LLMs from scratch in the Chinese community — baby-llama2-chinese (218M, 63.4B tokens), ChatLM-mini-Chinese (0.2B T5, 10.23M dialogues), and Steel-LLM (1.12B, 1T tokens, 8 months) — comparing corpus strategy, tokenizer decisions, and community ecosystem. Honest evaluation included: baby-llama2 scored a bottom-ranking 21 in MiniMind's side-by-side test; ChatLM has the strongest knowledge (62) but weak coding.
Four coding agent TUIs all follow the same two rules: tools maintain constant height (one-line summary, never jumping from 0 to N lines) and never auto-collapse (only user-initiated expand/collapse). looplane violates both.
OpenELM uses layer-wise scaling to shift parameters toward layers near the output; with 1.08B parameters and 1.5T tokens it beats OLMo 1.2B (+2.36% on the LLM360 average) despite OLMo training on 3T tokens. MiniCPM trains multimodal small models from scratch with a three-stage unfreezing recipe (Resampler first, vision encoder next, everything unfrozen last); MiniCPM-V 4.5 reaches sub-30B SOTA on VideoMME with only 8B parameters, and 4-bit quantization squeezes fp16's 16–17GB memory footprint down to about 5GB for phones.
A tour of Karpathy's three teaching repos: nanoGPT (2022, ~300 lines each to reproduce GPT-2 124M), llm.c (pure C/CUDA training), and nanochat (2025-10, one speedrun.sh from tokenizer to WebUI). $100 and 4 hours on 8×H100 buys a chatty model; GPT-2-grade capability is now down to about 2 hours and $48.
LitGPT (Lightning AI, ~13,600 stars, Apache 2.0) rewrites 20+ mainstream LLMs — Llama 3, Qwen2.5, Phi 4 — from scratch as single-file, no-abstraction implementations, with a full pretrain / finetune / evaluate / serve CLI. TinyLlama (1.1B parameters, 3T tokens) was trained on this codebase. This post breaks down how it differs from MiniMind, how to actually use it, and where it stops.
MiniMind is an open-source project for training LLMs from scratch: a 64M Dense model and a 198M-A64M MoE model that run the entire chain — Pretrain → SFT → LoRA → DPO → PPO/GRPO/CISPO → Agentic RL — in ~2 hours on a single RTX 3090 at roughly 3 RMB (~$0.40). Every core algorithm is implemented natively in PyTorch with no high-level wrappers.
Switzerland's Apertus (8B/70B, 15T tokens, 1,000+ languages) filters opt-outs and personal data before training to satisfy the EU AI Act; Japan's LLM-jp consortium shipped LLM-jp-4 (12T tokens) in April 2026, claiming wins over GPT-4o and Qwen3-8B on standard benchmarks. Both prove that from-scratch training outside the English sphere is a data-governance problem, not a technical one — plus a note on RWKV-7 as the non-Transformer alternative.
"Open-source LLM" is a spectrum: weights-only (Llama), weights plus data (most fully open projects), or data order, intermediate checkpoints, and training logs all released (LLM360 K2, OLMo 3's model flow). This piece unpacks the two projects that pushed transparency furthest: OLMo 3 shipped the first fully open 32B thinking model in November 2025, and K2 is the first 65B-class model whose checkpoints even include optimizer states.
The hands-on installment of the series: rent an RTX 3090 on RunPod (Secure Cloud $0.5/hr, Community Cloud $0.22/hr), follow the MiniMind README through pretrain (~1.21h) + SFT (~1.10h), spend roughly $0.55–1.50 USD total, and chat with your own 64M model trained from scratch in the terminal.
Open-source projects have pushed the cost of training an LLM from scratch absurdly low: MiniMind runs the full PreTrain-to-RL pipeline for about $0.4 (2 hours on a single RTX 3090), while at the other end OLMo 3 and LLM360 K2 publish everything — data, code, and stage-by-stage checkpoints of 65B models. This series walks the whole project spectrum from $0.4 to 65B in 11 articles, and flags where the map is biased.
Training from scratch only makes sense in three cases: you want to learn how training works, you have 10B+ clean tokens no open model has seen, or you need a fully transparent training process for research. Otherwise fine-tuning or RAG is almost always cheaper. This post collapses the series' main routes into one cost ladder and a decision tree.
YuLan-Mini is a 2.4B open-source model from Renmin University's AI Box lab, trained on 48 A800 GPUs with only 1.08T tokens — scoring 37.8 on MATH-500 and 64.0 on HumanEval, beating Qwen2/Qwen2.5 peer models trained on 7T–18T tokens at math and code. What's public isn't a slogan: per-phase data mixes, pre-annealing optimizer states, and even W&B logs of the ablation studies.
Traditional document parsing runs a fixed pipeline regardless of input, but contracts, financial reports, and technical manuals each need different strategies. Agentic Parsing lets LLM agents observe a document and dynamically choose tools — AgenticOCR parses only the regions that matter (70%+ visual token savings), and ParseBench shows even the best method scores only 84.9% across 2,000 enterprise pages. No silver bullet.
ColPali renders each PDF page as an image, generates patch-level multi-vector embeddings with a vision-language model, and retrieves via MaxSim late interaction. On table-heavy financial PDFs, recall jumps from 62% to 84% — no OCR, no chunking. The tradeoff: ~100× storage, GPU required, no BM25.
Small chunks give precise embeddings but lack context; big chunks have complete context but diluted embeddings. Hierarchical Chunking builds multi-level indexes (2048→512→128 tokens) with an Auto-Merge algorithm: leaf nodes match precisely, and when hit density exceeds a threshold the parent node is returned to the LLM instead. HiChunk shows a 12.7% evidence recall improvement; LlamaIndex and Haystack have it built in.
AI boosted task output by 34%, but code review time surged 441% and measured delivery actually slowed 19%. A four-round research survey maps the current landscape: deterministic guardrails (hooks) vs probabilistic ones (prompts), clean-context review, self-improving feedback loops, specification-driven development, AI test quality crisis (high coverage but median 53% mutation score), and the Replit agent fabricating test results.
Standard RAG retrieves one set of documents per query, but real questions often need reasoning across 2-4 documents. IRCoT pioneered interleaved retrieval-reasoning, PAR²-RAG beats IRCoT by 23.5% accuracy on four benchmarks, and CompactRAG compresses LLM calls down to just two.
Self-RAG trains four reflection tokens (Retrieve / IsREL / IsSUP / IsUSE) into an LLM, letting it decide on-the-fly whether to retrieve, whether results are relevant, and whether its own output is grounded. ICLR 2024 Oral (top 1%), the 7B model beats ChatGPT and Llama2-chat + RAG on multiple QA benchmarks. The catch: it requires fine-tuning—no API-only models.
Markdown-KV format achieves 60.7% LLM comprehension accuracy vs 44.3% for CSV — a 16-point gap from format alone. But retrieval and comprehension have different optimal formats: metadata prepend + row-wise key-value is the current best combination for table RAG.
Routines created by Claude itself (created_via: meta_mcp) prompt for connector approval on every call. User-created ones (created_via: http_api) don't. Fix: recreate the routine via the RemoteTrigger API.
Looplane now grades effects as read/modify/modify_execute/execute and fails closed on unclassified tools. Native MCP tools default to execute unless trusted read-only metadata lowers them. Approval events still land in events.jsonl first and grants can scope to one change set or backend; general command rules and universal sandbox coupling remain unfinished.
looplane now ships a bounded tool-program DSL: read-only programs support list/read/search/diff, repeat, and if_contains; modify/check transactions receive whole-transaction approval and roll back touched paths on failure. This is not arbitrary JavaScript/Python code mode, and transaction execution is not parallel.
Mature-agent compaction must handle triggers, complete-turn cut points, and recovery. looplane now has an 85% high-watermark, automatic compaction, a deterministic native-loop fallback summary, persisted checkpoints, and workspace-context reinjection. Cross-runtime fallback, model-quality summaries, and live-provider long-session validation remain open.
omp and claude-code provide cross-session memory while the other references mostly rely on instruction files. looplane now has an explicit remember/list/inject baseline: typed JSONL memories enter prompts across sessions, but retrieval is scope-and-recency only, with no semantic ranking, deduplication, forget command, or automatic extraction.
All five projects combine allow/ask/deny decisions, compound-command inspection, and fail-closed behavior. looplane now has a deny-first classifier, critical floor, shell segmentation, timeout-deny, configured allow/deny rules, and visible policy reasons. Broader syntax coverage and live interactive validation remain open.
I'm building my own Python coding agent called looplane. This series dissects the source code of five mature projects — pi, oh-my-pi, opencode, codex, and claude-code — topic by topic, while also comparing them with Looplane's current TUI, external CLI runtimes, local gateway, usage/OTel/session tooling, and Cloudflare slice. Every post follows a fixed five-part structure: design problem → how five projects do it → looplane's choice → academic grounding → improvement roadmap, with evidence cited at file#symbol level.
Hooks govern control flow, skills inject knowledge, and plugins package both. looplane now has opt-in deny-only project hooks, a bounded SKILL.md loader, exact enabled_skills selection, plugin manifests/install/list, and external-runtime projection. Input rewriting, full lifecycle coverage, remote registries, and a mature marketplace remain open.
looplane can now inject repository diagnostics and open-file state into the next model turn, push typed IDE context over WebSocket, package a VS Code bridge, and supervise long-lived LSP subprocesses through ManagedLspServer. Language-specific initialize/didOpen/didChange adapters and live-editor validation remain open.
An MCP client must handle transports, tool refresh, approvals, and credential boundaries together. looplane now supports allowlisted stdio, Streamable HTTP/SSE, tools/resources/prompts, tools/list_changed, OAuth metadata/PKCE, and a 0600 credential store. A real authorization-server E2E and MCP-specific confirmation UX remain open.
looplane now has static ModelRole/ModelRoute candidates, opt-in aliases such as --model @cheap, cross-provider fallback, and a no-tool reviewer lane that runs after verification. Role inheritance/override rules and automatic summarizer, parser, or scout routing remain open.
Wrapping an SDK directly buys you three walls within months: usage fields that don't agree, error semantics tied to SDK exception types, and tool-call formats that change per provider. All five reference projects separate 'wire protocol' from 'provider identity' as independent dimensions. Looplane goes further with pydantic canonical contracts (Message/ToolCall/Usage/ModelTurn) plus six protocol adapters, forces the OpenAI SDK's built-in retries to zero, and routes every failure through a classified ProviderErrorKind before any retry policy sees it. Its provider table is deliberately copied from pi's packages/ai — lineage, not coincidence.
An OS sandbox is the kernel boundary beyond path policy. looplane now ships a fail-closed CommandSandbox: sandbox-exec on macOS, Landlock plus seccomp on Linux, and exit 126 when containment cannot be proven. Coverage still focuses on verification commands, and external CI confirmation remains open.
Looplane's prompt is now `m3-exact-edit-v4`: the version persists into artifacts; core/tool/interaction/runtime/instructions/skills/workspace/memory are composed as stable or dynamic sections; and positive/negative examples cover replace_text, unified diffs, and direct replies. Unit tests pin the structure, while live-eval coverage still needs expansion.
Intermittent NVIDIA NIM 500s exposed Looplane's early gap: classified errors with no retry consumer. SDK retries are now disabled; the harness gives each candidate up to five attempts with jittered exponential backoff and capped Retry-After handling, then can move to an explicitly configured fallback model. Both model.retry and model.fallback enter the event log.
looplane now connects events.jsonl to a deterministic reducer, CLI timeline, canonical JSON, SDK replay, and safe event-point forks. Forking never replays prior tools or model calls; provider/live-runtime validation, redaction, and richer replay hooks remain open.
Mature subagents need roles, bounded fan-out, narrowed permissions, and a result contract. looplane now has native named-role schedules, parallel fan-out, child allowed_paths constrained by the parent, unsafe execution disabled by default, and parent-approved transaction proposals. Persistent background lifecycles, recursion trees, and automatic worktree merging remain open.
looplane now has CostBreakdown, an explicitly estimated static GPT-5-family price table, per-lane usage/cost, and OTel cost fields. Unknown models still show tokens without invented dollars; broader pricing coverage, authoritative external-CLI bills, and live billing reconciliation remain open.
Looplane's core surface has grown from seven tools to nine with a read-only `tool_program` and rollback-capable `tool_transaction`; search prefers ripgrep and arbitrary shell remains absent. Native MCP tools join only from allowlisted servers and default to execute approval without trusted read-only metadata.
Looplane's disposable clone and SafePathPolicy protect the source repo. `--sandbox-checks` can now wrap verification commands with macOS sandbox-exec, Linux bubblewrap, or Landlock, while Cloudflare provides a separate bounded Sandbox slice. Network policy, external-runtime coverage, and production hardening are not yet consistent across those backends.
looplane now has a Cloudflare Durable Object run resource with async creation, status/cancel/artifacts, live NDJSON, and Last-Event-ID SSE; remote approvals use a separate short-lived capability. Python also provides an attach client and a stateful conversation WebSocket. Production deployment, cross-runtime parity, and full multi-tenant hardening remain unverified.
The first observation of 'What course articles do you have?' was contaminated by an old cache entry. A real cache miss retrieved all four university maps but spent 51.169 seconds across three Writer and Critic passes; after catalog-specific retrieval and review fixes, one uncached production observation passed q21 in 26.821 seconds.
Ask AI first extracts intent, complexity, and 1–4 search terms. It then routes across metadata, BM25, Vectorize, and RRF; a retry adds Critic gaps and disables the first-pass-only BM25 short circuit.
Ask AI indexing runs in two production stages: source-hash changes update D1, post chunks, and FTS5 first; embedding checkpoints and a delete queue then let Vectorize catch up asynchronously. The two stores do not share one transaction.
Ask AI splits one question across the UI, `/api/chat`, Planner, Research, Writer, Validation, Critic, and Related stages. Answer text, displayed sources, and related-reading cards come from separate paths with separate gates.
Ask AI keeps golden contracts, offline fixtures, live SSE output, and production observations separate. A passing fixture proves harness reproducibility; public sources can measure expected-source recall, but they do not expose hidden ranked chunks or establish model-graded faithfulness.
One Ask AI request leaves five different evidence surfaces: public SSE, Langfuse traces, D1 logs, semantic cache, and a hidden shadow run. They expose different data, and no single surface reconstructs the complete retrieval context.
Ask AI finding a post does not mean the UI should display it as a source. An answer must pass deterministic Markdown and URL validation, then the Critic's relevance, intent, and grounding checks; if either gate fails, source cards are withheld.
Writer sees the first 8 candidates for a factual query or 12 for a recommendation by default. Citations must use an exact `source_url` from that set, and weak or empty retrieval triggers an instruction to abstain rather than fill gaps from model knowledge.
Agent Memory is a Cloudflare private beta service for letting agents remember users, teams, projects, and task context across conversations. It fits facts, events, instructions, and tasks; RAG documents, product data, files, and audit logs should still live in AI Search, Vectorize, D1, or R2.
Cloudflare Agents turns an agent session into a durable runtime: each agent instance has stable identity, local SQLite, WebSockets, scheduled work, recoverable execution, and tools. It is not just a chat example; it composes Workers, Durable Objects, AI models, Browser, Sandbox, AI Search, and MCP into a deployable agent app.
An AI app should not put conversations, artifacts, memory, retrieval documents, locks, and eval traces into one store. D1 fits queryable product data, R2 fits large files and artifacts, Durable Objects fit named coordination and per-session state, and Agent Memory / AI Search / Vectorize handle memory and retrieval.
AI Gateway is the control plane for AI calls: one layer for logs, analytics, cache, rate limits, retry/fallback, BYOK, and Unified Billing. In Workers, use env.AI.run(..., { gateway }); with external SDKs, change the baseURL or provider-native endpoint.
The Cloudflare AI Stack series covers the infrastructure around AI apps: where models run, how gateway control works, how RAG is built, how agents keep running, how memory is governed, and how browser, sandbox, secrets, data, and observability fit into a product.
Secrets Store is Cloudflare's open beta account-level secret store, currently integrated with Workers and AI Gateway. It fits provider API keys, BYOK keys, and secrets reused across Workers; per-Worker secrets still work, but the governance scope is different.
Vectorize is Cloudflare's vector database. AI Search is the right starting point for a managed RAG pipeline; Vectorize is the better fit when you need control over chunking, embeddings, metadata filters, hybrid retrieval, reindexing, and fallback behavior.
GPUtw.ai makes sense as a short-rental GPU learning tool: start with Jupyter, Ollama, or ComfyUI, then try LoRA/QLoRA on a small model. It is not a large foundation-model training platform, and the first run should verify deployment, billing, and data retention with a small budget.
Keenable.ai positions itself as search infrastructure for AI agents: a 100B+ document index, Search/Fetch APIs, MCP/CLI entry points, 100K free monthly requests, and keyless public endpoints. It is worth tracking, but the 100B+ index, latency, and quality claims are still mostly company-provided; NEEDLE is open, but needs external reruns and human review.
screenshot-to-code is not a one-shot screenshot-to-HTML tool. Its core is a 30-step Agent Loop with 7 tools — extracting real assets from screenshots, self-verifying with Playwright, and running 4 models in parallel so users pick the best output. 74,500+ GitHub stars, MIT License.
Warp's self-improving agent pattern is not about dumping every mistake into a prompt. A base skill does the work, humans leave feedback in GitHub or Slack, an improver skill turns repeated signals into a small diff, and humans review the PR before the next run inherits it.
TinyFish provides four web APIs for AI agents: Search, Fetch, Agent, and Browser. Search and Fetch are permanently priced at $0 with no credit card requirement, making them a practical default layer for RAG and document retrieval.
Four philosophies of ADE workspaces: arul28/ADE's Brain+Lane, Superset's 100-agent IDE, Herdr's Rust-native runtime, and Kadro/Orca's pane and Fleet angles — compared in one table with a decision tree for when to pick a workspace over an Omnigent-style control plane.
Omnigent moves governance off prompts into a Server-side Policy engine: Python functions returning allow/deny/ask, a three-layer stack with cost budgets and tool caps, plus Omnibox OS-native isolation via bwrap/seatbelt and egress credential brokering — compared with five peer governance stacks.
The same model can score 20 points apart on different harnesses, 32% of SWE-bench Pro verifier judgments were found to be wrong, and DeepSWE's 113 tasks make most models score zero. This guide decodes six major coding benchmarks — what they test, which are easy to game, and which ones you should care about.
CS50 AI's OpenCourseWare edition publishes seven weeks of lectures, slides, notes, and twelve Python projects with autograder feedback, plus a free CS50 Certificate if you score at least 70% on every project. The catch: weeks 0–5 still use the Spring 2020 recordings; only Week 6 (Language) was re-recorded, in 2023.
The same model can cost 2-5× more depending on the channel. Direct API is simplest, aggregators (OpenRouter) are most flexible, cloud platforms (Bedrock/Vertex) suit enterprises. This post compares actual August 2026 prices across six channels with a decision tree.
meta-harness means two things: Databricks' control plane and Stanford's outer-loop optimizer. This post uses a four-layer model (MCP/ACP/Runtime/meta-harness) to place Omnigent, Zed ACP, Vercel HarnessAgent and Cloudflare Flue.
MIT 6.7960 Deep Learning (Fall 2025) publishes all 21 lecture decks as public Dropbox PDFs, and most required readings map to free textbook chapters; but the five problem sets are released only through Gradescope, and solutions plus recordings live behind Canvas login. This guide covers how the three instructors split the course, a topic map of all 21 lectures, textbook-based substitutes for lectures, and where outside self-learners realistically stop.
Nearly every frontier open-source model in 2026 is MoE: Ornith 35B activates only 3B to beat 31B dense models, MiniMax M3 uses 456B total but 45.9B active to hit SWE-bench Pro 59%, DeepSeek V4 runs 1.6T total with 49B active. This post explains why MoE dominates coding and agentic benchmarks using four case studies.
The same Polly task — parallel git worktrees plus cross-vendor review — implemented four ways: Omnigent YAML governs at the Server layer, LangGraph controls flow with a StateGraph, CrewAI assembles roles quickly, and Goose ships a desktop Recipe, compared on tokens, latency, and maintainability.
Allen AI's OLMo is the only language model family that fully publishes weights, training data (Dolma, 9.3T tokens), training code, all intermediate checkpoints, and evaluation tools. OLMo 3's 32B Think model hits 96.1% on MATH — and you can use OlmoTrace to trace any output back to the exact training data that produced it.
Databricks' open-source Omnigent wraps Claude Code, Codex, Cursor, Pi and custom agents in a Runner/Server + Omnibox sandbox, adding three-layer Policies and shareable persisted Sessions so you can swap models and harnesses with one-line changes — 9.3k stars, still alpha.
'Open-source' in AI doesn't mean what it means in software. MIT and Apache 2.0 let you do almost anything; the Llama License requires a separate deal above 700M MAU; old Gemma terms let Google change rules unilaterally (Gemma 4 switched to Apache 2.0). This guide maps what you can and can't do by license type.
Open-source models now match closed-source on coding benchmarks, but self-hosting isn't just picking a model — vLLM handles high-concurrency production serving, SGLang is 29% faster on prefix-heavy workloads, Ollama is the local dev default, and llama.cpp runs on the least hardware. A100 cloud rentals run ~$1.4-2.2/hr; self-hosting breaks even at roughly 100M tokens/month.
Three non-big-lab teams used different RL post-training strategies to produce benchmark dark horses in 2026: Ornith's self-improvement loop (GRPO), Nous Research's DataForge + Atropos execution-reward RL, and MiniMax's massive-scale RL across 200K real environments. Different strengths, but one shared proof point: post-training RL matters more than pretraining scale.
Models don't read words — they read tokens. A Chinese character is typically 1-2 tokens; an English word is 1-3. The context window is the token limit per request. Inference is using a model; training is teaching one. What you do every day is inference.
Models don't understand text — they only understand numbers. Embeddings map each token to a vector of several hundred dimensions, where semantically similar words end up close together in vector space. This is the shared foundation behind search, RAG, and classification.
Benchmark scores in model releases have three common traps: cherry-picking (only showing wins), contamination (test data leaking into training), and saturation (when everyone scores 90%+, the benchmark stops being useful). The most manipulation-resistant signal is Chatbot Arena's Elo ranking — real humans, blind voting, uncontrolled questions.
Data changes often and you need citations → RAG. Need consistent style or want to run on a small device → fine-tuning. In practice, many production systems use both: fine-tune a small model that speaks your domain language, then use RAG to supply up-to-date facts.
A model uses loss to know how wrong it is and gradients to know which direction to adjust. Gradient descent repeats three things: compute loss, compute gradients, update parameters. The learning rate controls step size — too large and you overshoot, too small and training takes forever.
Every time a model predicts the next token, it assigns a probability to every candidate word. A loss function measures how far that probability distribution is from the correct answer — the further off, the higher the loss, the more the model knows it got it wrong. Cross-entropy is the standard formula; perplexity is its human-readable translation.
A 70B model needs ~140GB VRAM in FP16, but 4-bit quantization shrinks it to ~35GB. With llama.cpp's partial CPU offloading, it can run on consumer hardware. GGUF naming conventions (Q4_K_M, Q5_K_S) tell you the precision-size tradeoff. KV cache is why long conversations slow down.
Scaling laws show that loss decreases predictably with more parameters, data, and compute — following power-law relationships. The Chinchilla paper's key finding: most models were too large and undertrained. Given the same compute budget, training a smaller model on more data produces better results. This reshaped the entire industry's training strategy.
You don't need to become a researcher to understand AI models systematically. This series starts from what you can see (tokens, context windows) and works up to self-hosting open-source models — 18 articles covering everything you need to choose models, read benchmarks, and estimate costs.
Models charge by tokens, not characters. The BPE algorithm starts from individual bytes and repeatedly merges the most frequent adjacent pair to build a vocabulary. English 'understanding' might be 1-2 tokens, but Chinese '理解' could take 2-3 — same meaning, higher cost.
Every LLM goes through three training stages: pre-training reads the internet to learn language, SFT uses example conversations to learn the format, and RLHF uses human preferences to learn what a good answer looks like. The gap between a base model and a chat model is what the last two stages do.
The core of the Transformer is self-attention: for each token, the model computes how relevant every other token is, then takes a weighted sum. This lets the model reach across distance to figure out that 'it' refers to 'cat' not 'mat' — and is the foundation for how it handles long documents.
In 2025 RAG stopped being 'retrieve once, generate once.' Search-R1 trains models to search autonomously in multiple turns with RL, REX-RAG/AlignRAG add policy and alignment branches, OpenAI Deep Research productizes the loop, and MCP generalizes retrieval into unified tool invocation. This post unpacks the design philosophy, trade-offs against ten generations, and when to adopt the new paradigm.
Apple is giving App Store Small Business Program developers free access to AFM 3 models on Private Cloud Compute if their apps have fewer than two million first-time downloads. The five-model family includes the sparse 20B-parameter AFM 3 Core Advanced, which activates only 1–4B parameters on-device, and AFM 3 Cloud Pro on Google Cloud NVIDIA GPUs, refined with outputs from Gemini.
BytePlus ModelArk Coding Plan offers Lite ($10/month) and Pro ($50/month) subscriptions covering models such as DeepSeek-V4, GLM-5.2, and Seed-2.0 in tools including Claude Code and Cursor. Lite includes about 24,000 requests per month; Pro includes five times as many.
pi's loop is a double while-loop wrapped in an EventStream; claude-code's source openly says stop_reason is unreliable and uses tool_use blocks observed during streaming as the sole continue signal; codex models a turn as a cancellable SessionTask and records sessions with a dedicated rollout crate. looplane chose an ordering — manifest first, JSONL second — that turns Ctrl-C into verified resumption instead of a rerun. All evidence cited at file#symbol level.
Mature coding-agent CLIs have converged on the same conventions: positional prompt, -p means print, exec is headless, resume is a first-class command, -C changes directory; looplane inherits this vocabulary directly, driving learning cost close to zero.
LLMs break unified diffs on bookkeeping: wrong hunk counts, hallucinated context lines. The five reference projects split into two camps — simplify the diff grammar (Codex drops line numbers), or drop diffs entirely (Claude Code/Pi/OpenCode exact replace); OMP goes further by binding read state into the format via hash anchors. looplane took the minimal-intervention path: keep the guarded apply_patch, add a zero-fuzzy replace_text, and its qwen3:4b eval went from stable failure to 5/5.
Every mature coding agent ships a machine interface: codex has `exec --json` plus a full app-server JSON-RPC protocol, claude-code has `-p` with stream-json, and pi/opencode/omp each expose a JSON event stream. Wrapping these CLIs as your backend is the fastest path to subscription-backed coding — but they own their agent loop, their login, and their permission model. looplane's answer: let the external CLI fully own its loop while looplane holds only three things — an isolated working copy, patch audit, and final verification. One runtime never impersonates another.
The ecosystem treats /v1/chat/completions as the lingua franca, but your providers don't all speak it. The five reference projects split into three camps: pi and OpenCode make the client speak every dialect natively so no gateway is needed; OMP builds a real protocol translator (foreign wire → neutral context → provider adapter, no raw passthrough); Codex and Claude Code run proxies that translate nothing and exist purely to force traffic through a controllable path. Looplane copies OMP's boundary but narrows it to one wire in, one out: strictly parse OpenAI Chat into a canonical contract, then dispatch to any ModelProvider — and along the way hit a cross-event-loop client-close bug whose lesson is that provider lifecycles belong to the ASGI lifespan, not the signal handler.
The biggest problem when an agent enters CI is approval: no TTY, nobody to click approve. The five reference projects converge on two strategies — delegate permission decisions to the calling program (claude-code's control protocol), or replace approval semantics entirely (codex defaults to Never plus sandboxing, opencode auto-rejects). looplane keeps one AgentRunner loop and injects a different ApprovalPolicy: headless uses HeadlessApprovalPolicy, which never reads stdin so it cannot hang the pipeline, and denies EXECUTE by default — fail closed.
A blank config file drives people away; a bad credential discovered too late drives them away faster. All five mature agents treat setup as a first-class state, and looplane adds the step most of them skip: verify the key right after saving it.
After an agent run finishes, 'the model said it's done' is not evidence. Codex splits traces into a manifest + JSONL + payloads bundle, omp mirrors on-disk files into SQLite, pi indexes native session files with runs.jsonl. Looplane picked the strictest option: six fixed files per run, the run is incomplete if any is missing, and patch review reads changes.patch—not anyone's verbal claim.
Five external CLIs expose five different machine interfaces: JSONL event streams, JSON-RPC handshake, HTTP API, ACP, stream-json. The right way to support them is not one interface that pretends they're identical — it's a narrow runtime boundary plus an honest capability matrix. Availability means installed, not authenticated; protocol drift fails closed.
A local sandbox limits the blast radius of an agent on your machine; a cloud sandbox is about moving code safely onto someone else's machine. All five mature projects solve the first problem; only looplane actually deployed the second. Lessons from production: mocks can't catch SSE framing, green CI can't catch a stale wheel, and cleanup paths deserve timeouts just as much as success paths.
All five agents store sessions as append-only JSONL plus some form of single-writer protection, but crash recovery lives in the details: pi repairs torn tails, codex reopens and retries after write failures, and looplane picked a 'manifest first' ordering that reduces the only crash window to one repairable slot. This post dissects each project's write ordering and fail-closed conditions, all cited at file#symbol level.
Small models don't fail at reasoning first — they fail at format stability: tool-call JSON, diff hunk arithmetic, and context budgets all break. The mature harnesses build evals on real model behavior (pi's model-backed evals, OMP calibrating benchmarks from real session logs, Codex even relaxing its parser for weaker models). looplane picks the narrowest but hardest path: one fixture, five real Ollama runs, a manifest declaring exactly which files and patch fragments count as success — and M2's failure kept verbatim as evidence. Never pass mock off as E2E; never spin partial success into full passes.
A CLI tool pays its startup cost on every invocation, and performance optimization without a baseline means no regression protection. codex uses daemon reuse and skill snapshot caches; claude-code splits its entrypoint into dynamic imports plus a built-in startup profiler; opencode and omp each maintain lazy-loading discipline; pi does none of it and leans on Bun being fast. looplane is Python — slow by birth — so it applies the full discipline: lazy imports, single-flight disk cache, background controller prewarming, and hyperfine paired benchmarks wired to a CI gate that fails on >10% regression.
The five reference projects split into three camps on subscription auth. Codex and Claude Code implement OAuth only for their own official clients and store tokens in the OS keyring. pi and OMP directly reuse Claude Code's client ID to implement Pro/Max OAuth — technically feasible, but Anthropic's docs explicitly bar third parties from offering claude.ai login without approval. OpenCode removed its bundled Pro/Max plugins entirely, the cleanest policy precedent in the ecosystem. Looplane's rules: own your grant, never scrape another CLI's credentials, accept third-party OAuth only when the provider clearly supports it, and never copy or forward credentials.
An agent's two dependencies — the LLM and external CLIs — are both non-deterministic, but mature projects separate 'the moving parts' from 'the shape of the boundary': codex fakes the Responses API with wiremock plus a scripted SSE server and pins its TUI with insta snapshots; opencode built a VCR-style http-recorder package; pi splits model-backed evals from unit tests into two vitest configs; omp wraps its edit benchmark itself in unit tests. looplane stacks four layers against external CLIs: unit tests, fake-CLI contract tests, recorded-stream integration proofs, and Textual pilot TUI tests. The methodology in one line: record real non-deterministic output, then make deterministic assertions about it.
Mature coding agent TUIs never print the event stream directly — they build a typed projection layer first and update it in place. looplane took three steps (full-screen composition, runtime-first dual modes, removing the Ask/Agent split) before two old constraints — non-streaming output and resume-without-replay — were truly lifted.
None of the five reference projects enforces 'all declared verification commands pass' at the harness level: pi leaves verification to the model, OpenCode and Codex put it in the system prompt, Claude Code uses a separate adversarial verifier subagent but as a soft contract, and only OMP's cleanse actually runs checks from harness code. looplane takes the hardest path: if files changed, every declared verification command must pass before terminal_reason=verified; with no changes, checks don't rerun (no_changes). Whether to verify is decided by code, not by the model.
None of the five mature coding agents use Python — pi/opencode/claude-code run on TypeScript, codex rewrote TS into Rust, omp bolted ~80k lines of Rust native crates onto its hot path. looplane still chose Python; the costs are startup performance and packaging, compensated by lazy imports, uv, and Cloudflare Sandbox.
Same 'knowledge graph + retrieval' label, three different bets: Microsoft GraphRAG v3.1.2 pays indexing cost for global summarization, LightRAG cuts cost with dual-level retrieval and incremental updates, HippoRAG 2 turns RAG into growing associative memory via PPR — this guide splits the trade-offs by component with four query modes, indexing pipelines, and a selection matrix.
Anthropic Contextual Retrieval uses an LLM to prefix each chunk with 50-100 tokens, cutting failure rate from 5.7% to 1.9% with rerank at ~$1.02/1M tokens; Late Chunking encodes the full 32K-window document first then mean-pools by chunk boundaries for zero extra LLM cost — the trade-off is window, latency, and update shape.
Unsloth is the fastest, most VRAM-efficient local LLM fine-tuning tool — 2× training speed and 70% less VRAM. In 2026 it added a Desktop app that bundles inference, training, image/video generation, web search, and agent integration into a complete local AI workstation.
2021 was the year Transformers decisively entered computer vision: Swin Transformer won the ICCV Best Paper award, DINO showed that a self-supervised ViT could learn object segmentation without labels, and NeRF grew from one paper into an entire subfield. CVPR and ICCV both moved fully online because of the pandemic, yet the work published that year shaped architectural choices across computer vision for years to come.
2021 was the year diffusion models surpassed GANs, self-supervised learning made theoretical breakthroughs, and reinforcement learning confronted weaknesses in its evaluation methodology. NeurIPS received a then-record 9,122 submissions, ICLR’s Score-Based Generative Modeling paper became a theoretical foundation for the diffusion ecosystem, and ICML delivered substantial work on optimization theory and the dynamics of self-supervised learning.
2021 marked NLP’s shift from fine-tuning an entire model to adapting only a small fraction of its parameters. Prefix-Tuning at ACL, LoRA on arXiv, and Prompt Tuning at EMNLP all appeared that year; ACL Rolling Review launched; and the Findings track established itself as a second publication channel.
2021 was a dividing line for major AI conferences. Transformers spread from NLP throughout computer vision and time-series research, self-supervised learning became the most common cross-conference theme, and a diffusion model won an ICLR Outstanding Paper award before anyone realized it would displace GANs. Meanwhile, GNNs and federated learning reached historic peaks in paper volume before beginning to decline.
2022 marked computer vision’s turn from recognition toward generation. Latent Diffusion Models appeared at CVPR and led to Stable Diffusion; NeRF research jumped from 25 papers in 2021 to more than 50 at CVPR alone; ConvNeXt mounted a compelling counterattack for CNNs; and ECCV in Tel Aviv set a record with 157 oral papers.
2022 was the year diffusion models took center stage, Chinchilla scaling laws rewrote large-model training, and Chain-of-Thought turned reasoning into an ability that prompts could elicit. NeurIPS passed 10,000 submissions; three of its 13 Outstanding Papers directly concerned diffusion; and Chinchilla and data pruning both challenged the belief that bigger was always better. On the eve of ChatGPT’s release, every required piece fell into place at that year’s conferences.
2022 marked NLP’s shift from demonstrating model capabilities toward aligning and controlling them. InstructGPT brought RLHF into the mainstream, Chain-of-Thought showed that prompts could unlock reasoning, and Flan 2022 matured instruction-tuning methodology. ACL and NAACL adopted ARR as their sole review path, exposing infrastructure and reviewer-load problems. ChatGPT launched at year-end and rewrote the rules of NLP research.
2022 was a turning point at major AI conferences. Diffusion models moved from emerging to mainstream, with two NeurIPS Outstanding Papers; Chinchilla rewrote scaling laws; Chain-of-Thought showed that large models could reason; and InstructGPT used RLHF to teach language models to follow instructions. When ChatGPT launched at year-end, these academic topics instantly became global news.
In 2023, computer vision moved from seeing images to understanding, generating, and controlling them. Segment Anything turned segmentation into a general zero-shot capability, ControlNet made diffusion models precisely controllable, and 3D Gaussian Splatting challenged NeRF with real-time rendering. CVPR received more than 9,000 submissions and ICCV more than 8,000 as both conferences returned to in-person events.
In 2023, LLMs took over the machine-learning conference agenda. NeurIPS received more than 12,000 submissions; both Outstanding Papers addressed large models, while runner-up DPO became a practical alternative to RLHF within two years. DreamFusion opened the text-to-3D field, ICML spotlighted LLM watermarking and learning-rate adaptation, and the Mamba preprint emerged as the first serious architectural challenger to the Transformer.
2023 was the first full academic year after ChatGPT, and LLMs rewrote the NLP conference agenda. ACL's Best Papers examined humor understanding and the propagation of political bias; an EMNLP Best Paper explained in-context learning through information flow; and the HackAPrompt competition paper also won an EMNLP Best Paper award, signaling that security research had entered the mainstream. The year's largest shift was from asking how to make models more accurate to asking how we can tell when a model is misleading us.
2023 was the first year in which LLMs comprehensively rewrote the AI research agenda. DPO received a NeurIPS Outstanding Paper Runner-Up award, ReAct became an ICLR Oral, and hallucination grew from a marginal term into a major track at every conference. Meanwhile, 3D Gaussian Splatting swept through computer vision after its SIGGRAPH debut, Mamba emerged at the end of the year to challenge the Transformer attention monopoly, and publication volume for traditional NLP pipelines began a clear decline.
In 2024, 3D Gaussian Splatting took over 3D reconstruction, video generation moved from research toward products, and vision-language models spread into specialized domains. CVPR received a record 11,500-plus submissions; its Best Papers were Google Research's Generative Image Dynamics and the UCSD/Google collaboration Rich Human Feedback for Text-to-Image Generation. ECCV gave its Best Paper award to Columbia's Minimalist Vision with Freeform Pixels, an unconventional return to the physics of optics.
ML conference submissions exploded in 2024: NeurIPS received a record 15,671 papers, while ICML and ICLR passed 9,000 and 7,000. Research shifted from training ever-larger models toward spending inference compute more intelligently, making test-time compute scaling the year's defining new direction. VAR beat diffusion with next-scale image prediction, Rectified Flow became the theoretical foundation for Stable Diffusion 3, and ICLR gave its inaugural Test of Time Award to the original VAE paper.
NLP conferences redefined themselves under LLM dominance in 2024. ACL made open science its annual theme, and four of its seven Best Papers probed fundamental limits of language models. EMNLP turned toward multilingual and cross-cultural work, with Best Papers spanning speech representations and gradient interpretability. ACL and EMNLP received more than 10,000 submissions combined, but the deeper anxiety was what remains of NLP when LLMs can perform nearly every traditional NLP task.
The defining conference keywords of 2024 were agents, alignment, multimodal LLMs, and inference-time compute. The LLM share at five major conferences doubled again after its sharp 2023 rise; agent-related terms grew 4.3 times; and diffusion models graduated from an emerging topic to a second generative-AI pillar alongside LLMs. Traditional task-oriented NLP continued to contract, while GANs almost disappeared from top venues.
2025 was a two-conference year for computer vision, with CVPR and ICCV both taking place. CVPR received a record 13,008 submissions; Best Paper VGGT turned 3D reconstruction from iterative optimization into feed-forward inference. ICCV's Marr Prize went to BrickGPT, which generates brick structures from text that can actually be assembled. 3D Gaussian Splatting displaced NeRF, video generation moved toward products, and flow models began replacing diffusion, completing several paradigm shifts in one year.
ML conferences broke every submission record in 2025 and pushed peer review to its limit. NeurIPS received 21,575 papers and used more than 20,000 reviewers; ICML passed 12,000 for the first time, and ICLR reached 11,565. Reasoning and agents were the strongest trends. One NeurIPS runner-up, the conference's only perfect-score paper, challenged whether RLVR creates new reasoning ability. Awards for Alibaba Qwen's Gated Attention and a mechanistic theory of neural scaling laws showed a community moving from scaling at all costs toward understanding why scaling works.
NLP conference submissions nearly doubled in 2025: ACL received 8,360 papers and EMNLP 8,174. China-based first authors exceeded 51% at ACL, and DeepSeek's Native Sparse Attention won Best Paper. The deeper story was an identity crisis: an ACL president said 'ACL is not an AI conference,' a quantitative study asked 'Has ACL Lost Its Crown?', and EMNLP faced questions about what still distinguished it from ACL or NAACL.
The two strongest signals at AI conferences in 2025 were reasoning papers jumping from 47 to 216, a 4.6-fold rise, and agent-related terms exceeding 150 papers with 4.3–11-fold growth. Diffusion moved from breakout topic to infrastructure; RAG became a mainstream enterprise architecture with unusual coverage across all five conferences; state-space models and world models began tracing the early 2020–2021 path of Vision Transformers. Pure prompt-engineering papers encountered reviewer fatigue.
Publishing at a top conference as an independent researcher is possible, but the numbers are harsh: single-author papers have fallen to a single-digit share, the average author count has risen from 3 to 5, and the top 20 institutions account for 35-50% of authorships. Andreas Madsen spent eight months working without pay, earned an ICLR Spotlight, and still ended up returning for a PhD. This article examines real cases, evidence of review bias, and viable paths for researchers without a large lab behind them.
An AI conference paper passes through anonymized submission, format screening, reviewer bidding and assignment, independent scores from 3-4 reviewers, an author rebuttal, AC/SAC/PC decisions, and camera-ready revision—a process lasting about 4-5 months. ACL-family conferences add ARR's rolling-review model, in which review comes before the author commits the paper to a venue.
A paper submitted to a major conference can follow three very different routes: the Main Track is the highest-threshold formal publication, Findings is the ACL family's companion venue for solid work that misses the main program, and NeurIPS created the D&B Track specifically for datasets and evaluation methodology. Their review standards, prestige, and career signals differ enough that understanding the route matters before writing the paper.
The institutional map of top AI conferences is being rapidly redrawn. Industry labs dominate frontier model development—nearly 90% of notable models came from industry in 2024—but academia remains the largest source of highly cited research. Chinese universities went from challengers to nearly half of NeurIPS paper volume in five years, while OpenAI and Anthropic have nearly vanished from conference author lists. The decoupling of publication volume from research capability is the defining signal.
Stanford Marin pre-registers a paloma macro-loss of 2.04 with a 5-rung Scaling Ladder at 1% cost, then trains 535B-A23B on 11×GB200 in public with live W&B telemetry — 847 training buckets already show the most teachable frontier run.
There's no official certificate for being an 'AI top conference.' It's a community consensus built from four independent signals — CCF-A, CORE-A*, a high Google Scholar h5-index, and a low acceptance rate — and those four signals frequently disagree. ICLR being completely absent from CCF's list is a live example.
An academic-search pipeline cannot simply concatenate five APIs: use arXiv or PubMed for domain discovery, align OpenAlex and Semantic Scholar records through DOI, PMID, and arXiv IDs, then use Crossref and PubMed relationships to check the version of record, corrections, and retractions.
AG2 continues AutoGen's ConversableAgent model: agents collaborate through messages, while GroupChatManager selects the next speaker by round robin, manual choice, randomness, or an LLM.
These seven tools are not one product category: LangGraph, MAF, and Mastra emphasize durable workflows; CrewAI and AG2 emphasize multi-agent collaboration; Pydantic AI emphasizes typed Python agents; DSPy optimizes AI programs against data and metrics. Choose the control model first.
An agent should not send the user's sentence unchanged to every search service. Classify the need as exact lookup, keyword, semantic, or fielded search; move source, date, language, and field constraints into native provider parameters; then rewrite according to zero-result, overbroad, stale, or source-mismatch symptoms.
Amazon Bedrock is more than a reseller for multiple model APIs. It brings model invocation, IAM, Regions, Knowledge Bases, Guardrails, and CloudWatch into one AWS control plane. It fits teams already on AWS that value governance overhead more than the lowest token price.
Phoenix is an MIT-licensed open-source LLM observability and evaluation platform. It collects traces with OpenTelemetry and OpenInference, turns production failures into versioned datasets, compares prompt, model, or RAG changes in experiments, then writes code, human, and LLM evaluator scores back as annotations. It is not Arize AX, and self-hosting defaults require security work.
Authenticated browser state is not a convenience setting; it is a credential that can impersonate its owner. Use a dedicated low-privilege account and isolated profile, separate reading from reversible writes and high-risk transactions, and leave MFA plus final submission to a human.
Baseten puts custom-model packaging, GPU deployment, inference engines, autoscaling, and release workflows on one platform. Its value is not another OpenAI API, but retaining runtime control while operating less GPU orchestration.
Braintrust connects versioned datasets, immutable experiments, scorers, and production traces into one evaluation loop. Its value is not another score but the ability to turn production failures into offline tests. The company announced an $80 million Series B in February 2026; its customer list is company-reported.
Brave Search API exposes five endpoint families—Web, News, Images, Videos, and LLM Context—backed by Brave's own Web index and ranking models. Its core search is not merely a Google SERP wrapper.
Browserbase combines remote Chromium, persistent Contexts, proxies, and a Session Inspector in one control plane. It operates browser fleets; it does not decide an agent's next action. As of August 2026, the company reports more than 35 million monthly browser sessions and over 10,000 customers.
Cartesia's core is Sonic real-time TTS, Ink STT, and streaming inference. Although it offers the Line voice-agent platform in 2026, buyers must still separate the model layer from telephony orchestration and design consent, retention, and fallback for cloned voices.
Cerebras can dramatically accelerate generation on supported models, but agent latency still depends on prefill, tool I/O, model quality, and platform compatibility.
Chroma manages embeddings, documents, and metadata through collections; it embeds into Python locally, uses HNSW on a single node, and separates compute from storage with object storage, SSD caches, and SPANN in distributed deployments.
Anthropic interviewed 15 startups and distilled five Claude Code operating principles: everyone ships, automate the tedium, trust but verify, build for rebuilding, prototype to productionize. ClickHouse shipped 30% more features, Clay automated 100% of bug triage, Artemis Security hit 6,000+ PRs per week.
Kitesurf is a non-Chromium browser backend in Browser Run that remains in beta. It trades pixel compatibility, persistent authenticated sessions, WebGL, and full anti-bot behavior for low CPU and memory through Workers isolates, Rust/Wasm, and stateless components.
Cloudflare Sandboxes uses a Worker as the entry point, a named Durable Object as the control plane, and a Container inside an isolated VM as the execution plane. It fits Cloudflare-native fleets of ephemeral Linux workspaces, but persistence, security boundaries, and three layers of billing remain your responsibility.
Finishing 07-280 means more than reading 24 guides: produce a search engine, supervised-model comparison, CNN/GPT-2 experiments, and a small RL-plus-MCTS system before choosing 07-380, 10-301, or a specialist course.
07-280 is CMU's new Spring 2026 AI+ML core: 24 lectures and 12 main assignments move from heuristic search and CSPs to AlexNet, GPT-2, and AlphaZero. Its public material supports self-study, but complete recordings, Canvas checkpoints, Gradescope, and staff feedback remain unavailable.
Lecture 1 uses an alien autoencoder, the scope of AI and ML, and AI history to establish the course's coordinate system: an intelligent system turns inputs into representations and decisions under uncertainty.
Lecture 2 decomposes search into a problem, frontier, and priority: UCS uses paid cost, Greedy uses estimated remaining cost, and A* combines them as `f=g+h`; tree and graph search require different optimality conditions.
Lecture 3 turns a single path into a contingent plan: minimax faces an optimal opponent, alpha-beta skips branches without changing the root value, and expectimax replaces worst-case choice with probability.
Lecture 4 exposes structure through variables, domains, and constraints, then upgrades DFS with backtracking, forward checking, AC-3, MRV, and LCV; the goal is to prove failure earlier.
Lecture 5 formulates machine learning through `X → Y`, loss, risk, and empirical risk minimization: a training set only gives average observed loss, while the real objective remains generalization over an unknown distribution.
Lecture 6 recursively grows a tree from decision stumps, measures label uncertainty with entropy, and selects splits by `I(Y;W)=H(Y)-H(Y|W)`; this is computationally practical greedy ERM, not a global optimal-tree guarantee.
Lecture 7 applies ERM to linear functions and squared loss, moves from a one-dimensional slope to `argmin ||y-Xθ||²`, and derives the normal equation when `XᵀX` is invertible.
Lecture 8 moves from a one-dimensional parabola to vector gradients and compares batch GD, SGD, and mini-batches; the learning rate determines whether updates converge, oscillate, or diverge.
Lecture 9 models P(y=1|x) with a sigmoid instead of directly predicting 0 or 1, learns parameters with cross-entropy and convex optimization, and extends naturally to softmax regression.
Lecture 10 uses φ(x) to let linear models express nonlinear functions, then controls the resulting overfitting with train/validation/test separation, L1/L2 regularization, and model selection.
Lecture 11 expands a logistic unit into a multilayer network: linear layers produce z, activations produce a, and multiple neurons jointly learn a feature transform trained through a final loss.
Lecture 12 treats a network as a computation graph: the forward pass stores intermediates, the backward pass propagates upstream gradients, and local linear, activation, and softmax rules compute every parameter gradient efficiently.
Lecture 13 separates alignment into specification, distribution shift, oversight, and corrigibility, then uses benchmark selection, leakage, and post-hoc selection experiments to show why a final paper cannot audit an autonomous research workflow.
Lecture 14 replaces dense image models with local connectivity and parameter sharing, moving from convolution, stride, padding, and pooling to AlexNet, GPU data parallelism, ResNet skip connections, and BatchNorm.
Lecture 15 splits a pretrained model into representation g and task head h: freeze g and train only the head, or fine-tune some or all parameters at a smaller learning rate depending on data volume and source-target distance.
Lecture 16 starts from likelihood p(D|θ), uses i.i.d. to factor the joint probability and logs to turn products into sums; Bernoulli MLE yields sample proportions, conditional Bernoulli yields logistic cross-entropy, and Gaussian noise yields squared error.
Lecture 17 first decides how text becomes tokens, then uses N-grams to turn sequence probability into conditional probabilities estimated from corpus counts. Tokenization is the first design decision about what a model can see.
Lecture 18 truncates the chain rule with an N-gram Markov assumption, estimates probabilities from corpus counts, and contrasts greedy, categorical, and temperature sampling. The real bottlenecks are zero probability for unseen contexts and a fixed window.
Lecture 19 builds a minimal next-token model from two embedding matrices, dot-product similarity, softmax, and cross-entropy. Shared vector parameters replace the isolated count cells of an N-gram table.
Lecture 21 formulates stochastic sequential decisions as an MDP with known dynamics, defines value and Q-values through Bellman backups, and solves for an optimal policy with value or policy iteration.
Lecture 22 keeps the MDP structure but removes known transitions and rewards. TD learning updates value from one sample, and Q-learning uses an off-policy target to learn optimal action values directly.
Lecture 23 replaces a huge Q-table with Qθ(s,a): first derive a gradient update for linear features from squared TD error, then add replay data and a fixed target network to form DQN.
Spring 2026 Lecture 24 is MCTS, not Fall 2026 LLM post-training. It allocates simulations through selection, expansion, rollout, backup, and UCB, then connects policy/value heads and self-play to AlphaZero.
Lectures 1–12 form one decision pipeline: define states, moves, and objectives, then use heuristics, losses, regularization, and backpropagation to control an otherwise intractable search space.
Stage II uses HW8 and HW11 to test whether representation, computation graphs, training, transfer, and generation actually connect, rather than treating CNNs and Transformers as diagrams to memorize.
Stage III connects value, policy, bootstrapping, function approximation, and MCTS into AlphaZero: a network supplies priors and estimates, search improves decisions, and self-play creates the next training set.
Spring 2026 Lecture 1 focuses on neurons, perceptrons, connectionism, and the problem framing of deep learning. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 22 focuses on latent variables, the ELBO, the KL term, and the reparameterization trick. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 3 focuses on data distributions, hypotheses, losses, empirical risk, and their roles in generalization. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 4 focuses on gradients, learning rates, parameter updates, and the training of a linear neuron. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 5 focuses on computational graphs, the chain rule, local derivatives, and gradient reuse. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 6 focuses on non-convex loss surfaces, curvature, saddle points, and momentum's accumulated direction. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 7 focuses on the tradeoffs among full-batch, mini-batch, stochastic gradients, and second-order information. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 8 focuses on AdaGrad, Adam, regularization, BatchNorm, Dropout, and loss selection. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 9 focuses on local connectivity, weight sharing, convolution kernels, and feature maps. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 10 focuses on stride, padding, receptive fields, and multi-channel convolution. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 11 focuses on stacked convolutional architectures, feature hierarchies, and design tradeoffs. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 12 focuses on CNN training, architecture selection, and the end-to-end assembly of a vision model. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 13 focuses on sequence state, temporal unrolling, parameter sharing, and recurrent computation. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 14 focuses on backpropagation through time, gradient stability, and LSTM-style gated memory. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 15 focuses on variable-length input/output, unknown alignment, and the CTC objective. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 16 focuses on blanks, collapse rules, prefix probabilities, and approximate decoding. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 17 focuses on autoregressive factorization, conditional language models, and translation decoding. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 18 focuses on queries, keys, values, scaled dot-product attention, and the Transformer block. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 19 focuses on encoder/decoder structures, masks, residual paths, and architecture variants. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 20 focuses on scaled autoregressive models, training stages, inference, and capability boundaries. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 21 focuses on bottleneck representations, reconstruction objectives, dimensionality reduction, and representation quality. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 22 focuses on latent variables, the ELBO, the KL term, and the reparameterization trick. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 23 focuses on forward noising, reverse denoising, score or noise prediction, and sampling. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 24 focuses on the generator, discriminator, minimax objective, and training instability. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 25 focuses on message passing, aggregation, node representations, and permutation symmetry. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 26 focuses on states, actions, rewards, returns, values, and policy learning. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 27 focuses on associative memory, energy functions, fixed points, and pattern retrieval. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
Spring 2026 Lecture 28 focuses on energy-based probability models, stochastic units, the partition function, and learning difficulty. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.
CMU 11-785 Spring 2026 publishes official slides and YouTube recordings for all 28 content lectures, plus extensive bootcamps and recitations. Its HW1–HW4 specifications, starters, and evaluation still depend on Autolab, Piazza, and Kaggle.
Cognee is a data-to-memory pipeline: a relational store preserves sources and provenance, a vector store finds semantically similar content, and a graph store represents entity relationships, exposed through remember, recall, improve, and forget.
CS124 Winter 2026 opens by mapping a ten-week path from tokenization and classification to retrieval, speech, networks, and LLMs, while PA0 establishes the Jupyter environment used throughout the quarter.
Week 10 models the Web with anchor text, PageRank, and centrality; post-training, multilinguality, and speech belong only to a public final-deck outline labeled 2025, not the 2026 live narration.
Week 2 builds three layers: a token vocabulary with BPE, sequence comparison with dynamic-programming edit distance, and probability approximation with n-grams; PA1 turns regex and BPE into executable work.
Week 4 builds candidates with an inverted index, ranks them with tf-idf and cosine similarity, and then connects retrieved evidence to generation; PA3 exposes RAG's inspectable retrieval half.
Week 5's public materials support the distributional hypothesis, word embeddings, and cosine similarity; the paired Social NLP lecture is unrecorded and restricted, so concrete audit methods are labeled as author extensions.
Week 6 uses public neural-network slides for weighted sums, nonlinearities, loss, and backpropagation, then a public LLM/Transformer deck labeled 2025 for decoder-only architecture without treating it as the 2026 live transcript.
Week 7's public path is PA6a: implement causal self-attention, train a small Shakespeare Transformer, sample text, and compute perplexity; the live speech lecture remains an explicit source gap.
Week 8 sends text through TTS and back through STT, requiring error classification, formatting-loss analysis, and accent stress tests, while Lab 4 prepares Git collaboration for the team agent project.
Week 9 builds movie recommendations with item-item collaborative filtering, then packages recommendation, web search, databases, and memory as agent tools under API-budget and team constraints.
Lecture 3 decomposes neural-network training into computation graphs, local derivatives, and the chain rule: the forward pass computes a result; backprop accumulates gradients from the output so every parameter knows how to move.
Lecture 11 divides evaluation into what to test, how to measure it, and when the result stops being trustworthy. Benchmarks saturate or leak, prompts change scores, and an LLM judge remains a biased model.
Lecture 9 compares prompting, pruning, LoRA, prompt tuning, and adapters. Each asks the same question: how many parameters must change, and how much task-specific state must be stored, to adapt a large pretrained model?
Lecture 6 completes the Transformer picture with encoders, decoders, and cross-attention, then breaks the final project into formats, assessment, research topics, and data. A viable topic needs one explicit baseline and metric.
Winter 2026 Lecture 1 divides NLP into four eras: early exploration, symbolic systems, statistical machine learning, and deep/self-supervised learning. The point is not the dates but how each era redefined the language problem.
Lecture 15 is Been Kim's interpretability guest session, but the Winter 2026 site publishes no slides or agenda. This article does not invent lecture content; it maps the five official readings across concept discovery, agentic investigation, and new vocabulary.
Lecture 17 is Luke Zettlemoyer's multimodality guest session, but the site publishes no slides or agenda. Its official readings establish three routes: visual reasoning workspaces, early-fusion token models, and text autoregression with image diffusion.
The final lecture frames Open Questions in NLP 2026 as smart scaling: prolonged RL, Prismatic synthetic data, RL as pretraining, and open collaboration seek reasoning gains beyond adding parameters.
Lecture 8 explains how instruction tuning, preference data, and RLHF turn a pretrained model into an assistant, then derives DPO from winner–loser pairs. Every step converts human judgment into signal—and imports its biases.
Lecture 7 decomposes pretraining into scalable data, subword tokenization, three model objectives, and in-context learning. A general self-supervised objective yields reusable representations; downstream signals specify their use.
Lecture 10 moves from question answering and RAG into language agents, then decomposes them into reasoning and planning, memory, tools, data, and evaluation. An agent is an inspectable loop between a model and external state.
Lecture 12 shows that output policy is not a detail: greedy, beam, and sampling produce different text. It then moves from R1-Zero/R1 into PPO, GRPO, and DAPO, asking when longer reasoning actually helps.
Lecture 13 moves from inference efficiency to inference capability: speculative decoding drafts with a small model and verifies with a large one; on-policy distillation addresses drift; long context and test-time scaling spend inference resources.
Lecture 4 defines a language model as a next-word probability distribution, then uses an RNN to compress an arbitrarily long prefix. It also exposes recurrence's central cost: information and gradients travel one time step at a time.
Lecture 16 divides NLP's social impact into four questions: why models hallucinate, why AI-assisted creativity may homogenize output, how work is reorganized, and why value alignment cannot be reduced to one reward.
Lecture 18 is a John Schulman guest session. The official page gives only the title Tinker and LoRA Without Regret, date, and speaker—no slides, agenda, or readings—so this article records confirmed facts and unknowns only.
Lecture 14 moves from word, character/byte, and subword segmentation to BPE failures and cross-lingual fairness. A tokenizer determines sequence length, compute cost, and the units a model sees; it is not neutral preprocessing.
Lecture 5 moves from the long-range and sequential bottlenecks of RNNs to self-attention and the Transformer. It shortens information paths and enables parallel computation, at the price of quadratic attention and separately encoded position.
Lecture 2 moves from word2vec's prediction task, objective, and gradients to count-based vectors and evaluation. Meaning becomes a high-dimensional position learned from context, not a label retrieved from a dictionary.
SPINACH does not guess complete SPARQL in one shot. It searches entities and properties, inspects Wikidata entries and examples, executes small queries, and composes a final query under explicit action and stopping rules.
CHURRO represents full-page text, layout, and metadata in HDML, unifies multilingual historical data for a page-level VLM, and connects extraction to HistoryGenie for searchable, conversational archives.
The final lecture is not a complete LLM-training tutorial. It studies data efficiency under fixed data and abundant compute, revisiting epochs, batches, ensembles, self-training, and conditions for synthetic continued pretraining.
The [WikiChat paper](https://aclanthology.org/2023.findings-emnlp.157/) expands RAG into query formulation, retrieval, filtering, generation, claim extraction, renewed retrieval and verification, and removal of unsupported content—and evaluates retrieval separately from factuality.
Fall 2025 opens with computational thinking: reliability comes from decomposing retrieval, formal representation, verification, and generation into testable algorithms, not from one heroic prompt.
STORM uses perspective-guided questions, simulated interviews, and outlines to broaden research; Co-STORM keeps a person in the loop so discovering unknown questions and co-editing become part of the system.
SLIDERS induces a question-specific schema, applies semantic chunking and contextualized extraction, reconciles duplicate rows, and answers with SUQL instead of feeding every long document directly to one model.
ReactGenie annotates React components to expose data, actions, and views, parses composite voice commands into a DSL, and renders native graphical output against shared UI context.
The lecture parses patient records and trial criteria into SMT, retrieves candidates through a weaker propositional projection, and runs a solver on the reduced set. Reasoning is inspectable, but NL-to-SMT remains the main error boundary.
Automated qualitative coding defines event types and arguments in a codebook, then separates document classification, structured extraction, and entity linking. Constrained JSON fixes form, not expert judgment.
SUQL adds answer and summary functions over text to SQL. A semantic parser emits one hybrid query, while an optimizing compiler applies predicate pushdown, top-k pruning, and lazy evaluation.
CS224V splits task-agent evaluation into state updates and complete interaction: isolate the semantic parser, then test task completion, grounded queries, and valid actions with real users.
Genie Worksheets declare task capability as a form-like specification. A contextual semantic parser updates formal dialogue state while the runtime controls queries, actions, and responses.
Lecture 3 does not turn its survey of modern LLMs into a single best recipe. It finds a conservative consensus—pre-norm, RMSNorm, no biases, SwiGLU, and RoPE—plus a small set of deviations justified by inference cost or stability.
Lecture 4 studies two kinds of sparsity: linear/recurrent attention reduces sequence-length cost, while MoE activates only part of a model for each token. Both turn saved FLOPs into routing, balancing, communication, and kernel problems.
Lecture 14 moves raw documents through language, quality, and safety filtering; exact and near deduplication; and source mixing. Each stage reshapes model behavior, while synthetic instruction and agent trajectories extend the pipeline into executable environments.
Lecture 13 traces training sources through Common Crawl, Wikipedia, GitHub, arXiv, books, and open datasets. Technically accessible is not the same as licensed, and raw data is not training data; provenance must precede cleaning and mixing.
Lecture 12 moves from perplexity to exams, chat preferences, agents, reasoning, and safety. Every benchmark changes the capability definition, scaffold, judge, and contamination risk, so evaluation must first say whether it compares a method, model, or complete system.
Lecture 5 explains GPUs through SMs, warps, and the memory hierarchy, then unifies common optimization under low precision, fusion, recomputation, coalescing, and tiling. FlashAttention combines those principles for attention.
Lecture 10 separates prefill from decode: prefill parallelizes and is often compute-bound, while decode is sequential and commonly bandwidth-bound. GQA/MLA, quantization, speculative decoding, continuous batching, and PagedAttention reshape that cost.
Lecture 6 turns GPU principles into kernels: benchmark scaling across shapes, profile actual calls and time, then implement GeLU, softmax, reductions, and tiled matrix multiplication in Triton. Speed begins with measuring correctly.
Lecture 17 organizes CLIP/SigLIP, LLaVA, Qwen-VL, and Chameleon into three paths: contrastive encoders learn semantics, vision-encoder/projector/LM stacks provide understanding, and discrete image tokens enable generation. Resolution, token budgets, and modality balance constrain them all.
CS336's first lecture does not treat building a language model from scratch as reenacting every old technique. It separates mechanics, mindset, and intuitions, then uses BPE to show how raw bytes become trainable tokens.
Lecture 7 starts below FSDP APIs, building a communication language from broadcast, all-reduce, all-gather, reduce-scatter, and all-to-all before assembling data, tensor, and pipeline parallelism.
Lecture 8 moves from parallel primitives to system design: ZeRO progressively shards optimizer state, gradients, and parameters; TP, PP, SP, and EP split width, depth, sequence, and experts. Their composition must follow topology and dynamic activation memory.
Lecture 2 reduces model training to tensors, FLOPs, bytes, and time: use einops to track dimensions, arithmetic intensity and roofline analysis to identify bottlenecks, then trade compute for memory with gradient accumulation and activation checkpointing.
Lecture 16 moves from PPO to GRPO and RLVR. Math, code, and environment outcomes provide scalable rewards and avoid some preference-model overoptimization, but group-normalized advantages introduce difficulty and length bias while rollout infrastructure becomes the dominant cost.
Lecture 9 begins with log-log linear relationships between data and error, then uses scaling laws to compare architectures, optimizers, batches, and model-data allocations. The Chinchilla dispute shows how fitting methods, observed ranges, and deployment objectives change the answer.
Lecture 11 reads public recipes from MiniCPM, DeepSeek, Qwen, and Llama 3: hold most architectural ratios fixed, sweep learning rate and batch at small scale, then choose model/data allocation with IsoFLOPs. μP helps, but normalization, optimizers, and weight decay can break transfer.
Lecture 15 divides post-training into imitation and optimization. SFT extracts pretrained capabilities from instruction-response data; RLHF uses pairwise feedback to bridge demonstrations and preferences. PPO and DPO both inherit data bias, reward overoptimization, and mode collapse.
Daytona treats a sandbox as a long-lived computer that can start, pause, snapshot, and fork. It raised a $24 million Series A in 2026, while a Laude Institute case study reports 37,000 sandboxes in one week. It fits parallel evaluations and coding agents, but its core open-source repository is no longer maintained.
Deepgram combines streaming STT, LLM orchestration, turn detection, barge-in, and streaming TTS over one WebSocket while preserving paths for standalone speech models and bring-your-own LLM or TTS.
Dify puts models, Knowledge, visual Workflows, Agents, Plugins, and application APIs in one workspace; this guide builds a minimal Workflow that can be tested, published, and called through the API, then explains when an Agent is actually warranted.
DSPy replaces handwritten prompt strings with task Signatures, execution Modules, and Optimizers that compile better instructions and examples against a dataset and metric.
E2B combines Templates, Firecracker microVMs, and process, file, and network APIs into an agent execution layer. Its real selection advantage is preserving memory and processes across pause and resume, not merely providing another code interpreter.
ElevenLabs has expanded from a TTS vendor into the ElevenAgents platform: Scribe Realtime listens, Flash speaks, and the platform connects the LLM, turn-taking, tools, and telephony. The key choice is whether you need a voice model or the whole agent control plane.
Fireworks AI puts open-weight model evaluation, dedicated GPU deployments, and LoRA customization behind one API surface. Serverless fits low-volume starts, On-demand fits sustained traffic and custom models, while reserved capacity adds enterprise capacity guarantees.
Flowise uses Assistant, Chatflow, and Agentflow to cover simple assistants, single-agent systems, and multi-agent orchestration; however, its repository was archived in August 2026 and official EOL is scheduled for August 31, so new projects should not adopt it without a maintained fork and migration plan.
Galileo connects dataset experiments, LLM/code/Luna evaluators, production traces, and runtime guardrails. It fits enterprises that need observability plus intervention, but the old Protect surface is deprecated and evaluator models do not replace human calibration or application security.
In August 2026, it's not just five frameworks moving. Beyond OMP 2, Pi v2, Opencode 2, dsh, and Claude Code, three model makers — Google (Antigravity CLI), Meta (Muse Code), and xAI (Grok Build) — are building coding agents directly. Add Amp, Cline 2.0, and the Codex CLI Rust rewrite, and eight-plus frameworks are undergoing architecture-level changes simultaneously. Factor in 110+ total CLI tools, and H2 2026 is a divergence period for harness methodology. This article analyzes four architectural approaches, one shared direction, and one emerging trust crisis.
Haystack turns indexing, retrieval, generation, and evaluation into replaceable Components connected by directed-multigraph Pipelines; it fits Python teams that want RAG flows to be tested, versioned, and deployed as code.
Helicone is an open-source LLM gateway and observability platform: requests sent through its compatible endpoint automatically capture model, latency, tokens, cost, and custom properties, while managed credits or BYOK enable routing and fallbacks.
Hugging Face Hub is a collaboration layer for versioned models, datasets, and applications. Datasets handles data, Spaces runs demos, while Inference Providers and Endpoints provide managed inference.
Hyperbrowser packages Chrome sessions, proxies, stealth, profiles, and recordings behind managed Playwright and Puppeteer APIs. It fits agents that need to scale real-browser work quickly, while profile credentials, anti-bot compliance, and proxy bandwidth costs remain application responsibilities.
Jina Reader turns a known URL into LLM-friendly Markdown; production use still requires explicit rendering, scope, token-budget, validation, and fallback decisions.
LanceDB stores vectors, metadata, and multimodal source data in the Lance columnar format. Its OSS edition embeds in Python, TypeScript, or Rust processes; distributed Enterprise becomes relevant when the data or service outgrows one machine.
LangChain v1 provides a high-level agent loop through create_agent, runs it on LangGraph, and treats tools, structured output, and middleware as its extension boundaries.
LangSmith structures LLM applications as projects, traces, runs, and threads, then uses datasets, evaluators, and experiments to turn production failures into offline regression tests. It observes any LLM application and does not require LangChain.
Letta extends MemGPT's operating-system analogy but is not a standalone memory API. The runtime persists agent state, editable in-context blocks, conversation history, and external archival memory, while the model can actively curate memory through tools.
LiteLLM is not a model provider. It is a Python SDK and self-hosted proxy that normalizes 100+ LLM APIs, then centralizes routing, fallbacks, virtual keys, budgets, and observability at the gateway layer.
LiveKit models a voice agent as a server participant in a realtime media room, with AgentSession orchestrating STT, turn detection, LLM, TTS, and interruption. It raised a $100 million Series C at a $1 billion valuation in 2026. It fits products needing WebRTC, multiple client platforms, telephony, and swappable models, but self-hosting the media server does not self-host the entire AI pipeline.
Mastra is a TypeScript agent framework that combines agents, typed workflows, memory, MCP, tracing, and scorers in one Node.js development environment.
Mem0 sits between an agent and storage: it extracts durable facts from interactions, scopes them by user, agent, or run, and searches them before a later generation. Its appeal is a small API; its risks are extraction errors, stale memories, and authorization boundaries.
Milvus separates real-time ingestion, historical queries, index building, and persistence into independently scalable components. It fits large, continuously updated retrieval services, but smaller projects often pay too much operational complexity for that architecture.
Lecture 2 of the 2026 course addresses data where order changes meaning—text, audio, and time series—and connects directly to music generation in Lab 1.
Lecture 3 of the 2026 course moves from image tensors, convolution, and pooling to recognition systems, preparing for MNIST and face detection in Lab 2.
Lecture 4 of the 2026 course separates generative from discriminative tasks, organizes VAE, GAN, and diffusion objectives, and leads into Lab 2’s DB-VAE.
Lecture 5 of the 2026 course connects agent, environment, state, action, reward, and policy into an interaction loop, introducing credit assignment and exploration.
Lecture 6 of the 2026 course places deep learning in emerging applications and real constraints, emphasizing data, outputs, evaluation, and failure conditions.
Lecture 7 of the 2026 course starts from Asimov’s literary laws and examines modern safety protocols through traces, test data, and continuous evaluation.
Lecture 8 of the 2026 course uses the scientific-discovery loop to show how simulators, AI emulators, and experiments cooperate instead of reducing science to generic prediction.
Lecture 9 of the 2026 course starts with GPU memory pressure and moves through checkpointing, offloading, ZeRO, FSDP, and multiple forms of parallelism.
In the 2026 lab, part 1 classifies MNIST with dense and convolutional networks; Part 2 learns a facial latent distribution with a DB-VAE and changes training sampling.
In the 2026 lab, students build chat templates and generation with LFM2-1.2B, adapt style through LoRA, and combine OpenRouter with Opik for a judge workflow.
n8n is automation-first: a webhook, schedule, or application event starts a workflow, then an AI Agent may choose tools inside it; production still requires deliberate memory, approvals, credentials, execution data, and scaling architecture.
OpenRouter exposes many models and inference endpoints through an OpenAI-compatible API, with provider ordering, failover, BYOK, and zero-data-retention controls in one routing policy.
Volcano Engine's open-source OpenViking stores agent memory, knowledge, and skills as a viking:// virtual filesystem — browsable with ls, tree, and find. Three-tier loading (L0/L1/L2) averages just 550 tokens per retrieval, boosting LoCoMo memory accuracy from 24–57% to 80–83%.
Parallel Web Systems separates Search, Extract, and Task APIs into web-access layers with different latency and cost profiles, while Basis maps citations, excerpts, and confidence to output fields.
Patronus AI treats evaluators as reusable scoring units, then applies them to offline experiments and production traces. It suits teams that want managed hallucination, safety, and multimodal evaluators, but judge scores cannot replace human labels, deterministic tests, or real security validation.
pgvector is a PostgreSQL extension, not a standalone vector database. It adds exact and approximate vector search to the same data model, transactions, and operations stack, while leaving index tuning and horizontal scaling as PostgreSQL concerns.
Portkey sits between applications and model providers: one OpenAI-compatible endpoint adds routing, fallbacks, request logs, budgets, and guardrails, with an open-source gateway available for self-hosting.
Authorization must take effect before candidate generation, while ACLs, deletion events, and source versions must propagate to every derived index; freshness needs measurable event-time SLOs too.
The first private-corpus decision is not which vector database to buy. Define data classes, trust zones, policy enforcement points, and freshness SLAs so every index, model, and observability system receives only the minimum data it is allowed to process.
The repository has a 20-query Traditional Chinese/English golden dataset, but no document-level qrels, retrieval runs, raw latency data, or executable benchmark script. Reporting Recall@k, MRR, or nDCG as measured results would therefore be dishonest; this article defines the contract needed to run them reproducibly.
Private-corpus sync is not periodic refetching. It requires stable canonical IDs, source versions plus checksums for change detection, idempotent upserts, and tombstones that propagate deletion through every index.
Promptfoo combines prompts, providers, test cases, and assertions in YAML to produce repeatable local and CI evaluation matrices, with red teaming against the same targets. It lowers the testing barrier but does not remove output variance, LLM-judge bias, or hosted data-flow concerns.
Pydantic AI models an agent as Agent[Deps, Output]: dependencies, tool inputs, and final outputs are typed, and model results must pass Pydantic validation.
R2R packages document ingestion, hybrid search, knowledge graphs, RAG, Agents, and access controls behind a REST API; it fits teams that already own their product frontend and backend and need a retrieval service.
LlamaIndex and Haystack are code-first frameworks; RAGFlow and Dify are managed application platforms; R2R packages retrieval as an API service. Choose how much control your team needs over ingestion, retrieval, and operations before choosing a tool.
RAGFlow puts document parsing, human chunk review, retrieval tests, chat, and citations in one platform; it fits layout-heavy PDFs and tables, but carries more deployment weight and platform state than a Python library.
Runloop combines isolated microVMs, reproducible images, disk branching, credential proxies, and evals in one coding-agent platform; an official case study reports more than 10,000 concurrent Devboxes in one workload.
Sail Research lets each inference request declare a completion window, scheduling patient background agents on cheaper capacity, while Sailboxes provide persistent long-running execution environments.
Reliable citation is not appending URLs to an answer. Separate URLs, content copies, and source independence, then connect atomic claims to quote spans and snapshots through a rerunnable claim-source matrix.
SerpAPI primarily manages search-results-page retrieval and parsing: select an engine, receive a structured SERP, then handle location, pagination, asynchronous polling, and validation in your application.
Serper is a third-party Google SERP API: one POST request returns structured JSON such as organic, knowledgeGraph, and peopleAlsoAsk, but production code still needs optional-field validation, URL checks, retries, and source verification.
Slack Code moves AI coding agents from individual terminals into shared Slack channels where teams can see diffs, previews, and plans in real time. But it solves management's visibility anxiety, not engineers' productivity bottleneck — the real battle is over who becomes the agent control plane.
Lecture 1 of Stanford CS221 Autumn 2025 develops operational representations and algorithmic intuition through Overview: Defining Intelligence Under Resource Constraints.
Lecture 2 of Stanford CS221 Autumn 2025 develops operational representations and algorithmic intuition through Learning I: From Computation Graphs to Linear Regression.
Lecture 3 of Stanford CS221 Autumn 2025 develops operational representations and algorithmic intuition through Learning II: Linear Classification, Features, and Cross-Entropy.
Lecture 4 of Stanford CS221 Autumn 2025 develops operational representations and algorithmic intuition through Learning III: Deep Networks as Composable Computation Graphs.
Lecture 5 models search with states, actions, successors, and costs, then uses acyclic dynamic programming to show that an efficient algorithm still solves the wrong problem when state omits information needed by the future.
Lecture 6 of Stanford CS221 Autumn 2025 follows the official material on Search II: Priorities in UCS and A* and makes its assumptions and limits explicit.
Lecture 7 of Stanford CS221 Autumn 2025 follows the official material on MDPs I: Putting Uncertainty into State Transitions and makes its assumptions and limits explicit.
Lecture 8 of Stanford CS221 Autumn 2025 follows the official material on MDPs II: Learning Q-Values Without a Transition Model and makes its assumptions and limits explicit.
Lecture 9 moves from tabular RL to function approximation, derives REINFORCE with the log-derivative identity, and connects the derivation to the executable PyTorch implementation.
Lecture 10 extends single-agent search into adversarial game trees: expectimax averages chance outcomes, minimax takes the opponent's worst case, and alpha-beta removes irrelevant branches without changing the answer.
Lecture 11 first learns game values from experience with temporal-difference updates, then moves from sequential play to simultaneous games described by mixed strategies, minimax guarantees, and Nash equilibria.
Lecture 12 builds a joint distribution from random variables and factors, then uses Bayesian-network factorization to express conditional independence and make conditioning and marginalization executable.
Lecture 13 replaces costly exact inference with Gibbs sampling: resample one variable at a time from a conditional determined by its Markov blanket, then approximate query probabilities with sample frequencies.
Lecture 14 moves from maximum-likelihood counts and Laplace smoothing with complete data to EM, which alternates posterior responsibilities for latent variables with parameter updates.
Lecture 15 separates propositional syntax from semantics: model checking defines entailment through satisfying assignments, SAT finds witnesses, and inference rules must be judged for both soundness and completeness.
Lecture 16 compresses knowledge across objects with predicates, quantifiers, and functions, then derives conclusions through substitution, unification, and definite-clause forward inference while exposing termination and completeness limits.
Lecture 17 defines a language model as a chain-rule factorization of sequence probability, compares n-gram and neural conditional models, and shows how sampling, temperature, and evaluation shape generation.
Lecture 18 classifies AI's social effects as benefits, misuse, accidents, and structural harms, then connects fairness audits, research ethics, copyright, and platform terms to accountable institutional choices.
Lecture 19 uses the Economics of AI deck to connect compute, data, distribution, and organizational complements to GDP, labor, and ideas-driven growth.
Lecture 20 is Percy Liang's fireside chat on career and research, CS221 and Stanford, and AI's future, with every attribution tied to the official video and editorial synthesis kept separate from auto-caption uncertainty.
A slide-grounded reconstruction of Fall 2025 Lecture 1, covering Course map and tools, A common language for graph data, Hand-designed features and representation learning while documenting the classroom material unavailable to self-learners.
A slide-grounded reconstruction of Fall 2025 Lecture 2, covering Encoder-decoder view, Similarity and the objective, Random walks while documenting the classroom material unavailable to self-learners.
A slide-grounded reconstruction of Fall 2025 Lecture 3, covering From fixed embeddings to deep encoders, The message-passing framework, Aggregation and update while documenting the classroom material unavailable to self-learners.
A slide-grounded reconstruction of Fall 2025 Lecture 4, covering The GNN design space, Message, aggregation, and update, GraphSAGE while documenting the classroom material unavailable to self-learners.
A slide-grounded reconstruction of Fall 2025 Lecture 5, covering Graph-data augmentation, Feature and structural augmentation, Supervision and loss while documenting the classroom material unavailable to self-learners.
A Fall 2025 slide-grounded reconstruction of Lecture 6, covering What distinguishability means, The Weisfeiler–Lehman test, An upper bound for message passing while documenting unavailable classroom material.
A Fall 2025 slide-grounded reconstruction of Lecture 7, covering The perfect-GNN thought experiment, Three levels of standard-GNN failure, Identity-aware encoding while documenting unavailable classroom material.
A Fall 2025 slide-grounded reconstruction of Lecture 8, covering Self-attention and message passing, The scope of graph attention, Positional and structural encodings while documenting unavailable classroom material.
A Fall 2025 slide-grounded reconstruction of Lecture 9, covering Heterogeneous graph schemas, Relation-specific messages, R-GCN while documenting unavailable classroom material.
A Fall 2025 slide-grounded reconstruction of Lecture 10, covering Knowledge graphs and completion, Triple scoring, TransE and relation patterns while documenting unavailable classroom material.
A Fall 2025 slide-grounded reconstruction of Lecture 11, covering Graph formulation of recommendation, The matrix-factorization baseline, Message passing in NGCF while documenting the public-material boundary.
A Fall 2025 slide-grounded reconstruction of Lecture 12, covering Limits of the tabular pipeline, Mapping relational databases to graphs, Temporal entity graphs while documenting the public-material boundary.
A Fall 2025 slide-grounded reconstruction of Lecture 13, covering The multi-relational bottleneck, RelGNN composite message passing, Relation-specific aggregation while documenting the public-material boundary.
A Fall 2025 slide-grounded reconstruction of Lecture 14, covering The goal of relational foundation models, Zero-shot relational transfer, PRODIGY's prompt graph while documenting the public-material boundary.
A Fall 2025 slide-grounded reconstruction of Lecture 15, covering Limits of transductive KG embeddings, Entity-inductive link prediction, The relation graph while documenting the public-material boundary.
A Fall 2025 slide-grounded reconstruction of Lecture 16, covering Complementary gaps in LLMs and GNNs, Text-attributed graphs, The LLM as predictor or encoder while documenting the public-material boundary.
A Fall 2025 slide-grounded reconstruction of Lecture 17, covering From graph QA to agents, Multimodal retrieval in STaRK, Tool use and traversal while documenting the public-material boundary.
A Fall 2025 slide-grounded reconstruction of Lecture 18, covering The graph-generation problem and representation, Evaluating generation quality, GraphRNN's autoregressive factorization while documenting the public-material boundary.
The Fall 2025 conclusion studies roughly 315K GNN designs across 32 tasks: run a small set of anchor models, derive task similarity from rankings, and transfer the best designs from similar tasks.
Linear regression is more than a best-fit line: Chapter 1 connects squared loss to gradient descent, normal equations, maximum likelihood, and locally weighted regression.
Chapter 2 derives logistic loss from a sigmoid probability model, then contrasts it with the perceptron and extends it through softmax and Newton's method.
Chapter 3 uses exponential families, natural parameters, and link functions to place least squares and logistic regression inside one modeling template.
Chapter 5 replaces high-dimensional feature inner products with kernels, letting inner-product-based linear algorithms learn nonlinear functions without constructing the features.
Chapter 7 decomposes neural networks into composable modules and uses backpropagation and vectorization to explain how deep models can be trained efficiently.
Chapter 8 decomposes test MSE into irreducible noise, squared bias, and variance, then uses uniform convergence and VC dimension to explain when training performance transfers to new data. Double descent shows why parameter count is not a universal measure of complexity.
Chapter 9 presents three controls on generalization: explicit complexity penalties, optimizer-induced implicit regularization, and model selection on data excluded from training. MAP estimation then connects a Gaussian prior to an L2 penalty.
Chapter 10 introduces unsupervised learning through k-means: alternating updates make distortion non-increasing and numerically convergent, but do not guarantee a global optimum.
Chapter 11 starts from soft assignments in Gaussian mixtures, uses Jensen's inequality to construct the ELBO, interprets EM as alternating maximization over a variational distribution and model parameters, and extends the idea to VAEs through approximate posteriors and reparameterization.
Chapter 12 formulates PCA as geometric optimization: maximize projected variance along a unit direction to obtain the leading eigenvector of the covariance matrix. The top k eigenvectors give both maximum retained variance and minimum linear reconstruction error.
Chapter 13 models ICA as x=As: observations are unknown linear mixtures, and the goal is to estimate W=A^{-1} to recover independent, non-Gaussian sources. A Jacobian determinant enters the transformed density and leads to the Bell–Sejnowski likelihood update.
Chapter 14 starts with a fixed Gaussian noising Markov chain and learns to reverse each transition. The ELBO turns reverse-kernel matching into weighted noise prediction, while the continuous-time view explains reverse drift through the score ∇log p_t.
Chapter 15 compares linear probing, full fine-tuning, and LoRA—not only by trainable parameter count, but by representation movement, data needs, and memory cost.
Chapter 16 connects representation learning to systems: contrastive objectives shape an embedding space, semantic retrieval finds neighbors in it, and RAG passes retrieved context to a generator.
Chapter 17 runs from next-token loss through Transformers, KV caches, MoE, and SFT, connecting an LLM's objective and architecture to its inference costs.
Chapter 18 separates two levers for LLM reasoning: chain of thought adds test-time computation, while verifiable rewards and policy gradients train long-reasoning behavior.
Chapter 19 uses Bellman equations to turn long-horizon decisions into one-step updates, moving from value iteration in known MDPs to model learning and continuous-state approximation.
Chapter 20 exploits linear dynamics and quadratic objectives to solve LQR, then uses DDP for local nonlinearity and Kalman filtering with LQG for partially observed state.
Chapter 21 derives REINFORCE with the log-derivative trick, then uses reward-to-go, baselines, and PPO clipping to control policy-gradient variance and update size.
Steel packages Chromium sessions, CDP, proxies, stealth, and debugging behind an Apache-2.0 browser API. Its public repository has about 7,400 stars and it entered the Stripe Projects developer preview in 2026. Self-hosting fits development and data-control needs; Cloud addresses concurrency, managed proxies, CAPTCHA, recordings, and SLAs.
Together AI puts serverless APIs for open-weight models, dedicated GPU endpoints, batch inference, and fine-tuning on one platform, letting teams validate per token before moving to reserved deployment when traffic or customization justifies it.
I turned Anthropic's session-cost advice into a global skill, and the first version made the very mistakes it was meant to prevent: a description stuffed with trigger keywords, hard thresholds based on file counts and minutes, and 'protect the main context' conflated with 'spend fewer tokens overall'. Three rounds later the entry point is one page, details live in references, numeric thresholds became four judgment dimensions, and every claim from a draft post was checked against official docs.
Vapi connects phone and web audio, STT, LLMs, TTS, tool calls, and call observability in a managed voice runtime. Providers are swappable, but Vapi's realtime orchestration is not portable. In May 2026, the company reported one million developers and announced a $50 million Series B.
Vercel Sandbox isolates untrusted code in Firecracker microVMs and integrates with Fluid compute, Active CPU pricing, and Vercel OIDC. It fits agents already running on Vercel, but network defaults, memory billing, and persistence still require deliberate design.
Vertex AI is more than the Gemini API: it puts access to 200+ models, training, evaluation, deployment, and governance under one Google Cloud control plane. Since April 2026, its products and roadmap have moved into Gemini Enterprise Agent Platform, while the Vertex AI API, documentation paths, and many resource names remain in active use.
Extraction tools cannot be compared by HTTP 200s. The same 20 URLs must be scored for body text, headings, tables, code, links, metadata, noise, latency, and cost. This article publishes the corpus, adapter contract, and gates, but no winner without a same-version raw run across all four paths.
Zep does not merely vectorize chat history. It turns episodes into entities and facts with validity time, allowing new information to invalidate an old relationship without erasing history. Graphiti is the open-source framework; Zep adds managed scale and governance.
Agent Plugins 1.0 is a packaging format that bundles Agent Skills (markdown instructions) and MCP server configs into a single directory, loadable by ChatGPT, Cursor, GitHub Copilot, Kiro, and VS Code. It's not a new protocol — it's the wrapper above protocols. Vercel initiated it, OpenAI/AWS/Microsoft/Cursor co-authored it, and Google joined on launch day. Anthropic isn't on the governance board, but MCP is a core primitive of the spec.
AgentQL replaces brittle CSS and XPath selectors with queries shaped like the data you want: `query_data` returns structured values, while `query_elements` returns interactive Playwright locators. The public Starter plan lists 50 free API calls per month, but its payment, hard-stop, and remote-browser reset rules still need to be verified in Billing.
Apify is not a single crawler. It packages scraping programs as Actors, saves reusable configurations as Tasks, triggers them with Schedules, and delivers results through Datasets. It fits teams that do not want to operate queues, schedulers, and workers, but Actor fees, compute, proxies, storage, and transfer all draw from the same platform budget.
Browser Use combines browser state, model decisions, and actions such as click, type, and extract into a repeatable loop. The open-source package favors custom tools and execution control; Cloud manages browsers, profiles, proxies, and concurrent work.
changedetection.io is a web-change signal layer: narrow the monitored content, suppress noise, and notify downstream systems only when a meaningful change occurs. It is neither a search API nor a crawler replacement.
This site covers MCP thoroughly but has never written about the layer underneath it: when your agent acts for ten thousand end users reading their own Gmail, whose database holds those refresh tokens, who rotates them, who revokes them. Composio is currently the most complete answer — MIT-licensed SDKs, a commercial hosted execution and OAuth layer. It claims 1,000+ toolkits; the managed-auth page actually lists 121 with a Composio OAuth app and 96 that require your own credentials. New pricing effective 2026-08-15: 100K free tool calls, $29/mo Pro. This post takes the authorization model down to an operational level and draws the line between wiring up MCP servers yourself and buying an integration platform.
Chroma's controlled study shows that even when it fits, a full context degrades performance. Coding agent vendors have landed on seven different responses: compact, hand off, prune, defer loading, isolate, train it into the model, or change the unit of work. Amp removed /compact outright, Atlassian argues summarization should be a last resort, and Cursor's A/B test measured a 46.9% token reduction. The three real disagreements come down to what each team is measuring.
Crawl4AI handles retrieval after a URL is known: use JsonCssExtractionStrategy for stable DOMs, and switch to LLMExtractionStrategy only when extraction needs semantic judgment or must tolerate irregular layouts.
CrewAI (GitHub 57.4k stars, MIT, PyPI 11.6M weekly downloads) defines agents by role, goal, and backstory, then groups them into crews for collaboration. Unlike LangGraph's graph-first and MAF's workflow-first approach, CrewAI is team-first — you don't draw nodes and edges, you describe who's on the team and what each person does. It fully removed its LangChain dependency in late 2024 and is now a standalone framework. The commercial side splits into the open-source package and AMP, a managed platform adding visual building, deployment, tracing, and compliance.
Exa turns every indexed web page into an embedding and retrieves by vector similarity instead of keyword matching. Official pricing as checked on 2026-08-21: $7 / 1k requests for /search (first 10 results included), $1 / 1k pages for /contents, $12–15 / 1k for the deep tiers, with $20 in free credits for new accounts. This blog's CLAUDE.md puts Exa first among cloud fetch tools, 16 of its 38 skills reference it directly, and only four existing posts mention it in passing — with zero dedicated posts. This is that post.
Firecrawl puts single-page scraping, site discovery, whole-site crawling, and JSON extraction behind one API. Cloud removes browser, proxy, and worker operations; self-hosting gives infrastructure control, but not the complete Cloud feature set.
Free access is not one model: recurring allowances, balance top-ups, rate-limited access, one-time credits, and self-hosting have different steady-state costs.
Linkup separates search depth from response shape: start most agent queries with standard + searchResults, move to deep only for multi-step browsing, and treat the monthly $20 as a balance refill rather than a new $20 grant.
LlamaIndex (51,775 GitHub stars, MIT, verified 2026-08-21) has moved its center of gravity from indexing to Workflows: the standalone llama-index-workflows package pulls 2.81M weekly PyPI downloads, more than the 1.97M of the llama-index umbrella package itself. This post covers the core abstractions, the trade-off against hand-rolling a pipeline, and a hands-on test of its defaults on Traditional Chinese text — at the same chunk_size=1024, English fits 4,645 characters and Traditional Chinese only 1,332. Plus one fact you need before choosing: the TypeScript port is archived and unmaintained.
Meilisearch turns application data into fast, typo-tolerant full-text search; the hard parts are index settings, asynchronous tasks, Chinese tokenization, access filters, and tested recovery.
Microsoft merged Semantic Kernel and its own AutoGen into Microsoft Agent Framework, which hit 1.0 GA on 2026-04-02 for .NET and Python (Go is still public preview). The absorbed autogen-agentchat has not shipped since 2025-09-30. But AG2, the fork on the original authors' side, never merged — it shipped 1.0.2 six days ago, and `pip install autogen` gets you AG2, not Microsoft. This post covers MAF's abstractions, the migration clock, and how to read the tangle of names.
MIT 6.S191's 2026 edition publishes nine lecture videos, slides, three software labs, and solutions, making it an A3 self-study course. The supplied path still depends on Google/Colab, Comet, and OpenRouter for Lab 3, while unaffiliated learners do not receive MIT credit, project feedback, or API credits.
Modal is a per-second-billed serverless GPU platform that also treats agent sandboxes as a first-class primitive (company-reported: over 1 billion sandboxes launched, more than a third of revenue). The selection question isn't how convenient it is — it's your GPU utilization. Verified 2026-08-21: Modal's A100 80GB works out to $2.50/hr against RunPod's $1.59/hr for the same card, so above 64% utilization renting your own is cheaper. But on the same day, H100 SXM is $3.95/hr on Modal against $3.99 on Lambda — on that card the premium is gone.
Qdrant is not just a place to store embeddings: define the vector schema, index frequently filtered payload fields, then add dense+sparse queries, tenant boundaries, snapshots, and monitoring to build an operable retrieval service.
Scrapling puts HTTP, Playwright browsers, CSS/XPath extraction, and a Spider API behind one Python interface. Adaptive selectors save element properties and relocate a target by similarity after a layout change, but the output still needs validation.
SearXNG is a metasearch engine, not a crawler, and it does not own a web-wide index. Based on the official 2026.8.20 documentation, this guide covers Compose installation, settings.yml, engine selection, the JSON API, and empty-result diagnosis.
Three components, one job each: SearXNG finds, Crawl4AI reads, and a thin layer of your own code glues them together. Three defaults will stop you cold — `formats` only emits HTML, `secret_key` ships as the literal string `ultrasecretkey`, and turning on the limiter blocks your own code. Also, the official install path changed in March 2026, so most tutorials online point at a repository that is now archived.
Tavily and Exa are cloud-only APIs and can't be self-hosted. What you can assemble instead is SearXNG (269 upstream engines, 82 on by default) plus Crawl4AI (78.8k stars, Apache-2.0), and the ready-made Tavily-compatible wrappers are all still double-digit-star solo projects you should not depend on. But SearXNG has no index of its own, and running it from a datacenter IP gets you empty results — those two facts decide whether self-hosting is worth it.
CS124 is the first course in Stanford's NLP branch. Its textbook is Jurafsky's own Speech and Language Processing, free online, and all nine assignment repos are public. But a banner sits on the course homepage: it will not be taught at all in AY 2026–27. And the chapter numbers the syllabus points at no longer match the August 2026 textbook.
CS221 lays AI out along one axis, and reflex models — deep learning — sit in the lowest slot, with states, variables and logic above them. When Percy Liang took over in Autumn 2025 he replaced the slides with runnable Python and wrote 'Cut constraint satisfaction problems :(' into the source of the first lecture — yet ExploreCourses and Stanford Online both still advertise constraint satisfaction as a course topic. The project has gone from 20% of the grade in 2019 to extra credit only.
CS224N has kept every course website since 2000 online. In Winter 2019, Transformers were lecture 14, taught by a guest. In Winter 2026 they are lecture 5, and every lecture after that assumes you already know them. The machine translation assignment is gone; assignment 3 now has you code a decoder-only Transformer from scratch, with pytest suites that run on your laptop.
CS224U's teaching material isn't a slide deck — it's an Apache-2.0 GitHub repo holding the lecture notebooks, all three assignments, and the grading document for the final project. But the on-campus course has skipped three straight academic years since Spring 2023, and ExploreCourses briefly put it back on the books for Spring 2026-27, then dropped that section again by 29 September 2026. The official description still lists relation extraction and semantic parsing; the 2023 syllabus covers neither. And the data-loading cell in the first assignment breaks in a fresh environment today, on a Hugging Face compatibility change.
CS224V only became Agentic AI in the 2026–2027 catalog, and the rename changed nothing underneath: the course still translates natural language into formal semantics and constrains agents with SMT solvers and knowledge graphs instead of wiring frameworks together. Seven of the eleven mandatory readings come out of the instructor's own lab. Every slide deck is public, and the course site says outright that they are deliberately incomplete.
All six CS224W Colabs download and run today, and the first one needs only NetworkX — no PyG install at all. But the exam is 35% of the grade, the largest single piece, and it's an in-person closed-book sitting. The public recordings stop at 2021 and cover none of the current syllabus's second half: graph transformers, relational deep learning, LLM+GNN.
CS228's official prerequisite is a single line — 'basic probability theory and algorithm design and analysis' — with no named course. But ExploreCourses shows it was last offered in Winter 2024, and the next slot, Winter 2027, still has a blank instructor field. What a self-learner can actually get is cs228-notes: 16 chapters, complete, last touched in June 2025.
The three things you need to self-study CS229 run on three different clocks. The lecture notes are 278 pages and were recompiled in August 2026. The newest problem sets you can download are from summer 2020. The self-assessment Stanford Online tells you to attempt before enrolling is a PDF created in 2008. Seventeen lectures from spring 2026 are public, and the last three are mislabeled.
CS25 is Stanford's 1-unit seminar where attendance is the only homework and anyone can audit. Of the nine talks in the Spring 2026 season, the three worth your time are Albert Gu on the inductive biases of SSMs vs Transformers, Charles Frye on serving inference across thousands of GPUs, and Victoria Lin on what native multimodality still hasn't solved.
CS329Z is a new three-unit agent engineering course debuting at Stanford in Autumn 2026. Its first homework bans every agent framework: one chat-completion call plus code you write yourself, grown on a real corporate email archive from a RAG pipeline into an agent harness with tools, a terminal, memory and a human in the loop. DSPy is still in the lectures, but no longer in the homework. The course site lives in a public GitHub repo, and its commit log records every syllabus revision: three assignments cut to two, peer review grown into a fifth of the grade, and the project topic changed from fixed to open.
Of the seventeen regular CS336 lectures, only nine are executable Python programs; the other eight are PDF slide decks — and the split falls exactly along the two instructors. Assignment 1's handout carries eight 'Low-Resource Tips' for finishing it on a laptop. Assignments 2 through 5 carry none. The course page lists the hourly price of a B200; the handouts list how many B200 hours each problem needs.
Tavily exposes Search, Extract, Map, and Crawl through one web API for agents. The free plan includes 1,000 credits per month; basic, fast, and ultra-fast Search cost 1 credit each, while advanced costs 2.
vLLM is the de facto standard for self-hosted LLM inference (89,470 GitHub stars, verified 2026-08-21), built on managing the KV cache the way an OS manages paged memory. But the selection question isn't how fast it is — it's your GPU utilization. Using Red Hat's measured 793 output tokens/second, a fully saturated A100 costs roughly $0.70 per million output tokens; at 10% utilization that becomes $7, more than most cloud APIs.
A web retrieval benchmark must evaluate complete tasks, not HTTP 200s: 30 fixed cases across five failure strata and three live channels, measuring answers, citations, freshness, latency, cost, and unnecessary escalation. This article delivers the harness and gates, but no fabricated ranking while the three live channels remain unconfigured.
An agent should not open a browser for every web task: route first to Search or Fetch, then escalate on explicit signals such as status codes, weak content, JavaScript shells, authentication, or challenge pages, with retry, budget, cache, deduplication, and provenance constraints at every step.
Behavioral interviews aren't about improvisation — they're about a pre-prepared story library. AI Engineer behavioral interviews have unique focus areas: AI ethics (bias, fairness, privacy), technical decision impact narratives (why you chose this model/architecture), and experience driving ML projects across teams. Strategy: build 8-10 STAR stories, practice each until you can deliver it in under 2 minutes.
AI Engineer coding interviews aren't identical to SWE — beyond LeetCode medium, you'll face ML-flavored problems (implementing a tokenizer, writing a batch inference pipeline, handling sparse matrices). Strategy: practice LeetCode medium to 70% pass rate, then spend remaining time on numpy/pandas operations, data processing pipelines, and ML-related programming problems.
Deep learning interviews don't ask you to derive backpropagation — they test whether you can explain the design intuition behind architectures. High-frequency topics: CNN's locality and translation invariance, why the evolution from RNN to Transformer was necessary, self-attention computation and complexity, BatchNorm vs LayerNorm use cases, and common training tricks (learning rate scheduling, gradient clipping, mixed precision).
LLM Application Design is the hottest new interview topic in 2025-2026. Key focus areas: RAG pipeline chunking/retrieval/reranking design, agent tool-use and planning loops, context window management strategies, guardrails and safety design, and LLM application evaluation methods. Interviewers especially value whether you've hit real-world pitfalls.
ML fundamentals interviews don't test formula memorization — they test whether you can explain concepts intuitively and hold up under follow-up questions. High-frequency topics: the practical meaning of bias-variance tradeoff, the selection logic for L1/L2 regularization, why cross-entropy beats MSE for classification, SGD vs. Adam tradeoffs, and how precision/recall priorities differ by scenario.
The core of ML System Design interviews isn't choosing the model — it's how to turn a business objective into a system that's deployable, monitorable, and iterable. Interviewers want to see if you can: translate business goals into ML objectives, design data pipelines and feature stores, choose reasonable serving strategies, and plan monitoring and A/B testing.
MLOps interviews test whether you have experience pushing models to production. Key topics: ML pipeline CI/CD (how it differs from software CI/CD), model registry and version management, A/B testing design and pitfalls, inference scaling strategies (horizontal scaling, model compression, caching), and production monitoring and alerting design.
The dividing line in LLM interviews is whether you've actually used these things. High-frequency topics: BPE tokenization logic and multilingual challenges, pretraining objectives (CLM vs MLM), three levels of fine-tuning (full/LoRA/prompt tuning), RLHF workflow and failure modes, prompting as engineering practice, and the difficulty of LLM evaluation with current methods.
AI Engineer interviews go beyond ML — big tech emphasizes system design and coding, startups look for end-to-end delivery, and AI-native companies test LLM engineering depth. Strategy: identify your target company types first, then allocate prep time across six dimensions (ML fundamentals, system design, LLM applications, coding, paper reading, and behavioral).
Paper reading interviews don't test whether you've read that specific paper — they test whether you can quickly understand a new method and identify its limitations. AI-native companies (Anthropic, OpenAI) particularly favor this format. Strategy: practice reading a paper in 30 minutes and verbally stating contribution + limitation, build your own must-read list, and practice summarizing each paper in three sentences.
CS329A is built around the generation–verification gap: models can produce the right answer but can't tell which one it is. The conclusion the course draws about itself matters more — today's methods make models more consistent, not smarter. Nine lectures are public, out of twenty.
AIF-C01, MLA-C02, and AIP-C01 are not a difficulty ladder but three job-function slices: AIF tests AI business judgment, MLA tests whether you can put traditional ML, foundation models, and agentic workflows into production, and AIP tests whether you can integrate foundation models into a GenAI system. The MLA-C02 English beta is open for registration; general-availability dates and non-English versions remain unannounced.
Before comparing exam objectives there is one fact that outranks all of them: registration for the Claude certifications is open only to organizations in the Claude Partner Network — individuals cannot sign up. For those who clear that gate, four things actually decide the answer. CCAO-F ($99) does not count toward partner tier eligibility while the other three do. Claude Code is 20% of CCAR-F but only 3.1% of CCDV-F — the architect exam tests the tool far more heavily than the developer exam. CCAR-P is 28% non-technical (governance 14% plus stakeholder communication 14%), which nothing else in this series is. And CCAO-F is the cheapest but only 14% Prompting; Output Evaluation at 21% is its real spine. All four are valid 12 months, with retakes at 14 / 30 / 90 days and 4 attempts per rolling 12 months.
Across Microsoft's four AI/agent certifications, only AI-103 → AI-500 is an official ladder; everything else is positioning. Three forks decide it: whether you write Python (AI-103/AI-500 vs AB-620), whether you build or judge (AB-100 vs the rest), and whether you can actually start today — AI-500's four official learning paths currently 404, AB-620 has no practice test, AB-100 has a free one. For readers who prefer Chinese there is a fourth fork: AI-103 and AB-620 offer Traditional Chinese, AI-500 and AB-100 are English only. All four cost $165, expire after one year, and renew free but only inside a six-month window.
NVIDIA's generative AI line has four exams: NCA-GENL and NCA-GENM ($125 each, associate), NCP-GENL and NCP-AAI ($200 each, professional). Three decision inputs no other vendor forces on you. One: both professional exams still show 'Coming soon' next to Register, so any near-term plan is down to the two associates. Two: NVIDIA is the only vendor in this series whose official prep courses are all paid — real cost is exam fee plus courses, and the self-paced totals are $390 (NCA-GENL), $210 for only three of five courses (NCA-GENM), and $1,620 list price across NCP-GENL's five. Three: the official documents disagree with themselves — NCP-AAI's weights total 98% on the web page and 92% in the PDF, and two cells of NCP-GENL's web table carry misplaced text, one of it about OpenUSD. Lock-in also varies sharply: NCP-AAI is 7% NVIDIA-specific, NCP-GENL is 31% GPU and model-compression work.
The 2026 consensus for agent-built slide decks: outline-first, separate content from construction, then render to images and let a fresh-eyes subagent do visual QA. Anthropic's and OpenAI's official slides skills both converged on PptxGenJS plus a visual verification loop, and the research line (PPTAgent → PreGenie → DeepPresenter) points the same way. But two later corrections matter: PresentBench shows the widely cited PPTEval scores too generously, and SeaSlides argues the model should not write free-form HTML/SVG at all.
Governance carries more weight on these exams than most engineers expect — CCAR-P is 14% governance plus 14% stakeholder work (28% non-technical), CCAO-F is 15%, AB-100's deploy-and-govern block is 40–45%, AIF-C01 is 14% responsible AI plus 14% security/compliance/governance. But across all fifteen official exam guides in this series, not one names the EU AI Act, the NIST AI RMF, or ISO/IEC 42001; the only regulations any of them names are CCAR-P's GDPR, HIPAA, and FedRAMP. So the use of this post isn't memorizing frameworks for an exam — it's using the three frameworks as a skeleton to file six certifications' scattered governance objectives. The three split cleanly: the EU AI Act is law (fully applicable 2026-08-02, high-risk duties pushed to 2027-12-02 by the AI Omnibus), the NIST AI RMF is voluntary (GOVERN/MAP/MEASURE/MANAGE, and 1.0 is currently being revised), and ISO/IEC 42001 is a certifiable management system standard whose clauses sit behind a paywall — so this post uses only what ISO's own public page states.
The AIF-C01 exam guide moved to v1.1 on April 30, 2026, adding seven objectives — MCP, multi-agent patterns, context engineering, token-based pricing, and hallucination detection all became testable, with Bedrock AgentCore, Kiro, and Strands Agents joining in-scope services. Almost every summary online describes the older version. This guide builds on the official five-domain weighting, integrating AWS's four-step prep method and field-tested advice from 10+ candidates who passed. $100, 90 minutes, 65 questions (50 scored), pass at 700, valid 3 years — the only exam in this series offered in Traditional Chinese.
AIP-C01 tests integrating foundation models into production AWS applications: RAG, agents, security, cost, and evaluation. This guide follows the five official domains with a prerequisite self-check, four-step preparation plan, resource choices, and hands-on checkpoints, plus a ten-week schedule, an AI study prompt, and exam-day guidance. Start with a diagnostic, then close gaps through one RAG-and-agent project.
A complete study guide for Claude's official architect certification (CCAR-F): five domains weighted 27/18/20/20/15, four scenarios drawn from six, common anti-patterns, and hands-on preparation. Official specs are 60 items / 120 minutes / $125 / 12-month validity / 720 to pass; registration is limited to Claude Partner Network members, and on-time renewal is free and non-proctored.
CCAR-P is the most expensive and most senior of Anthropic's four exams ($175, 63 items, 120 minutes). Integration is the heaviest domain at 19%, but what really separates it from everything else in this series is the other two: Governance, Safety & Risk Management at 14% and Stakeholder Communication & Lifecycle Management at 14% — 28% combined on compliance, risk, discovery interviews, and delivery lifecycle rather than code. The guide names GDPR, HIPAA, and FedRAMP, and its Intended Audience explicitly excludes entry-level developers and anyone doing 'prompt writing without broader system design responsibility.'
CCAO-F is the cheapest of Anthropic's four exams ($99, 60 items, 120 minutes), aimed at people who work with Claude rather than build against it. The heaviest of its seven domains is Output Evaluation and Validation at 21% — spotting hallucinations, deciding when human review is required, and adapting outputs — with Governance, Risk, and Responsible Use at another 15%. Anthropic states plainly that it is not for developers building against APIs or designing agentic systems. One easily missed limitation: this credential does not count toward Claude Partner Network tier eligibility, while the other three do.
CCDV-F is the engineer's exam among Anthropic's four certifications. The official blueprint has eight domains, and the heaviest — Applications and Integration at 33.1% — is led by Claude Application Design (8.6%) and Software Engineering Foundations (7.4%), meaning a third of the exam is API mechanics and ordinary software engineering. The counterintuitive part: Claude Code is only 3.1% and Eval only 2.6%, while the sibling Architect exam gives Claude Code 20%. Official specs: $125, 53 items, 120 minutes, pass at 720, valid 12 months, registration limited to Claude Partner Network organizations.
Google PMLE, AWS AIF-C01 and AIP-C01, Microsoft AI-103 and AI-500, and NVIDIA NCP-GENL all test how to make a GenAI application fast, cheap, and reliable — and they form a three-rung ladder: AIF-C01 asks whether you know cost scales with tokens, AIP-C01 and the two Microsoft exams ask whether you can instrument and control it, NCP-GENL asks whether you can change the model and the hardware. Three different altitudes. NVIDIA works at the kernel and quantization layer (Model Optimization 17% + GPU Acceleration 14% = 31%, the heaviest single cost/latency block in the whole series), AWS and Microsoft at the application layer (three caching tiers, token caps, chargeback), and Google at the MLOps layer (CPU/GPU/TPU evaluation, data vs model parallelism, scaling serving backends by throughput). The shared core is eight levers, but each lever becomes a different question at each altitude. This post deliberately carries no prices and no hardware specs — that is the part of this topic that rots fastest.
Google's Professional ML Engineer exam guide was rewritten in 2026: Vertex AI is renamed Gemini Enterprise Agent Platform throughout, so older study material no longer matches the product names in the questions. This guide uses the official six-section weighting as its skeleton, listing what each section tests, which official materials cover it, and what to build — plus a study schedule whose reasoning is spelled out. Official specs: $200, two hours, 50–60 multiple-choice and multiple-select questions, two-year validity, 3+ years of industry experience recommended including 1+ year on Google Cloud.
`batch_runner.py` runs thousands of prompts in parallel into ShareGPT-format tool-calling trajectories, lets each prompt name its own container image, and resumes by matching prompt content rather than index. Two quality filters run before you see the data: samples with zero reasoning are discarded, and entries calling hallucinated tool names are dropped at merge time. This is why a research lab builds a personal agent — the agent is the data pipeline.
One gateway process fronts 30-plus chat platforms and denies every user not on an allowlist or paired by DM. The scheduler adds two unusual guards: pre-dispatch validation marks a misconfigured job `blocked_config` without making a single LLM call, and the model drift guard makes unpinned jobs fail closed when the global model changes — protecting you from an hourly job quietly following you onto a paid model.
Hermes install paths come in three support tiers: macOS (Apple Silicon), Windows 10/11, Linux/WSL2, and Docker are Tier 1; Termux and Nix are Tier 2; pip, brew, AUR, and Intel Macs are explicitly unsupported — fixes for those won't be merged. On the upgrade side, `hermes update` snapshots state first, then compiles nine critical files after the pull and hard-resets the checkout if any fail to parse.
Hermes Agent is Nous Research's MIT-licensed agent framework, built around a learning loop: it writes its own skills, curates its memory, and searches past sessions with FTS5. It ships `hermes claw migrate` to move you off OpenClaw — but OpenClaw was not replaced, and both projects are still moving. This is the series opener: what it is, how it differs, and when not to pick it.
Hermes memory has hard caps: 2,200 characters for MEMORY.md and 1,375 for USER.md, and an over-limit write returns an error instead of auto-compacting, forcing the agent to make room itself. Skills are maintained by a curator that runs every 7 days after 2 hours of idle, marks skills stale at 30 days and archives at 90 — but never deletes. The switches actually worth flipping are `memory.write_approval` and `skills.write_approval`, which stage the background self-improvement writes for review.
`hermes claw migrate` imports persona, memory, skills from four locations, model and provider config, platform tokens, and the approval allowlist — but secrets are never imported silently, and even `--preset full` requires an explicit `--migrate-secrets`. What can't move (cron jobs, plugins, hooks, the multi-agent list, deep channel config) isn't discarded but parked in `~/.hermes/migration/openclaw/<timestamp>/archive/` for manual work. Coming from Claude Code or Codex is a different command: `hermes import-agent`.
Hermes supports 40+ providers, and the consumer-subscription OAuth paths are where billing surprises live: Anthropic OAuth only spends Claude Max extra-usage credits, and Claude Pro can't use it at all. Auxiliary tasks default to `provider: auto`, meaning your expensive main model does compression and vision grunt work. The fallback chain is a one-shot switch per session, not continuous retry.
Approvals default to smart mode: an auxiliary model waves through low-risk commands, auto-denies genuinely dangerous ones, and escalates the uncertain cases to you. Neither `--yolo` nor `approvals.mode: off` can disable the hardline blocklist (`rm -rf /`, fork bombs, `dd` to a physical disk), and `approvals.deny` is its user-editable counterpart, evaluated before yolo. Upstream is explicit that the threat model is an honest-but-wrong agent, not an adversarial process.
Hermes can run commands on seven backends: local, ssh, docker, singularity, modal, daytona, and vercel_sandbox. The decisive trade-off isn't performance, it's approval — local and ssh run dangerous-command checks, the other five skip them entirely because the container is treated as the boundary. Also, Docker defaults to one long-lived container shared across sessions, not a fresh environment per conversation.
The Tool Gateway routes four tool categories — web search (Firecrawl), image generation (nine FAL models), TTS (OpenAI), and cloud browser (Browser Use) — through Nous infrastructure, replacing four signups with one OAuth. It's per-tool rather than all-or-nothing, and `use_gateway: true` overrides any direct key in your `.env` — the precedence rule people most often get wrong.
Attach enough MCP servers and the tool schemas alone eat your context — upstream's extreme example is Cloudflare's ~3,300 tools, whose names alone run about 32K tokens. Hermes answers with Tool Search: MCP and non-core plugin tools collapse into three bridge tools and schemas load on demand, while core tools never defer. Separately, plugins are disabled by default and only run when named in `plugins.enabled`.
AB-100 is the architect tier of Microsoft's agent line, weighted 25-30 / 25-30 / 40-45 with deployment and governance heaviest. Its outline runs on verbs like design, recommend, and propose — it tests judgment, not configuration. Three things to know first: the scope blurb on the official exam page is wrong (it is information-protection and DLP boilerplate, which I verified verbatim), so prepare from the study guide instead; it has a free practice assessment, the only one of Microsoft's three agent credentials that does; and the 15 associate certifications it lists are described as usable, not required. Official specs: $165, English only, pass at 700, one-year validity.
AB-620 is the low-code branch of Microsoft's agent certification line — it tests agent flows, adaptive cards, computer use, MCP tools, A2A, and Fabric data agents in Copilot Studio, not Python. The three skill areas weigh 30-35 / 40-45 / 20-25, with integration the heaviest. Official specs: $165, 120 minutes, pass at 700, one-year validity, and 13 languages including Traditional Chinese — the only localized exam of Microsoft's three agent credentials. It is generally available, but the practice assessment is not out yet.
AI-103 replaces AI-102, retired June 30, 2026, and the objectives were rewritten around Microsoft Foundry — prompt flow, Azure AI Studio, Azure OpenAI Service, and Azure AI Agent Service appear nowhere in them. The five skill areas weigh 25-30 / 30-35 / 10-15 / 10-15 / 10-15, with generative AI and agents the largest. Official specs: $165, 120 minutes, pass at 700, offered in Traditional Chinese, may include interactive components — and it is valid for only one year, though renewal is a free, open-book, unproctored online assessment.
AI-500 is a rare thing among the major clouds — an expert-level certification dedicated to multi-agent systems, weighted 15-20 / 30-35 / 20-25 / 20-25, naming Agent Framework, LangGraph, Hugging Face Transformers, MCP servers on Azure Functions / Logic Apps / API Management, A2A, Key Vault, and the AI Red Teaming Agent. Three constraints come first, though: it is still in beta (scores wait for rescoring), it requires AI-103 before you can take it, and the official training is not live — the four learning paths listed on the exam page all return 404 today, and the instructor-led course opens 2026-09-30.
Microsoft AI-500, AB-620, AB-100, NVIDIA NCP-AAI, and Claude CCAR-F all test multi-agent architecture, and they overlap on seven things: orchestration topologies, A2A and MCP, per-agent identity boundaries, three-layer memory, observability and agent replay, human-in-the-loop, and guardrails at four intervention points. But four vendors use four vocabularies for the same ideas, and each exam has objectives that don't transfer — Microsoft names four context-window failure modes nobody else names, 7% of NVIDIA's is locked to NeMo and NIM, and Claude tests SDK-level details like stop_reason. One correction along the way: Google PMLE's wall-to-wall 'Agent Platform' is a Vertex AI rename, not a multi-agent domain.
Pure image understanding has flattened out — four frontier models all clear 80% on MMMU-Pro within 3 points of each other. The real differentiation is video, long-document OCR, and realtime speech, each with a different leader. But the most useful lesson from assembling these rankings is that two credible sources named different Video-MME leaders more than 10 points apart — and that July and August each turned the field over again.
NCA-GENL is usually what a job posting means by 'NVIDIA Generative AI / LLM certification.' But the official blueprint diverges sharply from the name — Core Machine Learning and AI Knowledge 30%, Software Development 24%, Experimentation 22%, Data Analysis 14%, Trustworthy AI 10% — with LLM and RAG content scattered at bullet level rather than forming a domain, alongside spaCy, NumPy, Keras, and cross validation. The other thing to know first: NVIDIA's official preparation courses all cost money ($30–$500), making it the only vendor in this series without a free official learning path. Official specs: $125, 1 hour, 50–60 items, two-year validity, English only, pass/fail with no score reported.
NCA-GENM matches NCA-GENL on price, length, and level but not on emphasis: Experimentation rises to 25% (the heaviest), Core ML drops from 30% to 20%, and two new areas appear — Multimodal Data 15% and Performance Optimization 10%. The content covers U-Net, CLIP, diffusion models, multimodal loss functions, attention maps, and NVIDIA's Riva / NeMo / Triton / ACE SDKs. Watch the cost structure: two of the five recommended courses exist only as $500 workshops with no self-paced option, so a self-study path cannot cover the official set. Official specs: $125, 1 hour, 50–60 items, two-year validity, English only.
NCP-AAI is NVIDIA's professional-level agentic AI credential — $200, 120 minutes, 60–70 items, two-year validity. Two things come first: registration is not open (the Register button carries a 'Coming soon' label), and NVIDIA's own web page and PDF study guide disagree on the weights — Deployment and Scaling is 13% on the page and 5% in the PDF, Run/Monitor/Maintain is 5% on the page and 7% in the PDF, and the two versions total 98% and 92% respectively. Both are nvidia.com. This guide treats that as a range and an uncertainty rather than picking one.
NCP-GENL is NVIDIA's professional-level LLM credential — $200, 120 minutes, 60–70 items. What separates it from every other GenAI exam is where the weight sits: Model Optimization 17% plus GPU Acceleration 14% is 31% on quantization, distillation, pruning, distributed parallelism, and CUDA profiling — not on calling APIs. Two things first: the Register button says Coming soon, so you cannot sit it yet; and two description cells in the official weight table are corrupted — Fine-Tuning is described with OpenUSD data-interchange text and Model Optimization with deployment text. I verified both verbatim; the correct descriptions are in the official PDF.
Most people assume GenAI certifications are built around prompt writing. CCAO-F gives Prompting 14% while Output Evaluation gets 21%; CCDV-F gives Prompt and Context Engineering 11.0%. What actually gets tested is structured output, injection-resistant prompting, dynamic context injection, context compression and caching, prompt lifecycle governance, and proving a prompt change helped — closer to context engineering and software engineering than to writing craft. None of the ten asks you to write a prompt on the spot; they are all multiple choice, so explaining why beats having a feel for it. The single most useful line comes from CCAR-F: when business logic must be guaranteed, 'change the prompt first' is usually the wrong answer.
Four certifications genuinely test RAG and retrieval evaluation: AWS AIF-C01 (chapters 2 and 3 total 52%, covering RAG, vector stores, and FM evaluation metrics), AWS AIP-C01 (11 of the 27 skill points in its 31% Domain 1 sit in vector storage and RAG), NVIDIA NCP-AAI (Knowledge Integration 10% plus Evaluation and Tuning 13%), and Microsoft AI-500 ('multi-agent RAG architecture' inside its 30–35% Develop area). Google PMLE contributes exactly one LLM-as-a-judge objective, and Claude CCDV-F — the developer certification people most readily assume covers RAG — has no retrieval objective across its eight domains, with Eval at just 2.6%. Includes a same-vendor foundational-vs-professional comparison, a four-vendor terminology map, non-transferable objectives, and a practice project.
OpenClaw has 386k stars to Hermes Agent's 232k, yet Hermes passed it on OpenRouter daily tokens back on 2026-05-10 (224B vs 186B). The nine self-hosted agents that appeared this year aren't nine competitors — they're nine incompatible answers to one question. CVE-2026-44112 broke OpenClaw's own sandbox, and in the Meta alignment director's inbox incident there was no attacker at all: context compaction ate the safety instruction.
The Workers AI catalog currently holds 84 models. For general chat pick glm-4.7-flash ($0.06 / $0.40 per M, 131K context), for vision pick gemma-4-26b-a4b-it ($0.10 / $0.30, 256K), for cheap high-volume steps pick granite-4.0-h-micro ($0.017 / $0.112), and for embeddings pick qwen3-embedding-0.6b or bge-m3 (both $0.012 per M). This post is updated on a schedule.
A VS Code extension open-sourced by Microsoft employees that reads your local Claude Code / Codex / OpenCode session logs. The real payload is 45 Markdown rules: prompts under 30 characters, sending the next message within 15 seconds of receiving 20 lines of AI code, instruction files over 4,000 bytes — turning 'context engineering' into numbers you can argue with.
The course lists four techniques for directing agents: instruction files, hooks, commands, subagents. The instruction file is the only one loaded in full every startup, making it config rather than memory; hooks cover what instructions can't, because a rule can be ignored and a hook cannot; commands are the only one a human triggers. The course also marks just one and a half of seven task steps as human work.
Week 1 of CS146S is 'build Claude Code in 200 lines' plus a dissection of production system prompts. The agent loop really is that small. The course slides close with four things Claude does underneath, one of them being `<system-reminder>` tags scattered everywhere to stop the model drifting — which appears in no official documentation.
Factory breaks 'can an agent work in this repo' into eight pillars and five levels, and published real scores: CockroachDB L4 (74%), FastAPI L3 (53%), Express L2 (28%). The thesis is that agent readiness approximates the density of deterministic validation loops — linters, type checkers, tests are reward signals for agents.
The course measured AI SAST false positive rates at 50–100%, against 50%+ for traditional SAST — the genuinely new problem is nondeterminism: run the same prompt twice, get different results, and you can never answer "am I done scanning?" The course lists five agent attack vectors, one of which, intent breaking, attacks the agent's plan itself.
The Agent Skills spec fits in a sentence: a directory containing a SKILL.md. The real design is three levels of progressive disclosure — only name and description load at startup, the body loads on a match, bundled files load on demand. This site's own repo carries 35 skills and 7,893 lines of SKILL.md, and startup still costs only those 35 metadata pairs.
Google deployed AutoCommenter to tens of thousands of engineers and published the whole tuning process: suppressing 17 'technically correct but low-value' rules raised the useful ratio from 54% to 66%, with 80% set as the bar for the next rollout stage. Final comment-resolution rate landed around 40%. The bottleneck in AI code review was never detection — it's volume.
How an individual connects tools is a preference; how an organization does it is governance — who can touch what data, where keys live, whose budget it lands on. Anthropic's published record of ten internal teams contains a good indicator: security engineering accounts for 50% of all custom slash commands in the entire monorepo. Adoption doesn't spread evenly; it takes off first in teams that already build their own tools.
Background agents replace 'you watch it run' with 'it finishes and opens a PR.' Every vendor's design converges on the same parts: an isolated environment, external triggers (issues, Slack, Linear), and a PR as the output. The genuinely new problem is that you become the bottleneck — five agents finish at once, five diffs queue for you, and none of them know the others exist.
Fall 2026 compresses a full week of prompting into one bullet here and adds RePPIT (Research, Propose, Plan, Implement, Test) and MCP. Two RePPIT rules are worth stealing outright: always ask for exactly two proposals, and never let the instance that wrote the code review it. On the MCP side, Anthropic measured turning tools into code calls dropping 150,000 tokens to 2,000.
Stanford CS146S's Fall 2026 syllabus compresses prompting from a full week into a single bullet, drops the terminal and UI-generation weeks, and adds Agent Skills, Agent-Ready Codebases, Background Agents, and AI-Native Team. Grading moved too: the final project fell from 80% to 50%, with 30% now on open source contributions. This series reads all ten weeks.
The final session is 'self-running, self-improving software systems.' The parts all appeared in the previous nine weeks: deterministic validation loops, skills that can be written back, background agents, centralized governance. One easily missed proportion from the slides — coding is 30% of engineering time, and running it in production is the other 70%.
Researchers initially assumed neural networks are easy to fool because they're nonlinear. That was wrong — Goodfellow's 2014 paper argues the primary cause is their linear nature, and high dimensionality lets every tiny perturbation compound. The second half covers generative models: GANs' three pathologies, and why diffusion sidesteps two of them by adding noise and learning to remove it.
A BCG experiment found a jagged frontier: inside it, AI substantially improved consultants' work; outside it, AI made results worse — and people fell asleep at the wheel. The lecture also takes a strong position: avoid fine-tuning wherever possible, because by the time you're done tuning, the next model already beats your fine-tuned version.
Andrew Ng demonstrates error analysis on a deep researcher: columns are the pipeline stages, rows are 10 to 100 queries, you only look at the ones that went badly, and you mark each cell where something broke. The percentages don't have to sum to 100%. He says it takes three or four hours and saves weeks of going the wrong direction — and the fraction of people who actually do it is far below 100%.
The third reason Go can't be learned with supervision is the interesting one: the ground truth itself is ill-defined — the strongest human doesn't play their best moves every day, and even their best move isn't optimal. The last 20 minutes map RLHF fully back onto RL: the agent is the model being fine-tuned, the action is the next token, an episode is one full generation, and the reward is extremely sparse.
Andrew Ng walks a face-recognition door system through the entire project lifecycle, and the whole lecture has one thesis: speed. He gives teams a two-day deadline, on the reasoning that 'time spent preparing data should be commensurate with the time it takes to train the model once.' It closes on a line: my job is to build something that actually works, and that is not the same as building something that works on the test set.
CS230's second lecture derives embeddings through three case studies: day/night classification teaches you to use humans as a proxy for choosing resolution, trigger-word detection teaches you to manufacture a million training examples in three hours, and face verification walks you through designing your first loss function. The final step — from supervised triplets to self-supervised pairs — is why modern models can consume billions of unlabeled images.
Ask a model what a goose looks like to it and it draws a whole flock — because the labeled data tagged a flock as 'goose,' so it thinks the flock is the label. This lecture gives seven ways to open a CNN up, then says honestly: applied to transformers, even the frontier of this research only explains two layers.
CS230's first lecture is a course overview, but Andrew Ng spends most of it on three things: why scaling works, when prompting stops being enough, and why he thinks 'don't learn to code' is one of the worst pieces of career advice ever given.
I tested 10 open-source PDF parsing tools on four scanned NTU graduate entrance exams. VLM-based tools—Firecrawl, MinerU 3.4, and Marker v2—overwhelmingly beat conventional OCR on formulas and code, but installation was the real barrier: MinerU's old package name creates dependency hell, Marker's first model download takes 10 minutes, and PaddleOCR needs a separate engine. In practice, use RapidOCR for screening and MinerU or Firecrawl for close inspection.
The 2026-08-13 release added 363,246 lines and published the For You ranking weights for the first time: favorite 0.5, reply 5.0, report −234.0. But the weights are constants — the P(action) that actually decides order comes from a 2560-dim, 8-layer transformer.
Chroma tested 18 frontier models and all of them degrade as input grows — as a cliff, not a slope. Memory failures are usually retrieval failures in disguise. And the real cost of KV cache is bandwidth, not storage: every generated token reads the whole cache.
In November 2025 three frontier labs jointly broke all 12 previously proposed prompt-injection defenses. EchoLeak's payload passed Microsoft's own dedicated classifier. So the goal is not blocking every attack — it is surviving the ones that land, and that is harness work.
The line between workflow and agent is who decides the steps — the developer at design time, or the model at run time. By that definition most LLM systems in production today are workflows. Plus a usable test for choosing between RAG and an agent.
Salesforce's number from 20,000 deployments: 90% of the work on an agent happens after launch, the reverse of traditional software. Stripe merges 1,300 PRs a week with no human-written code, and credits the environment rather than the model.
MCP governs agent-to-tool, A2A governs agent-to-agent, Skills govern reusable knowledge. The test is whether the data changes: if it changes between calls you need MCP; if it's stable enough to write down, a skill file is simpler and has no runtime that can fail on its own.
Microsoft, OpenAI, Salesforce, Stripe and three others independently say the same thing: reliability comes from the engineering around the model. And 'give the deterministic parts back to code' has been shipped as a product four separate times — Agent Script, Procedures, runtime, blueprints.
Standard RAG gives a wrong answer when it retrieves the wrong chunk, and nothing in the system will notice. Agentic RAG adds a self-check, at the cost of the evaluator paradox: the ceiling on self-correction is whatever the evaluating LLM can judge about relevance.
Every AI certification an engineer can register for in 2026, listed one by one: AWS AIP-C01 / MLA-C01 / AIF-C01, Google PMLE, Microsoft AI-103 and AI-500, all twelve NVIDIA exams, Databricks, Snowflake, Oracle's Agentic AI track, IBM watsonx, Salesforce Agentforce, GitHub GH-300, Anthropic's four Claude exams, Taiwan's iPAS AI Application Planner, plus the governance and audit line (IAPP AIGP, ISACA AAISM / AAIA, CertNexus CAIP) — prices, validity, and registration gates all checked against vendor pages. Two things that hit your wallet: Google's PMLE exam guide has renamed Vertex AI to Gemini Enterprise Agent Platform throughout, making pre-mid-2026 study material worthless, and the iPAS intermediate certificate is valid 5 years, not permanently.
Firecrawl's open-source Rust conversion library turns 14 office formats (including legacy .doc / .ppt / .xls) into GFM at a 4.7ms median — 109× faster than Docling under the same timing basis. The trade-off: it does no OCR at all.
Scans and complex layouts leave you no choice but to infer structure with a model. But the technical gap between MinerU, Marker, and Docling is far smaller than the licensing gap — MinerU needs a separate license past $20M monthly revenue, Marker's model weights need payment past a funding threshold, and only Docling is cleanly MIT. Read the LICENSE before the benchmark.
The most common mistake in feeding documents to an LLM isn't picking the wrong tool — it's picking the wrong layer. Structure already in the file goes to the conversion layer (milliseconds); text without structure goes to extraction; only inferred structure needs parsing. anydoc's 4.7ms against Docling's 513.6ms is a 109× gap, and most people jump straight to the most expensive layer.
Digital-native PDFs already contain readable text — what's missing is structure, and heuristics can recover it. PyMuPDF, pdfplumber, pypdf, and Tika do this with zero GPU and zero inference cost. The biggest selection trap isn't accuracy; it's PyMuPDF's AGPL-3.0 license.
"Digital employee" isn't a technology — it's a pricing and accountability unit. Anthropic's Project Vend had Claude actually run three shops, and found the most effective intervention wasn't a smarter model but forcing it to follow procedures. Their words: "we rediscovered that bureaucracy matters." Gartner estimates only ~130 of the thousands of vendors claiming to be agentic actually are.
Every serious image-to-video model in 2026 runs latent diffusion on a DiT backbone, so visual quality is no longer a useful axis for choosing one. The real axes are native audio, self-hostability, and dollars per second. Three widely-repeated errors worth correcting: Sora's app shut down on April 26 and its API goes on September 24; Wan 2.7 is described everywhere as Apache 2.0 open weights but no first-party source has them; Veo 3.1 officially costs $0.40/s, not the $0.75/s that circulates on review sites.
There are four paths to a 3D model in 2026: AI generation (Meshy-6 / Tripo / Rodin Gen-2.5 / Hunyuan 3D), phone scanning, Text-to-CAD, and manual modeling. Picking wrong has concrete costs — AI-generated meshes can't be dimensionally edited, Rodin's STL exports usually need repair, and Meshy's free-tier assets are public. This guide selects by what the model is actually for, with current pricing from each vendor's own page.
From MarkItDown (175k stars, MIT) to curl_cffi (6k stars), a survey of 34 open-source tools for feeding data to AI. Categorized along five axes: whole-site crawling, AI browser agents, document conversion, smart extraction, and anti-detection infrastructure. The key to selection isn't which tool is best — it's scenario matching.
Uncle Bob's 4.18M-view post of 2026/7/23 isn't a manifesto — it's a reply to an engineer who started in 1983 asking whether needing to understand code psychologically makes him old-fashioned. And he doesn't skip the code entirely: his 6/1 four-stage pipeline post says 'I spot check the code,' with thresholds of crap ≤ 6 (convention is 30) and mutation runs that kill all survivors. Plus a breakdown of his open-sourced Acceptance-Pipeline-Specification and the three metric blind spots Grady Booch names.
The dominant paradigm in 3D generation in 2026 is video diffusion feeding feed-forward 3D reconstruction, and Lyra 2.0 is the flagship of that line. But three Best Papers at CVPR 2026 point at what comes next: SAM 3D brings foundation-model-scale object reconstruction, D4RT rebuilds dynamic 4D scenes in seconds from a unified transformer, and O-Voxel replaces Gaussians with structured latents. 3DGS still rules, but surface primitives are challenging it, and pixel-space diffusion is pushing back against latent space.
Every official course platform from OpenAI, Anthropic, and Google, plus Stanford CS146S/CS336, Elements of AI, Hugging Face, MIT 6.S191 and more — scraped page by page, then re-sorted into four tiers: AI-curious, vibe coding, shipping to production, and how models actually work. Also covers self-study repos still being updated in 2026 and browser-based platforms that need no local setup, filtered by last-commit date rather than star count. The conclusion: nearly all of it is free. What is scarce is not courses, it is the judgment to pick one. And tier four will not fix your tier three problem.
HeyGen's open-source HyperFrames defines video timelines with HTML data attributes, uses headless Chrome for frame-accurate seek-and-capture, then encodes via FFmpeg to MP4. 33k stars in 3 months, Apache 2.0, 21 agent skills — AI agents write HTML to produce video, no React needed.
Loop Engineering is the practice of designing systems that automatically prompt AI agents, rather than prompting them manually. Boris Cherny runs hundreds of agents, Addy Osmani coined the term, and Blake Crosley identified verification cost as the real bottleneck — this article covers primary sources, the five building blocks, applicability boundaries, and criticisms.
From the CLI tool kin3o to the CVPR 2026 paper OmniLottie — a survey of open-source approaches for converting text and images into Lottie animations, with performance benchmarks and selection guidance.
MUSE-Autoskill (2026) introduces a five-stage skill lifecycle framework. Self-created skills achieve 60.35% (+7.16%) on SkillsBench overall, and an impressive 87.94% on tasks where skill generation succeeds — surpassing the human-authored skill ceiling. This post synthesizes six arXiv papers to map the full landscape of skill evolution research.
Even with temperature=0, LLM outputs can still fluctuate by up to 15% in practice. To rigorously compare agent changes, you need a frozen golden set, at least 3 runs per query averaged out, LLM-as-judge blind evaluation (pairwise preference flip rate reaches 35%), and paired statistical tests -- not just running each version once and going by feel.
The industry has converged on using OpenTelemetry GenAI semantic conventions to turn every LLM call and tool call into a span. Detecting the three major failure modes then splits into three tracks: faithfulness + semantic entropy for hallucinations, framework-level symbolic guardrails for tool misuse, and max steps + action hash deduplication for infinite loops — all wired into a Final / Trajectory / Single-step three-layer evaluation framework.
Agent decision-making under resource constraints is bounded rationality reborn: Rational Metareasoning uses VOC rewards to save 20-37% of tokens, BATS proves that adding budget without budget awareness is futile, FrugalGPT cascades cut costs by up to 98%, and Speculative Actions reduce latency by 20%. The three constraints ultimately converge into a single Pareto curve, and the overarching trend is moving from humans tuning knobs to models making resource-rational decisions on their own.
Three seemingly distinct agent security problems — tool output injection, trust boundaries, malicious agents — share the same root cause: LLMs flatten instructions and data into a single token stream, making them architecturally unable to distinguish between the two. Understand this through-line and you can trace every attack from EchoLeak (CVE-2025-32711, zero-click) to the Morris II AI worm, and see why 'making the model behave' doesn't work — only architectural constraints (six design patterns, CaMeL) do.
Traditional RAG is a fixed pipeline of 'retrieve then answer.' Agentic RAG splits retrieval into three decision layers: when to retrieve (FLARE uses token probabilities; Adaptive-RAG uses a complexity classifier), what to retrieve (HyDE / RAG-Fusion / decomposition / Step-back), and how to fuse (RRF k=60 then cross-encoder rerank then compression -- Anthropic measured a -67% failure rate reduction). Key counter-intuitive insight: unnecessary retrieval hurts quality -- 'deciding not to retrieve' is a first-class capability.
Automatic prompt optimization (APO) has evolved from APE/OPRO to GEPA: replacing sparse rewards with linguistic reflection, winning over GRPO by ~6pp with 4-35x fewer rollouts. Meanwhile, tool descriptions are the overlooked prompt -- small wording changes can shift tool selection rates by 10x, and Anthropic's experiments show Claude self-rewriting tool descriptions outperforms human experts. These two lines are converging: eval-driven automatic optimization is eating hand-tuned prompts.
Inferring another's beliefs/goals/intentions from observed behavior is called Machine Theory of Mind. Three lineages: symbolic BDI, Bayesian inverse planning, and deep learning ToMnet. The biggest controversy in the LLM era is that GPT-4 still trails humans by >10 points on ToMBench — are high scores genuine reasoning or statistical shortcuts?
At 99% accuracy per step over 100 steps, the error-free completion rate drops to just 36% -- error compounding is a structural problem, not something prompt tuning can fix. Distributed systems' supervisor trees, bulkheads, circuit breakers, sagas, and durable execution can be mapped almost one-to-one into agent orchestration. But LLMs introduce a failure class that traditional systems never had -- semantic errors that don't crash -- which require Inspector agents (recovering 96.4%) and redundancy voting (MAKER: one million steps with zero errors) to address.
Cosine similarity and relevance systematically diverge across an entire class of scenarios: negation (most IR models score at or below random on NevIR), exact identifiers, numeric thresholds, and logical combinations (SoTA models achieve recall@100 < 20 on LIMIT) -- some of these hit the theoretical ceiling of the single-vector paradigm, and switching to a larger model will not help. Recommended remedy order: hybrid BM25 -> reranker (Anthropic measured -67%) -> upstream metadata routing -> domain fine-tuning / multi-vector.
As tools scale up, selection accuracy doesn't degrade gracefully — it collapses: 4 to 51 tools drops from 43% to 2%, 10 to 100+ drops from 78% to 13.62%. The root fix is to stop stuffing everything in at once — Anthropic's Tool Search Tool uses defer loading plus retrieval to cut 85% of tokens, pushing Opus 4.5 accuracy from 79.5% to 88.1%. Description quality has conditional payoff: negligible in simple scenarios, but correctness jumps from 44% to 50% in multi-tool chaining.
Traditional Chinese RAG retrieval failures are a three-layer stack: embedding granularity defects (BGE/GTE from 0.1B to 7B all mis-rank on simple queries like 'fried chicken'), Simplified Chinese / English corpus dominance causing local vocabulary drift ('premium', 'exclusion clause' alignment is unreliable), and MTEB Chinese benchmarks being Simplified Chinese making model selection signals misleading. The fix is architectural: OpenCC normalization -> hybrid + jieba segmentation -> reranker -> local fine-tuning last -- and the prerequisite for all of it is building a Traditional Chinese eval set first.
arXiv does not perform peer review, and roughly 2% of submissions are rejected. Quality judgment relies on external signals: top venue acceptance > institution + open-source reproduction > citation quality. Includes a 20-item practical checklist and a 2026 toolbox (PWC has shut down).
Making 'chunk and embed every uploaded file automatically' the default behavior means making a decision for the LLM that it could have made itself. From Self-RAG (2310.11511) and Adaptive-RAG (2403.14403) to AgenticOCR (2602.24134), the academic trajectory is pushing three layers of decision-making -- whether to retrieve, whether to parse, and how to chunk -- from the ingestion pipeline back to the agent at conversation time.
The hard part of LLM agents is not building function calling, skills, code interpreter, and document tools individually -- it is assembling them into a system that selects the right tool, writes code when needed, decomposes tasks, verifies results, and resists prompt injection. This post organizes the key papers into six engineering decisions: function calling reliability, tool/skill selection, code-as-action, multi-step planning, skill systems, and safety plus document generation.
A2UI is an agent generative UI protocol open-sourced by Google on 2025-12-15: agents send declarative JSON describing UI intent, and clients render it natively using their own component catalog whitelist, layered on top of A2A. It launched at format v0.8 and iterated to v0.9 within three months.
browse.sh, launched by Browserbase in May 2026, is two things: a browser skill catalog and the Browse CLI. The core thesis: the bottleneck for browser agents isn't reasoning — it's amnesia. By storing learned site-specific workflows as plain-text SKILL.md files, Autobrowse cut Craigslist task costs from ~$0.22 to ~$0.12 by their own metrics. Note: this has nothing to do with the 2018 Browsh text-mode browser.
CodeGraph uses tree-sitter to extract a codebase into a local SQLite/FTS5 knowledge graph, letting AI coding agents query the graph instead of scanning files. The official end-to-end benchmark (7 repos, median of 4 runs) averages 35% cost savings and 70% fewer tool calls -- but only if the agent actually walks the graph. Delegating exploration to a file-reading subagent that ignores CodeGraph turns it into pure overhead.
Reading papers is two problems stacked together: methodology (Keshav's three-pass method, 5-10 min / 1 hour / 4-5 hours) determines how to read, and tools (arXiv HTML, alphaXiv, NotebookLM, Connected Papers, Zotero) shorten the time for each pass. AI lowers the barrier to understanding; judging correctness always stays with the human.
An MIT-licensed open-source UI automation framework from ByteDance. UI actions rely solely on feeding screenshots to a vision-language model, with no DOM parsing. A single JS API works across Web / Android / iOS / desktop. The trade-offs: each step is slower and more token-expensive, and everything hinges on the model's grounding ability. Note that Midscene retired MCP after 1.9.8 in favour of Skills + CLI.
Claude has no docx_tool or pdf_tool -- it relies on bash + file tools, plus SKILL.md instructions and pre-installed libraries like pdfplumber / python-pptx inside the container, assembling file handling capabilities from three layers.
Anthropic shipped Claude Design on 2026-04-17. On 4-28, nexu-io/open-design went public -- same artifact-first loop, Apache-2.0, runs on the 16 coding-agent CLIs you already have. Two weeks from 0.1 to 0.7, 40k+ stars. A paradigm shift that flattens AI design tools from vertical SaaS into a skill bundle.
asgeirtj/system_prompts_leaks collects the raw system prompts of 40+ AI assistants, from GPT-5.5 and Claude Opus 4.7 to Gemini 3.1 Pro, with 40.3k stars, 461 commits, and an MIT license. The value isn't in obtaining secrets -- it's in turning vendors' implicit policies into comparable engineering material. What you should study is the design decisions, not the text itself.
Anthropic's 35-page startup handbook released 2026-05-14 reorganizes Idea/MVP/Launch/Scale around agentic AI. The most valuable takeaways are 'the easier it is to build, the more important validation becomes' and treating CLAUDE.md as the first MVP artifact. The part to discount: the Launch chapter puts compliance workstreams on Cowork -- but Anthropic's own docs say Cowork doesn't write audit logs.
AI agents can operate video generation tools through three approaches — Skills, MCP Connectors, and direct APIs. Choosing the right integration method matters more than choosing the right tool.
Stop stuffing all your tool descriptions into context at session start. Let the model write code, have the runtime execute it, and let tool definitions enter context only at the import line — Anthropic's GDrive→Salesforce example dropped from ~150K tokens to 2K, and Cloudflare's 2,500-endpoint schema shrank from 1.17M to 1K.
MIT research says 95% of enterprise AI pilots yield zero return. OpenAI and Anthropic announced multi-billion-dollar joint ventures in the same week, wholesale adopting the Forward Deployed Engineer model that Palantir has used for over a decade to bring AI into the enterprise battlefield.
In May 2026, OpenAI published its internal Codex deployment practices: sandboxes define technical boundaries, approval policies determine when to pause, Auto-review delegates approval decisions to a sub-agent instead of a human, and Managed configuration lets enterprise admins enforce policies top-down. The core philosophy: zero friction for low-risk actions, mandatory review for high-risk ones.
Spin up a local OpenAI-compatible endpoint at localhost:20128 that automatically routes requests from Claude Code / Cursor / Cline / Codex / Copilot through a Subscription → Cheap → Free 3-tier fallback to 40+ providers. Built-in RTK compresses tool_result (saving 20–40% input tokens), Caveman mode compresses output, OAuth auto-refresh, multi-account round-robin — install with npm install -g 9router and two commands.
Three vendors originally took three routes: Anthropic built an extension, OpenAI built its own browser, Google welded AI into Chrome. By August 2026 there are only two — OpenAI's Atlas stopped working on 9 August, with its capabilities folded back into the ChatGPT desktop app and Codex. The remaining split is 'live alongside Chrome' versus 'be Chrome'.
Stripe Minions says 'The walls matter more than the model,' but the case studies from four Silicon Valley companies never explained how to actually build those walls. This post breaks down the 15 walls we implemented in the daodao auto-dev agent: what each wall prevents, where the files live, and what the tradeoffs are. Tier 1 is mandatory, Tier 2 strengthens governance, Tier 3 is serious governance.
A PM checks a task card in Notion → the system syncs it to a GitHub issue → writes a plan → writes code → opens a PR for human review. This post explains what the system does, what it doesn't do, and why it's feasible now — written for people who don't write code.
Build a Notion task → GitHub issue → spec PR → code PR auto-dev agent from scratch. Using the daodao case as a template, this guide walks through every step — what to do, what to verify, and how to handle problems. Notion DB schema → bin/ scaffold → two Claude Code routines → cloud env vars → staging tests.
Anthropic open-sourced 12 financial-industry Agents and 11 MCP connectors. The real takeaway isn't the Agents themselves but the layered design of 'one prompt, two runtimes' and 'pure-file extensibility.'
5 rounds of consensus to write the plan, then team mode with 5 workers running 12 tasks in parallel — with plenty of pitfalls along the way. Writing it down for my future self and anyone else trying the same thing.
DeepSeek-OCR's paper is titled Contexts Optical Compression -- OCR is just the means; what it actually validates is that 'rendering text as images and feeding them to a VLM' achieves 10x compression at 97% accuracy. This is a qualitative shift for long-context LLM and RAG token costs.
For side projects, toy demos, and RAG prototypes, nobody wants to swipe a credit card on day one. This is a verified roundup of 40+ LLM inference providers still operating as of 2026/05, tiered by whether free resources auto-replenish or are one-time grants. Each entry notes credit-card requirements, supported models, paid starting prices, and catches. Chinese-origin providers including Zhipu GLM (permanently free), Doubao (2M tokens/day), Kimi, DashScope, and the Ollama local option are all included.
A Skill is a folder with a SKILL.md. Three-layer progressive disclosure lets Claude load details only when needed, eliminating the need to re-explain preferences every conversation.
Local Deep Research is a privacy-first deep research agent built on LangChain + LangGraph, integrating 20+ search engines and 30+ research strategies. Its flagship langgraph_agent_strategy takes the LLM-autonomous tool-calling approach, offering a fundamentally different paradigm from fixed-pipeline RAG graphs.
PageIndex skips chunking, embedding, and vector storage entirely. Instead it relies on LLM reasoning over a tree-structured table of contents the LLM itself wrote, reporting 98.7% on FinanceBench in its own vendor-run evaluation. It solves a different problem than vector RAG — finding the right section in a well-structured long document.
When using AI agents like Claude Code or Cursor, built-in WebFetch / WebSearch often gets blocked by Cloudflare, geo-restrictions, or rate limits. Connecting a search MCP server is the most direct fix. This post compares the options actually available in 2026.
Groq Console is the developer portal for Groq's in-house LPU chip, offering an OpenAI-compatible API, Playground, and free tier credits. Its selling point is running open-source models like Llama, Qwen, and DeepSeek at the fastest tokens/second on the market.
goose is an open-source AI Agent maintained by the Linux Foundation's AAIF, supporting 15+ LLM providers and 70+ MCP extensions, built with Rust as a Desktop App + CLI + API. It positions itself as a vendor-neutral, self-hostable alternative to Claude Code.
For running Traditional Chinese LLM workloads on Cloudflare Workers AI, the Gemma family follows instructions more reliably than same-tier Llama models. gemma-3-12b-it was marked deprecated on 2026-05-30; the current equivalent is gemma-4-26b-a4b-it: 256K context, Vision, Function calling, at $0.10 / $0.30 per M tokens.
Karpathy proposed the llm-wiki pattern in 2026, having LLMs proactively maintain a markdown wiki instead of running RAG from scratch every time. Over 100 open-source implementations now exist, ranging from local CLI tools to serverless Telegram bots.
On 2026/4/22 OpenAI launched Workspace Agents — powered by Codex, capable of long-running cloud execution, and integrating with Slack/Salesforce/Google Drive. They are the enterprise successor to Custom GPTs.
Using Weaviate Query Agent + ColQwen multi-vector model, a single prompt built a production-grade legal contract search system in 36 hours -- this post breaks down its architecture logic, technology choices, and what you actually need to watch out for.
Cloudflare ran a Multi-Agent Code Review system internally for 30 days — 131K reviews, median 3 minutes. This post breaks down their architecture and compares it with solutions from Anthropic, GitHub, CodeRabbit, Greptile, and others.
A detailed look at OpenAI's Codex agent loop design: how prompts are constructed, how multi-turn conversations are managed, how prompt caching prevents cost explosions, and how context window auto-compaction works.
OpenAI wrapped the Codex harness as a JSON-RPC over stdio App Server, enabling VS Code, JetBrains, Web, and desktop apps to share a single agent loop. Three core primitives: Item, Turn, and Thread.
An OpenAI internal team spent 5 months with 3 people and 0 lines of hand-written code, delivering a complete product using Codex. This article distills their core lessons on AGENTS.md design, repo-local knowledge bases, architecture enforcement, and entropy management.
Agentic Engineering isn't about making AI write code faster — it's about making software move through the entire delivery pipeline faster, by using multi-agent collaboration to compress cross-team coordination friction.
Agent memory isn't a plugin — it's part of the harness itself. Pick the right memory type, estimate data volume, then decide on the technology. And finally, figure out whether you actually own that memory.
AI models rationalize their own code when reviewing it. Using three different CLIs for independent review effectively catches blind spots -- this post covers the design philosophy and practical workflow patterns behind the approach.
Agentic AI is not just autocomplete — it is an AI system capable of autonomously executing multi-step tasks. This article breaks down the five phases of the SDLC, explaining where to plug in agents at each phase, how to progress from CLI tools to full-pipeline automation, and the most valuable external resources to track right now.
Encyclopedia of Agentic Coding Patterns catalogues 190 patterns to help you make the right software decisions in the age of AI-written code — and the book itself is autonomously written and maintained by an AI agent.
GitHub Copilot Coding Agent lets you assign an Issue to Copilot, which then automatically creates a branch, writes code, runs CI, and opens a PR — all inside a cloud sandbox. The key to success is setting up AGENTS.md; without it, the agent tends to go off track. Best suited for well-defined medium-sized tasks; requires Pro+ (1,500 premium requests/month) or Enterprise plan.
A six-layer deterministic pipeline that handles everything from URL ingestion to vector embedding automatically, filtering out garbage before it enters your RAG system through an eight-dimension scoring system.
MCP is not going away, but its effective scope is narrower than most people think. For local development, CLI and raw API almost always beat MCP. MCP's truly irreplaceable niche is the narrow gap of 'cross-agent shared local tool layer.'
Not everyone should use a coding agent to modify code directly. AI Native teams need interface specs, test-first development, monorepo, security guardrails, human-in-the-loop, and token budget controls. Building an agent platform layer on top of coding agents and clearly redefining developer roles is the right path forward.
Autoreason replaces the traditional critique-and-revise loop with a competitive multi-version evaluation mechanism (A/B/AB + blind Borda count), solving three structural problems in LLM self-refinement: prompt bias, scope creep, and lack of restraint.
An open-source coding agent reference implementation from Vercel Labs. A three-layer architecture separates the web UI, agent workflow, and sandbox VM — designed as a starting point for teams that want to self-host their own Claude Code or Cursor Background Agent.
Claude Octopus is a Claude Code plugin that simultaneously calls Codex, Gemini, Copilot, Qwen, Ollama, Perplexity, OpenRouter, and Claude to review the same code, using a 75% consensus threshold to catch single-model blind spots. It ships with 32 personas, 48 /octo:* slash commands, 51 skills, and a Dark Factory fully autonomous spec-to-code pipeline.
LLM Council is a local Web App Andrej Karpathy built over a weekend. It sends one question to multiple LLMs simultaneously, has them anonymously peer-review each other, and then a Chairman model synthesizes a final answer. Positioned as a small tool for comparing models while studying — 99% vibe coded with no plans for long-term maintenance — but the architecture itself is a minimal ensemble LLM implementation worth studying.
Claude Managed Agents is a beta service launched by Anthropic on 2026/04/08 that provides an agent harness plus cloud container sandbox, billed per token plus $0.08/session-hour. It suits long-running async tasks and is worth exploring if you don't want to build your own agent loop and sandbox.
Graphify uses tree-sitter AST to extract code structure, then applies LLM semantic analysis to documents and images, compressing an entire project into a queryable knowledge graph. It claims to save 71.5x tokens per query compared to reading raw files.
Claw Code is a from-scratch Rust rewrite of the Claude Code CLI, featuring 48K lines of code, 40 tools, and MIT licensing. Most remarkably, the entire project was built by multiple AI agents collaborating over just 5 days, surpassing 170K GitHub stars within a week of launch.
clawhip is a Rust daemon that routes AI coding agent events (commits, PRs, session status) to Discord / Slack, solving the observability problem of not knowing who is doing what when multiple agents run in parallel.
notebooklm-py reverse-engineers Google's batchexecute RPC protocol, letting you programmatically control NotebookLM via Python / CLI / AI Agent — including audio, video, slides, quiz generation and more.
oh-my-claudecode (OMC) adds 8 collaboration modes, 19 specialized agents, and cross-model orchestration (Claude + Codex + Gemini) on top of Claude Code, transforming a single-user CLI tool into a multi-agent development platform. Features include Deep Interview for requirement clarification, Smart Model Routing that saves 30-50% on tokens, and automatic rate limit recovery.
oh-my-codex (OMX) doesn't replace Codex CLI — it adds a structured workflow layer on top of it. From requirements clarification and plan generation to multi-agent parallel execution, four core Skills transform scattered prompt conversations into a trackable development process.
oh-my-openagent (OmO) transforms OpenCode from a single-LLM tool into a multi-model agent team — Opus as the workhorse, GPT-5.2 as the architect, Gemini for frontend, Sonnet for documentation lookup — all triggered to run in parallel with a single ultrawork keyword. With 48K stars, it is the earliest project in the UltraWorkers ecosystem to establish the multi-agent coding pattern.
An open-source Agent Harness framework from HKUDS (HKU Data Science Lab) that implements tool calling, skill loading, memory, permissions, and multi-agent collaboration as complete infrastructure, supporting Anthropic / OpenAI / GitHub Copilot API formats.
There are already 6,400+ .claude/agents/*.md files on GitHub. We dissected 4 representative projects — ChemistryTimes (content production pipeline), claude-sub-agent (document-driven development pipeline), agentic (Temporal.io DAG parallel execution), and vs-copilot-multi-agent (hook-enforced memory persistence) — plus ruflo's enterprise-grade swarm architecture, distilling 6 design patterns and 5 practical trends.
Top Silicon Valley companies are independently building internal AI coding agents that automate everything from a Slack message to a merged PR. This article deep-dives into architectures from Stripe, Ramp, Coinbase, and Spotify — including their 2026 growth numbers (Stripe 7,000+ PRs/week, Ramp 75% of merged PRs) — then expands to cover Google, Meta, Amazon, Uber, Shopify, PostHog, and more.
Andrej Karpathy proposed a framework for compiling personal knowledge wikis with LLMs — collect raw data, have the LLM compile it into .md wiki pages, run Q&A against the wiki, and file outputs back. This post compares three practical approaches: Karpathy's knowledge vault model, the community's experience vault model, and quidproquo's blog model.
After dissecting Claude Code's 18+ caching mechanisms, I found that you can't touch provider-level prompt cache, but embedding cache, tool result cache, and entity cache are not only within your reach — they deliver even better results. Includes a complete AgentCache interface design and per-tool TTL strategy.
Every one of Claude Code's 45 tools uses a prompt() method that dynamically adjusts based on user type, feature flags, and system capabilities. Applying this pattern to a ReAct Agent, tool descriptions are dynamically generated along three dimensions: orchestrator model capability, locale, and available tools. Small models automatically get few-shot examples; large models save tokens.
Claude Code runs from $20/mo Pro to $200/mo Max 20x. Quota is a rolling five-hour window with weekly limits on top, shared across Claude on web, desktop, mobile, and the terminal. When you run out you can switch to usage credits at standard API rates rather than stopping.
Cursor CLI brings the IDE agent to the terminal with an interactive TUI and headless mode, Plan/Ask/Agent modes, Cloud Handoff, and CI/CD integration. Billing now runs on two separate usage pools: Cursor's own models (Grok 4.6/4.5, Composer 2.5) and third-party models (Pro includes $20, Pro+ $70, Ultra $400).
The paying paths on Google's side: the individual free tier and Gemini CLI access on Google AI Pro / Ultra ended 2026/6/18, leaving individuals with Antigravity CLI or their own paid API key; enterprise licenses and Google Cloud are unaffected. The zero-cost starting option now belongs to someone else.
Kiro has five tiers: Free 50 credits, Pro $20/1,000, Pro+ $40/2,000, Pro Max $100/5,000, and Power $200/10,000, with add-on credits at $0.04. Auto mode mixes models to cut cost (the same task costs 1.3x credits via Sonnet), and the spec-driven flow turns vibe coding into traceable, structured development.
Codex rides your ChatGPT subscription (Free / Go $8 / Plus $20 / Pro 5x $100 / Pro 20x $200), and since 2026/4/2 billing is token-based credits. The model line is GPT-5.6 Sol / Terra / Luna; GPT-5.4 and 5.4 mini retire from ChatGPT-signed-in Codex on 2026/8/31.
OpenCode is a free, open-source TypeScript CLI agent (MIT, ~198K GitHub stars). It supports 75+ model providers including local Ollama, allows authentication via Copilot/ChatGPT accounts, and lets you switch models mid-session without losing context. There is also a desktop app and an official Zen gateway.
A comparison of six agent CLI subscriptions (Claude Code, Cursor CLI, Codex, Kiro, Antigravity/Gemini CLI, OpenCode) plus the multi-model routing pattern — cheap models for simple work, strong models for hard work. Nearly every one of these changed its billing in the first half of 2026; this version was re-verified on 8/18.
Comparing the NVIDIA DGX Spark, Apple Mac Studio M4 Ultra, ASUS Ascent GX10, MSI AI Edge, and more — helping you find the right local inference hardware.
With multi-model routing, 70% of simple tasks are directed to cheap models, and only 10-15% of complex tasks use flagship models — saving 40-85% on inference costs in practice. This article covers the architecture and implementation of five major open-source tools.
Agent CLIs are not smarter autocomplete tools -- they are AI agents that can read your codebase, execute multi-step tasks, and operate in real environments. Claude Code, Codex CLI, Gemini CLI, OpenCode, Aider, Pi, Kiro, Amp, Cursor CLI... the tools keep multiplying, but they all share a common set of design principles -- understanding these principles is how you actually get good at using them.
Sorted by GitHub Stars, a survey of 15 mainstream AI Agent frameworks in 2026 — their positioning, key features, and ideal use cases. Not a ranking — it's a map.
Use Claude Code as an orchestrator to chain Playwright screenshots, catbox.moe image hosting, Meta Graph API publishing, and Telegram notifications — generate and publish an IG carousel from a single sentence.
llama.cpp is the most widely used local LLM inference engine, implemented in pure C/C++. It supports CPU, Metal, CUDA, Vulkan, and other backends, and uses the GGUF quantization format to run multi-billion-parameter models on consumer hardware.
TurboQuant+ is an open-source implementation of a Google Research ICLR 2026 paper that uses PolarQuant + QJL two-stage quantization to compress the KV cache by 3.8-6.4x, enabling consumer hardware to run larger models with longer contexts.
The main on-device LLMs in 2026 are Gemma 3n, Qwen 3.5 Small, Llama 3.2, Phi-4-mini, Ministral 3, and SmolLM3. Sub-3B quantized models can hit 30-50 tokens/sec on phones with 8GB RAM, but RAM, thermal throttling, and context window remain hard constraints.
2026 Q1 saw a full-blown open-source model explosion: on the LLM front, GLM-5, Kimi K2.5, and Qwen3.5 caught up with closed-source models; Embedding and Reranker are dominated by Qwen3 and BGE; speech has Voxtral TTS and Whisper V3; image has FLUX.2; and video has Wan 2.2 rivaling Sora. This is the complete navigation map.
In 2025-2026, websites need to be readable not just by humans but by AI. From llms.txt and Schema Markup to GEO and RAG ingestion pipelines, this post maps out the complete technical landscape for turning your website into an AI-consumable data source.
A Harness is more than just an LLM wrapper. Tool Registry manages dynamic tool loading and selection, Guard System establishes a four-layer defense network, and Checkpoint-Resume enables long-running tasks to survive interruptions. These three patterns form the critical infrastructure of production-grade Agent systems.
A Skill is a prompt template you invoke manually. A Subagent is an independent agent that Claude routes to automatically. They look similar, but differ completely in trigger mechanism, tool isolation, and context management.
When AI agents can turn intent into a PR in minutes, the bottleneck in software engineering flips from 'planning what to do' to 'evaluating whether the output is correct.' Artifacts of the ticketing era — sprints, story points, backlog grooming — are collapsing to zero, replaced by review as the core practice.
The same model produces dramatically different results under different harness designs. Anthropic uses a dual-agent architecture, cross-session state files, and a GAN-inspired generator-evaluator loop to let Claude autonomously complete hours-long software development tasks.
Google outlined eight multi-agent design patterns: from the simplest Sequential Pipeline to the composable Composite Pattern. More complexity isn't always better — picking the right pattern matters more than stacking agents.
AI engineering has gone through three phases: Prompt Engineering (write better instructions) → Context Engineering (feed the right information) → Harness Engineering (design the entire working environment). Each evolution doesn't replace the previous one — it operates at a higher level of abstraction.
The agent loop is a serialized per-session run. The part worth studying is how it handles concurrency: an admitted run records an activeWriterRunId claim, every transcript write supplies expectedWriterRunId, and the commit transaction verifies the match — so a superseded run cannot commit stale data.
OpenClaw builds its own system prompt for every run; there is no runtime default prompt. What it builds is split by an internal cache boundary — the stable workspace prefix above, the per-turn channel context below — so backends with prefix caches can reuse the same prefix across channels.
SecretRefs keep credentials out of plaintext config, and the model-call chain sees process-local sentinels instead of the real value. But the docs say it plainly: this is not process isolation — the real value still exists in the same process's memory, and any plaintext file the agent can read bypasses the whole mechanism.
Cron is now called Automations (openclaw cron remains an alias), and automation spans six mechanisms. The core trade-off is one line: Automations give you exact timing and isolated execution, Heartbeat gives you full main-session context on a roughly-every-30-minutes cadence.
Standing orders grant an agent permanent operating authority for a defined program, written into AGENTS.md and injected into every session. They define what it may do; automations define when — and the automation prompt should reference the standing order rather than duplicate it.
Every enterprise channel is a plugin now, including Slack and Google Chat, which used to be built in. Slack has three transports — Socket Mode, HTTP Request URLs, and relay — and the docs say plainly that the first two have reached feature parity, so you pick by deployment shape, not by features.
Each channel has one gotcha that stops you cold: WhatsApp's login is QR-only and hard to do remotely, Telegram bots ship with Privacy Mode on so they never see group messages (and you must remove and re-add the bot after changing it), and Discord needs Message Content Intent or it receives nothing from servers.
The most interesting entry here is Reef — an end-to-end-encrypted side channel between OpenClaw agents owned by different people. Messages are sealed on your machine, screened in both directions by a pinned-model guard, and the relay operator can never read the content. It ships bundled.
OpenClaw supports 31 chat channels, but only WebChat lives in core — even Slack and WhatsApp are plugins you install. And group safety has two independent axes: allowlists govern who can trigger the agent, not which quotes and history the model sees. That second one is contextVisibility, and it defaults to wide open.
OpenClaw validates config strictly — one unknown key, a wrong type, or an invalid value and the Gateway refuses to start. It keeps a last-known-good copy, but neither startup nor hot reload restores it automatically; only doctor --fix does.
The Gateway binds to loopback by default, and binding anywhere else requires auth — that is enforced, not advised. Inside a detected container the effective default is auto, unless Tailscale serve/funnel is active, which always forces loopback.
Deploying OpenClaw to the cloud comes down to four decisions: where the Gateway binds, where state lives, who can reach it, and how you recover. Which platform you pick is the least important of them.
OpenClaw has six local install methods, and what separates them is not the command but whether you want reproducibility, isolation, or self-updating. The real blocker is package-manager lifecycle-script policy: both npm 12 and global pnpm installs block OpenClaw's build scripts by default.
OpenClaw's failover runs in two stages: rotate auth profiles within the provider, then fall back to another model. But what really governs behavior is who chose the model — a model you picked yourself with /model is strict, and its failure is reported rather than answered by some other model.
OpenClaw's hard requirement for a model is tool use plus a large enough context — onboarding only auto-suggests a local model when it confirms tool support and at least a 16K context window. The easier thing to get wrong is that provider, model, and agent runtime are three separate layers: an `openai/*` ref does not mean Codex.
The official provider directory now lists 60 entries. The most common failure when attaching a local model is writing Ollama's base URL with /v1 — that breaks tool calling, and the model starts emitting raw tool-call JSON as plain text.
An agent is a complete persona scope — its own workspace, auth profiles, model registry, and session store. But the isolation is not absolute: when a secondary agent's OAuth credential expires, OpenClaw reads through to the main agent's profile of the same id, and a workspace is only a default working directory, not a hard sandbox.
The best part of remote node execution is how approval binds: exec prepares a canonical systemRunPlan before approval, and once granted the gateway forwards that stored plan — not any later caller-edited command, cwd, or session fields — and re-validates the working directory before running.
"OpenClaw is a Gateway shell around Pi" is obsolete. The docs now say the built-in runtime id is openclaw, that pi is a legacy alias which normalizes to it, and that no external agent framework packages remain. The only Pi-related third-party dependency left is a terminal component toolkit.
Node is the required runtime because the canonical state store uses node:sqlite — Bun is only for installing dependencies. Windows changed the most: there is now a native Windows Hub companion app that installs without administrator privileges and can provision its own app-owned WSL distro for the Gateway.
The iOS and Android apps are nodes, not Gateways: they do not run the Gateway service, and Telegram or WhatsApp messages land on the Gateway rather than on the phone. The Apple Watch is the exception — because watchOS blocks generic low-level networking for ordinary apps, it uses signed HTTPS polling instead.
The official framing is to treat plugin installs like running code — ClawHub and the bundled catalog are trusted sources, while arbitrary npm, git, and local paths require --force in noninteractive installs. And verification means inspect --runtime, because a bare inspect is only a cold manifest check.
Sandboxing is governed by three independent settings: mode (when it applies), scope (how many containers), and backend (where it runs). The most common failure is an expectation gap — `tools.exec.host` now defaults to auto, so 'unset means sandboxed' is no longer true, and the security audit has a check specifically for it.
By default every DM lands in one main session, and group activity and background work report back into it. Memory is entirely Markdown on disk — the model only remembers what gets saved, with no hidden state. But if more than one person can DM your agent, DM isolation is something you have to turn on.
OpenClaw's security docs open by stating the scope: this is a personal-assistant trust model, one gateway per trusted operator. It explicitly is not a security boundary for mutually adversarial users sharing one agent — and a 'not vulnerabilities by design' list pins that down.
OpenClaw's browser is a separate agent-only profile, fully isolated from your personal browser. And web_search's return shape carries an externalContent.untrusted marker — search results are typed as untrusted external content at the type level.
exec is a mutating shell surface: disabling write, edit, and apply_patch does nothing to make it read-only. And since sandboxing is off by default, host=auto actually resolves to the gateway — if you really want the sandbox, say so explicitly and it will at least fail closed.
Skills load from six sources with the highest precedence winning on name collisions, and a per-agent list replaces rather than merges. Sub-agents get no session or message tools by default — they return plain text to the parent, and the right to speak to a human stays with the parent agent.
When the tool catalog no longer fits in the prompt, OpenClaw offers two answers: Code Mode shows the model only exec and wait and has it write small programs against a hidden catalog, while Tool Search keeps structured search/describe/call controls. Neither bypasses tool policy.
The official triage flow is seven commands and two minutes to a diagnosis. And the most common symptom — the assistant feeling limited or missing tools — is usually the tool profile: minimal allows only session_status, while coding is the default for new local configs.
The Control UI gained a session rail: it uses a utility model to produce a run digest and attaches a read-only companion thread, so you can ask what a session is doing without entering or interrupting the main agent run. Its contents never enter chat.history.
The model is the CPU, the harness is the operating system, and the agent is the application. No matter how powerful a model is, without a good harness it's just a demo. Phil Schmid argues that harness is the most critical infrastructure in AI engineering for 2026.
LangGraph models LLM workflows as directed graphs, solving the pain points of multi-turn iteration, conditional branching, and parallel execution that are difficult to handle with linear pipelines.
Langfuse is currently the most mature open-source LLM Observability platform. This post covers four core capabilities — Tracing, Prompt Management, Evaluation, and Datasets — showing you how to use them in real projects.
Context Engineering is the core concept that replaced Prompt Engineering in 2025: the focus shifted from 'how to ask' to 'what information to provide.' Delivering the right information at the right time into the context window is more effective than upgrading to a stronger model. This post covers the definition, four key strategies, practical techniques, and common failure modes.
Every AI tool has its own calling format, making integration costly. MCP (Model Context Protocol) is an open standard proposed by Anthropic that unifies the communication protocol between AI Agents and external tools/data sources, enabling tools to be reused across Agents.
RAG is read-only. Agent Memory lets AI not only read but also write and persist information. Three memory types: Procedural (behavior patterns), Episodic (temporal events), and Semantic (factual knowledge) form a complete cognitive memory system.
AI Agent is not a single technology -- it is an entire architecture system. This article is a systematic navigation: starting from the Agent Three Pillars (Context/Cognition/Action), through the three-stage evolution of AI engineering (Prompt -> Context -> Harness), to eight Multi-Agent design patterns and production-grade Harness infrastructure. Each topic links to a dedicated deep-dive article.
An AI agent is not a black box — it is built from three layers: what it knows (Context), how it thinks (Cognition), and what it can do (Action). Understanding these three layers is the key to grasping why agents are sometimes brilliant and sometimes go off the rails, and how to design a truly effective agent system.
A single RAG Agent handling all queries hits knowledge boundaries and performance bottlenecks. Multi-Agent RAG dispatches retrieval tasks to multiple specialized Agents, each with its own knowledge base and retrieval strategy, coordinated by a central Orchestrator that merges results.
Traditional RAG splits documents into small chunks for retrieval, but this causes information fragmentation. LongRAG leverages 100K+ token long-context models to retrieve larger document segments (entire sections or even whole documents), reducing fragmentation while maintaining retrieval efficiency.
Speculative RAG uses small specialist models to generate multiple answer drafts from different document subsets in parallel, then a large model verifies and selects the best answer in one pass. The paper reports +12.97 points accuracy and -50.83% latency on PubHealth — but that is the best cell in the table; other benchmarks gain far less.
Ollama wraps llama.cpp in a Docker-style CLI + REST API, letting you run LLMs locally with a single command. This post covers core concepts, installation, API, hardware requirements, Modelfile customization, and what this tool is — and isn't — good for.
RAG has evolved far beyond simple 'search + generate' into a technology ecosystem spanning ten generations — and since 2025 into an Agentic/Reasoning era. This article is a systematic navigation guide: from Naive RAG to Multi-Agent/LongRAG across ten generations, the post-ten Agentic Era (Search-R1/RL search, MCP, GraphRAG 3.x, vision-native retrieval), retrieval strategies, chunking, embedding, reranking, evaluation frameworks, observability, and cost optimization. Each topic has a dedicated deep-dive article.
vLLM uses PagedAttention to eliminate KV cache memory waste, combining continuous batching and prefix caching to become the most widely adopted open-source LLM inference engine today.
Building a chatbot is more than just calling an API. Conversation state management, memory mechanisms, streaming, guardrails, observability, and tech stack selection — every layer affects the user experience.
Good prompts aren't written in one go — they're iterated into existence. Start with the simplest prompt, test with real cases, classify error types, and make targeted fixes. This article covers the three-part System Prompt structure, reasoning framework selection, few-shot optimization, token budget management, and six common mistakes.
For complex multi-hop questions, a single RAG search isn't enough. Agentic RAG lets the LLM evaluate whether retrieved results are sufficient — if not, it rewrites the query and searches again, forming a ReAct loop.
Your choice of embedding model directly determines RAG search quality. BGE-M3's multilingual training, 1024-dimensional vectors, and matching Reranker make it a practical pick for Traditional Chinese RAG.
Chunks too large and retrieval loses precision; too small and you lose context; hit a table and retrieval falls apart entirely. Chunking is the most underrated part of RAG — pick the wrong strategy and no amount of downstream optimization will save you.
Bi-Encoders are too coarse, Cross-Encoders are too slow — ColBERT's Late Interaction finds the sweet spot: token-level comparison between query and document, but with document vectors that can be precomputed.
When you split a document into chunks, each chunk loses its place in the original document. Contextual Retrieval solves the isolated-chunk problem by generating a per-chunk context from the whole document and prepending it at index time.
Filters too strict and getting zero results? CRAG automatically relaxes them and retries — far better than letting the LLM hallucinate an answer from general knowledge.
Vector search similarity scores don't equal relevance. Cross-Encoders use pairwise comparison to reorder results and push the truly relevant documents to the top.
Vector search handles semantics; BM25 handles keywords. Combining them with RRF is what lets you handle both fuzzy queries and exact terms at the same time.
After each conversation, asynchronously extract likely user preferences and skill level, then automatically personalize search parameters on the next query — no manual setup required.
Ranking purely by relevance leaves you with five documents all describing the same route. MMR strikes a balance between relevance and diversity, and layering in popularity weighting makes results even more useful.
RAG doesn't have to be a rigid three-step process. It's a set of steps that can be dynamically enabled, skipped, or reordered. Pipeline as Code lets the system adapt its behavior without redeployment.
A single vector search on a complex query often misses relevant documents. Let the LLM rewrite the query into 3-5 sub-queries, run them in parallel, and recall improves significantly.
Climbing routes carry a ton of visual information (topos, wall photos) that text-only RAG misses entirely. Multimodal RAG makes images searchable and understandable.
Naive RAG works but has real problems. Advanced RAG patches those problems. Modular RAG rearchitects the whole system to be composable and configurable. Understanding all three generations is the key to understanding why modern RAG systems look the way they do.
For complex queries, have the LLM map out what information is needed and in how many steps — then execute that plan. More systematic than thinking on the fly.
"Adding a Cross-Encoder feels better" is not a scientific evaluation. A/B testing tells you whether a change actually works, how much it helps, and which query types benefit.
A RAG system needs data to answer questions, but data only accumulates as the system gets used. Cold-start strategy is what bridges the gap from empty to useful.
RAG system costs come from LLM tokens, Embedding APIs, and vector search. Every stage has room for cost reduction, but you need to verify that optimizations don't sacrifice too much quality.
No industry standard mandates one RAG evaluation tool. Measure retrieval, generation, and operations separately, then choose Promptfoo, RAGAS, DeepEval, or TruLens for the actual stack.
When a RAG system breaks, 90% of the time it's one of these 10 failure modes. Identify which one first, then apply the matching fix — far more effective than optimizing blindly.
The attacks RAG systems face go beyond the technical level — Prompt Injection and Jailbreak are real threats. Both inputs and outputs need independent protection layers.
Rolling your own traces is good enough, but open-source tools save you a lot of work. Langfuse, Phoenix, and LangSmith each have their niche — the right choice depends on your trade-offs around self-hosting, open source, and integration complexity.
The hardest part of a RAG system isn't building it — it's figuring out why a particular answer went wrong. Pipeline Tracing records every step's decisions and data so debugging has a clear trail to follow.
Search found the right documents, but the LLM's answers are still poor — often the problem lies in prompt design. System prompt structure, context formatting, and instruction placement all affect output quality.
LLM generation takes 3-5 seconds, and waiting for the full response before displaying it makes for a terrible experience. SSE pushes tokens as they're generated, reducing time-to-first-character from 5 seconds to under 1 second.
Limiting request count alone is not enough — a single long query can consume ten times the tokens of a normal one. Dual quotas (request count + token count) are what truly control costs.
RAG and Fine-tuning solve different problems. RAG gives the model new knowledge; Fine-tuning changes the model's behavior and style. In most cases you use both, not pick one.
BM25, vector search, HyDE, and Multi-Query each produce separate result sets -- how do you merge them sensibly? RRF uses ranks instead of scores, sidestepping the fundamental problem that scores from different systems are incomparable.
BM25 only recognizes words that appear in the query. SPLADE infers related terms and adds them to the search, gaining partial semantic capability while preserving the precision of keyword search.
Questions like 'how many routes did I complete this year' will never be answered well by RAG semantic search — querying the database directly is far more accurate. Let the LLM identify intent, extract parameters, and execute predefined SQL templates.
Vector database selection is more constrained by deployment platform than LLM selection. Determine your platform and scale requirements first, then evaluate features — don't just look at benchmarks.