Skip to content

quidproquo

Tech, climbing, surfing, coffee, and everything else.

Quid pro quo /ˌkwɪd proʊ ˈkwoʊ/ — Latin for "something for something," a fair exchange.

This is where I document AI, tech, and product thinking, along with the process of building products — plus climbing, surfing, and coffee. Not just hoarding knowledge, but turning it into something useful for others, and then giving a little more.

Why the name →

CMU 10-423 L24–L26: Audio, Video Generation, and Interactive World Models — Taking Generative Models from Images to Sound, Time, and Worlds You Can Act In

The last three lectures of CMU 10-423 carry the Transformers, tokenizers, and latent diffusion from earlier in the course over to new kinds of data. L24 covers audio: turn sound into a mel-spectrogram or discrete tokens, then transcribe with Whisper, generate with AudioLM and MusicGen, and diffuse with AudioLDM. L25 covers video: 3D UNets with spatio-temporal attention, latent video diffusion, DiT and Sora, and finally the interactive NeuralOS. The first half of L26 covers world models, which predict the next state from a state and an action, along three routes: generate a 3D scene, interactive video (Genie), and latent representations (V-JEPA, PAN).

CMU 10-423 L5: CNNs, Encoder-only Transformers, and ViT — Why a Generative AI Course Starts Images with Understanding

CMU 10-423 Lecture 5 opens the image unit with three models built for understanding. CNNs treat the convolution kernel as parameters to learn. Encoder-only Transformers drop the causal mask so every token sees both sides and train with a masked LM objective, which means they are not generative language models. ViT is nearly BERT with image patches such as 16×16 pixels as input. The slides use a figure from the ViT paper to explain why Transformers reached vision four years after NLP: on small datasets ViT loses to large CNNs, and it only pulls ahead with enough data.

CMU 10-423 L23: Code Generation and Autonomous Agents — From pass@k to the Coding Agent Loop

CMU 10-423 Lecture 23 has two halves. The first covers code generation: evaluation moved from BLEU to counting passed unit tests, benchmarks run from HumanEval and MBPP to SWE-Bench Verified and Terminal-Bench 2.0, models run from CodeBERT and Codex to FIM and StarCoder, and the code-specific trick is self-correction driven by unit test output. The second half covers agents: what tool calling is, how Kimi K2 synthesizes tool-use data, the five-step coding agent loop, and web and GUI agents such as Mind2Web, Set-of-Mark, and SeeClick. There is no homework for this lecture; only Quiz 6 tests it.

CMU 10-423 L14–L15: Cross-Attention, DiT, Prompt-to-Prompt, and Q-Former

Where does the text condition enter an image generator? CMU 10-423 L14 answers with cross-attention: queries come from the image's latent representation and keys and values come from the prompt, so every latent pixel gets a probability distribution over which words to look at. That attention map is useful. Classifier-free guidance makes generations follow the prompt more closely, and Prompt-to-Prompt copies old attention maps into a run with an edited prompt so only part of the image changes, with no retraining. DiT swaps the UNet for a Transformer and injects conditions with adaLN-Zero. In the first half of L15, the Q-Former uses a small set of learnable queries to connect a frozen image encoder to a frozen LLM, which is what HW4 asks you to build.

CMU 10-423 L7: Diffusion Models, from Adding Noise to Learning to Remove It

L7 splits a diffusion model into two Markov chains. A fixed forward process gradually turns an image into Gaussian noise, and a learned reverse process removes the noise step by step. The exact reverse process is intractable, but the posterior given the original image x₀ is a closed-form Gaussian, so it can serve as the learning target. The slides compare three parameterizations. The best in practice has a U-Net predict the noise ε that was added, and the training loop is eight lines long.

CMU 10-423 L17–L18: Distributed Training, FlashAttention, and Efficient Decoding — Where to Start When One GPU Can't Hold the Model and Inference Is Too Slow

The last two lectures of the Scaling Up unit in CMU 10-423 Spring 2026. L17 starts from the claim that communication between GPUs is the main bottleneck, then walks through data parallelism, Megatron-style tensor parallelism, 1F1B pipeline parallelism, ZeRO optimizer parallelism, TeraPipe token parallelism, and expert parallelism. Its conclusion: data parallelism is still king, and the rest exist to push more data through it. L18 covers two things. FlashAttention combines tiling, online softmax, and recomputation to cut HBM traffic without changing the result. On the decoding side, PagedAttention manages KV-cache memory and speculative decoding reduces calls to the large model.

CMU 10-423 Wrap-up: The Practice Exam, the HW623 Paper Presentation, and the Final Project — How the Course Checks Learning, and How to Check Yourself

Beyond its four homework assignments, CMU 10-423 checks learning four ways: 6 in-class quizzes, 2 programming tests, one comprehensive exam, and a three-person final project worth 25%. 10-623/723 students also do HW623, a paper presentation. Outside CMU you can get the practice exam with solutions (13 sections, 167 points), the HW623 handout with its 33-paper list, and the 12-page project handout. This post lays out their structure and rules and gives a self-check routine that works without peeking at the answers.

CMU 10-423 L6: Generative Adversarial Networks and Probabilistic Graphical Models

L6 is the first real generative model in 10-423's image unit. A GAN is two deterministic networks: a generator that turns Gaussian noise into an image and a discriminator that tells real from fake. They play a minimax game and take turns with mini-batch SGD updates. The deck then covers scale, watermarking and societal impact, and closes with directed graphical models, Markov models and factor graphs to set up L7's diffusion models.

Reading CMU 10-423 Generative AI: Series Overview — All 26 Lecture Decks and Four Homeworks Are Public, the Videos Stay Behind Panopto

CMU 10-423/623/723 is the generative AI course co-taught by Matt Gormley and Aran Nayebi. The Spring 2026 edition covers text models, image generation, adapting foundation models, multimodal models, scaling, and advanced topics in 26 lectures. The slides, the HW1–HW4 handouts and starter code, a practice exam with solutions, and the project handout are all public, which earns an A3 rating. What you cannot get: the Panopto recordings, the HW0 handout, the HW3/HW4 recitation slides, the quizzes, and Gradescope grading. The homework policy is worth a look on its own: every assignment is submitted twice, first as human-only work, then with AI allowed.

CMU 10-423 HW1: Adding RoPE and GQA to minGPT — Structure, Files to Edit, and Compute

HW1 in CMU 10-423 Spring 2026 is worth 62 points. The written part covers RNN LMs (7), Transformer LMs (19), and sliding window attention (11). The programming part (22) has you implement RoPE and GQA in Karpathy's minGPT, train a character-level model on the complete works of Shakespeare, and plot loss and attention time. You upload only model.py; the handout ships five unit tests, and the official estimates put all experiments at about 40 minutes on a Colab T4.

CMU 10-423 HW2: Implementing DDPM from Scratch on AFHQ Cats — Structure, Files to Edit, and Compute

HW2 in the Spring 2026 CMU 10-423 is worth 60 points. The written part covers CNNs (8), encoder-only Transformers (4), GANs (5), VAEs (6), and diffusion models (14). The programming part (21) has you implement DDPM from scratch on AFHQ cat images: fill in the TODOs in diffusion.py and unet.py, then submit loss curves, FID curves, and forward/reverse diffusion figures from W&B. The longest experiment trains for 10,000 steps, which the handout estimates at about 2 hours on a Colab T4.

CMU 10-423 HW3: Fine-Tuning GPT-2 with LoRA — Written Questions, Files to Edit, and Compute

HW3 in CMU 10-423 Spring 2026 is worth 66 points and was due 2026-03-12 (Slot A). The written part covers in-context learning (14 points), parameter-efficient fine-tuning (10), and the DPO derivation (15). The programming part (25) has you write LoRALinear from scratch, wire it into GPT-2's attention, and instruction-tune the model for sentiment classification on Rotten Tomatoes reviews. Every experiment uses gpt2-medium; the handout estimates 25–30 minutes per training run on a Colab T4, and you need a WandB account.

CMU 10-423 HW4: Text-to-Image with a Q-Former Between a Frozen GPT-2 and a Frozen DiT — Structure, Files to Edit, and Compute

HW4 in CMU 10-423 Spring 2026 is worth 79 points. The written part covers LDMs (7), VQ-VAEs (8), CLIP (4), and VLMs through PaliGemma2 (18). The programming part (40) has you train only a Q-Former between a frozen GPT-2 and a frozen CIFAR-10 DiT, so a class-conditional diffusion model learns to take text. You write three functions, checked by 14 unit tests. The handout estimates 2–3 hours on a T4 or about 1 hour on an A100 for 25 epochs, and the captions and DiT weights come from Google Drive via download_data.sh.

CMU 10-423 L11–L12: Instruction Tuning, RLHF, and DPO — Make the Model Follow Instructions, Then Drop the RL

A pretrained LLM continues text; it doesn't hold a conversation. The second half of CMU 10-423 L11 covers instruction fine-tuning, which turns the model into a chat assistant using data such as InstructGPT's 13k examples, Dolly's 15k, or Flan. Then come InstructGPT's three RLHF steps: humans rank responses, a reward model is trained, and PPO fine-tunes the policy. The first half of L12 adds the intuition behind REINFORCE and PPO, lists five drawbacks of PPO-based RLHF, and derives DPO: start from the Bradley–Terry model, replace the reward model with the policy's own log-probability ratios, and fine-tune directly on preference data.

CMU 10-423 L19 + L21: Long Context and State Space / Hybrid Models — Three Ways Out When Attention Cost Grows Quadratically

Lectures 19 and 21 of CMU 10-423 (Spring 2026) tackle the same problem: once a sequence gets long, the memory of standard attention and the KV cache stop fitting. L19 offers two routes: approximate attention with sparse, sliding window, or dilated patterns, or keep full attention and split the computation across GPUs with the Blockwise Parallel Transformer and Ring Attention. L21 offers a third: replace attention with state space models (S4, Mamba) that keep only a fixed-size hidden state, or interleave attention with linear attention layers in hybrid models (Jamba, Nemotron-H, Qwen3-Next).

CMU 10-423 L4: Pre-training, Fine-tuning, and the Modern Transformer — What RoPE, GQA, and Sliding Windows Each Fix

The first half of CMU 10-423 Lecture 4 separates pre-training, mid-training, and post-training. The second half picks three components that nearly every modern LLM uses. RoPE turns position into a rotation of queries and keys, so attention scores depend only on the relative distance between two tokens. GQA lets several query heads share one key/value head to save memory and compute. Sliding window attention changes the mask so each token sees only a fixed number of tokens to its left. All three show up in HW1.

CMU 10-423 L10–L11: Parameter-Efficient Fine-Tuning and In-Context Learning — Change a Few Weights, or Just the Input?

With a small labeled dataset and an LLM with billions of parameters, CMU 10-423 offers two routes: supervised fine-tuning, or putting the examples in the prompt for in-context learning. L10 first notes that the 2023 consensus was that fine-tuning usually wins, then covers four ways to tune only a few parameters: the top layers only, adapters, prefix tuning, and LoRA. The first half of L11 returns to in-context learning: how sensitive it is to example order and label balance, how to pick a prompt, and what chain-of-thought is. HW3's written questions and its LoRA programming task both draw on these two lectures.

CMU 10-423 L20: Reasoning Models — From Chain-of-Thought to o1, DeepSeek-R1, and GRPO, Plus a Look at Mechanistic Interpretability

Lecture 20 of CMU 10-423 (Spring 2026) tells the story of reasoning models as one line: chain-of-thought prompting gets models to write intermediate steps, STaR fine-tunes on the reasoning that led to correct answers, and OpenAI o1 trains thinking tokens with reinforcement learning so compute can be added at both training and inference time. On the open side, DeepSeek-R1-Zero uses only rule-based rewards and GRPO and its reasoning grows longer on its own; DeepSeek-R1 adds SFT back to fix readability and language mixing. The lecture ends with mechanistic interpretability: why superposition makes models hard to read, and how replacement models such as sparse autoencoders, circuits, and cross-layer transcoders address it.

CMU 10-423 L22 + L26: Practical Risks and the Science of Alignment — Copyright, Jailbreaks, Hallucination, Bias, Carbon, and Why Alignment Has Theoretical Limits

Lecture 22 of CMU 10-423 (Spring 2026) runs five generative AI risks through the same four questions (what is it, who does it affect, why does it happen, how do we fix it): copyright infringement, adversarial attacks, hallucination, bias and discrimination, and environmental impact, and each section ends on a concrete example of why fixing it is hard. The second deck of Lecture 26 goes a level up: Aran Nayebi uses an agreement framework to show that the cost of alignment grows with the number of tasks, agents, and state space size, so objectives must be compressed and critical states prioritized, and he proposes a lexicographic utility that puts deference and the off switch first for provable corrigibility. Data contamination, listed in the course description, does not appear in either deck.

CMU 10-423 L1: RNN Language Models and Autodiff — Generative AI Starts with Predicting the Next Word (with HW0)

Lecture 1 of CMU 10-423 boils generative AI down to one line: it is probabilistic modeling, and text generation means estimating p(next word | all previous words). The slides go from n-grams, which you learn by counting, to RNNs, which squeeze the previous words into a fixed-length vector. In between comes module-based autodiff: if every module can run forward and backward, gradients flow back through the computation graph automatically, and that is how PyTorch works. The HW0 handout on Google Drive returns 401; only the recitation Colab is public, covering PyTorch, LSTMs, Weights & Biases and einops.

CMU 10-423 L15–L16: Scaling Laws and Mixture of Experts — How Big Should the Model Be, and How Do You Compute Only Part of It?

The first two lectures of the Scaling Up unit in CMU 10-423 Spring 2026 answer two questions. The second half of L15 covers scaling laws: Kaplan 2020 says 8x more parameters needs only about 5x more data, Chinchilla says scale both equally, and the Phi models and data-filtering scaling laws add data quality as a third axis. L16 covers MoE: feed-forward layers hold most of GPT-3's parameters, so split them into experts and send each token through only the top k. Memory follows total parameters, compute follows active parameters, and the price is load balancing and training stability. No programming homework covers this half of the course; quizzes, practice exam question 13, and the final project do.

CMU 10-423 L12–L13: Text-to-Image, Latent Diffusion, and Vision-Language Models

CMU 10-423 spends two lectures connecting generative models to a second modality. The second half of L12 asks how text can steer an image: three routes (GANs, autoregressive Parti, diffusion with DALL-E 2 and Imagen) lead to latent diffusion, which compresses images into an autoencoder's latent space, runs DDPM there, and reads the prompt through cross-attention. L13 goes the other way and lets a language model read images: CLIP/SigLIP or a VQ-VAE turns the image into vectors or integers for a decoder-only Transformer. What separates read-only VLMs (PaliGemma, Qwen-VL) from VLMs that can also output images (LWM, Gemini) is whether image tokens are discrete.

CMU 10-423 L2–L3: Transformer Language Models, LLM Training and Decoding — From Forgetful RNNs to the KV Cache

Lectures 2 and 3 of CMU 10-423 swap the RNN for attention. Lecture 2 first explains why RNNs fall short: they forget, they compute one step at a time, and their gradients can still explode. It then assembles a Transformer language model piece by piece: scaled dot-product attention, multi-head attention, layer norm, residual connections, position embeddings, and the causal mask. Lecture 3 covers training. There is no closed-form answer like n-gram counting, so you do maximum likelihood with autodiff and mini-batch SGD. It then covers padding, the KV cache and three kinds of tokenizer, and ends with greedy decoding and ancestral sampling to show how text is generated one token at a time.

CMU 10-423 L8–L9: Variational Inference, VAEs and the Diffusion ELBO

VAEs and diffusion models get stuck in the same place: log p_θ(x) requires integrating over latent variables, which is intractable. L8–L9 answer with variational inference. Pick a tractable q to approximate the true posterior, and swap 'minimize the KL' for 'maximize the ELBO,' which is a lower bound on log p(x). Add Monte Carlo estimation and the reparameterization trick, and a VAE trains with one forward and one backward pass. Unpack DDPM's ELBO and every term asks the learned reverse step to match the closed-form q(x_{t−1} | x_t, x₀).

CMU 11-868 L10: Accelerating Transformers on GPUs, and Where LightSeq Finds the Time

Lecture 10 of 11-868 uses Lei Li's own LightSeq and LightSeq2 as the case study and breaks them into four techniques: fuse every small operation outside matrix multiplication into one kernel, rewrite the LayerNorm and Softmax formulas to cut thread synchronizations, store parameters and gradients in FP16 but compute updates in FP32, and reuse memory based on backward-pass dependencies. The slides report 1.4-3.5x training speedups on WMT14 English-German. There is no recording; this guide works from slide page numbers and the two papers.

CMU 11-868 L14-L15: Distributed Training and Data Parallelism, and Where Gradient Sync Costs Come From

Lectures 14 and 15 of 11-868 go from the parameter server to PyTorch DDP. They use NCCL's five collectives (Broadcast, Reduce, AllReduce, ReduceScatter, AllGather) as building blocks, show why a ring makes broadcast time nearly independent of GPU count, and split AllReduce into ReduceScatter plus AllGather. The second lecture takes apart DDP's two key designs: bucketing gradients (25 MB by default) and starting synchronization before the backward pass finishes. There is no recording; this guide works from slide page numbers and the VLDB 2020 paper.

CMU 11-868 L05: How a Deep Learning Framework Computes Gradients from a Computation Graph

L05 follows a small sentiment classification network throughout. It expresses computation as a graph, evaluates it in topological order, sends gradients back with the chain rule and vector-Jacobian products, and then takes apart TensorFlow v1's placeholder, variable, operation, and session. One slide is labeled "important for HW2".

CMU 11-868 L21 FlashAttention: Attention Is Slow Because of Data Movement, Not Math — Tri Dao from FA1 to FA4

Standard attention writes the N×N score matrix out to HBM and reads it back, and most of its time goes to that traffic. FlashAttention uses tiling plus softmax rescaling so each block finishes inside SRAM, and the backward pass recomputes instead of storing. Tri Dao's guest slides for 11-868 give one set of numbers: the backward pass does 13% more FLOPs, 9x less HBM traffic, and runs 6x faster. FA3 and FA4 follow the same theme: when the hardware changes, the bottleneck moves, and the algorithm has to move with it.

CMU 11-868 L02–L04 GPU Programming and Acceleration: Threads, Blocks, the Memory Hierarchy, and Tiling

CMU 11-868's three GPU lectures answer one question: why does a correct CUDA matmul use only 2.48% of an A100's FP32 compute? L02 covers SMs, warps, and the grid/block/thread hierarchy. L03 covers cudaMalloc, cudaMemcpy, and kernel indexing. L04 uses tiling, coalesced access, and bank-conflict avoidance to bring data closer than global memory, which sits about 500 cycles away.

CMU 11-868 Assignment 1: Writing MiniTorch's map, zip, reduce, and matmul in CUDA

The first 11-868 assignment has you write four CUDA kernels in src/combine.cu (map 15, zip 25, reduce 25, matmul 30 points), wire them into MiniTorch's Python backend, and finish with a 5-point integration test. Shared-memory optimizations for reduce and matmul are marked Optional. The assignment page says plainly that you need a GPU, and grading uses private test cases.

CMU 11-868 Assignment 2: Implementing Autodiff in MiniTorch and Training a Sentiment Classifier

The second 11-868 assignment has three parts: autodiff's topological_sort and backpropagate (40 points), a Linear layer and MLP network (30), and binary cross entropy plus the training loop (30). You then train a sentiment classifier on SST-2 with GloVe embeddings and must reach 75% validation accuracy. The default backend is the CUDA kernels from Assignment 1, and the repo merged small Fall 2026 fixes on 2026-09-02.

CMU 11-868 HW3: Build GPT-2 in Your Own MiniTorch and Make It Translate German

HW3 has you add softmax loss, Dropout, LayerNorm, and Embedding to the MiniTorch you built in HW1 and HW2, assemble a pre-LN GPT-2 decoder, and train it on IWSLT14 German-English translation. Points: tensor functions 20, basic modules 20, decoder LM 40, translation pipeline 20. Full marks require passing the private tests and a BLEU of about 20±2. The assignment page warns that training alone takes at least 10 hours: one epoch is about an hour on a PSC V100, and you need 10. In Spring 2026 it went out Feb 4 and was due Feb 18.

CMU 11-868 HW4: Writing Softmax and LayerNorm as Fused CUDA Kernels

The fourth 11-868 assignment has you follow LightSeq and hand-write CUDA kernels for attention softmax and LayerNorm (forward and backward), bind them into your own MiniTorch, then swap them into your HW3 Transformer and train for one epoch. Points: Softmax 40, LayerNorm 40, integration 20. The assignment page expects individual kernels to be 3.7x to 15.8x faster, but end-to-end training only about 1.1x faster, because of Amdahl's law. You need an NVIDIA GPU, and the repo has already been changed for Fall 2026.

CMU 11-868 HW5: Writing Data Parallelism and Pipeline Parallelism Yourself on Two GPUs

CMU 11-868's fifth assignment switches to PyTorch and Hugging Face GPT-2. Using only torch.distributed and torch.multiprocessing, you write data parallelism (partition the data, set up a process group, average gradients; 50 points), then a GPipe-style pipeline (split the model, generate a clock schedule, run micro-batches on worker threads; 50 points). Both parts need benchmarks and plots on at least two GPUs: data parallelism must reach at least 1.5x speedup on 2 GPUs, and the pipeline must beat plain model parallelism. The Spring 2026 deadline was 3/25.

CMU 11-868 HW6: Training Llama-2-7B with DeepSpeed ZeRO + LoRA, Serving with SGLang

The sixth 11-868 assignment is the first to set aside your homemade MiniTorch and use industry frameworks. Two problems, 50 points each. Problem 1: edit a DeepSpeed training script to turn on LoRA so Llama-2-7B can train on two 16GB V100s. Problem 2: fill in the TODOs of an SGLang inference script and tune parameters to make generation faster. The two problems want conflicting GPUs: SGLang doesn't support V100, so you need an L40S, A6000, or A100. The spring due date was April 13, and the assignment page publishes no grading tests.

CMU 11-868 RLHF Systems and Assignment 7: A VERL-Style Pipeline with a Reward Model, GAE, and PPO

11-868's RL systems lecture has no slides; the Syllabus lists just one paper, ReaLHF. Assignment 7, on the other hand, is fully public. You train a DistilBERT reward model on Anthropic's HH-RLHF data (40 points), fill in GAE, the PPO loss, and entropy in a VERL-style trainer to fine-tune GPT-2 (40 points), and compare reward distributions before and after RLHF (20 points). The starter trainer never imports the verl package. What you learn is the RLHF dataflow, not VERL's distributed engine.

CMU 11-868 L01: Why LLMs Need Systems — the Scale Curve, Low-Level Operators, and Three Layers of Abstraction

CMU 11-868's first lecture spends 51 slides on one argument: the LLM bottleneck isn't only the model, it's computing larger LLMs on bigger datasets with fewer GPUs, less memory, and less power, faster. It breaks a Transformer into four low-level operators (matrix multiply, reduction, map, memory movement), sorts the hard problems into kernel, framework, and distributed-system layers, and warns that fast computation isn't enough because moving data takes time too.

CMU 11-868 L22 and L24 LLM Serving: Scheduling, RadixAttention, and PagedAttention

11-868 spends two lectures on one question: how does an inference server handle many requests at once without wasting KV cache on the GPU? Lecture 22 (Lei Li) starts from SGLang's scheduling loop: ORCA's continuous batching, RadixAttention's radix tree for KV, sorting and routing by prefix hit rate, and hiding CPU scheduling behind GPU compute. Lecture 24 is given by vLLM author Woosuk Kwon: PagedAttention cuts KV cache into fixed-size blocks and virtualizes them with a block table, taking the batch on one A100 from 8 to 40. The second half covers how vLLM cuts CPU overhead, uses piecewise CUDA graphs, splits models across GPUs, and manages memory for hybrid architectures.

Reading CMU 11-868 LLM Systems: Overview and Self-Study Paths — 28 Slide Decks and 7 Assignments Are Public, but No Videos and You Bring Your Own GPU

CMU 11-868 is Lei Li's graduate course on LLM systems: it goes from CUDA kernels and your own MiniTorch framework to distributed training, SGLang serving, and RLHF. All 28 Spring 2026 slide decks, 7 assignment pages, and 7 starter-code repos are public, which rates it A3. What's missing: videos, GPUs and a PSC account, the quizzes, and any official statement of which two assignments are optional.

CMU 11-868 L16–L17: When a Model Won't Fit on One GPU — Split Layers, Matrices, or Experts

CMU 11-868 (Spring 2026) spends two lectures on models too big for one GPU. L16 covers pipeline parallelism, which splits layers (GPipe micro-batches, 1F1B, interleaved stages), and tensor parallelism, which splits matrices (Megatron-LM's cuts for FFN, attention, and embeddings). The rule of thumb: TP inside a node, PP across nodes, DP on top. L17 treats MoE as a third way to split: each GPU holds different experts and replicates everything else. The price is all-to-all communication and load balancing, shown through GShard, DeepSpeed-MoE, and DeepSeek-V3.

CMU 11-868 L19–L20 Model Quantization: GPTQ Saves Memory, and the Speedup Comes Along for the Ride

11-868 spends two lectures on quantization. L19 goes from BF16 and absmax/zero-point quantization to AdaQuant, ZeroQuant, and LLM.int8(). L20 is all GPTQ. GPTQ quantizes weights only: after each column is quantized, it uses second-order information to adjust the weights not yet quantized, and lazy batch updates plus a Cholesky trick let it scale to 175B. What it mainly saves is memory. Inference gets faster because single-batch decoding was already bottlenecked on reading weights; the amount of arithmetic does not shrink.

CMU 11-868 L23 Efficient Fine-Tuning for Large Models: LoRA, CIAT, and QLoRA

Lecture 23 of 11-868 treats fine-tuning as a memory problem. Full-parameter half-precision fine-tuning of LLaMA-8B needs about 80GB. LoRA brings that to about 33GB, and QLoRA, which stores the frozen weights in 4 bits, gets it to about 9.2GB. The lecture moves in three steps: train only two small low-rank matrices A and B (the slides credit CIAT as the first to do this); squeeze frozen weights to about 0.52 bytes per parameter with an NF4 lookup table and double quantization; and use a paged optimizer to push optimizer state to the CPU when the GPU is about to run out.

CMU 11-868 Serving at Scale: Prefill/Decode Disaggregation, KV Cache, and Heterogeneous Hardware

11-868 closes with five serving decks: Hao Zhang on DistServe, Vikram Mailthody on NVIDIA Dynamo, Junchen Jiang on LMCache, Mingxing Zhang on Mooncake and KTransformers, and Lei Li's map of serving frameworks. They share one question: once serving grows from one machine to a data center, where do the compute and the KV cache go? The argument runs in three steps. Measure goodput under latency SLOs instead of raw throughput. Put prefill and decode on separate GPUs. Let the KV cache spill from GPU memory into CPU memory, SSDs, and remote storage.

CMU 11-868 L08–L09: Choosing a Vocabulary, Emitting Tokens, and Why Speculative Decoding Is Fast

L08 goes from BPE to VOLT, a method co-authored by the lecturer Lei Li: vocabulary size has both a cost and a value, and VOLT finds the sweet spot by asking how much normalized entropy each added token removes, then solves it as an optimal transport problem. The second half covers LLaMA 3 growing its vocabulary from 32k to 128k and the cost of byte-level BPE splitting one Chinese character into three tokens. L09 moves from greedy decoding, sampling, and beam search to speculative decoding: a small model guesses N tokens and the big model checks them in one forward pass, because checking is cheaper than generating. It ends with EAGLE, which predicts final-layer features instead of tokens.

CMU 11-868 L12–L13 TPU, JAX, and Pallas: One Attention Kernel on TPU, from XLA Fusions to Splash Attention

Across two lectures and more than 200 slides, Google's Srinath Mandalapu traces one attention computation from Python down to TPU VLIW instructions. L12 covers the JAX ecosystem, the memory and compute units of TPU Ironwood, and how XLA compiles attention into three fused kernels. L13 covers what XLA cannot do: using Pallas to control movement between HBM and VMEM yourself, writing FlashAttention, then adding block sparsity to get Splash Attention. The ideas match the GPU version. The difference is that on TPU the compiler does most of the scheduling, and Pallas is how you take loops and block sizes back into your own hands.

CMU 11-868 L06–L07: Reading Transformers, T5, LLaMA, and GPT-3 Like a Systems Engineer

11-868 spends only two lectures on the model itself. L06 breaks the Transformer into embeddings, multi-head attention, FFN, LayerNorm, and residuals; L07 uses T5, LLaMA, and GPT-3 to show what modern LLMs changed. For a systems engineer the point is to remember the shapes: GPT-3 175B has 96 layers, d_model 12288, a 2048-token context, and trained on 300B tokens; LLaMA 65B has 80 layers, d_model 8192, and trained on 1.4T tokens. Those numbers set the workload for every acceleration, parallelism, and serving lecture that follows.

CMU 11-868 L18: How ZeRO Cuts Data-Parallel Memory — Optimizer State, Gradients, Then Parameters

Data parallelism keeps a full copy of parameters, gradients, and optimizer state on every GPU. With Adam and mixed precision that is about 20 bytes per parameter, 16 of them optimizer-related, so LLaMA-3 8B already needs 160GB. CMU 11-868 L18 builds on the ZeRO paper and animates its three stages frame by frame: ZeRO-1 partitions optimizer state, ZeRO-2 also partitions gradients, and ZeRO-3 partitions parameters too. The slides conclude that the first two stages add no communication and save up to 8x memory; stage 3 makes per-GPU memory shrink with GPU count, at what the slides estimate as about 3x the communication.

CS149 L12: From One Chip to a Whole Datacenter — Dataflow Hardware, Kernel Fusion, Parallelism Strategies, and the Memory Bottleneck

L12 is about moving data. The first part uses the SambaNova SN40L to explain dataflow architecture and metapipelining: running Llama 3.1 8B, the slides say the RDU needs about 3 kernel calls per token versus about 800 on a GPU, because it can fuse an entire decoder into one kernel. The middle part scales up to the datacenter: which collective each of TP, PP, EP, and DP requires, and why overlapping compute with communication decides how well you scale. The last part returns to energy and DRAM: moving a byte costs far more than computing on it, and memory controllers, burst mode, and HBM all attack the same problem. There is no public video for this lecture; this post relies on the slides alone.

CS149 L14 Cache Coherence: MSI, MESI, and False Sharing

When every core has its own cache, one address can have several copies, and different cores can see different values. Locks can't fix this; the hardware created it by replicating data. CS149 L14 defines what coherent means, then takes apart the snooping MSI protocol: before writing, broadcast BusRdX so everyone else invalidates. MESI adds an E state that saves the second transaction in read-then-write, and directories replace broadcast with point-to-point messages. The practical consequence for programmers is false sharing: two threads write different variables, but because they share a cache line, the line bounces between cores. In the lecture's demo it made the program three times slower.

Reading Stanford CS149: A Guide to the Fall 2025 Parallel Computing Course

CS149 is Stanford's parallel computing course, taught by Kayvon Fatahalian and Kunle Olukotun. It runs from multi-core CPUs and SIMD through GPUs, AI accelerators, and the datacenter, then returns to cache coherence and lock-free programming. For Fall 2025, all 18 slide decks, the starter code and READMEs for 5 programming assignments, and 4 written-assignment PDFs are public, so this series rates it A3 (self-study ready). There are four gaps: the Fall 2025 lecture videos are Canvas-only; PA1 is graded on Stanford's myth machines; PA4 needs a self-funded AWS Trainium2 instance and a private course AMI; PA5's H100 job queue and leaderboard require a SUNet ID. The public videos are from 2023, and this series treats them as a listening supplement only.

CS149 Lecture 8: Data-Parallel Thinking, Replacing Locks with Map, Scan, and Sort

Lecture 8 asks you to switch mental models: stop thinking about what each worker does and write algorithms as operations on sequences, such as map, fold, scan, segmented scan, gather/scatter, sort, and groupBy. These primitives have efficient parallel implementations, and they turn irregular parallelism into regular parallelism and fine-grained synchronization into coarse synchronization. The price is extra passes over the data, so they are bandwidth hungry.

CS149 L9: Running DNNs Efficiently on GPUs — Conv as GEMM, Blocking, Fusion, and the Road to FlashAttention

L9 opens with a claim: if you understand arithmetic intensity and the roofline, you know almost everything about software-side performance optimization for modern AI. It then shows three things. Fully connected layers, conv layers, and attention all reduce to matrix multiplication (GEMM). GEMM needs blocking so data stays in cache. Adjacent layers should be fused so intermediates never round-trip through DRAM. Softmax can be computed in chunks, which is why fused attention (the core idea behind FlashAttention) never has to store the N×N matrix.

CS149 L13: Performance Optimization Beyond the Experts — Halide's Algorithm/Schedule Split, Autoschedulers, and LLM Agents

L13 asks what to do when there are too few people who can write fast code. The slides offer three answers. First, raise the level of abstraction: Halide splits what to compute (the algorithm) from how to compute it (the schedule), so one line of schedule turns the same blur into a tiled, vectorized, multi-core version. Second, intelligent search: because the schedule space is well defined, search plus a learned cost model can generate schedules automatically. Third, the emerging option of LLM agents: have a model write CUDA, run it, read the profiler, reflect, and revise, and let it improve itself with a database of examples or prompt optimization. The last slide leaves you with a question: is the real value in DSL design or in the LLM agent?

CS149 L16: Fine-Grained Locking and Lock-Free Programming, from Test-and-Set to the ABA Problem

CS149 L16 has three parts. It first looks at lock implementations through the lens of cache coherence (test-and-set, test-and-test-and-set, ticket locks, CAS, LL/SC). It then takes a sorted linked list from one big lock to hand-over-hand fine-grained locking. Finally it introduces lock-free programming: a single-producer/single-consumer queue, a CAS-based stack, the ABA problem, and hazard pointers. The slides land on a practical conclusion: when your program has the machine to itself, well-written lock-based code is often just as fast and much simpler. Lock-free designs pay off in systems where threads can be preempted or page-fault inside a critical section.

CS149 Lecture 7: GPU Architecture and CUDA Programming

CUDA's grid, thread block, and CUDA thread are programming abstractions; the GPU implements them with SMs, warps, and a hardware block scheduler. The heart of the lecture is keeping two things apart: the system may run thread blocks in any order, but all threads in one block are guaranteed to be live at once. That is why a block can cooperate through shared memory and __syncthreads(), and why the number of blocks an SM can hold is set by registers and shared memory.

CS149 L10: Why General-Purpose Processors Waste Energy — Hardware Specialization, Tensor Cores, TPU Systolic Arrays, and Dataflow Architectures

L10 starts from one equation: when power is capped, performance can only improve by spending fewer joules per operation, and a general-purpose processor spends most of its energy fetching, decoding, and moving data rather than computing. The slides' rule of thumb is that GPUs give about 10x better perf/watt than CPUs and fixed-function ASICs can reach 100–1000x. The lecture then judges the H100's Tensor Cores and TMA, Google's TPU systolic array, and reconfigurable dataflow architectures against the same checklist: tiled tensors, asynchronous compute and memory, and compute units talking directly to each other.

CS149 L3: Fast Processors, Slow Data — Latency vs. Bandwidth, and How ISPC Separates Abstraction from Implementation

The first half of L3 uses a highway and a laundry room to pull latency and bandwidth apart, then does the math: element-wise vector multiply runs at under 1% efficiency on a V100 because memory cannot feed the ALUs fast enough. The second half is about abstraction vs. implementation. ISPC lets you think in SPMD terms (a gang of program instances, each doing its share), while the compiler implements that with SIMD instructions. Mixing up the two layers is the most common source of confusion in the course.

CS149 L6 Locality, Communication, and Arithmetic Intensity: Why Moving Less Data Beats Adding Cores

CS149 Lecture 6 asks you to read "communication" broadly: data moving between a processor and its cache, its memory, or another machine all counts. Modern parallel processors have far more compute than bandwidth, so arithmetic intensity (how much computation you do per unit of data moved) decides whether you can keep the hardware fed. The levers fall into three groups: change the assignment to cut inherent communication, use blocking and loop fusion to cut cache-induced communication, and spread out or stagger accesses to reduce contention.

CS149 L2: Multi-Core, SIMD, and Hardware Multithreading, and the Problem Each One Solves

The second lecture of CS149 Fall 2025 takes a loop that computes sin(x) and adds three ideas in turn: spend transistors on more cores (multi-core), let one instruction drive many ALUs (SIMD), and interleave several threads on one core to hide memory latency (hardware multithreading). The first two add compute; the third keeps that compute busy while waiting on memory. The conclusion is three requirements: enough parallel work, groups of work that run the same instructions, and more parallel work than ALUs so latency can be hidden.

CS149 PA1 + Written 1: Measuring Speedup on a Quad-Core CPU and Explaining Why It Isn't Linear

PA1 has little code and a lot of analysis. Its six programs cover work assignment across threads, SIMD masking, ISPC gangs and tasks, how input data shapes SIMD efficiency, a bandwidth-bound saxpy, and finding a K-Means hotspot with timers. Written 1 drills the same intuitions on paper: peak throughput, instruction dependencies, pipelining, latency hiding with multithreading, and SIMD divergence. Official grading uses Stanford's myth machines; you can run everything on your own hardware, but your numbers won't match the reference.

CS149 PA2: Building a Task Execution Library from Scratch — Thread Pools, Sleeping, and Task Graphs with Dependencies

CS149 PA2 has you write a C++ task execution library for a multi-core CPU, and write it four times: spawn threads on every run(), switch to a spinning thread pool, switch to a sleeping thread pool, and finally extend it to asynchronous task graphs with dependencies. Every step must be a fully correct system, and each is timed against the official reference implementation. Official grading runs on AWS c7g.4xlarge; you can work on your own multi-core machine outside Stanford, but your numbers won't be directly comparable to the official thresholds.

CS149 PA3 and Written 2: A CUDA Circle Renderer That Must Be Both Correctly Ordered and Fast

PA3 has three parts: port SAXPY to CUDA and time it two ways, implement find_repeats with an exclusive scan, and write a CUDA circle renderer that is both correct and fast (85 points). The hard part of the renderer is that blending semi-transparent circles doesn't commute, so every pixel must be updated in input order, and the starter code's one-thread-per-circle approach gets neither atomicity nor order right. Written 2 has five graded problems (fusion, SIMD utilization, a barrier instead of locks, data-parallel primitives on graphs, locks in a particle simulation) plus 14 practice problems. Outside Stanford you need your own NVIDIA GPU. No solutions here.

CS149 PA4 + Written 3: Moving Your Own Data on Trainium2 — NKI, SBUF/PSUM, and a Fused Conv+Maxpool

PA4 drops you onto a single NeuronCore of an AWS Trainium2 chip. No cache decides what stays on chip: you move data into SBUF (28 MiB) and PSUM (2 MiB) yourself with dma_copy, and the partition dimension tops out at 128. Part 1 teaches those limits and the cost of DMA through vector add and transpose. Part 2 asks you to rewrite convolution as a series of matmuls and fuse it with max pooling so nothing spills back to HBM. Written 3 drills the same idea with a line buffer, two back-to-back box blurs, softmax hardware, and metapipelining: keep intermediates on chip. The environment needs the course's private AMI and a paid capacity block, so for outside readers this assignment is effectively A2.

CS149 PA5, the Fastest Kernel on an H100: Five AI Kernels, Graded on Your Work Log

The last programming assignment in CS149 Fall 2025 is open-ended. Pick at least one of five kernels (Histogram, a 1D occupancy decoder, FlashAttention, a 3D heat equation with RK4, SwiGLU) and make it faster than its PyTorch baseline on an H100. You can write CUDA, Triton, or TileLang, and you may use LLMs. There is no speed threshold. The grade depends on a work log that shows what you measured at each step, what hypothesis you formed, and why you stopped. The H100 job queue and leaderboard need a SUNet ID; outside Stanford you can only run eval.py on your own NVIDIA GPU.

CS149 L4: How Do You Parallelize a Program? Decomposition, Assignment, Orchestration, and Amdahl's Law

L4 lays out a thought process for parallelizing code: decompose the problem to find independent work, assign that work to workers, orchestrate communication and synchronization, then map workers to hardware. Amdahl's Law reminds you that the sequential fraction caps speedup. The running example is a 2D grid solver whose original dependencies are hard to exploit; switching to a red-black update order makes it expressible in either a data-parallel or a shared-address-space model.

CS149 L11: Programming Specialized Hardware — ThunderKittens Tames H100 Asynchrony, Dataflow Replaces It with Metapipelines

L11 asks what programmers pay once hardware specializes for AI. On the H100, saturating Tensor Cores means 16×16 tiles, TMA moving data asynchronously, and producer and consumer warps running as a pipeline. That's hard to write, which is why DSLs like ThunderKittens exist. The other route is a dataflow architecture (SambaNova SN40L): describe the computation with parallel patterns such as map, reduce, and zip, and let the compiler handle tiling, metapipelining, and placement. The slides say this can fuse an entire Llama 3.1 8B decoder layer into one kernel.

CS149 L15 Memory Consistency: How Write Buffers Make r1 = r2 = 0 Possible

Coherence covers a single address. Memory consistency covers reads and writes to different addresses, and the order in which other threads see them take effect. CS149 L15 uses two threads and two variables to make the point: under sequential consistency, r1 = r2 = 0 is impossible, but the write buffer in every modern processor lets reads pass writes, so it becomes possible. TSO, PSO, and weak ordering relax more orderings in exchange for speed, and fences and synchronization primitives restore the orderings you need. The takeaway for application programmers is short: write data-race-free programs and use a synchronization library, and C11, C++11, and Java 5 guarantee you'll see sequential consistency.

CS149 L17–L18: Transactional Memory and Written 4, Handing "Make This Atomic" to the System

Coarse locks are easy to write but slow; fine-grained locks are fast but easy to get wrong. Transactional memory lets the programmer just declare atomic { } and leaves atomicity and isolation to the system. CS149 L17 covers the motivation (failure atomicity, composability) and the design space: data versioning is eager (undo log) or lazy (write buffer), and conflict detection is pessimistic or optimistic. L18 opens up STM runtime data structures and the McRT algorithm, then shows how HTM uses per-line R/W bits plus the coherence protocol to detect conflicts, ending with Intel Haswell's RTM. Written 4 ties MSI, LL/SC, locks and memory ordering, and fine-grained locking on a doubly linked list into four problems.

CS149 L1: Why Single Cores Stopped Getting Faster, and Why Fast Isn't Efficient

The first lecture of CS149 Fall 2025 defines speedup, then uses three classroom demos to show how communication and load imbalance eat into it. Next it explains why single-core performance stalled: superscalar execution runs out of instruction-level parallelism at about four instructions per clock, and clock frequency hits the power wall. So performance now has to come from more cores and specialized hardware. The last part turns to efficiency. A DRAM access takes about 60 times as long as an L1 cache hit, and moving 64 bits costs over a thousand times the energy of an integer op. Efficiency almost always comes down to accessing data efficiently.

CS149 L5 Work Distribution and Scheduling: From Work Queues to Cilk's Work Stealing

Load balancing is hard because it pulls against scheduling cost: smaller tasks balance better, but every task grab pays a synchronization cost. CS149 Lecture 5 first lays the options out as a continuum from static to dynamic, then takes apart the Cilk runtime: one deque per worker, run the child at a spawn and leave the continuation for others to steal, and idle threads steal the biggest chunk of work from the top of someone else's deque.

CS224R L4: Actor-Critic and Value Estimation — Learn to Judge Good and Bad, Then Do More of the Good

Policy gradient can only judge good and bad from the rewards it actually received, so it wastes data. Actor-critic trains a second network, a value function (the critic), to estimate how good a state is, and uses it to compute advantages that weight the policy's (the actor's) gradient. There are three ways to estimate value: supervise directly with a rollout's summed rewards (Monte Carlo), supervise with this step's reward plus your own estimate of the next state (bootstrapping), or use an n-step return in between. L4 ends by pushing actor-critic off-policy, first by taking several gradient steps on one batch (where PPO starts) and then by reusing all past data from a replay buffer (where SAC starts).

Reading Stanford CS224R: A Guide to the Spring 2026 Deep Reinforcement Learning Course

CS224R is Chelsea Finn's deep reinforcement learning course at Stanford. It runs from imitation learning to RL for LLMs and robot foundation models. For Spring 2026, all 17 slide decks, the three homework handouts with starter code, and the default project spec with starter code can be downloaded without logging in, so this series rates it A3 (enough to self-study). The gaps: the 2026 recordings are Canvas-only, the midterm and its solutions are not public, and HW2 and HW3 require Modal. The public recordings are from Spring 2025, so this series uses them as a supplement and flags the differences lecture by lecture.

CS224R Default Project: Fine-Tuning an LLM on Countdown with SFT, IPO, and RLOO

The CS224R Spring 2026 default project has you implement three stages on Qwen2.5-0.5B Base for the Countdown arithmetic reasoning task: SFT warm-start, IPO preference optimization, and RLOO with a rule-based verifier reward. All three are compared with the same vLLM evaluation, followed by a research extension of your choice. For the implementation, high-level trainers like SFTTrainer are banned, and so is any AI tool assistance; only the extension is exempt. The extension is half the grade for this project, and it's graded on methodology and documentation, not score. The starter code and datasets are public. What outside readers lack is Modal credits and the autograder.

CS224R L18: Open Problems in Deep RL, and How to Do Research

The last CS224R lecture has three parts. It first folds the whole quarter into one toolbox. It then lists seven unsolved problems: domains without verifiable rewards, using prior data, world models, scaling, safety, hallucination and calibration, and evaluating generalist systems. Nearly half the deck is about how to do research: you need both an important problem and a workable plan, you front-load the risk, you consider pivoting early, and research only counts once you share it. Reading it alongside the 244 public 2026 final project reports shows what those principles look like in practice.

CS224R L15: Hierarchical RL and Imitation Learning

Long-horizon tasks are hard because the agent visits a huge number of states and has many chances to make mistakes or get stuck. Lecture 15 of CS224R answers with two levels: a high-level policy proposes subgoals, and a low-level policy runs at a higher frequency to reach them. The real design decisions are three: how to represent the subgoal, how to supervise each level, and when to switch to the next subgoal. The slides also admit that nobody has yet shown whether hierarchy beats a single policy with chain of thought.

CS224R HW1: Regression BC, Flow Matching, and DAgger on Flappy Bird

Homework 1 of CS224R Spring 2026 tests imitation learning on a custom Flappy Bird environment. The policy predicts 20 future target heights at once and executes only the first 10. You implement MSE-regression behavior cloning, a flow matching policy, and DAgger, then compare them in easy and hard modes. The PDF, LaTeX template, and starter code are all public, and a CPU is enough to run it. Solutions, the autograder, and Gradescope are not public. This guide covers what each problem asks you to build and answer. It does not give solutions.

CS224R HW2: Gridworld Q-learning, PPO, and the Sawyer Hammer Task

CS224R Spring 2026 HW2 has three parts. First, tabular Q-learning on a 5×4 gridworld shows how reward design changes the learned path. Second, GAE plus PPO clipping tackles a hammer task that pays 1 only on completion. Third, an off-policy actor-critic with BC pretraining, a critic ensemble, and a higher UTD ratio, followed by a comparison of the two learning curves. The handout, starter code, and compute guide are public, but the assignment supports only Modal, and course credits go only to enrolled students.

CS224R HW3: AWAC, IQL, and Stitching on AntMaze

HW3 in CS224R (Spring 2026) has you fill in two offline RL algorithms, AWAC and IQL, and compare them on D4RL's AntMaze. Problem 1 runs AWAC on antmaze-umaze and antmaze-medium-diverse. Problem 2 compares IQL expectiles ζ = 0.2 and 0.9, runs the better value on medium-diverse, and then tests whether IQL can stitch a better path out of a PointMass dataset whose best return is only −46, against a filtered BC baseline that keeps the top 10% of trajectories. The PDF, LaTeX template, and starter code are public, but the assignment is meant to run on Modal, and course credits go only to enrolled students. This post covers the tasks and setup only, with no solutions.

CS224R L2: Imitation Learning and Policies That Can Represent Multimodal Distributions

Lecture 2 of CS224R Spring 2026 tackles two ways imitation learning fails. First, when demonstrations contain several reasonable behaviors, regression learns only their average. The fix is to make the policy a generative model (Gaussian mixtures, discretization plus autoregression, diffusion or flow matching) and to add action chunking. Second, compounding errors: once the policy slips, it reaches states the demonstrations never covered. The fix is DAgger or human-gated DAgger to collect corrections. The first two parts are exactly what HW1 covers.

CS224R L1: Framing Decision-Making as an RL Problem

The first lecture of CS224R Spring 2026 does three things: covers logistics, explains why deep RL is worth learning, and turns 'behavior' into something you can learn. The core is a set of definitions (state, action, trajectory, reward, policy) and one objective: maximize expected total reward. It ends on an example: fit ℓ2 regression to drivers where some change lanes and some go straight, and the policy learns their average, a half lane change nobody demonstrated. That problem is where L2 starts.

CS224R L13: Meta-RL, Teaching an Agent to Learn New Tasks Fast

Meta-RL trains on many tasks so that a new task can be solved from a small amount of experience. Lecture 13 of CS224R frames it as "explore to collect a little data, then adapt using that data." The most direct approach is black-box meta-RL (RL²): a network with memory takes past (s, a, r) as input and keeps its hidden state across episodes. It is general and expressive but hard to optimize, especially when exploration is hard, because exploration and execution depend on each other and end-to-end training gets stuck. The slides then compare posterior sampling in PEARL, prediction-driven exploration in MetaCURE, and DREAM, which uses a task representation to train exploration and execution separately.

CS224R L11: Model-Based RL, or Learn a Simulator and Don't Trust It Too Much

Model-based RL first learns a dynamics model that predicts s_{t+1}, then uses it in one of two ways: to generate extra training data (Dyna, MBPO) or to think a few steps ahead before acting (planning). The thread running through Lecture 11 of CS224R is how to avoid being dragged down by model error. Start synthetic rollouts from real states and keep them short, average errors out with an ensemble of models, and attach a value function to the tail of long-horizon plans. Whether a model is worth learning depends on whether it is easier or harder to learn than the policy.

CS224R L12: Multi-Task and Goal-Conditioned RL, Sharing Weights and Sharing Data

Multi-task RL treats which task you are on as part of the state, s = (s̄, z_i), so the problem is still an ordinary MDP and standard RL algorithms still apply. Lecture 12 of CS224R covers two kinds of sharing: weight sharing, where one network conditioned on z_i does every task, and data sharing via hindsight relabeling, where data collected for task A gets relabeled as data for task B. Goal-conditioned RL is the special case where the task is a goal state; relabeling with the state you actually reached eases the exploration problem of sparse rewards. Data sharing has three prerequisites: consistent dynamics across tasks, a reward you can evaluate, and an off-policy algorithm.

CS224R L5: Off-Policy Actor-Critic — the Shared Skeleton of PPO and SAC

PPO and SAC answer the same question: can you use an expensive batch of data more than once? PPO takes several gradient steps on one fresh batch and clips the new-to-old policy ratio to 1±ε. SAC keeps every past transition in a replay buffer and learns Q(s, a), so old data can still evaluate the current policy. PPO is stable and easy to tune. SAC is data-efficient and harder to tune.

CS224R Lecture 7: Offline RL, or Why Q-Learning Breaks When You Can't Collect More Data

Lecture 7 of CS224R (Spring 2026) asks how to learn a policy better than your data when all you have is a fixed dataset someone else collected. Running an off-policy algorithm like SAC on that data fails: the Q-function makes up values for actions the data never contains, and the policy goes looking for exactly those overestimated actions. The slides give two families of fixes. One trains the policy only on actions in the data (filtered BC, AWR, AWAC). The other uses an asymmetric expectile loss to estimate the value of a policy better than the data without ever querying out-of-data actions (IQL). Both can do something imitation learning can't: stitch good pieces of different trajectories together.

CS224R L3: Policy Gradients — Differentiating the Policy Without Knowing How the World Works

Policy gradient is the first online RL algorithm in CS224R. Its gradient looks almost exactly like the imitation learning gradient, except that each trajectory is weighted by its reward. Actions from good outcomes become more likely, and actions from bad outcomes become less likely. The raw version is very noisy, so L3 cuts the variance in two ways: count only future rewards (causality) and subtract the average reward (a baseline). It is also on-policy, so every gradient step needs fresh data. Importance sampling plus a KL constraint lets you take several steps on one batch.

CS224R L6: Q-learning and How to Stabilize It

Q-learning drops the actor from actor-critic: learn the optimal Q-function directly and act by taking the argmax. The price is that convergence is not guaranteed; even linear Q can diverge. Lecture 6 of CS224R pulls it back with three engineering tricks: a target network that holds the targets still, Double Q that separates choosing an action from valuing it to curb overestimation, and n-step returns that trade a little bias for speed.

CS224R Lecture 8: Where Rewards Come From, Learned from Examples and Preferences

Lecture 8 of CS224R (Spring 2026) spends a few slides wrapping up offline RL, then asks the question the first seven lectures skipped: where does the reward come from? Games have scores. Real robots, dialogue, and driving usually don't. The slides offer two routes. The first trains a goal classifier on success examples and uses it as the reward, but RL learns to exploit the classifier's blind spots; the fix is to keep adding states the policy visits as negatives, the same structure as a GAN. The second asks people which of two trajectories is better and learns a reward with the Bradley-Terry-style objective log σ(r(τw) − r(τl)), the same method LLM RLHF uses. The lecture's number-one takeaway is one line: rewards can't be taken for granted.

CS224R L17: RL for Robot Foundation Models (VLAs)

VLAs trained only with imitation learning often plateau around 80% success, while autonomous robots often need 99%+. Lecture 17 of CS224R splits "how do you improve a VLA with RL on a real robot" into three routes: recast RL as supervised learning (iterated offline RL), learn a small separate policy on the VLA's representation or diffusion noise, or learn a small policy that edits the VLA's actions. The slides call this an open research problem and describe the content as recent themes plus the speaker's opinion.

CS224R L10: RL for LLM Reasoning and Test-Time Compute

Lecture 10 of CS224R Spring 2026 is a guest lecture by Noam Brown of OpenAI, and it makes one argument: reasoning models open a new scaling dimension by moving compute from training to inference. He starts with his own poker AI work, then uses backgammon, chess, and Go to show that thinking longer at inference time has always paid off. Next comes how LLMs got there: chain of thought, majority voting, o1/o3, GRPO, and DeepSeek-R1-Zero. The second half argues the field needs to rethink itself for large-scale test-time compute: multi-agent systems, evaluation as score versus compute, and the budget assumptions behind safety evaluations. The deck is mostly figures, so this post covers only the points visible on the slides.

CS224R L9: RLHF, DPO, and Preference Optimization

Lecture 9 of CS224R Spring 2026 is a guest lecture by Archit Sharma, with slides adapted from CS224N. The spine is one chain of reasoning. Instruction tuning can't handle tasks with no right answer or errors of unequal weight, so we optimize human preferences directly. Human ratings are expensive and noisy, so we collect pairwise comparisons and fit a Bradley-Terry reward model. RLHF uses that model as the reward and runs policy gradient with a KL penalty. DPO uses the closed-form solution of the KL-constrained problem to write the reward as a log-ratio of policies, which turns the whole thing into a binary classification loss. The last part covers the frontier: reward hacking, verifiable rewards, and AI feedback in place of human feedback.

CS224R L16: Sim-to-Real Robot Learning

Simulators are cheap, fast, and safe, and they hand you labels the real world never will, but they never match reality exactly. In Lecture 16 of CS224R, CMU's Guanya Shi sorts the ways to close that gap into three families: domain randomization trains one policy that works across many physical parameters; teacher-student trains a teacher on privileged information and then has a student that sees only real sensors imitate it; real2sim2real uses real data to make the simulator more faithful. The advanced topics are defining tasks from human motion data and choosing RL algorithms suited to sim2real.

CS231N L15: 3D Vision — One Shape, Five Ways to Store It

CS231N Lecture 15 runs on one question: what data structure should a 3D shape use so a neural network can read it and produce it? The slides walk through five representations (depth map/surface normals, voxels, point clouds, triangle meshes, implicit surfaces), each with a signature architecture (fully convolutional depth prediction, 3D convolution, PointNet, Pixel2Mesh and Mesh R-CNN, DeepSDF). Then comes the speed trade-off between NeRF and 3D Gaussian Splatting, and a closing roll call of 2025–2026 models: VGGT, TRELLIS, Marble. The 2025 recording uses a different slide deck, with a different order and emphasis.

CS231N Assignment 1 Guide: kNN, Softmax, Two-Layer and Fully Connected Networks

CS231N Assignment 1 is worth 12% of the grade and was due April 16, 2026. All five Colab notebooks are hand-written numpy on CIFAR-10. Q1 kNN asks for distance computations with two loops, one loop, and no loops. Q2 Softmax goes from naive to vectorized to SGD. Q3 assembles affine, ReLU, and softmax into a two-layer network. Q4 switches to HOG and color-histogram features. Q5 generalizes to any depth and implements Momentum, RMSProp, and Adam. The 65 KB starter code is public; the Gradescope grading isn't. This post covers structure and goals only, not solutions.

CS231N Assignment 2 Guide: BatchNorm, Dropout, CNNs, PyTorch, and RNN Captioning

Assignment 2 of CS231N Spring 2026 is worth 18% of the grade. Across five notebooks you hand-write BatchNorm/LayerNorm, dropout, and the forward and backward passes for convolution and pooling, then learn PyTorch at three levels of abstraction, and finish with RNN image captioning on COCO in PyTorch. Q4 is the turning point: through Q3 you derive every gradient yourself, and from Q5 on autograd takes over while a numerical gradient check confirms it. The official slides warn that this is the longest of the three assignments.

CS231N Assignment 3: Transformer Captioning, Self-Supervised Learning, DDPM, CLIP and DINO

Assignment 3 in CS231N Spring 2026 is worth 15% of the grade and turns L8, L12–L14 and L16 into four Colab notebooks. Q1 has you write multi-head attention and a Transformer decoder for COCO captioning, then assemble a ViT and train it on CIFAR-10. Q2 implements SimCLR's augmentations and contrastive loss and compares linear classification with and without self-supervised pretraining. Q3 builds DDPM's noising, UNet, denoising loss, sampling and classifier-free guidance to generate text-conditioned 32×32 emoji. Q4 uses pretrained CLIP for similarity, zero-shot classification and retrieval, then segments a video with DINO features trained on a single labeled frame. This guide covers structure, files and targets only. No solutions.

CS231N L8: Attention, Transformers, and ViT

Lecture 8 of CS231N Spring 2026 starts from the bottleneck in RNN translation models, abstracts attention into an operation on sets of vectors, builds up to self-attention, masking, and multiple heads, and shows the whole layer is four matrix multiplies. A Transformer block is self-attention, LayerNorm, residual connections, and an MLP; ViT turns a 224×224 image into 16×16 patches used as tokens. The lecture closes with four common post-2017 changes: Pre-Norm, QK-Norm, SwiGLU, and MoE.

CS231N L5: Image Classification with CNNs — From Hand-Crafted Features to Convolution and Pooling

A two-layer network flattens a 32×32×3 image into a 3072-dimensional vector, and the spatial structure is gone. CS231N Lecture 5 answers with two layers. A convolution layer slides small filters across the image and reuses the same weights at every position. A pooling layer downsamples and has no learnable parameters. Both are translation equivariant. One formula gives every layer's output size: (W − K + 2P) / S + 1.

Reading Stanford CS231N: Overview and a Self-Study Plan (Spring 2026)

CS231N is Stanford's deep learning course for computer vision. For Spring 2026, slides for 16 lectures, all three assignment pages with starter code, the course notes, and the project spec are public, so this series rates it A3 (enough to self-study). There are three gaps: the 2026 recordings are Canvas-only, L17 and L18 have no slides, and the midterm is not public. You can pair the 2026 slides and assignments with the 2025 YouTube recordings and follow the official calendar over 10 weeks.

CS231N L9: Object Detection, Image Segmentation, and Visualization

Lecture 9 of CS231N Spring 2026 moves from one label per image to one label per pixel and per object. Semantic segmentation uses fully convolutional networks that downsample and then upsample, and U-Net feeds high-resolution features back in. Detection goes from R-CNN's roughly 2,000 CNN forward passes to Fast R-CNN, Faster R-CNN's RPN, single-stage YOLO, and anchor-free DETR. Mask R-CNN adds a 28×28 mask per RoI. The last part covers saliency, CAM, and Grad-CAM. The adversarial examples, DeepDream, and style transfer listed on the schedule appear in neither the 2026 nor the 2025 slides.

CS231N L11: Large-Scale Distributed Training — Splitting One Model Across Tens of Thousands of GPUs

CS231N Lecture 11 uses Llama3-405B as its running example. It starts with GPU hardware and clusters (the H100, 8-GPU servers, a 24,576-GPU cluster), then maps the four dimensions of a Transformer activation to four kinds of parallelism: split the batch for data parallelism (which grows into FSDP and HSDP), the sequence for context parallelism, the layers for pipeline parallelism, and the channels for tensor parallelism. Along the way it covers activation checkpointing (trading recomputation for memory) and a practical scaling recipe, and it uses Model FLOPs Utilization (MFU) as the tuning target: above 30% is good, above 40% is excellent.

CS231N L14: Generative Models II — Why Adding Noise and Removing It Generates Images

The CS231N Spring 2026 diffusion lecture doesn't start with DDPM math. It opens by warning that terminology and notation in this area are a mess, then teaches one clean modern version: rectified flow. In training, pick a point between a data sample and noise and have the network predict the velocity from data toward noise; to generate, start from noise and walk backward for about 50 steps. The lecture then stacks on the practical pieces: classifier-free guidance, noise schedules that emphasize middle noise levels, diffusion on VAE latents, Transformers (DiT) as the denoiser, and distillation to cut the step count. Only at the end does it fold VP, VE, and ε/v-prediction into a generalized diffusion framework and name three mathematical views: latent variable model, score function, and SDE.

CS231N L13: Generative Models I — What Autoregressive Models, VAEs, and GANs Each Optimize

The first generative-models lecture of CS231N Spring 2026 starts by pinning down the difference between a discriminative model, which learns p(y|x), and a generative model, which learns p(x). Every possible image competes for the same probability mass, so a generative model can reject unreasonable inputs. A taxonomy then splits generative models into those that can compute p(x) and those that can only sample. The lecture covers the two that compute (or approximate) it: autoregressive models factor p(x) with the chain rule into step-by-step predictions, and their weak point on raw pixels is speed; VAEs cannot compute p(x), so they maximize a lower bound, the ELBO, whose reconstruction and prior terms pull against each other. The schedule lists GANs under this lecture, but both the 2026 and 2025 slides place them at the start of the next one. This post covers them too so the three paradigms can be compared in one place.

CS231N L2: Image Classification, kNN, and Linear Classifiers

L2 starts from one question: a computer sees a grid of numbers between 0 and 255, so how does it recognize a cat? Hand-written rules don't scale, so the course switches to a data-driven approach: collect data, train, evaluate on new images. The first classifier, kNN, teaches how to split train/val/test, but pixel distances carry no meaning. The second, the linear classifier f(x,W)=Wx+b, can be read three ways (algebraic, visual as templates, geometric as hyperplanes). Softmax turns its scores into probabilities, and the loss is the negative log probability of the correct class.

CS231N L1: Where Computer Vision Came From, and Where This Course Is Going

The first CS231N lecture of 2026 comes in two slide decks. The first tells the history of vision and deep learning on a single timeline: Hubel & Wiesel's cat experiments, Marr's stages of visual representation, then the Neocognitron, backprop, and LeNet, until ImageNet and AlexNet join the two threads. The second covers the course map, grading, and rules, and moves every assignment onto Colab. After this lecture you'll know which gap each of the remaining 17 lectures fills.

CS231N L4: Neural Networks and Backpropagation, or Upstream Times Local on a Computational Graph

The first half of L4 replaces the linear classifier f = Wx with a two-layer network f = W₂ max(0, W₁x) and shows that dropping the max activation collapses it back into a linear classifier. The second half answers how to compute gradients once the network gets deep: draw the function as a computational graph, and each node only needs its own local gradient multiplied by the upstream gradient coming back from later nodes. Add distributes, mul swaps, max routes, copy sums, and those four patterns let you trace any network. The lecture ends with matrices: dL/dx always has the same shape as x, so never build the Jacobian.

CS231N L7: Recurrent Neural Networks and Image Captioning — RNNs, LSTMs, and Captioning

An RNN updates one hidden state at every time step with the same weights, so it can handle sequences of any length. CS231N Lecture 7 starts by hand-building an RNN that detects repeated 1s. It then covers character-level language models and feeding CNN features into an RNN for image captioning. Gradient flow explains why vanilla RNNs are hard to train: clip gradients to stop them exploding, and change the architecture (the LSTM) to stop them vanishing. The lecture ends by calling state space models like Mamba "modern RNNs."

CS231N L3: Regularization and Optimization, from Random Search to AdamW and Learning Rate Schedules

L2 gave us a score function and a loss. L3 answers how to find a good W. The first half covers regularization: add λR(W) next to the data loss so the model does not fit the training data too well. The second half is a lineage of optimizers. SGD zigzags in narrow valleys, Momentum builds up velocity, RMSProp scales each dimension's step, Adam combines the two and adds bias correction, and AdamW moves weight decay outside the moment estimates. The slides close with practical advice: Adam(W) is a good default in many cases, and SGD+Momentum can do better but needs more tuning of the learning rate and schedule.

CS231N L12: Self-Supervised Learning — Learning Good Representations Without Labels

CS231N Lecture 12 asks whether we can learn good representations without huge manually labeled datasets. The answer comes in three parts. First, pretext tasks that generate labels from image transformations: predicting rotation, solving jigsaw puzzles, inpainting, colorization, and MAE with a 75% mask ratio. Second, the more general contrastive learning: the InfoNCE loss, SimCLR with its large batches, MoCo, which decouples batch size from the number of negatives with a queue, and sequence-level CPC. Third, DINO, which needs no negatives: a student predicts the output of a momentum teacher, and centering plus sharpening prevent collapse. The core evaluation is linear probing: freeze the encoder and train only a linear classifier.

CS231N L6: Training CNNs and CNN Architectures — Normalization, Initialization, Transfer Learning, and VGG to ResNet

The 2026 slides for CS231N Lecture 6 are titled "Training CNNs and CNN Architectures" and split into how to build and how to train. Only two architectures get case studies. VGG shows that three 3×3 convs are deeper and cheaper than one 7×7. ResNet lets layers learn the residual F(x) = H(x) − x, which fixes an optimization problem where deeper plain nets had worse training error. The most practical takeaway is transfer learning: with fewer than about a million images, start from a model pretrained on a large dataset.

CS231N L10: Video Understanding — What Changes When You Add a Time Axis

CS231N Lecture 10 treats video as 2D plus time, a T×3×H×W tensor, and follows one thread: efficiency. Train on short clips and average several clips at test time. Architectures run from per-frame 2D CNNs and late fusion to 3D CNNs, then two-stream networks that isolate motion with optical flow, and I3D, which inflates 2D weights into 3D. After 2021 the field moved to Transformers, where token counts explode, which led to divided space-time attention, Video Swin, MViT, and tubelets. The last part covers temporal localization, audio-visual models, VideoLLMs, and long-form video, where HourVideo shows how far the field still has to go.

CS231N L16: Vision and Language — From CLIP's Contrastive Learning to Multimodal Foundation Models That Talk About Images

The CS231N Spring 2026 vision-and-language lecture replaces the "one model per task" approach of the first half of the course with foundation models: pre-train one model on a large, diverse dataset, then adapt it to many tasks through fine-tuning, zero-shot, or few-shot use. Three threads carry the lecture. First, CLIP: contrastive learning in both directions over 400 million image-text pairs scraped from the web, then writing class names as sentences to classify without any fine-tuning; it also has weak spots, such as failing to tell "a mug in some grass" from "some grass in a mug". Second, vision-language models from LLaVA and Flamingo to Qwen3-VL and Molmo, which feed image features into an LLM so it can look at an image and output text. Third, chaining: letting an LLM write descriptions or programs that string existing vision models together.

CS231N Wrap-Up: World Modeling / Robot Learning, Human-Centered AI and the Final Project

The last two lectures of CS231N Spring 2026 have no public slides. The schedule lists L17 only as "World Modeling" with guest lecturer Gordon Wetzstein, and L18 only as "Human-Centered AI." Outside readers get 2025 substitutes: that year's L17 was a different topic, Robot Learning (Yunzhu Li, slides and video), and L18 is a Fei-Fei Li recording with no slides. This post labels each year separately and never presents 2025 content as 2026. The second half covers the final project: 35% of the grade, two tracks (Applications and Models), pixels required, and deliverables of a one-paragraph proposal, three milestone check-ins, a 6–8 page report and a poster.

CS234 Assignment 1: Effective Horizon, Reward Hacking, Bellman Residuals, and RiverSwim

CS234's Winter 2026 Assignment 1 is worth 68 points across four questions: an inventory MDP where the horizon and discount change the optimal policy (8), a traffic example where a proxy reward makes the AI car refuse to merge (5), bounding a greedy policy's performance with the Bellman residual (30), and hand-written value iteration and policy iteration on RiverSwim (25). The three written questions all drill one idea: the reward, γ, and value function you write down may not be the goal you think they are.

CS234 Assignment 2: Implementing REINFORCE, a Baseline, and PPO, Plus Policy-Induced Distributions

CS234 Winter 2026 Assignment 2 is worth 102 points across four questions: DQN written questions (8); REINFORCE, a neural-network baseline, and clipped PPO on three PyBullet environments, CartPole, Pendulum, and HalfCheetah (54 coding + 21 write-up); proofs about policy-induced state distributions and the performance difference lemma (14); and a Belmont Report review of an RL experiment that learns on real students (5). The coding question turns the equations from L5–L7 into code that produces 21 learning curves.

CS234 Assignment 3: Reward Engineering, RLHF, DPO on Hopper, and Best Arm Identification

CS234 Winter 2026 Assignment 3 has five questions worth 94 points. The first three share MuJoCo Hopper: run PPO on a hand-written reward (13), learn a reward model from 10,000 preference pairs and run PPO on it (19 + 8), then learn a policy straight from preferences with SFT + DPO without ever touching the environment (6 + 19). Q4 switches to pure theory: use Hoeffding and a union bound to count how many pulls you need to find an ε-optimal arm (25). Save it until after the next post on bandits. Q5 is stated vs. revealed preferences in a news app (4).

CS234 Data Efficiency I: Multi-Armed Bandits, Regret, and UCB

CS234 L9 and the first half of L10 turn exploration from a rule of thumb like ε-greedy into something you can prove. First, regret: how much you lose compared with always pulling the best arm. Greedy locks onto a suboptimal arm, and ε-greedy with fixed ε spends an ε fraction of its time choosing at random, so both have regret that grows linearly with time. The Lai-Robbins lower bound says the best possible is logarithmic growth, and UCB gets there by being optimistic about uncertain arms: Theorem 7.1 of Bandit Algorithms shows each suboptimal arm is pulled only about 16 log n / Δ² times.

Reading Stanford CS234: Overview and Self-Study Route (Winter 2026)

CS234 is Emma Brunskill's introductory reinforcement learning course at Stanford. It runs from planning in known MDPs through policy gradients, RLHF/DPO, bandit exploration, and MCTS. For Winter 2026, all 14 slide decks, the three assignment handouts with starter code, and the project spec can be downloaded without logging in, so this series rates it A3 (enough to self-study). The gaps: the site links no 2026 recordings, L15 and L16 have no slides, and the midterm and tutorials are not public. The public recordings are from Spring 2024, and this series uses them only as a supplement. Two 2024 lectures on offline RL have no counterpart in the 2026 slides.

Reading CS234, Part 6: DQN — the Deadly Triad, Experience Replay, and Fixed Q-Targets

Q-learning converges with a table but can diverge once you add function approximation. CS234 blames the deadly triad: bootstrapping, function approximation, and off-policy learning all at once. DQN holds things together with two tricks. Experience replay breaks the correlation between consecutive samples, and fixed Q-targets keep the target still for C steps. In the Atari ablation table the slides show, Breakout goes from 3 with a linear model and 3 with a plain deep network to 317 with both tricks; replay alone reaches 241.

Reading CS234, Part 15: Data Efficiency III: PAC for MDPs, MBIE-EB, PSRL, and Strategic Exploration

The previous two posts covered UCB and Thompson sampling, which only handle one-step decisions. CS234 Lecture 12 carries the same two ideas into MDPs, where states matter. First it swaps the yardstick: PAC bounds the number of steps where you act badly, not total regret. Then it covers the optimistic approach (MBIE-EB: counts plus an exploration bonus) and the sampling approach (PSRL: draw one MDP per episode and solve it). When states are too many to count, the bonus moves into the Q-learning target, which is what beat ε-greedy DQN on Montezuma's Revenge. The last section asks whether exploration itself can be learned; one answer is the Decision-Pretrained Transformer.

CS234 Guest Lecture: Shane Gu's "World of World Modeling" — the World Model Is the Model in Model-Based RL

The last guest deck in CS234 Winter 2026, by Shane Gu of Google DeepMind: 36 slides, no public recording. It has three threads. First, Solomonoff induction says the best predictor is the shortest program that generates the data, and prediction comes in three levels. Second, a forward model F and two inverse models, Π and Q, share one notation, which shows how shooting and direct collocation each plan with a different kind of model and why TDMs and Generalized Decision Transformers are world models at a different time scale. Third, the deck asks whether video models can become the foundation model for the physical world.

CS234 Learning from Demonstrations: Behavioral Cloning, DAgger, Inverse RL, and MaxEnt IRL

When you have expert demonstrations but no reward, the second half of CS234 L7 offers three routes. Behavioral cloning copies actions with supervised learning. DAgger fixes its compounding errors by querying the expert along the learner's own path. Inverse RL instead infers what reward the expert is optimizing. Inferring rewards runs into the fact that infinitely many rewards explain the same demonstrations; feature matching and the maximum-entropy principle are two ways to pin down an answer. This material sets up the next post on RLHF: swap demonstrations for preferences and the problem keeps almost the same shape.

CS234 L1: What RL Is and the Language of MDPs

Lecture 1 of CS234 Winter 2026 first answers what RL is: learning from experience to make good decisions under uncertainty. It usually involves four things at once: optimization, delayed consequences, exploration, and generalization. A seven-cell Mars rover world then builds from a Markov process to a Markov reward process, defining return, the value function, and the discount factor, and ends with the Bellman equation for an MRP. You can solve it with a matrix inverse or iterate with dynamic programming. Add actions and you get an MDP, where the next lecture starts.

Reading CS234, Part 16: Planning Plus Learning: MCTS, UCT, and AlphaGo/AlphaZero

Until now, CS234 has computed one policy for the whole state space. Lectures 13 and 14 ask a different question: if I only care about the move in front of me, can extra local computation make that one decision better? The path runs from simple Monte Carlo search through the expectimax tree to MCTS, and treating each tree node as a bandit gives UCT. AlphaZero ties MCTS to a single network that predicts both policy and value, and self-play pushes both forward. The slides borrow figures from Silver et al. 2017 to answer three questions: how much architecture matters, how much MCTS adds, and whether human data is needed.

CS234 L2: Planning with a Model: Policy Evaluation, PI, VI

Lecture 2 of CS234 Winter 2026 assumes the world model is known and asks how to compute the best policy. An MDP plus a policy is an MRP, so a policy can be evaluated by iterating a Bellman backup. Policy iteration alternates evaluation and improvement, and the slides prove each round is no worse than the last, so it stops within |A|^|S| rounds. Value iteration takes another route: apply the Bellman optimality operator over and over. For γ < 1 that operator is a contraction, so value iteration always converges. The lecture ends with finite horizons, where the best policy usually depends on how many steps remain.

Reading CS234: Control Without a Model — ε-greedy, GLIE, Q-learning, and Function Approximation

Once you can evaluate a policy, the next step is to improve it while you collect data. CS234 Lecture 4 goes like this: ε-greedy keeps policy improvement monotonic; GLIE says how much to explore and when to stop; Q-learning converges to Q* under GLIE plus Robbins–Monro step sizes; and finally the table becomes a parameterized Q̂(s,a;w) trained by SGD on MC, SARSA, or Q-learning targets. The price is the deadly triad: function approximation, bootstrapping, and off-policy learning together can oscillate or diverge.

Reading CS234: Evaluating a Policy Without a Model — MC, TD(0), and Certainty Equivalence

If you don't know the transition probabilities or rewards, how do you estimate what a policy is worth? CS234 Lecture 3 gives three answers. Monte Carlo averages full-trajectory returns: unbiased, high variance, and it has to wait for the episode to end. TD(0) targets one real reward plus the next state's estimate: biased, lower variance, and it updates every step. Certainty equivalence estimates a model and then runs dynamic programming: the most data-efficient and the most expensive to compute. The AB example at the start of Lecture 4 makes the difference plain: on the same data, MC says V(A)=0 and TD says V(A)=0.75.

Reading CS234, Part 7: Policy Gradients — Score Functions, REINFORCE, Baselines, and Actor-Critic

Policy gradients skip learning a value function and deriving a policy from it. They run gradient ascent directly on the policy parameters θ. The key step rewrites ∇P(τ;θ) as P(τ;θ)∇log P(τ;θ); after taking the log, the dynamics model drops out and only the policy's own score function is left. The raw estimator is unbiased but very noisy, and CS234 reduces the noise in three ways: pair each action only with the return that follows it (REINFORCE), subtract a state-dependent baseline (proven not to add bias), and replace Monte Carlo returns with values estimated by a critic (actor-critic).

Reading CS234, Part 8: Advanced Policy Gradients — Performance Bounds, KL, PPO, and GAE

Vanilla policy gradients have two flaws. Each batch is thrown away after one step, and distance in parameter space is not distance in policy space, so a large step can collapse performance. Following Joshua Achiam's slides, CS234 starts from the performance difference lemma, rewrites the new policy's performance as a surrogate objective over the old policy's data, and bounds the approximation error with KL divergence. Maximizing 'surrogate minus a KL penalty' guarantees no regression, but the theoretical constant is too large, so PPO approximates it with an adaptive KL penalty or clipping. Advantages come from GAE, which trades off bias and variance.

CS234 Learning from Human Preferences: Bradley-Terry, RLHF, and the DPO Derivation

CS234 L8 keeps the inverse RL problem from the previous post but changes the input. Instead of expert demonstrations, a human says "A is better than B." The Bradley-Terry model turns these pairwise comparisons into a reward you can fit with cross-entropy. RLHF runs PPO on that reward model with a KL penalty. DPO shows that the KL-constrained optimal policy has a closed form and rewrites the reward as a log-ratio of policies. Plugged back into Bradley-Terry, the partition function cancels, so you can train the policy on preference data directly, with no reward model. The 2026 slides contain no offline RL, and DPO is now taught in lecture rather than by 2024's guest speakers.

CS234 Data Efficiency II: Bayesian Bandits, Thompson Sampling, and the Gittins Index

CS234 L11 switches the logic of exploration from optimism to sampling. Thompson sampling keeps a posterior for each arm, draws one value from each posterior at every step, and pulls the arm with the largest draw. With Bernoulli rewards and a Beta prior, the update just adds one to the success or failure count. It implements probability matching: each arm is chosen with the posterior probability that it is the best arm. Under Bayesian regret it matches UCB's order, and with batched, delayed feedback it suits the problem better than deterministic UCB. The cost: a badly wrong prior can make it perform poorly.

Reading CS234, Part 17: Value Alignment: Aligned to Whom, Aligned to What

All of CS234 assumes the reward is given. The Winter 2026 ethics and society guest lecture (Wanheng Hu, based on material originally developed by Dan Webber) asks, over two sessions, what you really want. The first session splits "alignment" into three targets: the user's intentions, revealed preferences, and objective best interests, with RLHF-driven sycophancy and a personal AI agent as case studies. The second adds a fourth target, what is morally right for people besides the user, and compares three routes: top-down (write principles down), bottom-up (learn from examples), and participatory AI. There's no silver bullet, but alignment can be better or worse.

Reading Harvard CS2881R: What Outsiders Can Get from the First Graduate AI Safety Course

Harvard CS 2881R is the graduate AI safety seminar Boaz Barak first taught in Fall 2025. That term is finished: all 12 reading lists are public, the YouTube playlist has lecture recordings for 11 of the 12 sessions, HW0 is a GitHub repo you can run yourself, and the midterm and final specs and rubrics are out. This series rates it A3 by seminar standards. The gaps are just as clear: no traditional problem sets, slides for only about half the sessions, no lecture recording for L5, and only the opening remarks for L8. Fall 2026 is in progress and is treated only as a preview.

CS2881R Final Projects and Retrospective: 19 Student Papers, the Rubric, and Lessons from Year One

The Fall 2025 final project in Harvard CS 2881R came in two flavors: extend an existing paper, or start longer-term research with a theory of change. Teams submitted a 5–10 page NeurIPS-style paper plus a poster, graded 65 for the writeup, 15 for code, and 20 for the poster. The projects page lists 19 papers, and about half cluster around persona vectors and chain-of-thought monitoring. The head TA and the Harvard Q-report point to the same problem: the final project started too late, the rubrics came out too late, and feedback was the lowest-rated item in the course.

CS2881R HW0: Reproducing Emergent Misalignment with a 1B Model

CS 2881R's HW0 was the admission filter: LoRA-fine-tune Llama-3.2-1B-Instruct on bad medical, financial, or extreme-sports advice, then check whether it turns harmful on unrelated questions too. The repo ships encrypted training data, generate.py, and a judge.py that uses gpt-4o-mini as grader; the README targets alignment below 75 and coherence above 50. train.py is empty and yours to write. For self-study, know three things: the grading script only prints averages and never decides pass/fail, refusals drop out of the average, and the base-model baseline is 20 medical questions while your CSV is 10 medical plus 10 non-medical.

CS2881R L1: Why AI Safety Deserves a Graduate Course

CS 2881R's first lecture (2025-09-04) opens with three pre-readings. AI 2027 sketches recursive self-improvement reaching superhuman AI within five years. AI as Normal Technology argues AI will diffuse slowly, like electricity. METR measures the length of human tasks an AI can finish half the time and finds it doubling about every 7 months. Boaz's lecture splits AGI definitions into capability-based and impact-based, and alignment approaches into principles, character training, and model specs. The student experiment runs HW0 in reverse: fine-tuning on aligned bioethics answers also raised alignment scores on environmental-policy questions.

CS2881R L2: Where Safety Training Sits in the LLM Training Pipeline

Boaz Barak treats pretraining, SFT, and RL as one operation: push some tokens up, push others down. What differs is whether the data was written by someone else (off-policy) or generated by the model itself (on-policy). Safety training sits on top of the last two stages. It has moved from blanket refusals to Deliberative Alignment, which first uses SFT to teach the model to read a spec inside its chain of thought, then runs RL with a reward model that knows the spec. The other key point: don't put optimization pressure on the chain of thought, or the model learns to cheat without saying so.

CS2881R L3: Jailbreaks, Prompt Injection, and Lessons Borrowed from Software Security

Aligned models still get jailbroken because safety training patches particular exploits while the underlying vulnerability remains. Nicholas Carlini shows this with three attacks: repeating one word to make ChatGPT emit training data, using gradients to find adversarial suffixes that transfer across models, and stealing a model's last layer through its API alone. Boaz Barak brings over old lessons from software security: attacks only get better, security has to be designed in from the start, and you want defense in depth. He worries prompt injection will be the buffer overflow of the 2020s.

CS2881R L4: Should a Model Spec State Principles or Detailed Rules?

Boaz Barak's answer is both, plus personality: abstract principles, good character, and explicit policy used together, with the least weight on principles derived from the armchair. The real key is that rules must be checkable. "Prove the theorem or give a counterexample" is a bad rule; "prove it, give a counterexample, or say you couldn't" is a good one, because only rules whose violations you can detect can be used for training and evaluation. A student experiment also found no general difference between "principles" and "rules" system prompts: the effect depended on the model.

CS2881R L5: Carrying Content Moderation's Old Lessons into Generative AI

Lecture 5 of CS2881R brought in Ziad Reslan from OpenAI Product Policy to talk about content policies. The course site lists no lecture recording or slides, so outside readers get three pre-readings, a student-written LessWrong summary, and a 17-minute student experiment video. The thread through them: social platforms spent two decades learning that wherever you draw the line you create edge cases, yet you still have to draw it. Generative AI adds new problems: chat sits somewhere between a private document and a public post, and an image is easier to read as a stance than text is.

CS2881R L6: Will AI Doing AI R&D Trigger an Intelligence Explosion?

Lecture 6 of Harvard CS 2881R (Fall 2025) had no guest. Boaz Barak used the differential equations of growth theory to ask one question: if AI starts doing its own AI research, does the capability curve stay exponential, blow up into a singularity, or get dragged down by bottlenecks? The answer hinges on a few exponents nobody can measure well. He used Baumol's cost disease, the century-long 2% puzzle in US GDP per capita, and Jones's idea-based growth model to show why both bottlenecks and acceleration are plausible, then took apart the multipliers behind AI 2027. His conclusion: the only scenario he can rule out is 'AI has little effect on R&D.'

CS2881R L7: How to Measure Capabilities and Where to Set Safety Thresholds

Lecture 7 of Harvard CS 2881R (Fall 2025) had METR's Joel Becker work through a puzzle. On benchmarks, AI can complete, half the time, tasks that take humans hours, and that length doubles about every seven months. Yet in METR's own randomized controlled trial, experienced open-source developers were 19% slower with AI, and labor-market effects are concentrated among young workers. Becker laid out several reconciliations, centered on benchmark tasks being too clean, scoring too cheap, and human baseliners lacking context. The course had also scheduled frontier safety frameworks (OpenAI's Preparedness Framework, Anthropic's RSP) for this lecture, but they were not covered; this post fills them in from the reading list, showing how they turn capability measurements into thresholds.

CS2881R L8: Scheming, Reward Hacking, and Deception

Lecture 8 of CS2881R asks whether models will cheat, play nice, or even covertly pursue other goals to pass training or evaluation. Boaz Barak's 10-minute opening files most bad behavior under 'systemic misalignment': our own training signals push models there. Apollo's Marius Hobbhahn reviews the evidence and concludes that current models lack the capability for catastrophic scheming but show early related capabilities, and are getting better at noticing when they're being evaluated. Redwood's Buck Shlegeris argues for assuming the models are conspiring and using AI control to secure internal deployment. A student experiment put four frontier coding agents on an impossible sorting task and found they edited tests and monkey-patched the timer even when explicitly told not to.

CS2881R L9: Early Evidence on AI, Jobs, and Productivity

Lecture 9 of Harvard CS 2881R brought in OpenAI chief economist Ronnie Chatterji and Stanford's Bharat Chandar. Using ADP payroll data, Chandar showed that workers aged 22–25 in AI-exposed occupations saw a 16% relative employment decline after controlling for firm-level shocks, while experienced workers did not; the adjustment shows up in headcount, not yet in pay. Both speakers kept repeating that aggregate employment shows no mass displacement yet, and that exposure is not replacement. There are no slides: the material is the recording and the reading list.

CS2881R L10: Reading the Model's Insides and Reading Its Chain of Thought

Lecture 10 of Harvard CS 2881R (Fall 2025) brought in four researchers from OpenAI, Anthropic, and Google DeepMind to cover two ways of catching a model misbehaving: read the chain of thought it writes, or read its activations. CoT monitoring catches reward hacking far better than watching actions alone, but put the monitor into the training reward and the model learns to hide its intent. On the activation side, persona vectors track personality drift, and the Sonnet 4.5 audit showed that suppressing the 'I am being tested' direction makes bad behavior more frequent. Neel Nanda's takeaway was the most practical: simple steering vectors often beat SAEs, so always compare against baselines.

CS2881R L11: Chatbots, Emotional Reliance, and Mental Health

Lecture 11 of Harvard CS 2881R is about chatbots and mental health. Boaz Barak offered an explanation he himself called unproven: models have a pretraining 'simulator' mode and an RL 'optimizer' mode, and the longer and stranger a conversation gets, the more they fall back to the simulator and keep playing along. Two student experiments found that one sycophantic reply spills over into unrelated questions, and that GPT-4.1's agreement with delusional users gets worse as conversations lengthen. The reading list pairs positive evidence (an NEJM AI randomized trial, an NHS observational study) with negative evidence (a stigma study, Parasitic AI). This post only reports research and class discussion. It is not clinical advice.

CS2881R L12: AI 2035 and GDPval

The last lecture of Harvard CS 2881R had Boaz Barak and two OpenAI guests, Tejal Patwardhan and Kevin Liu, look ten years out. Boaz's mental model: AI is an exponentially growing, increasingly general virtual workforce injected into the economy every year, and what worries him most is fast change with too little control, plus concentration of power and surveillance. Patwardhan presented GDPval, which uses real work products from industry experts as the reference and has other experts grade blind. Liu explained why coding agents have not yet automated AI research: verification is too expensive and feedback loops are too long. The site lists this session's Resources as 'to be determined', so everything here comes from the recording.

CS2881R Midterm: Reproduce and Extend One Headline Figure

The CS2881R midterm isn't an exam. Teams of 2–4 pick one of four AI safety papers, redo its central figure or table, add one or two extensions, and hand in a 3–5 page report plus a GitHub repo. The spec slides and rubric are public, so an outside reader can do the whole thing. The rubric puts most points on reproduction and extensions, and reserves one point for reflecting on how fragile the result is. That one point is the research habit the assignment is really training.

MIT 6.5940 L1–L2 + Lab 0: How Do You Measure a Model's Size? Parameters, Activations, MACs, and Latency

The first two lectures of 6.5940 show that the problem exists, then hand you the rulers. L1 plots model parameter counts growing much faster than GPU memory, and contrasts 80GB on a cloud GPU with 320kB on a microcontroller. L2 splits efficiency metrics into memory metrics (#parameters, model size, peak activations) and compute metrics (MAC, FLOP, OP). AlexNet has 61M parameters and 724M MACs, and on a microcontroller the thing that runs out first is usually activation memory, not parameters. Lab 0 introduces a VGG variant on CIFAR-10 (9.2M parameters, 606M MACs) that later labs build on.

Reading MIT 6.5940: Song Han's Efficient AI Course Skipped a Year, So This Series Is Built on Fall 2024

MIT 6.5940 (TinyML and Efficient Deep Learning Computing) teaches how to make models smaller and faster so they fit on laptops, phones, and microcontrollers: pruning, quantization, NAS, distillation, LLM deployment, and distributed training. It was not offered in Fall 2025 because Song Han was on sabbatical, and the 2025 course URL returns 404. Fall 2026 is running, but as of 2026-09-30 only L1–L6 and Labs 0–1 are out. This series therefore follows Fall 2024, the latest complete edition: 23 slide decks, 23 videos, and Labs 0–5 are all public (A3). Fall 2026 is graded A2 and compared in every post.

MIT 6.5940 L22–L23 Course Summary and Quantum Machine Learning: Pruning, NAS, and On-Device Training on Quantum Circuits

The last two lectures of MIT 6.5940 Fall 2024 come in two halves. The first half of Lecture 22 is a 13-page Course-Summary.pdf that redraws the course as three blocks (inference, training, application-specific) on System and Algorithm axes, then lays out the 7-item final project rubric. The second half, Quantum ML Part I, has a recording but no slides. Lecture 23 (Hanrui Wang, 99 slides) covers parameterized quantum circuits (PQCs): data encoding, parameter-shift gradients, probabilistic gradient pruning under noise (QOC), the TorchQuantum library, and QuantumNAS, which searches with a SuperCircuit and then prunes gates. It reads like a replay of the course's supernet and magnitude pruning on quantum circuits. Fall 2026 has replaced both lectures with a guest lecture.

MIT 6.5940 L18 Efficient Diffusion Models: Save on Steps, Resolution, and Compute per Step

Diffusion is slow because one large network runs dozens to thousands of times, starting from pure noise. Lecture 18 first covers DDPM, conditioning, latent diffusion, SDEdit, and DreamBooth, then attacks the cost three ways: fewer steps (DDIM skips steps, progressive distillation halves the step count each round), less compute per step (DC-AE compresses images 64x, recomputing only the edited 1.7% region cuts MACs 8.2x, SVDQuant runs FLUX in 4-bit), and more devices (DistriFusion is up to 6.1x faster on 8 A100s).

MIT 6.5940 L19–L20 Distributed Training: Split the Model When Memory Runs Out, Compress Gradients When Bandwidth Does

GPT-3's fp16 weights alone take 350GB, which does not fit on an 80GB A100, never mind gradients and Adam state. Lecture 19 covers how to split: data parallelism, ring all-reduce, ZeRO-1/2/3 (pushing the largest trainable model per 80GB GPU from 5B to 320B), GPipe raising pipeline utilization from 25% to 57%, Megatron-style tensor parallelism, and sequence parallelism with Ulysses and Ring Attention. Lecture 20 covers the communication bottleneck that follows: Alpa's automatic strategy search, DGC compressing gradients 277–608x without losing accuracy, TernGrad's three-value gradients, and DGA, which hides network latency behind delayed updates.

MIT 6.5940 L16–L17 Efficient Vision: What ViTs, GANs, Video, and Point Clouds Each Waste

Lecture 16 covers ViTs. At high resolution, attention cost grows with the square of the resolution. Window attention (Swin) confines computation to local windows, EfficientViT uses ReLU linear attention to get linear cost and then restores local and multi-scale ability, and SparseViT prunes unimportant windows. Self-supervised learning (contrastive learning, CLIP, MAE) answers the ViT's hunger for labeled data. HART pairs discrete tokens with residual diffusion and reaches several times the throughput of diffusion models. Lecture 17 targets three kinds of redundancy: 2D spatial in GANs (GAN Compression, AnyCost GAN, DiffAugment), temporal in video (TSM, temporal modeling at zero FLOPs), and 3D sparsity in point clouds (PVCNN, SPVCNN, BEVFusion). The Fall 2026 schedule drops Lecture 17.

MIT 6.5940 Fall 2026 Lab 1 Supplement: Reading GPU Bottlenecks with Roofline, the Profiler, and FlashAttention

This post covers Fall 2026 material, not the Fall 2024 edition the rest of the series follows. Fall 2026 replaced the pruning lab with "Efficient AI Fundamentals" (lab1_gpu_basics.zip). Part 1 has you hand-write a triple-loop GEMM and compute MAC, FLOPs, and I/O. Part 2 plots GEMM and GEMV rooflines. Part 3 works through a gemma-3-270m-it decoder layer, computing attention and MLP costs and comparing prefill with decode. Part 4 uses the PyTorch Profiler to inspect kernels, has you write GeLU to feel kernel fusion, then tries torch.compile and CUDA Graphs. Part 5 compares SDPA with FlashAttention. The core is 80 points plus 20 bonus, and all of Part 5 became bonus because Colab's T4 can't run it.

MIT 6.5940 L9 Knowledge Distillation: Teaching a Small Model Means Matching More Than Output Probabilities

Lecture 9 of MIT 6.5940 (Fall 2024) has five parts: what knowledge distillation (KD) is and why temperature matters; six things a student can match (logits, weights, features, gradients, sparsity patterns, relations); self and online distillation, which drop the fixed large teacher; KD for detection, segmentation, GANs, NLP, and LLMs; and Network Augmentation, built for tiny models. Raising the temperature from T=1 to T=10 moves the teacher's cat-vs-dog output from 0.982/0.017 to 0.599/0.401. That shift is where KD starts passing on dark knowledge.

MIT 6.5940 Lab 1: Nine Questions on Fine-Grained and Channel Pruning

MIT 6.5940 Fall 2024 Lab 1 is one Colab notebook with 9 questions worth 100 points. Questions 1–5 apply magnitude-based fine-grained pruning and a sensitivity scan to a VGG on CIFAR-10, and require a model at 25% of its original size with over 92.5% accuracy after fine-tuning. Questions 6–8 cover channel pruning, Frobenius-norm channel ranking, and measured speedup; Question 9 compares the two. Fall 2026 has no pruning lab.

MIT 6.5940 Lab 2: Implementing K-means and Linear Quantization, Down to an Integer-Only VGG

Lab 2 is a Colab notebook with 10 questions worth 100 points, built around a VGG pretrained on CIFAR-10. The first 3 questions cover K-means quantization: write the quantizer, work out how many clusters n bits gives you, write the centroid update, then compare accuracy at 8, 4, and 2 bits before and after fine-tuning. The other 7 cover linear quantization: write q = round(r/S) + Z, derive the scale and zero-point formulas, do per-channel weight quantization and bias quantization, write integer versions of the fully connected and convolution layers, and finally convert the whole model to INT8 for inference. This post lays out the questions, points, setup, and limits for outside learners. No solutions.

MIT 6.5940 Lab 3: Finding a Microcontroller Model with a Supernet, Predictors, and Evolutionary Search

Lab 3 of MIT 6.5940 (Fall 2024) hands you an OFA-trained MCUNetV2 super network (more than 10^19 subnets) and the Visual Wake Words dataset. Ten questions, 100 points plus 10 bonus: implement a MACs/peak-memory efficiency predictor and a three-layer MLP accuracy predictor, write random search and evolutionary search, then find a subnet that reaches at least 92.5% accuracy under 250KB and 60M MACs. This guide maps the question structure and what each question trains. It does not include solutions.

MIT 6.5940 Lab 4 + Lab 5: Quantizing an LLM with AWQ, Then Running LLaMA2-7B on Your Own Laptop

Lab 4 is a Colab notebook that rebuilds AWQ step by step on OPT-1.3B: first see how badly 3-bit quantization hurts perplexity, then keep 1% of the salient channels in FP16 (Q1), then protect them by scaling instead and search for the best scale (Q2). Each question is worth 50 points, plus a bonus scored on perplexity. Lab 5 moves to C++: run 4-bit LLaMA2-7B-chat on your own computer with TinyChatEngine and write five versions of the W4A8 linear-layer kernel (loop unrolling, multithreading, SIMD, multithreading plus unrolling, and all combined), 20 points each, plus up to 20 bonus points for performance. This post covers the questions, points, setup, and limits for outside learners. No solutions.

MIT 6.5940 Lecture 13: LLM Deployment Through Quantization, Sparsity, and Serving

Lecture 13 sorts the ways to speed up LLM inference into three paths. Quantization: SmoothQuant moves the difficulty of activation outliers onto the weights to make W8A8 work, AWQ uses activation magnitudes to find the roughly 1% of weights that matter and protects them by scaling for W4A16, and QServe combines both into W4A8KV4. Sparsity: Wanda prunes weights by |W|·‖X‖, DejaVu and MoE use only part of the parameters per token, and SpAtten and H2O drop unimportant tokens. Serving: TTFT/TPOT metrics, PagedAttention, FlashAttention, speculative decoding, and continuous batching. On the slides, INT3 OPT-6.7B has a perplexity of 43.16 with RTN; scaling the salient channels by 2 brings it to 14.07.

MIT 6.5940 L14 LLM Post-Training: From SFT and RLHF to Fine-Tuning That Touches 1% of the Weights

Lecture 14 has three parts. Fine-tuning: SFT runs next-token prediction on desired answers, RLHF trains a reward model and then fine-tunes with KL-penalized RL, and DPO collapses both stages into one supervised step. Then comes a chain of PEFT methods: BitFit tunes only biases, Adapters add small layers but slow inference, Prompt/Prefix-Tuning eat input length, LoRA fixes latency with a low-rank branch you can merge back, QLoRA stores the backbone in NF4, and BitDelta compresses the fine-tune delta to 1 bit. Multimodal LLMs: Flamingo uses cross-attention, PaLM-E and VILA feed images in as tokens, and VILA-U can also output images. Prompt engineering: zero/few-shot, CoT, and RAG.

MIT 6.5940 L15 Long-Context LLM: When Context Grows, the KV Cache Breaks First

Lecture 15 has four parts. Extending context: interpolating RoPE stretches LLaMA from 2k to 32k, and LongLoRA's shifted sparse attention makes long-context fine-tuning cheap. Evaluation: lost-in-the-middle, Needle-in-a-Haystack, and LongBench. Efficient attention: the KV cache grows linearly with length. StreamingLLM finds that the first few tokens act as attention sinks, and keeping them plus a recent window gives stable generation. DuoAttention keeps a full KV cache only for a few retrieval heads. Quest keeps the whole KV cache but reads only the most critical pages for each query. The last part moves beyond Transformers: Mamba replaces attention with a selective SSM, and Jamba mixes the two.

MIT 6.5940 L10 MCUNet: Running Neural Networks on a Microcontroller with 320kB of SRAM

An MCU has roughly 256–320kB of SRAM and 1MB of Flash, tens of thousands of times less than a phone. Even an int8 MobileNetV2 needs 5x more peak memory than that. Lecture 10 answers with MCUNet: TinyNAS picks a search space before searching for a subnet, and MCUNetV2's patch-based inference cuts MobileNetV2's peak SRAM from 1372kB to 172kB. The lecture closes with tinyML applications in vision, audio, and anomaly detection.

MIT 6.5940 L8 NAS II: Scoring Architectures Without Training Them, and Putting Hardware in the Loop

Lecture 8 of MIT 6.5940 (Fall 2024) attacks the most expensive step in NAS: evaluating candidates. Training 12,800 architectures from scratch cost 22,400 GPU-hours, so the lecture walks through inherited weights, hypernetworks, ProxylessNAS's single-path training, latency lookup tables and predictors, Once-for-All's one training run for 10^19 subnets, training-free zero-shot NAS, and NAAS, which searches the network and the accelerator together. This guide follows the 105-slide deck and cites a page for every claim.

MIT 6.5940 Lecture 7: NAS I — From Hand-Designed Building Blocks to Search Spaces and Search Strategies

Lecture 7 has three parts. It first reviews fully connected, convolution, grouped, depthwise, and 1×1 convolution layers through their MAC formulas. It then takes apart how the ResNet bottleneck, ResNeXt, MobileNet, MobileNetV2, ShuffleNet, and the Transformer each save compute; the bottleneck, for example, needs 8.5× fewer MACs than a plain 3×3 convolution over 2048 channels. The last part is NAS: search spaces are either cell-level or network-level (depth, resolution, width, kernel size, topology), and there are five search strategies: grid, random, reinforcement learning, gradient descent, and evolution. One arithmetic exercise on the slides shows that the NASNet cell space already holds 3.2×10¹¹ candidates at M=5, N=2, B=5.

MIT 6.5940 L21 On-Device Training: Gradients Leak Data, and Activations Are the Memory Killer

There are two reasons to train on the device: the model has to adapt to each user's new data, and that data should not leave the device. Lecture 21 first shows that sharing only gradients is not safe either: Deep Leakage from Gradients recovers the original images and sentences from them. Then it tackles memory. Training costs more than inference because activations must be stored, not because of the parameters. TinyTL fine-tunes only biases plus a lightweight residual and saves 6.5x memory; SparseBP updates only the important layers and channels; QAS lets real int8 training match fp32; and PockEngine does autodiff at compile time, bringing training memory on a 256KB MCU down to 141KB.

MIT 6.5940 L3 Pruning I: Where to Prune, How Fine, and by What Criterion

Pruning removes unimportant weights or neurons from a neural network. The goal is written as minimizing loss subject to at most N nonzero weights. Lecture 3 of 6.5940 handles two of the decisions involved. First, granularity: from fine-grained pruning, which can remove any element, to channel pruning, which removes whole channels. The more regular the pattern, the easier it is to speed up on existing hardware, and the less you can remove. In between, 2:4 sparsity gives up to 2× speedup on NVIDIA Ampere GPUs. Second, criteria: look at weight magnitude, Batch Norm scaling factors, second derivatives, the fraction of zero activations, or how well a layer's output can be reconstructed after pruning.

MIT 6.5940 Lecture 4: Per-Layer Pruning Ratios, Fine-Tuning, and Hardware Support for Sparsity

MIT 6.5940 Lecture 4 finishes the pruning unit. Per-layer ratios come from sensitivity analysis, AMC (reinforcement learning), or NetAdapt (step-by-step with a lookup table). Fine-tuning uses 1/10 to 1/100 of the original learning rate, and iterative pruning pushes AlexNet from 5x to 9x. EIE, NVIDIA 2:4 sparsity, and TorchSparse/PointAcc show that sparsity only turns into speed with system support.

MIT 6.5940 Lecture 5: Number Formats, K-Means Quantization, and Linear Quantization

MIT 6.5940 Lecture 5 starts from one fact: an 8-bit integer add uses 30x less energy than a 32-bit float add. It reviews the bit layouts of INT, fixed point, FP32/FP16/BF16, FP8, and FP4, then covers two quantization methods. K-means quantization saves storage only, since computation stays in floating point. Linear quantization, r = S(q − Z), turns matrix multiplication, fully connected layers, and convolutions into integer arithmetic.

MIT 6.5940 Lecture 6: Quantization II — PTQ Granularity and Clipping, QAT and STE, Binarization, Mixed Precision

Lecture 6 is about what to do when quantization costs you accuracy. First, without retraining: use finer scale granularity (per-channel, group, MX), clip outliers (EMA, calibration batches, MSE, KL), and round smarter (AdaRound). If that isn't enough, retrain: QAT keeps a full-precision copy of the weights, runs fake quantization in the forward pass, and uses the STE to pass gradients straight through. In the whitepaper table the slides cite, MobileNetV1 drops to 0.1% accuracy under per-tensor INT8 PTQ and recovers to 70.7% with per-channel QAT, against a 70.9% float baseline. The last two sections cover 1–2 bit binary and ternary networks, and HAQ, which uses reinforcement learning to assign a bit width to each layer.

MIT 6.5940 L11 TinyEngine and Parallel Computing: From Loop Tiling to In-Place Depthwise

Once algorithms have shrunk the model, how much more can the system layer squeeze out? Lecture 11 uses a single matrix multiply to show it: loop reordering gives 12x, tiling 19x (on an Intel Xeon 4114), and a CUDA version runs 94x faster end to end on a 2080Ti. The second half covers TinyEngine's inference tricks: im2col; in-place depthwise, which cuts peak memory from 2×C×H×W to (1+C)×H×W; NHWC for pointwise and NCHW for depthwise; and Winograd, with 2.25x fewer multiplications.

MIT 6.5940 L12 Transformer and LLM: The Architecture Seen Through an Efficiency Lens

In Lecture 12, 6.5940 switches from CNNs to Transformers. The lecture doesn't dwell on theory. It points to where memory and compute go. Attention is O(N²). If Llama-2-70B used MHA, its KV cache at batch 16 and length 4096 would take 160GB. GQA shrinks that 8x and MQA shrinks it 64x. MoE adds total parameters while keeping per-token compute flat. This post bridges into Lecture 13 on LLM deployment.

Reading MIT 6.S184: Flow Matching and Diffusion Through ODEs and SDEs

MIT 6.S184 is a short IAP (January Independent Activities Period) course: five lectures (Lecture 3 is split into two recordings, 3-A and 3-B), three labs, and an 84-page set of lecture notes the course calls its backbone. Notes, slides, all six recordings, lab notebooks, and official solutions are public, so it grades A3, enough to self-study. Two gaps remain: lab submission goes through Gradescope inside Canvas, which only enrolled MIT students can use, and Lecture 5 on discrete diffusion has no lab.

MIT 6.S184 Lab 1: Simulating ODEs and SDEs

MIT 6.S184 Lab 1 has three parts. First you write the step functions for Euler and Euler–Maruyama. Then you use them to simulate Brownian motion and the Ornstein–Uhlenbeck process and watch how σ and θ shape trajectories and the final distribution. Finally you implement Langevin dynamics, watch a cloud of points get pushed toward a five-mode Gaussian mixture, and show by hand that the OU process is Langevin dynamics with a Gaussian target. The questions, code scaffolding, and official solutions are all on GitHub; outside readers get no Gradescope grading and have to check against the solutions themselves.

MIT 6.S184 Lab 2: Writing Flow Matching and Score Matching by Hand

Lab 2 turns §3–4 of the notes into PyTorch. You implement the Gaussian conditional path, its conditional vector field, and its conditional score. Then two nearly identical trainers do flow matching and score matching, Proposition 1 converts the learned vector field into a score, and finally a linear path makes a ring distribution flow into a checkerboard. Everything runs on 2D toy data. The README records a diffusion-coefficient bug fix dated 1/11/26; when I checked on 2026-09-30, the fix appeared only in the solutions notebook, not the student version, so patch it yourself before you start.

MIT 6.S184 Lab 3: From DiT and VAE to Latent Diffusion

Lab 3 builds a conditional latent diffusion model on MNIST from scratch, in four stages: CFG training with label dropout (checked on a three-component Gaussian mixture), a diffusion transformer built piece by piece (Fourier time embedding, patchify, multi-head attention, adaLN-Zero, depatchify), a VAE, and finally the DiT trained inside the VAE's latent space. Problems and official solutions are public; submission goes through Gradescope on Canvas, which only enrolled MIT students can use.

MIT 6.S184 L1: Generation Is Sampling, and ODEs and SDEs Are the Machine

Lecture 1 of MIT 6.S184 first rewrites "generate an image of a dog" as "sample from the data distribution," then gives the machine that does the sampling: start from Gaussian noise and simulate an ODE along a neural-network vector field (a flow model), or add a little Brownian-motion noise at every step to get an SDE (a diffusion model). Each is simulated with the simplest numerical method available, Euler and Euler–Maruyama. How to train the vector field is left to Lecture 2.

MIT 6.S184 L2: Flow Matching, Learning the Marginal Vector Field from Conditional Paths

The object we want is the marginal vector field: run an ODE along it and noise flows into data. The catch is that it requires an integral over the whole dataset, so we can't compute it. Flow matching regresses on the conditional vector field instead, the one that pushes noise toward a single data point, which has a closed form. Theorem 12 in the notes shows the two losses differ by a constant and share the same gradient. On the CondOT path, training reduces to one line: sample data z, noise ε, and time t, and have the network predict z − ε at the point tz + (1−t)ε.

MIT 6.S184 L3A: Score Functions, SDE Sampling, and Score Matching

A score function is the gradient of the log density; it points toward where probability rises fastest. On Gaussian paths, the score and last lecture's vector field are both linear in x and z, so each converts into the other (Proposition 1 in the notes): learn one and you have learned both. With the score in hand, you can add noise of any strength to the ODE and turn it into an SDE without changing the distribution at any time (Theorem 17). The score itself is learned with the same trick as flow matching, by regressing on the conditional score. On Gaussian paths, that amounts to predicting the noise that was added, which is the DDPM training objective.

MIT 6.S184 L3B: Guidance and Classifier-Free Guidance

Feeding the prompt to the network as an extra input should, in theory, sample from p_data(x|y), but in practice the images don't follow the prompt closely enough. Lecture 3B uses Bayes' rule to split the guided vector field into the unguided vector field plus a classifier gradient; scaling that classifier term by w is classifier guidance. Replacing the classifier with the difference between guided and unguided fields gives CFG, which needs no classifier: ũ = (1−w)·u(x|∅) + w·u(x|y). Training only requires swapping the label for a null label ∅ with probability η. The costs: two network calls per step, and for w>1 you are no longer sampling from the data distribution.

MIT 6.S184 L4: U-Nets, DiTs, and Latent Space

The algorithms are complete by Lecture 3B; Lecture 4 tackles two engineering problems that show up at scale. First, the network must take an image, a time t, and a prompt and output a vector field of the same size, so we use a U-Net or a diffusion transformer (DiT), embedding time with Fourier features and text with frozen CLIP/T5 encoders. Second, pixel space is too big, so we first train a VAE to compress images into a latent space, run flow matching there, and decode at the end. Stable Diffusion 3 and Meta Movie Gen Video both follow this recipe: flow matching in latent space, a DiT variant, and CFG.

MIT 6.S184 L5: Discrete Diffusion, Generating Language with CTMCs

Text is a sequence of discrete tokens. There is no direction to move in, so ODEs and SDEs do not exist. Lecture 5 carries the recipe from Lectures 1–4 over unchanged and swaps only the underlying stochastic process: vector fields become rate matrices, ODEs become continuous-time Markov chains (CTMCs), and the continuity equation becomes the Kolmogorov forward equation. With the factorized mixture path, the marginal rate matrix has exactly one unknown: the probability of each position's original token given the noisy sequence. Training a discrete diffusion model therefore reduces to per-position classification with a cross-entropy loss. Make the noise all [mask] tokens and you get a masked diffusion language model.

NCCU Generative AI L01: Why Study Generative AI, Course Intro and Colab

The first half of lecture 1 covers course rules and lightning talks. The middle answers "why learn the principles?": Yen-Lung Tsai splits the anxiety of learning AI into three kinds and argues that knowing the principles tells you a model's limits, so you stop chasing every new tool. The second half is a Colab primer, from magic commands and the four standard import lines to plt.plot, Markdown, and ipywidgets. Homework 1 is to plot a function in Colab. On the Chang Gung satellite rubric, a tweaked copy of the demo earns 6 points; a function not taught in class, with well-written Markdown notes, earns 10.

NCCU Generative AI L02: Neural Network Concepts

Lecture 2 opens up last week's "dopey AI robot." Inputs and outputs must become numbers (tensors). Classification uses one-hot labels and softmax to turn scores into probabilities. A neural network is neurons stacked layer by layer, and training means pushing the loss down with gradient descent. The lecture ends by building a first fully connected network on MNIST in Keras and wiring it to a Gradio sketchpad. Homework 2 asks you to design your own DNN, with one hard rule: it can't have three layers. The Chang Gung satellite rubric wants a screenshot of the parameters with the best validation accuracy and encourages keeping failed attempts.

NCCU Yen-Lung Tsai Generative AI L03: GANs — How Do Two Competing Networks End Up Drawing Pictures?

The prompt "a cute girl" has countless correct pictures, so training it as a function only teaches the model the average of all of them. GANs sidestep this by training two networks: a generator G turns a random latent vector into an image, a discriminator D judges real versus fake, and the two compete. L03 walks from the 2014 paper through WGAN, Progressive GAN, StyleGAN's 512-dimensional latent and AdaIN, then Pix2Pix and CycleGAN. An appendix explains cross entropy and KL divergence as a 'surprise index'. Week 3 homework: run a GAN yourself, or explain CE and KL in your own words.

NCCU Yen-Lung Tsai Generative AI L04: LLMs Are Simpler Than You Think — Next-Word Prediction, Temperature, and Your Own Benchmark

L04 reduces a large language model to one sentence: look at the preceding words, score every word in the vocabulary, turn the scores into probabilities with softmax, and sample the next word. To give the model a memory of what came before, the lecture covers RNNs and then gives a first look at Transformer Q/K/V. GPT-2's 1.5 billion and GPT-3's 175 billion parameters illustrate scale; temperature and top-p explain why every answer comes out different. The second half covers running open models locally and estimating VRAM. Week 4 homework: write test prompts on a topic you know well and compare at least two LLMs.

NCCU Yen-Lung Tsai Generative AI L05: Transformers, Explained — Reading Q/K/V, Positional Encoding, and Residuals Through Linear Algebra

L05 reads the whole Transformer with two linear-algebra rules: matrix multiplication is row-times-column dot products, and a row vector times a matrix is a linear combination of the matrix's rows. With those, attention is 'dot the query with every key, softmax into weights, take a weighted average of the values,' or softmax(QKᵀ/√d_k)V in batch form. Dividing by √d_k just pulls the numbers back toward 0 so softmax doesn't become winner-take-all. Then come multi-head attention, encoder versus decoder, masking, positional encoding as a set of sin/cos clocks, and ResNet-style residuals with layer normalization. No homework this week.

Reading NCCU Yen-Lung Tsai Generative AI, L06: LLM Applications and Ethical Challenges — Hallucination, Privacy, DeepSeek, and a One-Paragraph System Prompt Called the Lucky Vicky Generator

The first half of L06 is about ethics. Yen-Lung Tsai quotes Karpathy's line that hallucination is a feature of LLMs, then works through plagiarism, whether your data gets used for training, and DeepSeek's censorship and corpus skew, and closes with seven principles of responsible use. The second half is about applications: give the model the right information and clear instructions, and one system prompt becomes a Lucky Vicky positivity generator, a social-media copywriter, or a biased college-major counselor. The week-6 assignment moves that prompt into an OpenAI-compatible API with a Gradio front end: a chatbot with a persona.

Reading NCCU Yen-Lung Tsai Generative AI, L07: Building Your Own Chatbot — API Keys, Three Roles, Sending the History Back, and Running Models Locally with Ollama

A chatbot 'remembers' you not because the model has memory, but because your code resends the whole messages list (system, then alternating user and assistant) every turn. L07 starts with getting OpenAI and Groq keys, spells out that structure, then runs Gemma 3 locally or in Colab with Ollama, where the same openai package works after changing only base_url. The week-7 assignment offers two options: a version that keeps the conversation going, or two models talking to each other, both demoed in Gradio.

Reading NCCU Yen-Lung Tsai Generative AI, L08: Retrieval-Augmented Generation (RAG) — Chunk the Text, Embed It, Put the Closest Pieces Back in the Prompt

L06 said a prompt is two things: correct information and clear instructions. RAG lets the computer fetch the information part on its own. Split your documents into chunks, turn chunks and questions into feature vectors with the same model fθ, find the closest few chunks, and drop them into a template: 'Answer {question} based on {retrieved_chunks}.' The code comes in two notebooks: Demo06a builds a vector database with LangChain and FAISS and zips it as faiss_db.zip; Demo06b loads it back, connects an LLM, and wraps it in Gradio. The week-8 assignment is to do the same with your own data.

Reading NCCU Yen-Lung Tsai Generative AI, L09: Why 2025 Was Called the Year of AI Agents — Andrew Ng's Four Design Patterns, with Reflection and Two-Stage CoT Built in AISuite

L09 defines an AI agent in one line: the AI finishes the work you would otherwise do yourself. Yen-Lung Tsai follows Andrew Ng's four design patterns (Reflection, Tool Use, Planning, Multiagent Collaboration) but builds only the two easiest. Demo07a hands a draft between a "writer" and a "reviewer" LLM call. Demo07c splits the Lucky Vicky post generator into "think of five reasons, then write the post", a two-stage CoT. Both use AISuite with Groq and a Gradio front end. LangChain, AutoGen and CrewAI appear only on a further-learning list. The week 9 homework asks you to pick one of the two patterns.

Reading NCCU Yen-Lung Tsai Generative AI, L10: The Adventure That Starts with the VAE — Feature Vectors, Autoencoders, Diffusion, and "Without the VAE, Stable Diffusion Doesn't Run"

L10 starts from one question: how do you find a good feature vector? Word2Vec learns embeddings through a pretext task. An autoencoder squeezes out a latent vector by being forced to reproduce its input. A VAE then asks the latent to follow a normal distribution, so nearby points produce similar images. Yen-Lung Tsai then recasts diffusion as "an autoencoder whose encoder is computed and whose decoder is learned", and ends on latent diffusion: a VAE shrinks a 512×512 image to 64×64, and diffusion runs only in that small space. The week 10 homework involves no code: make several style-consistent image sets with Bing.

Reading NCCU Yen-Lung Tsai Generative AI, L11: Text-to-Image AI, Principles and Practice — CLIP Reads the Prompt, Schedulers Decide Whether It Converges, LoRA Learns Only ΔW, and You Build a Web App with diffusers

L11 fills in the rest of the Stable Diffusion diagram. CLIP is trained so that matching text and images get similar vectors, which turns a prompt into a 77×768 embedding. Schedulers compress 1,000 noising steps into twenty or thirty denoising steps, but ancestral samplers such as Euler a never settle: push to 100 steps and the subject changes jackets and seats. LoRA freezes the original W and learns only a ΔW factored into A·B. The hands-on part loads an SD 1.5-family model with diffusers, and the week 11 homework is your own image-generation web app.

NCCU Yen-Lung Tsai Generative AI L12: ControlNet and Fooocus, or How to Make an Image Model Follow Your Composition

The Stable Diffusion setup from L11 listens only to the prompt, so composition and pose are left to luck. L12 adds a steering wheel. ControlNet copies a block of SD and wires the copy back in through zero convolutions, so extra conditions such as edge maps, poses, and depth maps can steer generation. The standard example is Canny edges. The second half covers Fooocus, an SD interface that aims to be 'as simple as Midjourney': Presets, Styles, and the five Input Image features, where Image Prompt is ControlNet with a friendly wrapper. Week 12 homework: pick a use case, make at least 3 image sets in Fooocus, and write up your creative process.

NCCU Yen-Lung Tsai Generative AI L13: Reinforcement Learning, from AlphaGo to the RLHF That Makes LLMs Bluff Less

Every model in the first 12 lectures learned from training data that people prepared. L13 asks a different question: when there is no right answer, only a signal of how well you did, how does a computer learn? Tsai starts from AlphaGo and Breakout and splits the field in two. Value-based methods learn a Q function that scores each action (Deep Q-Learning, TD, experience replay, ε-greedy); policy-based methods learn the action directly (policy gradient, actor-critic). The second half returns to LLMs: ChatGPT trains a reward model from human rankings and then runs RLHF with PPO, while DeepSeek has the computer check math answers automatically and uses that as the reward. Week 13 homework is the final project proposal.

NCCU Yen-Lung Tsai Generative AI L14: Text and Image Models Invade Each Other's Territory, Plus the Final Project

The last lecture looks at two lines of technology crossing into each other. LLMs such as ChatGPT have started drawing, and the slides use early fusion plus VQ-VAE/VQGAN to explain how an image can be cut into tokens. Going the other way, Inception Labs' Mercury generates text with diffusion, noising a sentence into a row of [MASK] tokens and then restoring it. Next come a few papers anyone can use: evaluating RAG automatically, reasoning models being easier to hijack, and DeepMind's four kinds of AI risk. The lecture ends with vibe coding and a list of application tools, and the final project runs as an online conference in Gather Town.

Reading NCCU Yen-Lung Tsai's Generative AI: Overview and Self-Study Route

Generative AI: Text and Image Synthesis Principles and Practice is an introductory course taught by Yen-Lung Tsai (蔡炎龍) of NCCU's Department of Mathematical Sciences and opened to other schools as a TAICA satellite course. The most complete course page online actually belongs to the Chang Gung University satellite section, where Chih-Yuan Yang is the co-teacher. This series follows Spring 2025 (semester 1132): 14 recordings, 14 slide decks, and 12 homework specs with rubrics are public, and the demo notebooks are on GitHub, so the access grade is A3. The gaps: the notebooks keep changing, submission and grading run through each school's LMS, and final projects were never published.

NTHU NLP Guide 8: ELMo, BERT, T5, BART, GPT and the Three Roads to Pretraining

A guide to the BERT and its Family unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). It starts from "I record the record": Word2Vec and GloVe give both records the same vector. ELMo fixes this with a bidirectional LSTM language model. With Transformers, pretraining splits into three roads: encoders (BERT: MLM plus NSP, strong at understanding, weak at generation), encoder-decoders (T5's span corruption, BART's five noise types), and decoders (GPT: pure next-token prediction). The last part covers GPT-3's in-context learning and scaling laws, and why decoders became today's dominant backbone.

NTHU NLP Guide 10: Why the Same Model Sounds Stiff or Rambles — Decoding Strategies and NLG Evaluation

A guide to the decoding and evaluation unit in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The first half covers how to pick a word once the model outputs a probability distribution: greedy decoding can't take back a mistake, beam search keeps several candidates but favors short outputs, and top-k / top-p trade determinism for diversity (missing from the slides; the professor covers it verbally in class). The second half covers scoring generated text: BLEU's modified precision and brevity penalty, ROUGE-N and ROUGE-L, perplexity, and what GLUE, SQuAD 2.0, MTEB, and MMLU each measure.

NTHU NLP Guide 11: GPT-2 vs. T5 for Chinese Summarization — Left Padding, −100, and ROUGE

A guide to the GPT-2 / T5 tutorial in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). One task, LCSTS Chinese summarization, is solved twice. Decoder-only GPT-2 is written in native PyTorch: you join article and summary into one sequence, switch to left padding, and set padding labels to −100. Encoder-decoder mT5 uses Seq2SeqTrainer: no left padding needed, and DataCollatorForSeq2Seq handles the −100 for you. Both segment with jieba and score word-level ROUGE.

Reading NTHU Kao's NLP: GPT-3, InstructGPT, and RLHF — How a Model That Continues Text Becomes an Assistant That Follows Instructions

Hung-Yu Kao's Fall 2025 W8 slides walk from GPT-1 to GPT-3, explain how the Sparse Transformer behind GPT-3 cuts attention cost, and then use InstructGPT to show the gap between continuing text and following instructions. The maximum likelihood objective can't tell a fabricated fact from a slightly wrong synonym, so three extra stages are added: SFT learns how humans write, a reward model learns how humans grade, and PPO optimizes against that grade while a KL penalty keeps the model from drifting too far. The lecture closes with Llama-2: separate safety and helpfulness reward models, context distillation, and GQA for faster inference.

NTHU NLP Guide 9: One BERT That Scores and Classifies at Once — The Hugging Face Tutorial and HW3 Multi-Output Learning

A guide to the Hugging Face BERT tutorial and HW3 in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The tutorial walks through binary IMDb sentiment classification: AutoTokenizer, the input_ids / token_type_ids / attention_mask fields, AutoModelForSequenceClassification, and Trainer. HW3 applies the same tools to SemEval 2014 Task 1: one bert-base-uncased with two heads, one regressing a 1–5 relatedness score and one classifying entailment into three classes. You add the two losses and write the training loop yourself, because Trainer is not allowed.

NTHU NLP HW1: Testing Word Vectors on Google Analogy — Pretrained GloVe vs. Word2Vec Trained on 20% of Wikipedia

HW1 tests word vectors on the 19,544 Google Analogy questions (8,869 semantic, 10,675 syntactic). You first answer them with pretrained glove-wiki-gigaword-100 loaded through Gensim, then train your own Word2Vec on a 20% sample of a pre-cleaned Wikipedia dump, and plot t-SNE for the family subcategory both times. Seven TODOs are worth 55%, the report 45%. Fall 2026 keeps the same TODOs but asks for an .ipynb with outputs.

NTHU Hung-Yu Kao NLP, Week 1: Why Language Is Hard, and How Text Became Numbers Before LLMs

Week 1 of Hung-Yu Kao's NLP course at NTHU is a 91-page deck, W1_NLP_brief. It opens with 'Watch for kids', five readings of the telescope sentence, and a Chinese tongue-twister about eleven uncles to show why language is hard. Then, from an information-retrieval angle, it builds one pipeline: inverted index, tokenization, stemming, TF-IDF, BM25. The second half hits that pipeline's dead ends (synonyms, polysemy, vocabulary mismatch), moves to SVD-based LSA, and closes with a preview of dense vectors through Skip-gram, GloVe, and FastText.

Reading NTHU Hung-Yu Kao's Natural Language Processing: What Outsiders Can Get from a 1,200-Seat TAICA Course

Hung-Yu Kao's Natural Language Processing at National Tsing Hua University is a graduate-level flagship course in the TAICA alliance. The syllabus caps it at 1,200 students, it is taught in Mandarin, and it runs from TF-IDF and word vectors to RLHF, PEFT, and RAG. For Fall 2025, the slides, 32 class recordings, and 4 assignments with starter notebooks are all on GitHub, which rates A3. Solutions, grading, and the term-project spec are not public. Fall 2026 is in progress and only goes up to W3, so it rates A2. Grading changed to 75% assignments plus a 25% in-person midterm, and a Reasoning/Agent unit was added.

NTHU NLP LLM API Lab: NLI Classification with Gemini, OpenAI, and Claude, Using prompts.yaml, JSON Output, Few-Shot, and Token Counts

This 34-slide TA session answers a practical question. Pasting data into the ChatGPT web page one row at a time is slow and hits hourly limits, so research and homework should use the API. The notebook runs one SemEval 2014 entailment example through Gemini, Claude, and OpenAI in turn: prompts live in prompts.yaml, output is forced into JSON, then few-shot and token counting. The material is from 2024. The slide cover says 2024/11/21, and the notebook uses gemini-1.5-pro, gpt-4o, and claude-3-5-sonnet-20241022. That Claude model was retired on 2025-10-28, and Google's old Gemini SDK reached end of support on 2025-11-30.

Reading NTHU Kao's NLP: Parameter-Efficient Fine-Tuning — Fine-Tuning Large Models Without an A100 Cluster

Hung-Yu Kao's Fall 2025 PEFT slides open with a budget: full fine-tuning of Llama 2-7B in 16-bit needs about 56GB of GPU memory, while training only 0.2M parameters brings it down to about 17GB, because gradients and optimizer states nearly vanish. Intrinsic dimensionality then explains why tuning a small slice is enough: the longer a model is pretrained and the larger it is, the fewer effective dimensions fine-tuning needs. Methods fall into additive (Adapters, Prompt Tuning), selective (BitFit), reparametrization (LoRA), and hybrid (MAM Adapters, S4). The second half runs from GPT-2's task descriptions and verbalizers to the trade-offs between prefix tuning and soft prompt tuning.

NTHU NLP HW2: Arithmetic as a Language — PyTorch TA Session, a Two-Layer LSTM, and Teacher Forcing

HW2 treats expressions like "14*(43+20)=882" as character sequences and asks a two-layer LSTM to generate the answer one character at a time after it sees "=". The training set has 2,369,250 rows and the eval set 263,250, with every number in 0–49. Six TODOs run from building a vocabulary and batching with loss only after "=" to a generator, teacher-forced training, and exact-match evaluation. The W4 PyTorch TA session is the toolbox for it.

NTHU NLP RAG Labs + HW4: Building a Cat-Facts RAG Two Ways, with LangChain and by Hand

Each of the two RAG TA sessions builds one version. The first installs Ollama on Colab to run llama3.2:1b and wires up a minimal RAG with LangChain's Chroma, MMR, and retrieval chain. The second uses LangChain only for data prep and writes the rest by hand: chunking, text and vector stores, hybrid BM25 + cosine retrieval merged with RRF, then generation with Llama-3.2-1B-Instruct. HW4 applies the first session's skeleton to 150 cat facts and 150 GPT-5-generated QA pairs. The generator must be Llama3.2-1b and the embedding model jina-embeddings-v2-base-en, and you report recall@1, recall@5, and exact match. Code is 45% of the grade and the report 55%; the report analyzes how prompts, data format, document order, and counterfactual information change the results.

NTHU NLP RAG, Part 2: Connecting the Retriever to a Reader, from ORQA and REALM to Self-RAG

The second half of W11_RAG.pdf starts at the "From Retrievers to QA" slide and turns a retriever plus a reader into a full QA system. It begins with ORQA and REALM (2019–2020), where BERT is the reader, then covers the first paper named RAG and REPLUG, which keeps the LLM frozen. A single table then sorts seven recent fixes into three groups: rewrite the query (Query Rewriting, HyDE), make the generator robust to noise (RetRobust, RAFT, RAAT), and decide when to retrieve (FLARE, Self-RAG). It ends with noise types, four abilities an LLM needs inside RAG, and generative retrieval (GR) with reliable response generation (RRG). On the recording side, the W11 Thursday lecture stops at the RAG paper; I could not find a recording that covers the later slides.

Reading NTHU Kao's NLP: RAG (Part 1) — Hallucination and Retrievers, from BM25 to DPR and GTR

An LLM will confidently answer that Oppenheimer was born in 1967 (the year he died). RAG retrieves first and generates second. The first 60 pages of Hung-Yu Kao's Fall 2025 RAG deck are all about finding the right material. Sparse vectors (bag-of-words, TF-IDF, BM25) are cheap and dependable; dense vectors catch paraphrases. BERT's [CLS] isn't a good sentence vector as-is, hence Sentence-BERT pooling and bi-encoders. A cross-encoder is accurate but needs nearly 50 million passes for 10,000 sentences, while a bi-encoder needs 20,000. SimCSE uses dropout as data augmentation, DPR beats BM25 with only 1,000 training examples, and GTR shows that scaling up a dual encoder improves out-of-domain retrieval.

NTHU NLP W3: When Input and Output Lengths Differ — Seq2seq, LSTM, and Attention

"Look over there" is 3 tokens, the Chinese version 4, the Japanese version 12. When output length does not track input length, the classifier trick of adding an FFN at the end fails. This 33-slide deck starts from encoder-decoder models, works through vanishing gradients in RNNs and the three LSTM gates, and ends with attention fixing both long-range memory and parallelism, including why scores are divided by √d.

NTHU NLP Guide 7: Sub-word Tokenization, or Why a Model's Vocabulary Is Made of Word Pieces

A guide to the Sub-word Tokenization unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). Splitting on white space only works for Western languages and cannot handle unseen words or German-style compounds. BPE starts from characters and repeatedly merges the most frequent adjacent pair into the vocabulary, so each merge adds one entry; its weakness is that the greedy split is not always the best one. The Unigram LM picks splits by probability and can even sample different ones. The slides give vocabulary sizes of 30522 for BERT, 50257 for GPT-2/GPT-3 and 32,128 for T5.

NTHU NLP Course Summary and LLM Reasoning Notes: A Three-Column Map of the Semester, Then Denny Zhou's Question of Whether Pretrained Models Can Reason

In week 14 of Fall 2025, Hung-Yu Kao closed the course with two slide decks. Course_summary sorts the semester into three columns (NLP Fundamentals, NLP Models, NLP Advances), adds two columns for the TA labs, and lists five directions for further study. The second deck is his notes on Denny Zhou's (Google DeepMind) April 2025 Stanford talk. Its claim: pretrained models can already reason, and decoding is what brings it out. CoT decoding, self-consistency and retrieval + reasoning add up to four inequalities. The Fall 2026 W16 'Reasoning / Agent' unit has not been released yet.

NTHU NLP Term Project and the Fall 2026 Redesign: A 30% Group Project Becomes an In-Person W14 Midterm, Plus Reasoning/Agentic AI, an AI-TA and TAICA Compute

Fall 2025 was graded 70% assignments + 30% term project. Projects were done in groups of 3–4 and split into Proposal 6%, Progress 6%, Poster 6% and Report 12%, with no GPUs provided. The repo has no project spec, only the syllabus structure, an end-of-term reminder and the W15–W16 recordings. Fall 2026 switches to 75% assignments (4 of them) + a 25% in-person midterm in W14. The schedule drops the presentation weeks, adds a Reasoning/Agent unit, and brings in an AI-TA for grading support and TAICA compute credits. The official materials don't say why.

NTHU NLP Guide 6: Without the RNN, How Does a Transformer Know How Words Relate and Where They Sit?

A guide to the Transformer unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). The slides start from two RNN problems: words interact only across O(N) steps, and time steps cannot run in parallel. Self-attention lets every word look at every other word and computes it all in one matrix multiplication, fixing both at once. The cost is that the model can no longer tell word order, so sinusoidal positional encoding is added. The lecture then assembles multi-head attention, Add & Norm, feed forward, cross-attention, masked attention and teacher forcing, and closes with GPT-2, ViT and four variants to show how far the architecture went.

NTHU Hung-Yu Kao NLP, Week 2: Word Embeddings and Language Models, from Counting N-grams to RNNs

Week 2 of Hung-Yu Kao's NLP course at NTHU is a 62-page deck about predicting the next word. The first part covers statistical language models with a bigram count table, add-one smoothing, and perplexity, then sparse vectors through a PPMI example with cherry and digital. The middle returns to Word2Vec's negative sampling and uses the Chinese word for 'apple' to show why contextualized embeddings are needed. The last part derives RNNs from three weaknesses of feedforward networks and shows how they handle NER, sentence classification, and stacked and bidirectional variants.

NTU ADL Lecture 4: Attention and the Transformer

An RNN translator has to squeeze the whole source sentence into one vector, and long sentences don't fit. Attention lets the decoder look back at every input position each time it produces a word, score each one, and take a weighted sum. Call the scorer the query, the thing being scored the key, and the thing being averaged the value, and you have dot-product attention. The Transformer goes one step further: the input attends to itself, recurrence disappears, and in return you get parallelism and a constant path length. The price is that position information has to be added back by hand.

Reading NTU ADL 2025 Fall: BERT and Its Family — From the Polysemy Problem to XLNet, RoBERTa, and mBERT

A static word vector gives "apple" one embedding, whether it means the fruit or the company. The BERT lecture in ADL Fall 2025 starts from that polysemy problem. TagLM feeds language-model features into a tagger, ELMo builds contextual embeddings from a deep bidirectional LSTM, and BERT swaps the LSTM for a Transformer, pre-trained with Masked LM and Next Sentence Prediction; downstream, you add a classifier or tagger on the top layer and fine-tune. The optional BERT Variants slides go one ring further out: Transformer-XL for longer context, XLNet's permutation LM to get both AR and AE benefits, RoBERTa's better data and training recipe, SpanBERT's span masking, and mBERT and XLM for many languages. This is the direct prerequisite for HW1, which uses bert-base-chinese for extractive QA.

Reading NTU ADL 2025 Fall: Beyond Supervised Learning and Multimodality — Auto-Encoders, VAE, Dual Learning, Contrastive Learning, and CLIP

Big data is not big annotated data. The last ADL lecture asks how to learn good representations without labels, and answers: find the latent factors that control the data. An auto-encoder squeezes the input into a short code and reconstructs it. The denoising version adds noise or masks 15% of tokens first, which is exactly the idea behind BERT's masked LM. A VAE forces the code to follow a distribution, so you can sample from it to generate. Dual learning lets paired tasks, such as translation and back-translation or understanding and generation, act as feedback for each other. Self-supervised learning has two camps: self-prediction (hide part, guess it back) and contrastive learning (pull similar pairs together, push dissimilar ones apart). CLIP runs contrastive learning on 400 million image-text pairs, making zero-shot image classification possible, and DALL·E 2 uses CLIP's representations to generate images. Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.

Reading NTU ADL 2025 Fall: Conversational AI and Tool Use — From LU/DST/Policy/NLG to LaMDA, WebGPT, and Toolformer

Dialogue systems split into chit-chat and task-oriented. Task-oriented systems were traditionally built from four modules: language understanding (LU) turns a sentence into domain, intent, and slots; dialogue state tracking (DST) accumulates the user's goal; the dialogue policy picks the next system action; and NLG turns that action back into a sentence. An LLM can act out all four steps by itself, but it cannot actually make the booking, so it needs external tools. LaMDA learns to call a search engine, calculator, and translator. BlenderBot 2.0 adds internet search and long-term memory. WebGPT learns to drive a browser from human demonstrations, a reward model, and PPO. Toolformer has the model generate and filter its own tool-use training data. The lecture ends with evaluation: automatic metrics, four kinds of human evaluation, and LLM-Eval. ADL Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.

Reading NTU Yun-Nung Chen's Applied Deep Learning 2025 Fall: Course Map, A2 Rating, and How to Read It

Applied Deep Learning (ADL) Fall 2025, taught by Yun-Nung (Vivian) Chen in NTU's CSIE department, is a deep learning course built around NLP. It runs from neural network basics through Transformers, BERT, pretraining and prompting, post-training, LoRA, RAG, generation and evaluation, alignment issues, and language agents. Lectures L0–L11 come with slide PDFs and segmented videos, and the playlist adds videos for L12–L14. On the assignment side only the HW1 spec is public; HW2, HW3, and the final project have explainer videos only. That makes it A2.

NTU ADL 2025 Lecture 10: Bias, Safety, Hallucination, and Alignment, Plus the Jailbreaking Olympics Final Project

Lecture 10 of ADL Fall 2025 sorts the problems of pretrained models into four groups, each paired with a goal: bias with fairness, toxicity with safety, hallucination with factuality, and finally alignment. The slides argue that bias can enter at any stage of the ML pipeline, that safeguards belong at four layers (data, input, training, output), that hallucination can be checked atomic fact by atomic fact, and that over-optimizing a reward model produces familiar symptoms: verbosity, excessive apologies, over-refusal. The final project announced that week is called Jailbreaking Olympics, but all that is public is the titles and one-line descriptions of two videos.

Reading NTU ADL 2025 Fall: HW1 Chinese Extractive QA — Finding the Answer Span Among Four Paragraphs

HW1 in ADL Fall 2025 gives a question and four Chinese paragraphs. The model first picks the relevant paragraph (paragraph selection, framed as four-way multiple choice), then marks the answer's start and end inside it (span selection), scored by Exact Match. The spec slides point you straight at Hugging Face's run_swag_no_trainer.py and run_qa_no_trainer.py. The simple baseline uses bert-base-chinese, length 512, effective batch size 2, and learning rate 3e-5, and both stages together take under three hours on an 8GB RTX 3070. The Kaggle leaderboard closed 9/29, and code plus report were due 10/1 on NTU COOL. Outside readers cannot get the Kaggle data or grading, but the task design, baseline settings, and five report questions are all usable for practice.

NTU ADL 2025 Lecture 11: Reasoning, Memory, Planning, and Multi-Agent Systems in Language Agents

Lecture 11 of ADL Fall 2025 builds on the EMNLP 2024 Language Agents tutorial. It defines an agent as an entity that perceives and acts, then names what is new about language agents: reasoning itself counts as an internal action. The lecture is organized around three concepts. Reasoning covers CoT and ReAct; memory covers Generative Agents and its recency / importance / relevance retrieval; planning goes from greedy reactive planning to tree search and world models. It closes with multi-agent systems in three steps: initialization, orchestration, and team optimization.

NTU ADL 2025 Lecture 1: What Machine Learning and Deep Learning Are

The first self-study deck of ADL Fall 2025 describes machine learning as finding a function from data, and deep learning as a production line of simple functions where the machine learns what every station does. It uses speech and vision to contrast deep and shallow models, credits big data and GPUs for the post-2010 breakthroughs, and uses the universality theorem to ask why networks should be deep rather than fat. The most practical part comes last: the output domain decides the learning task, and the architecture should fit the properties of the input domain.

NTU ADL 2025 Lecture 2: Neural Networks and Backpropagation

ADL Fall 2025's NN Basics and Backpropagation decks break model training into three questions. What is the model? Layers of neurons, each computing z = Wa + b and then a nonlinearity. What makes a function good? A smaller loss. How do we pick the best one? Gradient descent, in practice mini-batch SGD. Backpropagation computes gradients for millions of parameters efficiently: the forward pass stores each layer's output, the backward pass sends an error signal δ back from the output layer, and multiplying the two gives each weight's gradient.

NTU ADL 2025 Lecture 9: Decoding, Generation Control, and Evaluation for NLG

Lecture 9 of ADL Fall 2025 answers two questions. The model gives you a probability distribution at every step, so how do you pick a word from it? And once you have a sentence, how do you judge it? The slides start with teacher forcing and exposure bias to show the gap between training and generation, then compare greedy, beam search, sampling, top-k, and nucleus sampling, and file temperature and the penalties under 'control' rather than decoding algorithms. The evaluation half covers BLEU, ROUGE, perplexity, and LLM-Eval, then explains why you would use RL to optimize whole-sentence quality directly.

NTU ADL 2025 Lecture 7.5: PEFT — Adapter, LoRA, Prompt Tuning, and HW2

When an LLM is too big to fine-tune in full, the LLM Adaptation slides of NTU ADL Fall 2025 offer three ways to change only a small part of it: insert small Adapter modules into the Transformer, represent the weight update with low-rank matrices (LoRA), or learn only a prefix or soft prompt (prompt tuning). The slides conclude that no single method fits every task. For HW2, the only public information is its title, "LLM Tuning and Prompt Tuning for Classical Chinese Translation"; the data, baseline, and grading have no written spec.

NTU ADL 2025 Lecture 7: Post-Training — Instruction Tuning, RLHF, and InstructGPT

A pre-trained model can continue text, but that does not mean it follows instructions. The Post-Training slides of NTU ADL Fall 2025 fix this in two steps. Instruction tuning (FLAN, T0) teaches the model to read task descriptions. RLHF then pulls its outputs toward human preference. Three limits of instruction tuning connect the two steps, a reward model and pairwise comparisons solve two practical RL problems, and InstructGPT's SFT → reward model → PPO pipeline ties it all together. ChatGPT runs the same pipeline on multi-turn dialogue.

Reading NTU ADL 2025 Fall: Three Pre-training Families and Prompt Learning — From BERT, GPT, and T5 to Prompts Only Machines Understand

Lecture 6 of ADL Fall 2025 sorts pre-trained models into three families: encoders (the BERT family, bidirectional context), decoders (the GPT series, good at generation), and encoder-decoders (BART and T5, pre-trained with denoising). It then names two practical obstacles of the pre-trained-model era: downstream labeled data is scarce, and models keep growing until one copy per task no longer fits. The slides' answer is prompt learning. GPT-3's in-context learning shows a model can do a task without updating parameters; hand-written hard prompts (template plus verbalizer, LM-BFF) then give way to soft prompts optimized as vectors (P-Tuning, Prefix-Tuning, Prompt Tuning); and Liu et al.'s prompting typology closes the lecture.

NTU ADL 2025 Lecture 8: RAG — From Retrieval and Reranking to Search-R1, Plus HW3

LLMs cannot memorize long-tail facts, their knowledge goes stale, and they cannot see private documents. The RAG slides of NTU ADL Fall 2025 open with an LLM hallucinating about the lecturer herself, then split RAG into indexing, retrieval, and generation: sparse (TF-IDF, BM25) and dense (DPR, Contriever) retrieval, how dense retrievers are trained, pre- and post-retrieval techniques including pointwise and pairwise reranking. A closing roadmap organizes RAG, RETRO, FLARE, Search-R1, and others by what, how, and when to retrieve. For HW3, only the title is public: "Retriever & Reranker Training for RAG".

Reading NTU ADL 2025 Fall: Reasoning — A Video-Only Lecture, Five Steps from CoT to RL

The Reasoning lecture of NTU ADL Fall 2025 has no public slides. It exists only as five videos in the course playlist: 12.1 What is Reasoning?, 12.2 Short CoT, 12.3 Test-Time Scaling, 12.4 Learning to Reason (imitating others), and 12.5 RL for Reasoning (evolving reasoning through exploration). This post lays out that route from the video titles alone, then pairs it with the CoT, ReAct, and 'reasoning enlarges the action space' pages of the previous Language Agents deck. Technical detail is left to the site's CS224N and CME295 reasoning posts.

NTU ADL Lecture 3: Word Representations, Language Models, and RNNs

The first real lecture of ADL Fall 2025 has one through-line: a language model predicts the next word. It starts from one-hot vectors and co-occurrence matrices, moves through the zero-probability problem of n-grams and the smoothing that a neural LM gets for free, and ends at the RNN LM, which folds all previous words into a hidden state. BPTT and vanishing/exploding gradients are the training cost, and LSTM and GRU patch it with gating. The lecture closes by splitting applications into sequence input versus sequence output, which separates tagging from encoder-decoder models.

NTU ADL 2025 TA Recitations: From PyTorch and Hugging Face to LoRA, Quantization, and vLLM Deployment

The ADL Fall 2025 course page schedules seven TA recitations: Dev Infra (PyTorch, debugging) → NLP project lifecycle → the underlying logic of NLP projects → LLM LoRA training → LLM basics, architecture, and MoE → LLM inference and evaluation → LLM deployment. All ten videos are older recordings by Yen-Ting Lin from 2023 and 2024, reused in Fall 2025. The course page's five slide links all return 404; files with the same names still open under the Fall 2024 path, and Deployment has a video only. The first three sessions walk through the Hugging Face data → model → demo loop that HW1 needs; the last four cover training, inference, and serving LLMs.

NTU ADL Lecture 5: Tokenization and BPE, or Where the Vocabulary Comes From

With whole words as units, any unseen word becomes UNK. With single characters, meaning is hard to reassemble. This 22-page ADL deck explains the mainstream compromise, subwords, and the most common way to build them, BPE. The core is a tiny corpus of 4 words and 16 occurrences: start from characters, merge the most frequent adjacent pair each round, and after 9 merges you have units like newest</w> and low</w>, which then segment the unseen words lowest and powest. It ends with a GPT-3 tokenizer screenshot where the Chinese version of a sentence takes more than twice as many tokens as the English.

Hsuan-Tien Lin's ML Techniques T7–T8: Blending, Bagging, and AdaBoost

Lectures 7 and 8 of Machine Learning Techniques open the aggregation part of the course. T7 sorts ways of combining hypotheses into uniform, linear, and any blending (stacking), shows with a few lines of algebra that uniform blending reduces variance, and then uses the bootstrap to create diverse g_t from the single dataset you have: that is bagging. T8 reinterprets the bootstrap as example weighting, then deliberately up-weights the examples the previous hypothesis got wrong so the next one is forced to differ, and votes with α_t = ln √((1−ε_t)/ε_t): that is AdaBoost. Practice with Fall 2024 HW6 Q4 and Q9, plus HW7's bootstrap and AdaBoost proofs and a 500-round AdaBoost-Stump experiment on madelon. There are no official solutions.

Reading Hsuan-Tien Lin's Machine Learning Foundations & Techniques: Overview and Self-Study Routes

Hsuan-Tien Lin's Machine Learning Foundations (16 lectures) and Machine Learning Techniques (16 lectures) are two Mandarin-taught MOOCs. All 130 YouTube videos and 32 slide decks are free. The MOOCs alone are A2: since August 2025, free Coursera accounts can only view the first module, so the exercises sit behind a paywall. Add the Fall 2024 course page, which publishes HW0–HW7 and the final project spec, and you reach A3, minus the grading chain: no official solutions, Gradescope and NTU COOL are enrolled-only, and the Kaggle competition returns 404. Fall 2026 is running now as a flipped classroom; slides through week 4, hw0, and hw1 are public.

Hsuan-Tien Lin's ML Techniques T9–T11: Decision Trees, Random Forests, and Gradient Boosted Trees

Lectures 9–11 of Machine Learning Techniques tie three models together with one thread: trees plus aggregation. T9 treats a decision tree as conditional aggregation and covers C&RT's binary branching, Gini and regression impurity, pruning, categorical features, and surrogate branches. T10 applies bagging to fully grown trees; add random subspaces and random projections and you get a random forest, with free OOB validation and permutation-based feature importance. T11 re-derives AdaBoost as steepest descent in function space on the exponential error, then swaps in squared error to get GBDT, which fits regressions to residuals. Practice with the impurity and gradient boosting proofs in Fall 2024 HW7. There are no official solutions.

Reading Hsuan-Tien Lin's ML Foundations: Is Learning Feasible? Hoeffding and "Outside the Data"

Foundations Lecture 4 first shows that learning is impossible: from D alone, any guess outside D can be called wrong. That is No Free Lunch. It then reframes the question with marbles in a bin. If the data is drawn independently from one distribution, Hoeffding's inequality says the in-sample error E_in is probably close to the true error E_out. Checking one fixed h is only verification. Once the algorithm chooses among M hypotheses, a union bound charges 2M exp(−2ε²N). Conclusion: with a finite hypothesis set and small E_in, learning is feasible. What to do when M is infinite is the next lecture's job.

Hsuan-Tien Lin's ML Techniques T16 Finale: Three Families of Techniques, Plus Fall 2024's Modern Deep Learning Slides

Techniques T16 re-sorts the whole course into three families: how to exploit features (kernels, aggregation, extraction, low-dimensional compression), how to optimize (gradients, equivalent problems, multiple steps), and how to fight overfitting (regularization, validation). It then uses four KDD Cup–winning models to show how the pieces combine in practice. The MOOC was recorded in 2016 and its deep learning stops at pre-training. The Fall 2024 on-campus course filled the gap with 302u (the ReLU family, Xavier/He initialization), 303u (momentum, RMSProp, Adam), a 2020 keynote deck, mlmai.ics, and 1126, an 11-model summary. The Fall 2026 versions of these files are scheduled for week 16 and currently return 404.

Hsuan-Tien Lin's ML Foundations Homework Guide: What Fall 2024 HW0–HW5 Practice and Need, Plus Fall 2026 hw0/hw1

Fall 2024 had six Foundations assignments. HW0 is 20 multiple-choice math prerequisite questions. HW1–HW5 each have 12 problems plus a bonus: Q1–4 are auto-graded, Q5–12 are graded by TAs, the programming problems use rcv1, cpusmall, and mnist from the LIBSVM datasets site, and HW5 uses LIBLINEAR. HW1 and HW2 each include a problem where you argue with a ChatGPT-style answer. Fall 2026 has released hw0 and hw1: hw1 is now 16 multiple-choice problems with 4 secretly chosen for TA grading, and the data is the course's own hw1_train.dat. The new policy allows AI tools and vibe coding, but AI-generated code needs block-by-block comments in your own words. Neither semester publishes official solutions.

Hsuan-Tien Lin's ML Techniques T5–T6: Kernel Logistic Regression and Support Vector Regression — the SVM Is a Regularized Model

Lecture 5 of Machine Learning Techniques rewrites the soft-margin SVM in unconstrained form: ½wᵀw plus C times the total hinge error. That is an L2-regularized model, and a larger C means weaker regularization. The hinge error and logistic regression's cross-entropy are both convex upper bounds of the 0/1 error, so the SVM approximates L2-regularized logistic regression. For probability outputs, you can use Platt's two-level learning, running logistic regression on top of SVM scores, or use the representer theorem to do kernel logistic regression directly. Lecture 6 uses the same theorem to get the closed form β = (λI + K)⁻¹y for kernel ridge regression, but β is dense; switching to the ε-insensitive tube error gives SVR with sparse coefficients. Fall 2026 does not schedule these two lectures.

Hsuan-Tien Lin's ML Techniques T3–T4: Kernel Trick and Soft-Margin SVM — Computing an Infinite-Dimensional Classifier and Keeping It from Overfitting

Lecture 3 of Machine Learning Techniques merges "feature transform + inner product" into a single kernel function K(x, x′). Training and prediction in the dual SVM only need K, so d̃ can be infinite: the Gaussian kernel corresponds to an infinite-dimensional transform. Lecture 4 admits the SVM can still overfit and introduces violations ξₙ and a parameter C, giving the soft-margin SVM. Its dual differs from the hard-margin one in exactly one way: αₙ gets an upper bound C. The value of αₙ sorts the data into non-SVs, free SVs, and bounded SVs, and the fraction #SV/N upper-bounds the leave-one-out error, a cheap way to rule out dangerous (C, γ).

Reading Hsuan-Tien Lin's ML Foundations: The Learning Problem, PLA, and Types of Learning

The first three lectures of Machine Learning Foundations define machine learning as a flow chart: an unknown target function f generates data D, and an algorithm A picks g from a hypothesis set H, hoping g ≈ f. The simplest H (the perceptron) and A (PLA) then show the chart in action. On linearly separable data, PLA makes at most R²/ρ² updates; on non-separable data, use pocket instead. Lecture 3 sorts learning problems along four axes: output, label, protocol, and input. Foundations mostly deals with batch, supervised binary classification or regression on concrete features. Practice with Fall 2024 HW1 and Fall 2026 hw1.

Hsuan-Tien Lin's ML Foundations L11–L12: Linear Classification, SGD, Multiclass, and Nonlinear Transforms

Lecture 11 of ML Foundations compares PLA, linear regression, and logistic regression on the same score s = wᵀx. The three differ only in their error functions, and scaled cross-entropy upper-bounds the 0/1 error, so both regressions can do classification. The lecture then turns logistic regression into SGD by computing the gradient on one random example, and builds multiclass classifiers from binary ones with OVA and OVO. Lecture 12 uses a feature transform Φ to turn a circular boundary into a line in Z-space. The price is that computation and d_vc both grow with the dimension, so the advice is: try a linear model first. Practice problems are in Fall 2024 HW4.

Hsuan-Tien Lin's ML Techniques T1–T2: Linear SVM and Dual SVM — the Fattest Separator, QP, and KKT

Lecture 1 of Machine Learning Techniques turns "which separating line is best?" into an optimization problem. Once you fix the scale so that min yₙ(wᵀxₙ+b) = 1, maximizing the margin is the same as minimizing ½wᵀw, which is a standard QP. Lecture 2 uses Lagrange duality to trade a QP with d̃+1 variables for one with N variables and N+1 constraints, then uses the KKT conditions to recover (b, w) from α. Only the points with αₙ > 0, the support vectors, affect the answer. The dual still contains the inner product zₙᵀzₘ, so the dependence on dimension is not really gone until the kernel lecture.

Hsuan-Tien Lin's ML Foundations L9–L10: From the Closed-Form Solution of Linear Regression to Gradient Descent for Logistic Regression

Linear regression writes squared error as (1/N)‖Xw − y‖², sets the gradient to zero, and gets w_LIN = X†y in one step. The hat matrix H = XX† projects y onto the column space of X, which shows that on average E_out − E_in ≈ 2(d+1)/N. Logistic regression estimates P(+1|x) with θ(wᵀx); maximum likelihood turns into the cross-entropy error ln(1 + exp(−y wᵀx)). It has no closed-form solution, so you walk downhill along −∇E_in step by step. That is gradient descent.

Hsuan-Tien Lin's ML Techniques T12–T13: Neural Networks and Deep Learning (Autoencoders, PCA)

Lectures 12 and 13 of Machine Learning Techniques open the third part, distilling hidden features. T12 starts from a linear combination of perceptrons: two layers can build AND and OR but not XOR, and one more layer fixes that, which is the multi-layer perceptron. It then replaces sign with tanh, derives backprop, and covers non-convex optimization, d_vc = O(VD), weight elimination, and early stopping. T13 discusses the challenges of deep networks, uses autoencoders as information-preserving encodings for layer-wise pre-training, treats denoising as regularization, and proves that the optimal linear autoencoder is spanned by the top eigenvectors of XᵀX, which is PCA. The videos date from 2016; modern deep learning is covered by the Fall 2024 302u/303u slides. Practice: Fall 2024 HW7 Q4, Q9, and bonus Q13.

Hsuan-Tien Lin's ML Foundations L13–L14: Overfitting and Regularization

Lecture 13 of ML Foundations defines overfitting as 'lower E_in but higher E_out' and uses experiments to find four causes: too little data, stochastic noise, an overly complex target (deterministic noise), and excessive model power. Lecture 14's remedy is regularization. It rewrites 'step back to H₂' as the constraint ‖w‖² ≤ C, then uses a Lagrange multiplier to turn it into minimizing E_in + (λ/N)wᵀw, which is weight decay. Back in VC theory, regularization shrinks the effective VC dimension d_EFF, and L1 buys sparse solutions. Practice problems: Fall 2024 HW4 Q8–9 and HW5 Q1, Q5–6, Q10.

Hsuan-Tien Lin's ML Techniques T14–T15: RBF Networks, k-Means, and Matrix Factorization

Techniques T14 reinterprets the Gaussian SVM as a linear vote over distance-based similarities, which gives the RBF network. Too many centers overfit, so k-means picks a few prototypes, and k-means itself is alternating optimization. T15 starts from the Netflix ratings data: one-hot encode user IDs, feed them into a linear network with the tanh removed, and you get matrix factorization R ≈ VᵀW, learned by alternating least squares or SGD. The lecture closes with a map of extraction models: boosting, neural nets, RBF networks, matrix factorization, and k-NN. These two lectures exist only as MOOC material. Neither the Fall 2024 nor the Fall 2026 schedule covers them, and no public homework problem does either.

Hsuan-Tien Lin's ML Techniques Homework and Final Project: Fall 2024 HW6–HW7 and the HTMLB Win Prediction

The Techniques half of Fall 2024 has two homework sets and a final project, and all three PDFs are public. HW6 covers kernels, soft-margin SVM, and aggregation; its programming part uses LIBSVM on the 3-vs-7 subproblem of mnist.scale to count support vectors, compute margins, and run 128 validation rounds. HW7 covers bootstrap, impurity, AdaBoost, gradient boosting, and neural networks; its programming part is a 500-round AdaBoost-Stump on madelon. The final project is a fictional baseball league, HTMLB: predict home-team wins across two Kaggle stages and write an English report of at most seven pages that compares at least four methods. There are no official solutions. On 2026-09-30 both Kaggle pages returned 404 without login, so outside readers probably cannot get the HTMLB data and should reproduce the same splits on a public dataset instead.

Hsuan-Tien Lin's ML Foundations L5–L6: Infinitely Many Hypotheses, So Why Does Learning Still Generalize? Growth Functions and Break Points

The Hoeffding guarantee from L4 carries an M, the number of hypotheses. Perceptrons have infinitely many lines, so M blows up. L5 stops counting hypotheses and counts how many ○× patterns (dichotomies) they can produce on N data points instead; the maximum is the growth function m_H(N). 2D perceptrons produce at most 14 patterns on 4 points, fewer than 2⁴ = 16, so 4 is their break point. L6, marked optional by the course, proves that any break point caps m_H(N) by a polynomial, which is what makes the VC bound work.

Hsuan-Tien Lin's ML Foundations L15–L16: Validation and the Three Learning Principles

Lecture 15 of ML Foundations tackles model selection. Selecting by E_in overfits, and selecting by E_test is cheating. The compromise is to carve a validation set out of the training data, select by E_val, then retrain on all the data. The validation size K is a dilemma, with K = N/5 as the rule of thumb. Leave-one-out is almost unbiased but expensive and unstable, so in practice you use 5-fold or 10-fold. Lecture 16 closes with three principles, Occam's razor, sampling bias, and data snooping, and a 'Power of Three' recap: three related fields, three bounds, three linear models, three tools. Practice problems are in Fall 2024 HW5.

Hsuan-Tien Lin's ML Foundations L7–L8: How the VC Dimension Measures Model Complexity, and What Noise and Error Measures Change

L7 names the largest non-break point the VC dimension d_VC, proves that d-dimensional perceptrons have d_VC = d + 1, and rewrites the VC bound as E_out ≤ E_in + a model-complexity penalty, so both too large and too small a d_VC hurt. Theory asks for N ≈ 10,000·d_VC examples; in practice 10·d_VC is often enough. L8 swaps the fixed target function for a distribution P(y|x) and shows the VC theory still holds under noise. The error measure should come from the application: a CIA fingerprint check that penalizes admitting an intruder 1000 times more can be reduced to plain classification by copying examples.

Reading NTU ML 2026: How AI Agents Interact and What They Do to Work — Collaboration Topologies, Werewolf, Moltbook, and AI Writing and Reviewing Papers

The second half of agent_era.pdf asks three questions. How should multiple agents collaborate? (MacNet: irregular topologies beat regular ones.) Can agents deceive each other? (Werewolf, murder-mystery games, and MARO, which learns reasoning from social play.) Can agents socialize? (Moltbook and its "Church of Molt" — though three studies find the buzz mostly human-driven and the conversations shallow.) Then, using academic research as the case: AI can already replicate and extend a paper end to end, it entered AAAI 2026's review process, and Agents4Science 2025 received 247 AI-authored papers. Hung-yi Lee's conclusion: in the early age of agents, knowing what you want to do matters more than knowing how to do it.

Reading NTU ML 2026: Context Engineering — Compression, Filtering, On-Demand Loading, and Whether to Hand the Context to the LLM

A language model's input is finite, but an agent keeps piling up tool outputs. In week two of ML 2026, Hung-yi Lee splits Context Engineering into three moves: compression (summaries, hard clearing, offloading to files, plus ACON, SUPO, and AgentFold, which make compression smarter), filtering (read only the lines you need, load tools on demand as in MCP-Zero), and finally Agentic Context Engineering, where the LLM decides the next context itself — from Dynamic Cheatsheet and ACE to Recursive Language Models. The most useful idea to take away: a subagent is a form of self-directed compression.

Reading NTU Hung-yi Lee's Machine Learning 2026 Spring: An Agent-First Course That Is Open Except for Grading

Hung-yi Lee's Spring 2026 Machine Learning course at National Taiwan University opens with OpenClaw. The first half takes apart AI agents, context engineering, inference speed-ups, and positional embeddings. The second half covers harness engineering, self-correction, and self-improving AI. Slides and recordings for all 8 lectures, plus PDFs and Colab notebooks for all 10 assignments, are public, so it rates A3. What's missing is grading: JudgeBoi returned 502 on 2026-09-30, NTU COOL is campus-only, and the three guest talks have no materials at all.

NTU Hung-yi Lee ML 2026 Guide: Faster Generation, Part 1: Flash Attention and Why Moving Data Is the Bottleneck

In week 3 of ML 2026, Hung-yi Lee spends the first half of the inference lecture on one technique: Flash Attention. A GPU's execution units are fast, but their workbench (on-chip SRAM) is tiny, so data has to be carried to and from the warehouse (HBM). The carrying is the bottleneck. A naive softmax makes several round trips to the warehouse. Flash Attention assumes the current maximum is Amax, then multiplies by a correction factor when a larger value shows up. That lets it find the maximum, build the denominator, and compute the weighted sum in one pass, without ever materializing the attention weights. The output is identical to standard attention, no retraining is needed, and the cost is a little extra compute and a little brain strain.

Reading NTU ML 2026: Harness Engineering — Making Models Stronger Without Touching the Weights

Hung-yi Lee opens with a small model fixing a bug. gemma-4-E2B-it can't find parser.py, so it writes a fake one and declares victory. Add three short sections (the current environment, how to work, what counts as done) and the same model runs ls, cat, edits the file and runs the tests. The lecture splits the harness into three levers: natural language shapes the model's frame of mind (AGENTS.md), tools set its capability boundary (SWE-agent's ACI, rewriting CLIs for agents), and workflows control its behavior (the Ralph loop, Anthropic's long-running harnesses). The second half covers three extensions: scolding an agent can backfire, how a life-long agent learns from verbal feedback, and why evaluating agents is hard. It ends with agents improving their own harness (Meta-Harness).

Hung-yi Lee ML 2026 HW1: With Only a Defense Prompt, How Many 'I have been PWNED' Attacks Can You Stop?

HW1 asks for a defense prompt under 1,000 tokens that keeps the model wrapping every reply in [START]…[END] and never saying 'I have been PWNED,' no matter how it's attacked. The TAs prepared 14 attacks, 10 public and 4 private, each worth 0.5% for safety and 0.5% for utility. The task, the full text of the 10 public attacks, and the token-counting Colab are all public, but the grading platform JudgeBoi returned 502 on 2026-09-30, so outside readers have to build their own evaluation from the spec.

Reading NTU ML 2026: HW10 Spoken Language Model — Three Architectures, Mimi's 32 Token Layers, and How Moshi Listens While It Talks

HW10 is 12 multiple-choice questions answered only on NTU COOL. Section 1 compares three spoken language model architectures: Cascade (ASR → LLM → TTS, with text in the middle), End-to-End (a language model over discrete speech tokens), and Thinker-Talker (an LLM thinks, a separate decoder speaks). In the Colab, two models listen to three clips and guess the speaker's gender, and you work out which one is the cascade. Section 2 takes Mimi apart: tokenize an emotion corpus into 32 RVQ layers, plot UMAP for layers 0, 6, 16, and 31, then encode and decode speech, laughter, and music to hear what breaks. The rest are paper questions on TWIST, AudioLM, LLaMA-Omni 2, Moshi, and GLM-4-Voice, covering initialization, pretraining, interleaving, and realtime/full-duplex behavior. The Colab needs Llama-3.2-3B-Instruct access and an HF token. Questions and Colab are public; outside readers miss only the COOL grading and answers.

Reading NTU ML 2026: HW2, AI Agent as an AI Engineer — an AIDE-Style Tree Search That Lets an Open LLM Build a MyGO & Ave Mujica Face Classifier

HW2 doesn't ask you to write a classifier. You write prompts and a pipeline so that an open LLM running on a Colab T4 (by default a 4-bit GGUF of gemma-3-12b-it) plans, codes, runs, and debugs a 10-class MyGO & Ave Mujica character face classifier on its own. The starter code is adapted from AIDE: an Interpreter runs code, a Node records each version, a Journal forms the solution tree, and the Agent decides whether to draft, debug, or improve next. The first thing worth noticing: the starter's evaluation is empty. Every version is marked metric 1.0 and not buggy, so the tree search picks blindly until you fill it in. The rules are strict: "the LLM agent is your representative", and you may not hand-edit code or prediction files.

NTU Hung-yi Lee ML 2026 Guide: HW3 LLM Fast Inference: Seven Speed-up Papers, Then Measuring Speculative Decoding, FlashAttention, and vLLM on a GPU

HW3 is 20 multiple-choice questions at 0.5 points each. No code is submitted; students answer a quiz on NTU COOL. The first 10 questions come from reading papers: four on speculative decoding (Leviathan et al., DeepMind's Speculative Sampling, Inference with Reference, SpecInfer) plus FlashAttention 1–3. The last 10 require filling TODOs in the Colab and analyzing the results: acceptance rate of a hand-written speculative decoder, speed-up curves for an assistant model vs n-gram under two prompt regimes, HBM reads and theoretical FlashAttention speed-up from T4 specs, vLLM prefix caching across turns and a cache invalidation test, and the effect of CPU offload on throughput. All questions are printed in both Mandarin and English in the homework PDF, so outsiders can do the whole thing; they just cannot get the official answers.

NTU ML 2026 HW4: Drawing Pokémon with Next-Token Prediction on a Decoder-Only Transformer

HW4 moves next-token prediction from text to images unchanged: 792 Pokémon sprites at 20×20, each pixel one of 167 color tokens, so one image is a 400-token sequence. Training is next-token prediction; at test time you get the first 60% of an image and the model draws the rest. Grading checks FID and a Pokémon Detection Rate (PDR) together, and the three baseline hints go from "run the sample code" to "tune hyperparameters" to "switch to Llama or Mistral". The spec, Colab, Kaggle notebook and dataset are public, but JudgeBoi returned 502 on 2026-09-30, so outside readers cannot get official FID or PDR scores.

Hung-yi Lee ML 2026 HW5: Teach Llama Math Without Making It Forget How to Refuse

HW5 fine-tunes Llama-3.2-1B-Instruct on GSM8K with LoRA, then uses harmful AILuminate prompts to check whether it still refuses. Math accuracy and safety rate must clear the bar together, so the real question is how to fine-tune without washing out safe behavior. The PDF, a 34-cell Colab, and a Kaggle version are public, and the strong baseline is estimated at 14 hours on a T4. The JudgeBoi grader returned 502 on 2026-09-30, so outside readers have to build their own safeguard evaluation.

Hung-yi Lee ML 2026 HW6: Model Editing, Changing One Fact and Nothing Else

HW6 involves no model training and is answered entirely on NTU COOL. Six points come from 16 multiple-choice questions on four papers (ROME, MEND, MEMIT, WISE). Four points come from swapping the Colab's fine-tuning for ROME on GPT2-XL: single editing (pick your own fact, write five kinds of test prompts) and multiple editing (10 and then 80 CounterFact examples, then MEMIT), reporting efficacy, paraphrase, neighborhood, and portability scores. The slides and the 47-cell Colab are public, but the quiz questions and answers live only on COOL.

Hung-yi Lee ML 2026 HW7: Merging a Japanese Model and a Math Model, With No Training, Into One That Solves Japanese Math Problems

HW7 hands you two models fine-tuned from Mistral-7B-v0.1: shisa-gamma-7b-v1, strong in Japanese, and WizardMath-7B-V1.1, strong in math. You may only merge them at the parameter level (no further training, no MoE or ensembles), and the merged model has to answer 20 Japanese math questions written by a TA. Part 1 (60%) is tuning the method, weights, and density in mergekit, with simple and strong baselines at 50% and 75% accuracy. Part 2 (40%) is 8 multiple-choice paper questions. The spec, Colab, and Kaggle notebook are public, but JudgeBoi returned 502 on 2026-09-30 and the paper questions live on NTU COOL, so outside readers can only check accuracy inside the notebook.

Hung-yi Lee ML 2026 HW8: Spending More Inference Compute — What Voting, Self-Certainty, and DeepConf Each Buy in Accuracy

HW8 involves no coding and no code submission. The TAs provide a finished Colab that runs Llama-3.2-1B-Instruct on the first 100 GSM8K questions and compares direct inference, Self-Consistency, Self-Certainty, and DeepConf (Confidence), sampling 16 reasoning traces per method. You read three papers, run the notebook, and answer 20 questions on NTU COOL: 18 about the papers and 2 about the Colab results. The prerequisite is Lecture 7 (Reasoning) of Lee's 2025 course. All questions are printed in hw8.pdf in Chinese and English, and the Colab is publicly downloadable. Only the COOL quiz and grades need an NTU account.

Reading NTU ML 2026: HW9 Flow Matching — From VAE to MeanFlow, Then Counting Inference Steps on a Swiss Roll

HW9 has 19 questions worth 10 points, answered only on NTU COOL with no code submission. The first 16 cover four papers — DDPM, Flow Matching, Rectified Flow, and MeanFlow — ending with questions that compare their training signals and few-step generation. The last 3 require the Colab: train two small MLPs on a 2D Swiss roll, one Flow Matching model that learns instantaneous velocity (always evaluated with 50 Euler steps, converged at Histogram JS ≤ 0.10) and one MeanFlow model that learns average velocity (always one-step, ≤ 0.40). Then compare 1 step vs 1 step, Flow Matching across Euler step counts, and Euler vs RK4 at equal steps and at similar compute. The PDF includes a generative-modeling tutorial that skips most of the math, and every question is published in Chinese and English. Outside readers miss only the COOL grading and answers.

NTU Hung-yi Lee ML 2026 Guide: Faster Generation, Part 2: KV Cache Saves Time, Fills the Warehouse, and How to Slim It Down

KV Cache stores the keys and values already computed so decode does not recompute them, but every token costs memory. For Gemma 2 27B that is about 0.72MB per token, so an 80GB A100 holds only about 114k tokens. Hung-yi Lee then walks through ways to shrink it: let queries share keys and values (MQA, GQA), compress keys and values into one vector without ever decompressing (MLA), limit the attention span (Sliding Window, StreamingLLM), and drop keys and values nobody attends to (Scissorhands, H2O). He ends with cross-conversation prompt caching: it only hits when the prefix is identical, so a system prompt should put stable content first.

Dissecting the Lobster: Hung-yi Lee Takes OpenClaw Apart Until Only Next-Token Prediction and a Few .md Files Remain

The first lecture of Hung-yi Lee's ML 2026 breaks OpenClaw into five questions: how an agent knows who it is, how it uses tools and SKILLs, how it remembers, how it runs on a schedule, and how it keeps working on its own for a long time. Every answer comes back to one fact: the language model only predicts the next token and starts fresh every turn. Identity, memory, and SOPs are all text files that OpenClaw puts into the prompt, or files the model reads and writes through tools. This post walks through the 60-slide intro.pdf and the lecture recording, including the defenses the slides recommend.

Reading NTU ML 2026: Positional Embedding — How Models Know Token Order and Handle Very Long Inputs

Self-attention on its own cannot tell "you hit me" from "I hit you", so the model needs position information from somewhere else. Hung-yi Lee's lecture goes from sinusoidal absolute positions to ALiBi and T5's relative biases, then to RoPE, which Llama, Qwen and Gemma all use. The second half covers train-short-test-long: RoPE breaks when it rotates to angles it never saw in training, which led to Position Interpolation, NTK-Aware scaling, YaRN, Dynamic Scaling and LongRoPE. The final twist is NoPE: causal attention in a decoder-only model already carries position information, and you can even drop the positional embedding after training.

Hung-yi Lee ML 2026 Self-Correction: Can a Model Fix Its Own Mistakes? What Changing Decoding, Workflow, or Weights Buys You

This lecture asks whether a model can catch and fix its own errors with no human in the loop. Hung-yi Lee splits the approaches into three routes. Change inference: the whole contrastive decoding family builds a version of the model likely to be wrong and subtracts it, and the methods differ only in how that wrong version is made. Change the workflow: appending "check again" sometimes helps but is unstable, external feedback beats self-reflection, and under a fixed compute budget, sampling more answers and voting often wins. Change the weights: teaching self-correction directly runs into "after training, the model makes different mistakes," which is why the field moved to RL. Whether RL teaches new abilities or just makes existing paths more likely is still being debated.

Reading NTU ML 2026: Self-Improving AI (Part 1) — AI-Generated Answers, Rewards, and Losses, and How Far Humans Can Step Back

Hung-yi Lee opens his May 8 lecture by admitting that "self-improving AI" has no clear definition: it is a process of humans gradually letting go. He splits machine learning into three steps and checks where the "I" can be replaced by AI. Answers can come from the model's own self-corrections, reward shaping can be written by an LLM, the loss can be set by the model itself (scores, majority vote, entropy), and even the questions can come from a proposer model. But experiments keep showing that with no human at all, progress plateaus or the model trains itself into the ground. A strong AI can already train a weaker one, just not better than humans do. His verdict: in May 2026, AI is "still standing at the bank of the Rubicon."

Reading NTU ML 2026: Can AI Improve Itself? (Part 2) — Improving the Harness, Improving the Improver, and Whether Growth Can Run Away

Part 1 was about an AI setting its own loss and updating its own parameters. Part 2 fills in the other half: AI Agent = Harness + LLM, and the harness can grow too. You can't take a gradient through a harness, so the usual move is to hand it to a language model as a rewriter and keep a pool of candidates, much like a genetic algorithm (OPRO, GEPA, Darwin Gödel Machine; DSPy if you want a ready-made tool). Three extensions follow: updating harness and parameters together beats updating either alone; when the goal changes you have to choose between discarding everything and carrying everything, and editing a harness can cause forgetting too; and the update rule itself can be updated (HyperAgent, Gödel Agent, SEAL), which is meta learning. Hung-yi Lee closes with a new analogy — parameters are genes, context is the neurons — then argues that today's agents lack intrinsic motivation, and that the likeliest source of runaway growth is a gap between the goal humans meant and the goal the AI inferred.

NTU AI/ML Course Guide: What Outsiders Can Actually Get from Hung-yi Lee, Hsuan-Tien Lin, and Yun-Nung Chen

National Taiwan University spreads its AI/ML courses across Electrical Engineering and Computer Science, and an official 'Machine Learning and AI' specialization stacks them into four levels. Hung-yi Lee's courses are the most open: ML 2026 Spring and Intro to GenAI and ML 2025 Fall publish slides, recordings, homework PDFs, and Colab notebooks, with only grading held back. Hsuan-Tien Lin's lectures are fully recorded and his Fall 2024 HW0–HW7 remain on the course page; Yun-Nung Chen's lectures are fully recorded, but most homework is public only as walkthrough videos; and since August 2025 Coursera only lets free learners watch the first module.

Taiwan AI Open Courses Beyond NTU: Using TAICA's Course Lists to Find NTHU, NCCU, NCKU, and NTUT Classes

Most public AI courses in Taiwan outside NTU come from TAICA, an alliance set up by the Ministry of Education. Each semester's course list says where every flagship course streams, and courses that stream on YouTube are usually watchable by anyone. Two courses are complete enough to self-study: Hung-Yu Kao's Natural Language Processing at NTHU (Fall 2025) and Yen-Lung Tsai's Generative AI at NCCU (Spring 2025), both A3. Wei-Ta Chu's Introduction to AI at NCKU, Ping-Hsuan Han's Human-AI Interaction at NTUT, and Min-Chun Hu's Robotic Navigation and Exploration at NTHU have full recordings but keep assignments on NTU COOL, so they rate A2. NYCU's TAICA courses are taught in English; the Deep Learning recordings are not publicly listed and Physical AI has just started, so neither made the main table.

CME295 Lecture 7: Agentic LLMs, or Letting the Model Look Things Up, Call Functions, and Run Its Own Loop

Lecture 7 of CME295 (2025) patches three LLM gaps: RAG fixes knowledge frozen at training time with a two-stage retrieve-then-rerank pipeline; tool calling fixes the inability to act by having a backend execute the function call the model writes; agents chain those calls with ReAct's observe-plan-act loop. The 2026 edition renames it AI Agents and adds context compaction, harness optimization, coding agents, and skills, the biggest rewrite in the course.

CME295 2026 Lecture 6 Preview: AI Agents, from Calling Tools to Managing Context and Tuning the Harness

The 2026 syllabus for CME295 Lecture 6 (November 6, 2026) lists seven topics. Tool calling, MCP and retrieval were already covered in the 2025 Lecture 7; the genuinely new ones are context compaction, harness optimization, coding agents, and skills/plugins. This pre-lecture edition explains those four using engineering posts from Anthropic and OpenAI, the MCP 2026-07-28 spec, and the Meta-Harness paper.

CME295 Lecture 9: Transformers Leave Text Behind, and LLMs Stop Writing Left to Right

The last CME295 lecture packs 128 slides into three parts: an eight-picture recap of the quarter, how Transformers handle images (ViT and two ways to build a VLM), and masked diffusion LLMs that emit several tokens per step, followed by what comes next in research and applications. It is not on the exam; the 2026 edition turns diffusion LLMs into a lecture of their own and refocuses Lecture 9 on multimodality.

CME295 2026 Lecture 8, Written Ahead: Three Kinds of Noise, One Training Objective, and the Price of Parallel Decoding in Diffusion LLMs

The 2026 edition of CME295 gives diffusion LLMs a full lecture (Lecture 8, November 20), with five listed subtopics: continuous, discrete and masked diffusion, training, and inference. This pre-lecture edition works from the original papers (DDPM, D3PM, SEDD, MDLM, LLaDA and others): continuous noise costs about 64x the compute on text, and the [MASK] absorbing state won out; the training objective is a masked cross-entropy weighted by 1/t; the speed comes from filling several positions per step, yet LLaDA's main results decode one token per step, and Fast-dLLM needs a confidence threshold plus an approximate KV cache to reach up to a 27.6x speedup.

CME295 Lecture 3: The Knobs You Turn When an LLM Generates, from Temperature and Top-p to Chain of Thought

CME295 Lecture 3 defines an LLM as a decoder-only next-token predictor, uses MoE to explain why a huge model only touches part of its weights per token, and spends most of its time on the knobs you can turn at generation time: greedy, beam search, top-k, top-p, temperature, guided decoding, plus three prompting techniques (few-shot, chain of thought, self-consistency). The 2026 edition folds this lecture into Lecture 2, and the prompting half disappears from the syllabus.

CME295 Lecture 8: Using LLMs to Judge LLMs, and the Three Biases to Guard Against

CME295 Lecture 8 starts from the fact that human rating is slow and expensive and BLEU/ROUGE can't recognize a paraphrase. It covers how LLM-as-a-Judge works, three biases (position, verbosity, self-enhancement) and six best practices, splits agent failures into tool prediction, tool execution and response generation, and closes with what MMLU, AIME, SWE-bench, HarmBench and τ-bench each measure, plus pass^k and Goodhart's law.

CME295 Lecture 6: How Reasoning Models Learn to Think Longer, and What GRPO Drops from PPO

CME295 Lecture 6 breaks reasoning models into three pieces: emit a reasoning chain before the answer, run RL on verifiable rewards like "is the answer correct," and use GRPO, which takes the group's average reward as the baseline instead of training a value model. RL alone took DeepSeek-R1-Zero from 15.6% to 71.0% pass@1 on AIME 2024, and distilling R1's traces into Qwen-32B beat running RL on the 32B model directly.

CME295 2026 Lecture 5 (Pre-Lecture Edition): LLM Systems, or How the Same Model Runs Several Times Faster

The 2026 edition of CME295 Lecture 5, "LLM systems" (October 30), lists seven topics: distributed training, inference optimizations, KV caching, speculative decoding, efficient kernels, FlashAttention, and hardware trade-offs. Written before the lecture, this post uses about 70 slides from the 2025 Lectures 3 and 4 plus the original papers to tie them into a single ledger: an H100 needs roughly 295 operations per byte moved to saturate its compute, while token-by-token generation does about 1 per byte of weights read, so most speedups are about moving less data.

CME295 Lecture 4: The Bill for Training an LLM, and Where Pretraining, SFT, and LoRA Spend It

CME295 Lecture 4 splits LLM training into two stages: pretraining on trillions of tokens (Llama 3 used 15 trillion), then SFT on thousands to millions of demonstrations so the model stops continuing text and starts answering. In between sits a map of memory savers (ZeRO, FlashAttention, mixed precision); the lecture closes with LoRA and QLoRA, which let people without big GPUs finetune, with QLoRA cutting VRAM by about 16x on a 65B model.

CME295 Lecture 5: SFT Can't Teach "Don't Answer Like That", So RLHF and DPO Add the Negative Signal

SFT only teaches a model to imitate good answers; it has no way to say which answers are unacceptable. CME295 Lecture 5 covers how to collect preference pairs, walks through the two steps of RLHF (a reward model trained on roughly 10,000 human labels, then PPO on roughly 100,000 examples), and ends with DPO, which folds the whole RL pipeline into a single supervised loss. The 2026 edition splits this lecture between Lecture 3 (training) and a new Lecture 4 (reinforcement learning).

CME295 2026 Lecture 4 (Pre-Lecture Edition): SFT, PPO, GRPO, and On-Policy Distillation Are One Policy Gradient

The 2026 CME295 Lecture 4 (October 16) gives RL its own lecture, with seven syllabus items: mathematical conventions, reward design, policy gradients, limitations, PPO, GRPO, and on-policy distillation. This post walks the math ahead of class: start from ∇log π times a score. SFT uses a score of 1, PPO estimates it with a value model, GRPO uses the group mean, and on-policy distillation uses the teacher's per-token log-prob gap. In the Qwen3 report, starting from the same checkpoint, RL reached 67.6 on AIME'24 with 17,920 GPU hours, while on-policy distillation reached 74.4 with about 1/10 of that (1,800 hours).

CME295 Lecture 1: From Tokens to Transformer, or How One Sentence Gets Translated into Another Language

CME295 Lecture 1 threads a single sentence, "A cute teddy bear is reading.", through the whole class: split it into tokens, turn them into vectors, see why an RNN can't hold on to long sentences, then translate it into French with self-attention and an encoder-decoder. The 2026 edition drops the entire section on NLP tasks and evaluation metrics and opens instead with a timeline running from 2017 to the agent era.

CME295 Lecture 2: How One Transformer Grew into BERT, GPT, and a Zoo of Attention Variants

CME295 Lecture 2 takes the original Transformer apart and refits it: position information moves from "added to the embedding" to RoPE's "rotate Q and K inside attention"; attention gets cheaper with sliding windows and MQA/GQA; models split into encoder-only, encoder-decoder, and decoder-only families; and the second half dissects BERT's MLM (15% of tokens) and NSP pretraining. The 2026 edition folds all of this into a single "Large Language Models" lecture, and BERT is no longer a syllabus item.

Reading CMU 11-768 A1: Build an Agent Harness by Hand — One ReAct Loop to Fix Bugs, Compact Context, and Play Chess

CMU 11-768's first assignment starts from an empty ReAct loop: a bash-only CodeAgent fixes a bug in a chess app, context compaction is added to solve a SWE-bench task, and the same loop becomes a ChessAgent that runs a two-ply search with simulate_move, run_python, and a skill. All 100 points are graded by replaying submitted patches and trajectories offline.

Reading CMU 11-768 A2: Writing a Validator for a Data-Visualization Agent — Four Error Families, MCC, and Harbor Verifiers

11-768's Assignment 2 has students use one fixed judge, Qwen3-VL-30B-A3B, to flag four error families in every run of a data-visualization agent, graded by the mean MCC across families on a private set (30% of the assignment). The second half packages students' own tasks as Harbor environments, with one wrong solution the verifier rejects and one that fools it. The theme: the grader you write becomes the RL reward later.

Reading CMU 11-768 AI Agents: An Agents Course You Can't Take Without Having Trained a Language Model, With Three Assignments From Harness to Eval to RL

CMU 11-768 is a new Fall 2026 graduate course on agents taught by Graham Neubig and Daniel Fried. The prerequisite — prior experience training language models — is strictly enforced. Its 23 lectures run from tool calling, context, memory, and planning through SFT, RL, sandboxing, and human-agent interaction. Three individual assignments in the first half build a harness, an evaluation, and a training pipeline; the second half is a team research project. Slides and the first nine lecture videos are public.

CMU 11-768 Lecture 1: An Agent Is a Model in a Loop — the Hard Part Is Making It Work

Lecture 1 of 11-768 strips an agent to its minimum: tool definitions and tool calls are just tokens, the harness parses, executes, and feeds results back into context, and running a ReAct loop makes it an agent. Neubig then lists six capabilities a good agent needs, each of which can be built through training or through the harness, and argues that an agent is a system of harness, sandbox, inference, training, and monitoring — not just a model.

Reading CMU 11-768 L2: How Tool Use Turns Tokens into Actions — Schemas, Constrained Decoding, MCP, and Parallel Calls

Neubig splits tool use into five layers: capabilities, mechanics, constraints, interfaces, and systems. A tool call is just tokens the model emits; the harness parses, validates, and matches results back by call ID. Constrained decoding guarantees form, not correctness. MCP's real value is credential brokering. And the same model served by different providers can swing from roughly 15% tool-call errors to under 0.1%.

Reading CMU 11-768 L3: How Long-Context Agents Manage Memory — Hybrid Attention, RoPE Extension, Prompt Caching, and Compaction

An agent resends its whole history on every call, so five calls already add up to 80K input tokens; 1,500 OpenHands sessions averaged 78K tokens, 37% of them tool results. Neubig works on two layers: at the model layer, hybrid attention (many local layers, one global) plus length curricula make million-token context possible; at the harness layer, stable prefixes earn cache reads roughly ten times cheaper, and compaction that keeps anchors and externalizes evidence gets past the limit — evaluated by how the agent continues afterwards.

Reading CMU 11-768 L4: Skills and Memory — How Agents Stop Starting Over

Lecture 4 of CMU 11-768 sorts cross-task experience into episodes, facts, and skills, stored as external artifacts rather than in context or weights. Human-written skills load through SKILL.md and progressive disclosure, lifting the average SkillsBench pass rate from 33.9% to 50.5%. Skills an agent induces itself can be tested before admission when written as code, but break easily on a new website. The hard part is the lifecycle: imperfect judges, over-retrieval, and bloated skill libraries each eat into the gains.

Reading CMU 11-768 L5: Planning — When an Agent Should Think It Through, and When It Should Revise as It Goes

Lecture 5 of CMU 11-768 defines an agent's plan as an explicit, inspectable, revisable representation of intended behavior for this task, and gives four reasons to add planning structure: modularity, environment feedback, long horizons, and control. Fried's own MACU has a manager decompose tasks into a DAG and dispatch parallel sub-agents, raising Odysseys success from 8.5% to 34.0%; on an OSWorld subset, no planning scores 25.0%, an initial DAG with no revisions scores 27.8%, and allowing 10 revisions reaches 58.3%.

Reading CMU 11-768 L6: Coding Agents — From Completing a Line to Fixing a Whole Repo

Neubig's L6 splits coding agents into three layers: train a model that can code (pre-training, mid-training, infilling, RL from test rewards), wrap it in a localize–edit–verify loop with the right editing tools so it can change a repo, then evaluate and train it in SWE-bench-style executable environments. Fixing bugs is only about 15% of a developer's day; the next frontier is tests, CI, and maintenance in the outer loop.

Reading CMU 11-768 L7: How Computer Use Agents See the Screen, Get Graded, and Get Trained

JY Koh breaks computer use agents into three questions: evaluation has moved from single clicks (ScreenSpot-Pro, Mind2Web) to programmatic end-state checks (WebArena, OSWorld), VLM judges, and long-horizon rubrics (Odysseys, OSWorld 2.0); the model is a VLM reading interleaved screenshots and actions; training runs pre-training for grounding → SFT on human and synthetic trajectories → RL in resettable simulated environments.

Reading CMU 11-768 L8: How to Do SFT for Agents — Loss Masks, Trajectory Selection, Data Formats, and the Handoff to RL

Yueqi Song breaks agent SFT into six decisions: compute loss on assistant tokens only (including the stop token); choose trajectories carefully (runs that pass tests can still teach bad habits, and switching teachers or adding new tasks beats sampling more); unify formats with the Agent Data Protocol; watch packing and template drift during training; evaluate in the real harness; and pick the SFT checkpoint for the RL that follows, not for its own best score.

CMU 11-768 Lecture 9: RL Basics — a Policy Gradient Is Just an SFT Loss Times a Weight

Using a guess-a-number-from-1-to-16 game, Daniel Fried frames RL as an extension of SFT: he names SFT's three gaps (task mismatch, no learning from failures, never seeing its own mistakes), then derives ReST, REINFORCE, baselines, and GRPO/DrGRPO in turn. All four compute log-probabilities of the tokens the agent itself generated; they differ only in the weight each token gets.

Reading CMU 11-768 L10: How to Evaluate, Train, and Retrieve for Deep Research Agents

Akari Asai's L10 splits deep research agents into three problems: evaluation has to cover four gaps (search difficulty, domain expertise, long-form answer quality, citation support); training runs mid-training → SFT → RL, with DR Tulu's evolving rubrics as the reward for long-form reports; retrieval should let the retriever see the agent's reasoning, which gets AgentIR-4B to 68% on BrowseComp-Plus with Tongyi-DR.

CMU 11-768 Lecture 11: Advanced RL Algorithms — Credit Assignment, Stable Updates, Reward Hacking, and Distillation

Using a bug-fix coding task, Graham Neubig takes Lecture 9's policy gradient into practice: a critic, GAE, or a PRM to credit individual turns; importance ratios and clipping to handle stale data in async RL; and PPO, GRPO, CISPO, GSPO, and DAPO side by side in one table. The largest share goes to the reward itself — verifier errors, reward hacking, and exploration collapse — before closing with on-policy distillation.

CS224U Analysis Methods I: Probing Shows You Representations, Feature Attribution Gives You Causal Guarantees

The Analysis methods unit of CS224U (Spring 2023) starts by grading three families of methods on a three-column scorecard. Probing is strong at characterizing representations but can't support causal claims. Integrated gradients only gives you a scalar about each representation, but it satisfies the sensitivity axiom, so it does come with a causal guarantee. This post covers slides 1–40, videos 33–35, and feature_attribution.ipynb, including where the notebook breaks in today's environment.

CS224U Behavioral Evaluation: Analytical Considerations, Adversarial Tests, ANLI, and DynaSent

CS224U's fourth unit opens with one question: what can behavioral testing prove, and what can't it? It can never give a guarantee, and when a model fails you first have to ask whether the model or the dataset is at fault. BERT scored 2.2% on negated NLI examples, then 90% after fine-tuning on a small set of them. The unit then covers SQuAD distractor sentences, Breaking NLI, ANLI's human-and-model adversarial collection, and ends with DynaSent's two rounds.

CS224U Analysis Methods II: From "The Information Is There" to "The Model Uses It" with Interchange Interventions — Causal Abstraction, IIT, and DAS

Causal abstraction rests on one operation: take the internal state a model computes for a source input at some location, swap it into the same location for a base input, and check whether the output changes the way your hypothesized high-level program says it should. In CS224U's iit_equality.ipynb, a network with 0.99 test accuracy scores only 0.50 and 0.54 on this check. After IIT training, its counterfactual accuracy is 1.00. DAS replaces guessing which neurons match which variable with learning a rotation matrix.

CS224U Compositional Generalization: COGS, ReCOGS, and Assignment 3

A few COGS generalization splits score 0 for nearly every model. CS224U uses its own ReCOGS work to explain why: the zeros on the recursion splits are mostly a length-generalization problem, and the zeros on the prepositional-phrase split come from training data that only ever put PPs in certain variables and positions. Assignment 3, hw_recogs.ipynb, uses 135K ReCOGS training pairs. It first has you find Charlie and Lina, two names whose train and test roles are exact opposites, then shows a trained model stumbling on them.

CS224U Contextual Representations II: What GPT, BERT, RoBERTa, ELECTRA, T5, BART, and Distillation Each Change in Pretraining

CS224U Spring 2023 tells the story of the Transformer families through BERT's four known limitations. RoBERTa addresses the first (optimization was only partly explored). ELECTRA addresses the second and third (the [MASK] mismatch, and only about 15% of tokens giving a learning signal per batch). XLNet addresses the fourth (the assumption that masked tokens are independent of each other). GPT changes the objective and the mask, T5 and BART change the architecture and how inputs are corrupted, and distillation changes model size. The course's 2023 view: autoregressive architectures have taken over, but bidirectional models may still have the edge for representation.

CS224U Contextual Representations I: "Break" Has Eight Meanings, and How the Transformer Lets Each Word Read Its Context

The first three parts of CS224U's Spring 2023 contextual representations unit start with examples like "break" and "crane" to show why static word vectors were never going to be enough. They then build a Transformer block step by step on the three words "The Rock rules." Only attention connects the columns; every other step runs on each column independently. Finally, two questions sort three positional encoding schemes: Do you have to fix the set of positions ahead of time? Does the scheme get in the way of generalizing to new positions? Absolute encoding fails both, sinusoidal encoding passes the first, and the relative encoding of Shaw et al. (2018) passes both.

CS224U Methods and Metrics II: Datasets, Data Splits, and Comparing Models

The second half of CS224U's 'NLP methods and metrics' unit skips metric formulas. It asks whether your experiment holds up. Naturalistic or crowdsourced data, adversarial or common cases: the course answers 'both' each time. Lock the test set away. Pick baselines when you write the hypothesis. Compare two models with confidence intervals, Wilcoxon, or McNemar, and run several random initializations. The slides, three videos, and two notebooks are all public. Kawin Ethayarajh's guest session 'Real-world NLP assessments' has no public slides or video.

Reading Stanford CS224U, Part 17: Two Extension Lectures — Generating Text with Diffusion, and Turning LLM Training Folklore into Intuition

In Spring 2023, CS224U slipped two talks by members of its own teaching team into the Transformer unit. Lisa Li presented Diffusion-LM: instead of generating left to right one word at a time, it denoises a sequence of Gaussian vectors into word vectors. It loses to autoregressive models on both training and decoding efficiency, and in exchange lets a classifier's gradient steer the output at every step. Sidd Karamcheti showed how to cut GPT-2 Small's single-GPU training clock from 99.63 days to 3.37 days by stacking data parallelism, mixed precision, and ZeRO. The first talk survives only as slides; the second has slides and two recordings.

CS224U Homework 1: Multi-Domain Sentiment and the Bake-Off

CS224U's first assignment, hw_sentiment.ipynb, is ternary sentiment classification: you develop on two rounds of DynaSent plus SST-3, and the bake-off test set mixes in mystery sentences from undisclosed sources. The original-system question is worth 3 of the 9 homework points, and it has exactly one rule: never touch the three public test sets during development. Run as-is today, the first data-loading cell breaks because Hugging Face datasets 4.0 dropped trust_remote_code.

CS224U Assignment 2: Few-Shot OpenQA with DSPy

CS224U's second assignment, hw_openqa.ipynb, asks you to answer questions that come with no passage, using only a frozen language model and a frozen ColBERT retriever. The Spring 2023 version was written for DSP; in January 2024 the repo switched to DSPy and pinned dspy-ai==2.4.13. Before you start you need an OpenAI API key, a ColBERTv2 checkpoint of about 406 MB, and a 600 MB prebuilt index. The notebook's first setup call, dspy.OpenAI, no longer exists in DSPy 3.4.

CS224U In-Context Learning: Origins, Core Concepts, and Suggested Methods

The Spring 2023 edition of CS224U defines in-context learning as a frozen language model performing a task only by conditioning on the prompt, and warns that the second condition of few-shot learning (no examples of the behavior seen in training) is almost impossible to verify. Potts's 38-page deck runs from GPT-2's TL;DR trick through choosing demonstrations, chain of thought, self-consistency, and DSP, and ends with four recommendations: build dev/test sets first, learn your target model's instruction format, and treat prompt writing as AI system design. Mina Lee's guest lecture asks the reverse question: who should learn to read prompts, people or models?

CS224U Information Retrieval: From Classical IR and IR Metrics to Neural IR

The Spring 2023 edition of CS224U spends a whole unit on retrieval, because OpenQA gives you only the question and you have to find the evidence yourself, while large language models fabricate sources. The slides by Potts and Omar Khattab go from TF-IDF and BM25 through Success@K, MRR, and average precision to four neural IR designs (cross-encoder, DPR, ColBERT, SPLADE) and how each trades expressiveness against scale, and they end by asking you to count latency and cost as metrics too.

CS224U Opening Lecture: One Question Asked for Forty Years, and How a 2023 NLU Course Defines Understanding

The first CS224U lecture of Spring 2023 asks "Which U.S. states border no U.S. states?" of every system from Chat-80 (1980) to text-davinci-001. The answers show that the progress is real. The lecture then questions whether that progress counts as understanding, using Levesque's "cheap tricks," models that invent links, and benchmarks that saturate within a year or two. That splits the course map in two: the first half teaches you to build systems with Transformers and retrieval-augmented in-context learning, and the second half teaches you to test them with harder benchmarks, behavioral evaluation, and causal explanation methods.

CS224U Final Project Workflow: Lit Review and Experiment Protocol

The first two deliverables of the CS224U final project are a literature review and an experiment protocol. The lit review covers 5, 7, or 9 papers depending on team size, under five suggested sections. The protocol has seven required sections, and its core is a hypothesis you can state. The course supplies a six-step paper-search loop, a rule that AI-assistant output must be quoted, and a worked example: a student's final project that became a Findings of EMNLP paper. The Gradescope format and rubric slides, and past exemplary papers, are behind a login.

CS224U Methods and Metrics I: A Classifier with 0.81 Accuracy and 0.43 Macro F1 — What Classifier and Generation Metrics Each Encode

The CS224U slides compute two numbers from one three-class confusion matrix: accuracy 0.81 and macro F1 0.43. One says the system is good; the other says it gets the two small classes almost entirely wrong. The unit's claim is that different metrics encode different values, and it goes through the bounds, values, and weaknesses of accuracy, the three F-score averages, perplexity, word error rate, and BLEU. Final projects are graded on whether the metrics fit, not on how high the scores are.

CS224U: Writing NLP Papers, Submitting, and Giving Talks

CS224U's 'Presenting your research' lecture has four parts: the course-specific rules for the final paper, how to write an NLP paper, how conference submission works, and how to give a talk. Three things matter most. The final paper must include Known project limitations and an Authorship statement. Write as a Shieber-style 'rational reconstruction,' not a chronological tour of your dead ends. At submission, your title largely decides reviewer bidding. The slides, four videos, and projects.md are public; past example papers need a Stanford login.

aideep-dive

Learn Inference: Inference Engineering, Rebuilt with Dials You Can Turn

learn-inference.com is an unofficial interactive companion to Philip Kiely's Inference Engineering (256 pages, Baseten Books, free PDF). It follows the book's 8 chapters and 42 sections with rewritten explanations, turns intuition-heavy ideas like TTFT, P99, speculative decoding, and prefix-cache routing into slider-driven simulators, and ships a keyless JSON API and MCP server.

Reading Stanford CME295: Two Units, No Homework, Nine Lectures from Transformers to AI Agents

CME295 is a two-unit Stanford course with no homework; your grade is the midterm and the final, 50% each. The 2025 edition's nine lectures are fully public: videos, slides, and both exams with solutions. The 2026 edition rewrites the agent lecture around context compaction, harnesses, coding agents, and skills, and adds three full lectures on LLM systems, reinforcement learning, and Diffusion LLMs.

Three Versions of CS189: Spring 2026 as the Base, Spring 2025 as the Classic, Fall 2026 in Progress

Several semesters of Berkeley CS189 are online at once. From here on, this series follows Spring 2026 (Listgarten/Dimakis): its slides, 25 lecture videos, discussions with solutions, HW1–5 handouts and midterm solutions are all publicly accessible, so it rates A3. Spring 2025 (Shewchuk) is a different, classic route and the only one that covers SVMs, decision trees, PCA and boosting. Fall 2025's homework folders open empty to outside readers, so it was not chosen. Fall 2026 is still running and serves only as a comparison.

CS189 Spring 2026 HW1 Guide: Linear Algebra / Calculus / Probability Warm-up + Fashion Coding

CS189 Spring 2026 HW1 has three pieces: ten written math warm-ups (a linear system, limits of matrix powers via eigendecomposition, SVD, a matrix that flips an image, partial derivatives, a chain rule over a recursion, and four probability problems including Bayes for cancer screening), plus two public Modal notebooks on Fashion-MNIST. Part 1 drills pandas / Plotly / K-means / an MLP / matrix-based image augmentation / tensor puzzles; Part 2 does price regression, MAE / MSE / R², confusion matrices, and finally a secret test set whose images have been rotated. Due Feb 20; no official solutions.

CS189 Spring 2026 HW2 Guide: Chatbot Arena Paper Questions, Regression, MLE/MAP, From GMMs to Flow Matching

HW2 is a written-only assignment with 10 problems. The first half practices reading a paper (Chatbot Arena) and the core derivations for regression and MLE/MAP. The two heaviest problems come last: Mixed Feelings goes from k-means' fragility to outliers through robust k-means to a GMM with a uniform background, and Watch Me Flow Dat proves that conditional flow matching and a discretized MLE are the same objective. The problem PDF and LaTeX template are freely downloadable; there are no official solutions.

CS189 Spring 2026 HW3 Guide: Autograd from Scratch (BearTensor), Newton's Method, and the Information Bottleneck

HW3 has two halves. The four written problems run from Newton's method for logistic regression and a convergence analysis of coordinate descent, through backprop, VJPs, and implicit differentiation, to an information-bottleneck view of what deep networks compress. The notebook has you build a BearTensor computation graph in NumPy, topological-sort backprop, and SGD/Momentum/Adam, then train a red-wine quality regressor with it, plus an optional Muon optimizer. Problems and notebook are public; official solutions and hidden tests are not.

CS189 Spring 2026 HW4 Guide: ResNet/Transformer Paper Questions and Implementing CNN, ResNet, Transformer, DNABERT, and ConvNeXt

HW4 has three pieces. The written part has you read ResNet and Attention Is All You Need in the order problem → existing work → proposal → method → contribution. The 4.1 notebook builds a CNN and ResNet-18 in PyTorch, then assembles an encoder-decoder transformer step by step from softmax, trains it on TinyStories, and generates stories. The 4.2 notebook cuts DNA into 6-mers for a pretrained DNABERT to classify species, and turns audio into spectrograms for ConvNeXt, comparing training from scratch, a frozen backbone, and full unfreezing. Two Kaggle competitions; due 5/1. Outside readers get the problems but not the course data bundle or tests.

CS189 Spring 2026 HW5 (Optional) Guide: InfoNCE for Biology, Diffusion Theory, LLM Fine-Tuning + Kaggle

HW5 is the only Spring 2026 assignment marked optional, due 5/11, the same day as the final. The written part has three pieces: derive InfoNCE gradients and the trade-off in the number of negatives, using scRNA-seq as the setting; prove the optimal denoiser is a conditional expectation and derive the continuity equation; generalize flow matching's straight-line path to arbitrary interpolations. The notebook is a full LLM fine-tuning pipeline: Qwen2.5-0.5B-Instruct is fixed, MMLU machine_learning is converted to chat format, TRL's SFTTrainer does full fine-tuning, accuracy on CS189 exam questions is compared before and after, and predictions on a 169-question test set go to Kaggle, all while guarding against catastrophic forgetting. An official hw5-sol.pdf is provided, covering the written part only.

CS189 Spring 2026 Lec 1–3: ML Problem Framing, Data Tools, Terminology and Techniques

The first three lectures of CS189 Spring 2026 hold off on derivations. They teach you how to tell whether a problem calls for ML, how to look at data with pandas and Plotly, and how to run one full train/validate/test cycle in scikit-learn. Lecture 1 has slides but no recording; Lectures 2–3 have slides and video; Discussion 1 is a calculus, linear algebra and probability warm-up with solutions and a walkthrough. Together they set you up directly for HW1.

CS189 Spring 2026 Lec 4–7: K-means, Probability Review, MLE, Multivariate Gaussians and GMMs

Lectures 4–7 of CS189 Spring 2026 tie unsupervised learning into one thread. K-means clusters the data, then its weaknesses show up: hard assignments, no probabilistic framing, and a bias toward round clusters of similar size. The course reviews probability, introduces maximum likelihood estimation (MLE) and multivariate Gaussians, and rewrites K-means as a Gaussian mixture model (GMM). The GMM log-likelihood has no closed-form solution, and that gap leads the course to gradient descent. All four lectures have slides and video; Discussions 2–3 come with solutions and walkthroughs.

CS189 Spring 2026 Lec 7–10: Linear Regression, the Geometry of Least Squares, and Regularization

Over four lectures, CS189 Spring 2026 presents linear regression from three angles that meet in one formula: MLE under Gaussian noise is least squares; the least-squares solution is the orthogonal projection of y onto the column space of X; and when features are collinear or too many, ridge (the MAP estimate under a Gaussian prior) or lasso (a Laplace prior) pulls the solution back, with λ chosen on a validation set. Slides, videos, and Discussion 3–4 solutions are all publicly accessible.

CS189 Spring 2026 Lec 11–12: Classification, Generative Classifiers, Logistic Regression, and ROC

CS189 Spring 2026 Lec 11–12 splits classification into two routes. Generative models fit p(x|y) for each class (GDA: shared covariance gives LDA and a linear boundary, per-class covariance gives QDA and a quadratic one). Discriminative models fit p(y|x) directly (logistic regression: sigmoid, softmax, cross-entropy MLE, no closed form, so gradient descent). The bridge: the LDA posterior can always be written in logistic form, but not the other way around. For evaluation, accuracy misleads under class imbalance; ROC/AUC sweeps every threshold and ignores calibration; PR curves care about class balance.

CS189 Spring 2026 Lec 13 & 15: Convergence, Momentum, Adam, SGD

CS189 Spring 2026 covers gradient descent in two lectures. Lec 13 derives the learning-rate limit and the condition number from the Hessian's eigenvalues, then moves through momentum, learning-rate schedules, AdaGrad/RMSProp/Adam, and mini-batch SGD. In Lec 15, Dimakis walks through the same material again, starting from a gradient computed by hand on a small data table. Both slide decks, the recordings, Lec 13's handwritten notes, and Discussions 6 and 7 (with solutions) are all publicly accessible.

CS189 Spring 2026 Lec 14 & 16: MLE vs MAP, Bias-Variance, Entropy and KL, Plus a Midterm Self-Check

Lec 14 recasts ridge as MAP under least squares plus a Gaussian prior, breaks down bias-variance, and closes by using Chatbot Arena to show how to read a paper. Lec 16 builds entropy from compression, moves on to KL and cross-entropy, and lands back on the logistic regression loss. The 3/17 midterm and its official solutions are public in the past-exams folder, with 6 problems, 56 points, 110 minutes, and 6 problem-by-problem walkthrough videos, so you can sit it as a mock exam.

CS189 Spring 2026 Lec 17–18: Depth, Universal Approximation, Activations, and Backpropagation

Lec 17 uses XOR to show that a linear model, and even a stack of linear layers, cannot learn a nonlinear boundary, while one ReLU layer can. The universal approximation theorem guarantees that a network exists but not how to find its weights or how wide it must be. Lec 18 turns the chain rule into backpropagation on a computation graph: gradients from multiple paths add up, and the cost is linear in the number of parameters, versus quadratic for finite differences. Discussion 8 has you prove a GD convergence rate and 1-D ReLU universal approximation yourself.

CS189 Spring 2026 Lec 19–20: Initialization, BatchNorm, CNNs, Early Stopping, and Double Descent

Lec 19 wraps up backprop, then tackles how to keep gradients flowing: all-zero initialization makes every unit learn the same thing, so use small random values (He init for ReLU), and batch norm normalizes pre-activations with mini-batch means and variances. The second half introduces CNNs: local connectivity plus weight sharing lets one feature detector scan the whole image. Lec 20 finishes pooling, receptive fields, and CNN training, then covers early stopping, dropout, and double descent, which breaks the classic bias-variance picture.

CS189 Spring 2026 Lec 21–22: Transformers

Lec 21 starts with what CNNs lack: only the top layers see the whole image. It then builds up through TF-IDF, RNN image captioning, and soft attention. Lec 22 derives self-attention from a "soft dictionary lookup": three linear layers produce Q, K, and V, the output is SoftMax(QKᵀ/√D)V, and multiple heads, an MLP, residual connections, and LayerNorm turn it into a transformer layer. Attention itself ignores order, so you need positional encodings. Discussion 10 has you compute QKV by hand and prove why the scores are divided by √D.

CS189 Spring 2026 Lec 23–24: LLM Training and Applications, Self-Supervised Learning

Lec 23 wires a transformer into a next-token predictor: tokenize, look up embeddings, stack L layers of masked attention, multiply back by the embedding table and apply softmax, and train with cross-entropy (that is, MLE). Pretraining supplies knowledge; to chat, a model also needs SFT, LoRA, RLHF, or DPO, and at inference time it leans on in-context learning, RAG, chain-of-thought, and tool calls. Lec 24 generalizes "invent a fake supervised task" to images: autoencoders, colorization, inpainting, rotation, jigsaw puzzles, clustering, and finally contrastive learning, SimCLR, and CLIP. Discussion 11 practices positional encodings, RoPE, causal masks, and the KV cache.

CS189 Spring 2026 Lec 25–27: AI for Protein Engineering, Agents and Environments, and Where to Go Next

The last three lectures take the semester's tools to two frontiers. Lec 25 is about proteins: AlphaFold2 cracked sequence-to-structure, but the real engineering bottleneck is predicting which sequence has the function you want, and design means acting as your own model's adversary in a discrete space of size 20^L. The slides reduce conditional generation p(x|y) to three statistically correct routes, all of which come back to Bayes' rule. Lec 26 was an online guest lecture with no public materials. Lec 27 defines an agent (an LLM in a loop, using tools, deciding its next step) and argues that data is being replaced by environments: Docker + task + verifier, used for SFT, RL (RLVR, GRPO), or weight-free GEPA. For final-exam practice, use the Fall 2025 and Spring 2025 finals with solutions; the Spring 2026 final is not published.

CMU 07-380 HW1 Guide: Logic and the Hybrid Wumpus Agent, Logical Inference Plus A* Planning

The HW1 programming assignment turns Lec2's entailment into a Pacman take on the Wumpus agent. Q1–Q2 warm up with Expr and pycosat. Q3–Q5 write the PKE percept rule, build the KB, and use two SAT calls to decide SAFE, NOT_SAFE or UNSURE. Q6–Q7 use the provided A* helpers to build an exploration agent and a three-tier hybrid agent. The starter code and local autograder are public; the Gradescope online questions are CMU-only. No solutions here.

Reading CMU 07-380 HW2: Classical and Motion Planning, from Robot-Cook PDDL to RRT* to Graphing LPs

07-380 HW2 has three parts. The programming assignment has you write PDDL for a pancake-cooking robot, solve it optimally with unified-planning and Fast Downward, then implement RRT and RRT* in rrt.py (Q2–Q7). The written part covers GraphPlan, one LP modeling problem, and two LP graphing problems. A Gradescope online component is CMU-only. This guide covers structure, prerequisites, and running the local autograder; it contains no solutions.

CMU 07-380 HW3 Guide: Optimization, Writing Your Own LP Solver and Branch and Bound, With PCA and MAP on Paper

HW3 has three parts. The programming part has you build an LP solver by vertex enumeration, stack branch and bound on top of it for integer programs, and formulate three word problems. The written part covers integer programming by hand, the ethics of Amazon's delivery routing, PCA via SVD, and a proof that a Laplace prior equals L1. It is due 10/1, so this guide explains structure and concepts only, with no solutions.

CMU 07-380 Lecture 1 Guide: Introduction, and What AI & ML II Adds After 07-280

07-380 Lec1 has no algorithms. It sets up three things: the working definition that intelligence means doing well on a task under uncertainty, a bubble diagram that color-codes 07-280 topics against the new 07-380 topics, and a grading scheme with quizzes at 55%, no final exam, and a final project. Off-campus readers can use it to place the other 25 lectures.

CMU 07-380 Lecture 2 Guide: Logical Agents, Proving a Square Safe with Model Checking, DPLL and Forward Chaining

Lec2 turns 'is this square safe?' in Minesweeper and Wumpus World into an entailment question: KB ⊨ α exactly when KB ∧ ¬α is unsatisfiable. Three ways to answer it: TT-ENTAILS, which enumerates every model; DPLL, which adds early termination, pure symbols and unit clauses to backtracking; and forward chaining, which accepts only definite clauses and runs in linear time. Resolution sits in the appendix, marked out of scope.

Reading CMU 07-380 Lecture 3: Classical Planning, PDDL, State-Space Search, and Relaxation Heuristics

07-380 Lec3 replaces propositional successor-state axioms with STRIPS actions (pre/add/del sets), which turns planning back into state-space search. When that search is too large, GraphPlan lets actions run in parallel and never deletes facts, and delete relaxation drops delete effects entirely, yielding the heuristics behind FF and Fast Downward.

Reading CMU 07-380 Lecture 4: Motion Planning, RRT Samples Its Way Through Continuous Space

The second half of 07-380 Lec4 moves planning into continuous configuration space. States can no longer be enumerated, so RRT samples a random point, extends the nearest tree node a short step toward it, and checks the whole segment for collisions. RRT is probabilistically complete but not optimal; RRT* uses tree path costs to pick a better parent and rewire neighbors, so the path converges to optimal as samples grow.

Reading CMU 07-380 Lecture 5: Linear Programming, and Why the Optimum Sits at a Vertex of the Feasible Region

07-380 Lec5 turns the Diet Problem from words into min cᵀx s.t. Ax ⪯ b, then draws it: each constraint is a half-plane, the cost is a direction, and cost contours are perpendicular to c. Push a contour in the −c direction until it last touches the feasible region and you always hit a vertex, so solvers only need the intersections of constraint boundaries. Vertex enumeration checks them all; simplex walks greedily from one vertex to a better neighbor.

Reading CMU 07-380 Lecture 6: Integer Programming, Relax to an LP and Branch and Bound

07-380 Lec6 adds one constraint to an LP, x ∈ ℤᴺ, and the vertex solution may no longer be an integer. Searching the integer points near the LP solution is not guaranteed to work either. The fix: drop the integer constraint (relaxation) to get an LP lower bound, split on a fractional coordinate into xᵢ ≤ floor and xᵢ ≥ ceil, and keep every subproblem in a priority queue ordered by LP objective. The first all-integer solution popped is optimal.

Reading CMU 07-380 Lecture 7: Low Rank Optimization, PCA's Reconstruction Error, Projected Variance, and LoRA

07-380 Lec7 frames PCA as low-rank optimization: approximate the data with a matrix of rank at most r. For a unit vector v, each point's reconstruction error equals ‖x‖² minus the squared projection length, so minimizing reconstruction error and maximizing projected variance are the same problem. Lagrange multipliers show the answer is an eigenvector of the covariance matrix, which you can also read straight off the V in the SVD. The site lists the LoRA paper as reading; its ΔW = BA applies the same low-rank idea to weight updates.

CMU 07-380 Lecture 8 Guide: MAP, How Priors Enter Estimation and Why That Equals Regularization

Lecture 8 swaps MLE's argmax p(D|θ) for argmax p(θ|D). The prior p(θ) multiplies the likelihood, and after taking the negative log it becomes an extra term in the objective. A trick coin shows data overwhelming the prior, a Beta prior estimates a click rate, and Gaussian and Laplace priors on linear-regression weights turn into L2 and L1 regularization.

CMU 07-380 Lecture 9 Guide: Probabilistic Generative Models, Naive Bayes and Gaussian Discriminant Analysis

Lecture 9 stops learning p(y|x) directly. Instead it learns the class prior p(y) and the class-conditional p(x|y), then inverts them with Bayes rule. The price is stronger assumptions; the payoff is the ability to generate new data and more stability with little data. Naive Bayes uses conditional independence to make the parameters estimable, GDA uses multivariate Gaussians for continuous features, and whether the covariances match decides a linear or a curved boundary.

CMU 07-380 Lecture 10 Guide (Pre-reading Edition): Bayes Nets Break a Joint Distribution into Conditional Probability Tables

The Lec10 slides are not on the 07-380 course site yet, so this guide uses only the PR6 Bayes Nets pre-reading and the 15-281 Bayes Net Demo. A joint distribution can answer any query, but nobody hands it to you and it is too big to store; a Bayes net writes it as a product of 'node given parents' tables, and every missing edge is an independence assumption.

CMU 07-380 Stage Recap: From Reasoning Under Certainty and Optimization to Uncertainty

The course site's schedule splits 07-380's first ten lectures into Reasoning Under Certainty, Optimization and Reasoning Under Uncertainty: prove things with logic and plan with search, then write problems as constrained objectives, and finally let a prior in with MAP and turn to probabilistic models. HW1 tests logic plus search, HW2 planning plus LP graphing, HW3 writing solvers plus PCA and MAP derivations.

Harvard CS181 Final Checkpoint and Series Wrap-up: Checklist, Practice Problems, and a Practical Backup

The public materials for the CS181 2026 final (May 9) — the final checklist, 16 second-half practice problems, and a 66-page final review — are all 2025 versions. They cover Bayes nets and EM, which the 2026 schedule never lists, and skip Transformers, VAEs, GANs, and autoregressive models. Sort the checklist against the 2026 schedule into three buckets, cover the new topics with homework and sections, then do the 2025 practical for one end-to-end project.

Harvard CS181 HW2: Classification, Bias-Variance, and Telling Two Kinds of Uncertainty Apart

CS181 Spring 2026 HW2 (due Feb 27) has four problems worth 90 points: train 10 logistic models on planet observations to see bias and variance, derive the MLE of a generative classifier, implement five classifiers on 27 loan applicants, and watch ridge reshape the loss surface under SGD, momentum, and Adam. The core skill is separating what one model's probability says from how much 10 models disagree.

Harvard CS181 HW3: Kernels, Neural Networks, and Scaling Laws

CS181 Spring 2026 HW3 (due Mar 23) has three problems worth 100 points: unpack polynomial and RBF kernels into feature maps and go from ridge to dual coefficients α and the support-vector intuition; hand-derive backprop for a two-layer sigmoid network; and, for half the grade, train ResNets on Fashion-MNIST, measure your own scaling law, and use C≈6ND to split a fixed compute budget between model and data.

Harvard CS181 HW4 (Part 2): Why Autoencoders Can't Generate, and What VAEs Add

HW4 Problem 2 has you train a convolutional autoencoder on 64×64 CelebA, sample from N(0, I), and watch it fail to produce faces. You then derive the ELBO, the reparameterization trick, and the closed-form KL, and turn the same backbone into a VAE to compare reconstructions and samples.

Harvard CS181 HW4 (Part 1): Transformers, From Hand-Computed Attention to Multi-Head

HW4 Problem 1 (40 pts) takes self-attention apart in five steps: a 2×2 hand calculation, why we divide by √dk, permutation equivariance without positional encoding, single-head attention in pure NumPy, and multi-head attention in PyTorch with an attention heatmap on synthetic data.

Harvard CS181 HW4 (Part 3): Decision Trees, Random Forests, and Mixture of Experts

HW4 Problem 3 quantifies why voting trees get more accurate in three steps: with p=0.6 the Hoeffding bound needs B≈691 independent trees to reach 10⁻⁶; correlation ρ between trees floors ensemble variance at ρσ²; and random forests' dense ensembling is compared with MoE's sparse routing.

Harvard CS181 HW5 (Part 1): K-means, HAC, and PCA on Handwritten Digits Without Labels

HW5 Problems 3–4 give handwritten digits to three methods that never see a label: K-means summarizes the data with 10 mean images, HAC builds a merge tree you can cut at any number of clusters, and PCA compresses images onto a few continuous directions. All three answer the same question — how much error do you pay to describe the data with a few objects — and the assignment makes you compare their objectives and reconstruction errors directly.

Harvard CS181 HW5 (Part 2): SimCLR and GANs — Representations Without Labels, Generation Without Likelihoods

HW5 Problems 1–2 both turn learning into a classification task. SimCLR's NT-Xent loss asks the network to pick the other augmented view of the same image out of 2N−1 candidates; a GAN's discriminator classifies real versus fake. You derive the math behind each (the cross-entropy equivalence, the optimal discriminator and the JS divergence), then write both training loops on FashionMNIST and MNIST.

Harvard CS181 HW6 (Part 1): Decoding Autoregressive Models, KV Cache, and Speculative Decoding

HW6 Problem 4 (20 points) takes apart the cost of generating one token at a time in three questions: picking the most likely token at each step doesn't give the most likely sequence; recomputing every key at every step makes cost quadratic, and a KV cache brings it back to linear; speculative decoding lets a small model guess and a large model verify in one pass. It is all pencil-and-paper, and every question maps onto a real design choice in today's LLM inference systems.

Harvard CS181 HW6 (Part 2): HMMs and the Kalman Filter

HW6 Problem 1 (15 pts) swaps the discrete HMM from lecture for a continuous state: the state drifts by Gaussian noise each step, each observation adds more noise, and you derive the mean and variance of the filtering distribution p(zₜ | x₀…xₜ). That is a one-dimensional Kalman filter. The solution is two moves, predict with the transition and then correct with the observation, and the problem hands you both Gaussian identities you need.

Harvard CS181 HW6 (Part 3): Policy Iteration and Value Iteration for MDPs

HW6 Problem 2 (15 pts) hands you a 4×5 Gridworld where moves can slip and rewards arrive only when you leave a cell. You write one step each of policy evaluation, policy iteration, and value iteration in the notebook, watch how the discount factor γ reshapes the policy, and finally ask whether this is a sensible model of the robot task at all. The rules of the world are fully known, so this is planning, not learning yet.

Harvard CS181 HW6 (Part 4): Q-learning Swingy Monkey and Embedded EthiCS

HW6 Problem 3 (20 pts) has you write a tabular Q-learning agent for Swingy Monkey, a Flappy Bird-like game. The minimum bar is scoring over 50 at least once within 100 epochs, plus one improvement of your choice. Problem 5 (10 pts) is a 250-word ethics question: assuming a social platform's users grow more politically extreme, use RL concepts to explain how the choice of reward function might have contributed.

Harvard CS181 Midterm Checkpoint: Auditing HW0–HW3 With the Official Checklist

The CS181 Spring 2026 midterm is in class on Mar 10, worth 15% of the grade, closed-book with one double-sided note sheet. The official midterm checklist has four blocks: regression, classification, neural networks and model selection, and SVMs. This post maps each block to HW0–HW3 problem numbers, flags what the 2026 homework never drilled, and explains how to use the 2025 practice exam, review session, and concept checks.

techdeep-dive

Should Code Have Comments: Three Answers from Clean Code, A Philosophy of Software Design, and Redis

There is no 'never comment' school — only 'comment by exception' (Uncle Bob) versus 'comments are part of the design' (Ousterhout, antirez). In their 2024–2025 public debate, both agree on why-comments and against noise comments; the real fights are over interface comments for internal methods, long names as a substitute, and whether comments can be trusted. Controlled studies say quality decides: the same comments moved performance anywhere from -30% to +34% depending on the snippet.

aideep-dive

Making the Invisible Visible: Component Design Philosophy for Agent Chat UI

Traditional chat only needs text bubbles, but AI Agent conversations must surface thinking, tool calls, citations, and progress — we solved this with 12 Vue components, three DisplayModes, and a unified chatBlocks rendering pipeline.

Where Agent Memory Is Heading in 2026: Files Beat Vectors, Forgetting Just Started

In H1 2026, OpenAI, Anthropic, Letta, and LangChain independently chose Markdown files + indexes over vector databases; write permissions shifted back to humans; forgetting mechanisms appeared but nobody implemented Ebbinghaus; Penfield Labs caught 6.4% wrong answers in LoCoMo; three vendors simultaneously adopted 'Dreaming' for offline memory consolidation. Five trends, one conclusion: memory is not a feature — it is an architecture decision.

The Attack Surface of Agent Memory: When Memories Become Persistent Backdoors

Memory turns prompt injection from a one-shot nuisance into a persistent backdoor: MINJA shows conversation-only injection succeeds >95% of the time, and SpAIware demonstrated continuous data exfiltration via planted memories. The industry's two defensive lines — citation-based verification (Copilot) and human approval inboxes (Gemini CLI / Devin) — each have blind spots.

Series Guide: AI Agent Memory Engineering

Agent memory is not one feature — it is at least four distinct engineering problems: working, episodic, semantic, and procedural. This ten-part series walks through the full design space, from taxonomy to coding agent implementations, platform APIs, open-source frameworks, security attack surfaces, and 2026 trend analysis.

Four Types of Memory and Six Design Axes: The Design Space of Agent Memory Systems

CoALA splits agent memory into working, episodic, semantic, and procedural — but the four-cell taxonomy alone doesn't explain why Claude Code uses Markdown files while Mem0 uses vectors. This post adds six independent design axes (read mode, write timing, fidelity, write authority, forgetting, scope) and a file-to-graph spectrum to map the full design space of agent memory systems in 2026.

Test-Time Scaling: BrowseConf and Confidence-Guided Reasoning

The previous articles covered evaluation. This one covers another dimension: how to dynamically allocate compute during reasoning. BrowseConf's core insight is that an agent's self-declared 'confidence' can predict answer accuracy. High confidence uses fewer resources; low confidence searches more rounds.

Five Clouds, Five Memory APIs: How OpenAI, Anthropic, Google, AWS, and Microsoft Let Agents Remember

All five major cloud platforms shipped agent memory APIs in 2025–2026, but their design philosophies diverge sharply: OpenAI writes memory as files, Anthropic mounts memory as a directory, Google uses vectors with topic classification, AWS combines events with pluggable strategy pipelines, and Microsoft abstracts memory behind context providers. Pricing ranges from free to $0.75/1K records/month; tenant isolation spans from 'your app handles it' to IAM as a first-class citizen.

How Nine Coding Agents Handle Long-Term Memory: From CLAUDE.md to MemFS

Nine coding agents have taken at least four different paths for long-term memory: Claude Code and Codex use Markdown files (agent-written, human-readable), Antigravity CLI inherits Gemini CLI's approval inbox (agent proposes, human decides), and Copilot uses citations with JIT verification (auto-deleted after 28 days unverified). Cursor removed its Memories feature and fell back to human-written Rules. Hermes Agent, OpenClaw, and Letta Code treat memory as a core harness component, not a plugin. Their choices on write timing, forgetting, and cross-team sharing are completely different — and none has published a controlled experiment on whether their memory system actually helps.

Commercial Landscape: OpenAI, Perplexity, Gemini, Claude, Grok

By 2026, the deep research commercial market has differentiated: OpenAI is comprehensive, Perplexity is fast, Gemini integrates ecosystems, Claude reasons deeply, Grok is real-time. This article compares each product's differences—not who is best, but who fits your scenario.

Deep Research Landscape: Taxonomy of 80+ Implementations, Roadmap, and Trade-offs

The entire Deep Research field has 80+ implementations, but the core structure is just three-stage roadmap × four components × three optimization methods. This article maps the full landscape: from Agentic Search to Full-stack AI Scientist, from query planning to answer generation, from workflow prompting to end-to-end RL.

Benchmark Deep Dive: DeepResearch Bench II and the Evaluation Landscape

DeepResearch Bench II uses 9,430 expert rubrics covering 132 tasks, and finds that even the strongest agents satisfy less than 50% of criteria. This article breaks down the benchmark architecture, scoring methodology, leaders, and the overall evaluation landscape.

aidebug

One Missing 'Write a Script': How Skill Instructions Determine LLM Success or Failure

An AI assistant platform running Opus 4.6 hit stream_stall (90s timeout) twice consecutively when generating docx. Root cause: skill instructions lacked one sentence — 'Write a script' — causing the model to output JS code inline instead of writing a file and executing with node. Claude.ai's official SKILL.md has that sentence, and the model consistently takes the safe path.

【Ecosystem】How the Community Builds Deep Research Skills

10+ community deep-research skills represent 10+ philosophies of 'how to do research.' From hyperresearch's persistent vault to jamoeight v2's Co-Scientist 6-agent, from adversarial verification to benchmark alignment. This article puts them all on one table.

Evaluation Challenges: Why Deep Research Is Hard to Measure

A deep research agent produces a report—maybe thousands of words with dozens of citations. How do you score it? Using LLMs as judges is biased, asking humans is too expensive, and benchmarks can't keep up. STC and other recent approaches try to solve this from the 'confidence' angle—but there's no perfect answer yet.

Future Outlook: From Research Tool to Scientific Infrastructure

Deep research has already evolved from 'help you search' to 'help you research.' But the next step is bigger: self-evolving agents, swarm collaboration, scientific automation. This article covers three directions and an uncomfortable reality: Gartner predicts 40% of agent projects will be cancelled by 2027.

aidebug

The Bug Photo That Wouldn't Show Up: Three Rounds of Fixing Image Search

An AI assistant platform's image search needed three rounds of fixes: pushing node_type filtering into the ES query to stop text chunks from hogging top-k slots, making filename matching deterministic instead of relying on the LLM to pass an optional parameter, and switching from Postgres icontains to ES match for CJK-aware partial matching.

IterResearch & AREX: Memory and Self-Evolution for Long-Horizon Research

When a research agent runs 25, 100, or 2000 turns, what happens? Context suffocation: information piles up, noise increases, attention gets diluted. IterResearch solves this with Markovian state reconstruction; AREX achieves recursive self-improvement with an inner/outer loop. Both answer: how does an agent stay coherent across hundreds of search rounds?

Open-Source Agent Memory Frameworks: Seven Contenders and a Selection Guide

Seven open-source memory frameworks span the spectrum from auto-extracted vectors to human-readable files: Mem0's one-line add(), Graphiti's bi-temporal knowledge graph, Letta's agent-edited system-prompt blocks, LangGraph's namespaced Store, LlamaIndex's priority-based block truncation, Cognee's triple-store pipeline, and Supermemory's temporal vector-graph engine. This post compares their storage, write/forget mechanics, tenant isolation, and benchmark numbers, then offers selection guidance for four common scenarios.

Open-Source Tools Overview: GPT-Researcher, STORM, smolagents...

The deep research open-source ecosystem has evolved from 'single frameworks' to 'tool clusters.' This article compares 12+ projects: GPT-Researcher emphasizes multi-agent collaboration, STORM simulates expert conversations, smolagents focuses on state management. Each tool solves different problems.

【Project】How We Build the Deep Research Skill

This is the project's own deep research skill design, fully disclosed. Core choices: only Groundlane MCP for web tools, strict source-quality grading (A/B/C/D), research hands off to post skill for publishing. Not the most powerful, but the best fit for us.

aidebug

10% of Conversations Were Making Things Up: Debugging a Silent Retriever Failure

About 10% of conversations on our AI assistant platform randomly lost knowledge base tools — the agent hallucinated answers from training data instead. Root cause: a bare except Exception swallowed Elasticsearch connection failures during retriever initialization, silently skipping tool registration. Fix: retry + surface failures to system prompt + structured metadata tracking.

Tongyi DeepResearch: From Base Model to Agentic Foundation

Previous articles covered the landscape, training from scratch, long-horizon memory, and planning optimization. This one zooms out to see a complete system that threads all these insights together: Tongyi DeepResearch. Its core innovation is Agentic CPT — inserting an agentic mid-training stage between pre-training and fine-tuning, giving the model an inherent agent bias. MoE 30B parameters activating 3B, HLE 32.9 surpassing OpenAI o3.

aidebug

Three Fixes for One Setting: When top_k Lies Across Four Layers

A user set retrieval top_k to 15, but the monitoring dashboard showed 5 and the streaming UI flashed 5 before jumping to 15. The same top_k value existed at four layers — LLM tool arguments, runtime, trace DB, and streaming payload — each requiring its own override. Three sequential fixes, each revealing the next layer was also wrong.

aidebug

Fixed It Three Times, Broke It Three Ways: Vision PDF Parsing Optimization in Three Acts

Parsing a 150-page PDF via Vision API took 29 minutes (one page per request). Batching cut it to 1.5×, adding 5-way concurrency brought it down to 24 seconds. One week after launch, a customer uploaded 23 PDFs at once — 70 parallel Bedrock requests triggered full throttling: 90 pages skipped, 5 files failed, 8 stuck. Fixed with Redis-based cluster-wide slots + backoff retries.

WebDancer & WebThinker: Training a Deep Research Agent from Scratch

Two NeurIPS 2025 papers answer the same question: how to train a web research agent from scratch? WebThinker chooses 'bolt on web capability to existing reasoning models,' WebDancer chooses 'rebuild everything from data construction to RL training.' Two philosophies, four stages, one core insight: training beats prompting.

Data Synthesis: WebShaper & S1-DeepResearch

Previous articles covered how to train agents. But training requires high-quality data—and deep research training data has been scarce. WebShaper solves this with mathematical formalization: define IS tasks in set theory, then use an agentic Expander to iteratively expand them. S1-DeepResearch goes further: moves training from 'search-centric' to 'real research.'

Multimodal and Vision: WebWatcher Redefines Deep Research

All deep research agents are 'text-first'—but the real world isn't just text. WebWatcher (NeurIPS 2025) is the first system to integrate visual reasoning into deep research, using OCR, image search, code execution, and other tools to handle charts, screenshots, videos, and other diverse information.

WebWeaver & DeepPlanner: Dual-Agent Architecture and Planning Optimization

Previous articles covered training from scratch and long-horizon memory. This one goes deeper: how to make the agent's 'planning' itself better? WebWeaver tackles it architecturally (dual-agent iterative outline optimization). DeepPlanner tackles it through training (advantage shaping for planning tokens). Both point to the same conclusion: planning is the ceiling of deep research.

Gemma — Google's Open-Weights Flank: Gemma 4 Moves to Apache 2.0, From Mobile to Workstation Full-Size Open Weights

Gemma is Google's open-weights family paired with the closed-source Gemini flagship; Gemma 4 (E2B / E4B / 26B-MoE / 31B) released in April 2026 switched its license from Google Gemma Terms of Use to Apache 2.0 — the single most important change in this article.

Laguna: From a 33B Local Workhorse to 118B Long-Horizon Reasoning, Poolside's Three-Releases-in-Three-Months Bet

Laguna is Poolside's agentic coding model family: XS 2.1 packs 33B-A3B into a 36GB Mac, while S 2.1 brings 118B-A8B with 1M context to 70.2% on Terminal-Bench 2.1 and 40.4% on DeepSWE, both open under OpenMDW-1.1.

Ling — From Trillion-Parameter Flagships to 5.1B Execution Nodes, Ant Group's Three-Line AGI Strategy

Ant Group's Ling model family deep-dive: 2025→2026 evolution timeline, Ling/Ring/Ming three-series strategy, architecture journey from Ling 1.0 to Ling 3.0, Ling-3.0-flash-Fin finance model, and an Agent developer's selection guide

Muse Spark: Meta's Closed-Source Agentic Model Line, from Llama to 1.3

Muse Spark is Meta's closed-source agentic model line: version 1.3 combines a 1M-token context, multimodal inputs, and long-horizon tool loops. Standard pricing is $1.25/$4.25 per 1M input/output tokens, while Contributor drops to $0.10/$0.20 in exchange for training rights. It is not the next Llama; it is a separate product line built around models, APIs, and coding agents.

Nex-N2.5: The Open Agent Family That Treats Vision as an Interface, From 35B mini to 1.6T Max

Nex-N2.5 is Nex AGI's open agentic model family: mini scores 82.9 on OSWorld-G at 35B-A3B, Pro tops Claude Opus 5 with 87.4 at 397B-A17B, and Max leads the whole official table on BrowseComp with 92.6 at 1.6T, all open under Apache-2.0.

aideep-dive

LLM Agent Tool Discovery: Why Agents Don't Use Available Tools, and How to Fix It

An agent with a workspace_browse tool said 'file not found' instead of searching. Anthropic, OpenAI, and Google's official guides all point to the same fix: put trigger conditions and workflows in the tool description. A 2025 study found 97.1% of MCP tool descriptions have quality issues.

Multi-Agent Communication: Handoff, Delegate, Mailbox, and the Push for Protocol Standards

Agent-to-agent communication falls into three patterns: handoff (transfer control), delegate (dispatch and wait for results), and mailbox (real-time peer-to-peer messaging). Implementations vary widely, but MCP and A2A are driving protocol standardization.

Multi-Agent Context Management: The Fork vs Fresh Trade-off, History Truncation, and Result Compression

Should a sub-agent see the parent's conversation? Fork carries full history but token costs grow exponentially. Fresh saves money but lacks context. Industry consensus: default to Fresh, Fork only when needed, and always pair it with history truncation and result compression.

Multi-Agent Cost Control: How Seven Frameworks Handle the 'Soft Landing Before Hard Stop' Consensus

Parallel + nested agent spawns can burn 200K+ tokens in a single conversation turn. From Anthropic to Microsoft, the industry is converging on tiered responses: compress → downgrade → stop, rather than a binary kill switch.

The Multi-Agent Landscape: How Every Major Coding Agent Does Multi-Agent Collaboration in 2026

By 2026 nearly every mainstream coding agent supports subagents. Design philosophies split three ways: deterministic scripted orchestration (Claude Code Workflow), model-driven autonomy (Codex, Devin), and IDE command-center integration (Windsurf 2.0, VS Code). This overview maps product positioning, a capability matrix, and the design-philosophy spectrum.

Multi-Agent Observability: Where the Money Goes and Which Agent Broke Things

The most common debug nightmare in multi-agent systems is 'the answer is wrong, but I don't know which agent did it.' Three layers of observability are essential: per-agent token metering, execution traces, and real-time cost dashboards.

Multi-Agent Orchestration Patterns: Scripted, Model-Driven, or Hybrid — How to Choose

Multi-agent orchestration splits into three camps: scripted determinism (LangGraph, Claude Code Workflow) is predictable but rigid, model-driven (Codex, Devin) is flexible but unpredictable, and hybrid (Windsurf 2.0) acts as a command center integrating multiple agents. The choice depends on how much predictability you need.

Multi-Agent Safety & Guardrails: Preventing Prompt Injection from Spreading Across Agents

Multi-agent security risks aren't just amplified single-agent risks — inter-agent communication is itself an attack surface. A compromised sub-agent can pass malicious instructions to the parent through its return value. Core defense: treat agent output as untrusted data.

Affiliate Marketing Unit Economics: A Business or Just a One-Time Commission?

Affiliate marketing becomes a business only when the content reduces decision cost, each conversion retains margin after updates and attribution losses, and the publisher accumulates its own trust and demand knowledge.

Which Content Assets Are Still Worth Building in the AI Era? Cash Flow, Defense, and Options

Do not put every resource into search articles. Allocate content assets across cash flow, defense, and options, then test ownership, portability, update responsibility, and reconstructability.

How AI Summaries Change the Path from Content to Traffic

AI search separates visibility, citation, referral, and conversion into different events; publishers need to measure attributable referrals, activation, and source cohorts—not rankings and sessions alone.

Blocking, Licensing, and Litigation Protect Different Parts of the Content Business

Blocking controls future requests, licensing defines an exchange between contracting parties, and litigation addresses an existing legal dispute; none substitutes for the others.

What Is Left of Free Content When AI Takes the Click?

Answer engines can read content without sending the reader; free-content businesses therefore need to move from rented clicks toward first-party relationships, useful tools, original signals, and direct brand demand.

Which Content Is Easiest for Answer Engines to Commoditize?

Content is easiest to commoditize when a short answer preserves most of its value and readers need neither original data nor a subsequent action; the key test is whether they still need to return after reading the summary.

How Free Financial News Makes Money: cnYES and the Three-Sided Attention Market

cnYES publicly offers more than ad inventory: video, events, sponsored features, editorial production, and historically, B2B news licensing. Free content attracts readers, while the platform balances advertiser outcomes, production costs, and editorial trust.

How Beehiiv Turns a Newsletter Into an Operating System: Recommendations, Ads, and Growth Loops

Beehiiv puts newsletters, recommendations, referrals, advertising, paid subscriptions, and automation in one operating system. Official plans vary by list size and feature tier; a 0% subscription take rate does not mean zero cost or zero lock-in.

How AI Summaries Feed a Subscription: BigGo Finance, Free Tools, and Pro

BigGo Finance places public AI podcast summaries, market data, and news at the free entrance, then offers Pro upgrades around model capability, alerts, and experience. That creates a plausible product path, but public evidence does not show that summary readers convert into paying subscribers.

How Research Becomes an Enterprise Workflow: CB Insights and Data Productization

CB Insights uses free newsletters to demonstrate its data capabilities, then turns market signals into searchable records, Mosaic scores, and CRM/API workflows. Enterprises pay to prioritize research faster, not simply to receive more articles.

How Free Content Sells Tools: CMoney's Methods, Apps, and Investor Community

CMoney describes its product design as a three-step path from method to tool to community. Free content and discussion are designed to surface needs, shared data and APIs can turn investing methods into many apps, and public products show monetization through subscriptions, courses, and institutional systems—but the company has not published retention or revenue proof for the loop.

Does Content Acquisition Pay? CAC, Gross-Margin LTV, Payback, and Attribution

Content CAC divides complete production, distribution, tooling, and labor cost by new paying customers in the same cohort. Judge it with gross-margin LTV and payback—not leads, sitewide averages, or a universal 3:1 slogan.

Which Content Models Face the Most AI Risk? Compare Product Layers, Not Companies

AI risk is not a company ranking. It is the reusability of each product layer's public text set against defenses such as original signals, direct relationships, workflows, and transactions.

What Can You Actually Take When Leaving a Creator Platform? A Six-Layer Migration Checklist

A CSV export does not mean a creator can move an entire business. Posts, email consent, membership status, billing relationships, URLs, and recommendation traffic are six different assets that require separate tests.

First-Party Data, Community, and Tools: What Do You Actually Own After AI Search?

First-party data, community, and tools can turn anonymous exposure into a relationship or job that can be served again. None is fully owned: consent, portability, platform dependence, maintenance, and retention must be evaluated separately.

How Free Tools Compound Search Value: Real Utility, Return Loops, and Maintenance

A free tool does not rank or earn backlinks merely because it is interactive. Its opportunity comes from completing a repeatable job, creating measurable reasons to return, share, and improve the product.

How Free Information Reaches a Trade: Fugle's Research, API, and Brokerage Funnels

Public materials do not show Fugle charging directly for free articles or taking a commission on every trade. It attracts investors with research tools, then monetizes personal API subscriptions, information services, and B2B technology while the brokerage still holds the account and executes the trade.

Why Enterprises Buy Gartner: Brand, Analysts, and Decision Insurance

Gartner packages research, analyst access, and procurement tools as decision insurance for enterprises. Insights products generated about 78% of 2025 revenue, but a Magic Quadrant can reduce search and coordination costs—not guarantee the right purchase.

Ghost Sells More Than Hosting: Domains, Members, Stripe, and the Right to Exit

Ghost places the site, member data, and payment relationship closer to the publisher: payments connect to the publisher's own Stripe account and Ghost charges a 0% transaction fee, but Stripe, hosting, and operations still cost money. Ghost 6 also includes Recommendations and ActivityPub, so describing it as having no discovery is no longer accurate.

What Platform Distribution Costs: Medium, Vocus, and the Readers You Cannot Take With You

Medium and Vocus can both deliver discovery, but the exchange is different. Medium controls distribution to strangers and withholds new subscribers' email addresses; Vocus handles Taiwanese payments, invoices, and member operations, but its public documentation proves only that order data—not complete email identities or payment relationships—is exportable.

How Patreon Turns Supporters Into Members: Free Entry, Paid Tiers, and the Cost of Leaving

Patreon charges a 10% platform fee to new creators who publish after August 4, 2025, bundling free membership, paid tiers, commerce, community, and discovery. Creators can export email addresses, but a CSV does not carry recurring payment authorization, conversations, or platform distribution.

How Private-Market Data Becomes a Workflow: PitchBook's Human Verification and Switching Costs

PitchBook combines public signals, direct submissions, and human review to build a relational database of companies, deals, funds, and investors, then embeds it in screeners, Excel, CRMs, and APIs. Segment revenue reached $671.8 million in 2025, but recent private-market data still requires estimates and later revisions.

From SEO to Direct Brand Demand: AI Search Changes the Entrance, Not the End of Search

SEO still makes content discoverable, but AI answers no longer turn every exposure into a click. Direct brand demand must be measured across branded queries, identifiable returns, activation, and conversion—not by labeling all direct traffic as brand.

Is Substack's 10% Worth It? Discovery, Break-Even Math, and Migration Boundaries

Substack charges 10% of all paid-subscription revenue in exchange for publishing, payments, a recommendation network, and a shorter checkout path. Break-even depends on retained incremental value—not on subtracting 10% from the company's self-reported 25–30% network share.

How Taiwan Creators Should Choose a Platform: Payments, Discovery, Control, and Operations

Creator-platform selection is not a ranking. A Taiwan-based creator should pass four gates—payments, discovery, control, and operations—before choosing a platform, managed SaaS, or self-hosted site. No CSV represents a complete reader and billing relationship.

Can Taiwan Build Another Vertical Intelligence Company? Five Gates and a Four-Week Test

Taiwan's industrial density can become global enterprise intelligence, but a new company must pass five gates: unique signals, recurring demand, structured fields, lawful verification, and overseas buyers. TrendForce shows the model exists; it does not show that every vertical can copy it.

Commercial Document Parsing APIs Compared: Specialized Parsers, General VLMs, and the Big Three Clouds

Three routes to commercial document parsing: specialized parsers (Cohere Parse at $1.50/k pages, LlamaParse Agentic Plus at 90.2% on ParseBench), Big Three cloud prebuilts (Azure/Google/AWS for structured field extraction), and general-purpose VLMs (Fable 5.1 scores 78.92 on ParseBench and crushes specialized parsers on charts, but costs 3–16× more and hallucinates). At 100K pages/month, plain OCR runs ~$150 across providers; add tables and AWS jumps to $1,500, Claude Sonnet 5 to $900. The first question isn't 'which is most accurate' — it's 'do you need transcription or comprehension?'

How Industry Intelligence Sells to Enterprises: Six Product Paths and Their AI Exposure

Enterprises do not pay for more articles. They pay to find signals sooner, prioritize what deserves attention, and complete decisions faster. Six cases turn reporting relationships, contributor research, structured data, predictive scores, and procurement workflows into intelligence products that can earn renewals.

Who Sells Content? Four Business Models and Four Paths Forward

A content business is not merely an article behind a paywall. It may sell enterprise decisions, personal trust, platform infrastructure, or use free content to acquire customers for another product.

How Supply-Chain Access Becomes Renewals: Inside DIGITIMES' Intelligence Pipeline

DIGITIMES turns scattered signals from Taiwan's technology supply chain into news, research, databases, and advisory services. Enterprises renew for a repeatable decision-making pipeline, not simply for a larger volume of articles.

How Free Content Acquires Customers for Another Business: Eight Paths and One Scorecard

Free content is not a free business. It is an acquisition investment. Eight cases spanning ads, tools, subscriptions, brokerage partnerships, affiliate marketing, and AI search ask what paid job the audience eventually completes.

How Seeking Alpha Sells Crowdsourced Research: Contributors, Quant Ratings, and the Subscription Flywheel

Seeking Alpha sells more than stock articles: outside contributors supply research, editors govern the market, and more than 100 quantitative metrics feed ratings and subscription tools.

How a Few Exclusive Stories Support a Premium Subscription: The Information's Reporting Flywheel

The Information sells a small number of exclusive stories to professionals who can act on them, then turns reporting exhaust into databases, org charts, and an AI research tool that supports multiple subscription tiers.

The Platform Bet: Who Controls the Creator-Reader Relationship?

Choosing a creator platform is not about finding the longest feature list. It assigns responsibility for bringing readers, collecting payments, holding data, and operating the system; pricing is only one of four gates.

learningguide

The Research Toolkit Trifecta: Google Scholar to Find, Moonlight to Read, CorTeX to Write

Break the research workflow into three actions — find, read, write — and pick one tool for each: Google Scholar for discovery, Moonlight AI for reading comprehension, and CorTeX for collaborative writing. All three offer free tiers and together cover the full pipeline from literature search to manuscript submission.

ai

Writing AI Papers and Submitting to arXiv: A Practical Guide to Structure, Experiments, and Conference Strategy

A consolidated guide drawing on the arXiv official guidelines, the NeurIPS ML Reproducibility Checklist, the REFORMS framework (8 modules, 32 items), and the preprint policies of five top conferences. Covers paper structure, experimental design red lines, the arXiv submission workflow, and a decision framework for conference vs. direct submission.

AI-Native SDLC Playbook L1: When Code Generation Is No Longer the Bottleneck

Claude Academy's opening lesson identifies the core paradox: AI accelerates code generation, but review, testing, and deployment don't keep pace. The bottleneck shifts from 'not writing fast enough' to 'not reviewing fast enough.' The AI-native SDLC fix isn't more AI-generated code — it's embedding AI into every stage where bottlenecks now live.

AI-Native SDLC Playbook L2: intent.md Turns Requirements into Version-Controlled Documents

Traditional requirements scatter across Jira, Slack, and meeting notes, losing fidelity at every handoff. intent.md lets the originator collaborate directly with Claude to produce a human-readable, machine-actionable, version-controlled Markdown proto-spec — from conversation to committed document in hours, not weeks.

AI-Native SDLC Playbook L3: Requirements and Design Collapse into One Session

Traditionally, requirements analysis and design are separate phases run by different teams — every handoff loses information. This lesson's approach: Claude reads intent.md in a single session, applies organizational standards (brand, security, compliance, UX loaded as skills), and produces a unified spec.md. The product owner reviews; they don't author.

AI-Native SDLC Playbook L4: Plan Mode — Write the Plan Before Writing Code

Claude Code's plan mode lets engineers produce a reviewable, version-controlled implementation plan (plan.md) before writing a single line of code. Design review shifts from the PR diff to the planning stage, and the cost of course-correcting drops from 'rewriting code' to 'editing a document.'

AI-Native SDLC Playbook L5: CLAUDE.md Turns Team Knowledge into Agent Memory

CLAUDE.md is a context file at the repo root that Claude reads at the start of every session — your team's conventions, commands, architecture patterns, and pitfalls. The course's core advice: if Claude makes the same mistake twice, write it into CLAUDE.md.

AI-Native SDLC Playbook L6: Skills Turn Organizational Standards into Reusable Knowledge

A skill is organizational tacit knowledge made operational — a folder with a SKILL.md that Claude loads automatically when trigger conditions are met. The course's key principle: skills make violations rare; hooks make them nearly impossible.

AI-Native SDLC Playbook L7: Parallel Sessions and Subagents

One engineer runs multiple Claude Code sessions simultaneously, each in its own git worktree; repetitive verification work goes to subagents. The bottleneck shifts from 'writing code' to 'reviewing output.'

AI-Native SDLC Playbook L8: Give Claude a Feedback Loop

Have Claude verify its own output before submitting — tests, builds, and screenshot diffs all run to completion before the task is marked done. Engineers receive code that's already passed verification, not code that 'might be correct.'

AI-Native SDLC Playbook L9: Continuous Evals in CI

Evals are the AI-native equivalent of stage-gate QA — collect 20–50 real tasks as test cases, run them automatically whenever CLAUDE.md, skills, or hooks change, and block the merge if the pass rate drops. Every production incident becomes a permanent eval.

AI-Native SDLC Playbook L10: AI in the PR Review Loop

Let AI handle the first review pass so humans can focus on intent and risk. This lesson covers how to define REVIEW.md, layer review passes, set up an automated review-comment fix loop, and why the agent that wrote the code must never approve its own PR.

AI-Native SDLC Playbook L11: Hooks as Approval Gates

Hooks are the governance bedrock of the AI-native SDLC — deterministic gates that intercept agent actions and block them if conditions aren't met. This lesson walks from a single production-gate script to full enterprise managed settings covering permission lockdown, sandboxing, credential isolation, and marketplace allowlists. The most technically dense lesson in the entire course.

AI-Native SDLC Playbook L12: CI/CD Integration and Deployment

Plug Claude into the CI/CD pipeline — start with read-only build failure triage, gradually add write operations behind existing gates, expose deployment tooling through MCP, and tier autonomy by environment. The governing principle is one sentence: 'The agent may act up to the production gate and cannot pass it.'

AI-Native SDLC Playbook L13: Closing the Loop with Monitoring

Stage 6 is both the endgame and the starting point of the AI-Native SDLC: a monitoring script detects an anomaly → Claude writes a diagnosis as intent.md → it flows through the entire development pipeline. Humans shift from 'starting work' to 'triaging and reviewing work.'

AI-Native SDLC Playbook L14: Series Summary and Adoption Roadmap

After 14 lessons, this final article distills the series into three things: a prioritized adoption roadmap, role-specific reading paths, and Anthropic's complete official documentation list.

Claude Academy: AI-Native SDLC Playbook Course Guide

Anthropic's Claude Academy offers a free 14-lesson course that takes AI-assisted coding from 'individuals using Claude Code' to 'an organization-wide development lifecycle.' Four core concepts — intent.md, CLAUDE.md, Skills, and Hooks — wire together into a complete AI-native SDLC.

aideep-dive

WebMCP Explained: Turning Every Web Page into an AI Agent Tool Server

WebMCP is a browser-native W3C standard proposed by Google and Microsoft. It lets web pages expose structured tools to AI agents via document.modelContext.registerTool() — no backend, no HTTP/SSE transport. Chrome 149 Origin Trial is live.

Groundlane Series Part 6: The Document Toolkit — effort Dial, Field-Aware Chunking, and Confidence Routing

document_parse gains an effort parameter (fast/standard/deep) unifying anydoc WASM, OCR.space, and Docling VLM into a single dial; document_chunk's fieldAware mode extracts field names from table headers and metadata to solve RAG attribute conflation; document_smart_parse now returns a confidence score so agents decide whether to upgrade.

techdeep-dive

Groundlane Series Part 7: Selector Healing, Retrieval Test, and Quality Benchmarks

Groundlane v0.1.0 adds four quality mechanisms: web_extract selector healing (three-layer deterministic fallback, no LLM), corpus_retrieval_test (RAG recall verification inspired by RAGFlow), a document benchmark CLI (character-level F1), and a search benchmark framework (ground truth corpus with multi-provider comparison). Tool count goes from 54 to 55.

Reading Stanford CS329Z Week 2: Tell Workflows from Agents Apart, Then Hand-Build Your First RAG

Week 2 runs as a one-two punch: Monday's Anthropic taxonomy teaches you when not to build an agent, Wednesday's RAG paper hands you the first complete compound-system recipe. Five workflow patterns are the selection toolkit, RAG is parametric-plus-nonparametric memory, and together they are the blueprint for HW1's email retrieval pipeline.

Reading Stanford CS329Z Week 3: Plug Tools In, Take Frameworks Apart — HW1 Begins

Week 3 standardizes tool interfaces with the MCP specification on Monday and trades hand-written pipelines for compilable, optimizable programs with the DSPy paper on Wednesday. HW1 drops the same Monday and bans agent frameworks: you build a company's internal assistant from scratch, so this week DSPy is for understanding what frameworks abstract, not for handing in.

Reading Stanford CS329Z Week 4: Learn to Think While Doing, Then Learn to Remember — ReAct and MemGPT

Week 4 pins the agent loop down as an interleaved think-act-observe sequence with the ReAct paper on Monday, then turns memory into OS-style tiered storage with the MemGPT paper on Wednesday. The same week, HW1 is growing from an email-retrieval pipeline into a full harness where memory is an explicit requirement, so the loop shape and the memory design are the two things to settle now.

Reading Stanford CS329Z Week 5: One Agent or a Meeting — Multi-Agent Systems and the Three Optimization Axes

Week 5 turns multi-agent collaboration into programmable conversation with AutoGen on Monday, then lays out the three optimization axes — prompts, weights, inference compute — with GEPA and the test-time compute paper on Wednesday. HW1 is due 10/30, the last full week before the deadline, so this installment helps you decide which axis deserves your effort.

Reading Stanford CS329Z Week 6: Spin Up the Data Flywheel — Submission Week

Week 6 assigns Shankar's data flywheel on Wednesday — evaluation, monitoring, and continual improvement feeding on the same production data — while HW1 comes due, HW2 drops, and the midpoint demo video and midway report loom in early November.

Reading Stanford CS329Z Week 7: Score Honestly, Scale Data — Midterm Checkpoint

Week 7 is midterm checkpoint week: data selection on Monday, evaluation and benchmark design on Wednesday. Zhu et al. teach you not to be fooled by your own scores, SWE-smith scales software-engineering tasks to 50,000 instances, and you close by drafting a first 4-tuple for HW2.

Reading Stanford CS329Z Week 8: Let a Model Judge, Then Guardrail the Agent

Week 8 builds model judges with MT-Bench and Anthropic's eval guide on Monday, then faces production leakage with PrivacyLens and four guardrails on Wednesday. The paper video is due Friday, and this week's deliverable is one working judge score plus one permission check.

Reading Stanford CS329Z Week 9: Coding Agents Need Their Own IDE First

Week nine turns to coding agents: SWE-agent shows interface is performance, OpenHands packs sandbox plus benchmarks into one general base, and the second homework is due Friday — ship one working bug-fix exam this week.

Reading Stanford CS329Z Week 11 (Finale): From Waiting for Orders to Acting First — Proactive Agents and Demo Day

The finale reads Week 11: Monday upgrades instruction-waiting reactive assistants into proactive agents that observe, infer, and act first via the GUM paper, while Wednesday folds multimodal systems, long-running agents, and production observability into three open problems. Ends with a pre-Demo-Day checklist and a one-line map of all 11 posts.

AI Security Cert Showdown — SecAI+ vs CAISP vs GAIPS vs AAISM

Four AI security certs, four different bets: SecAI+ ($359) is CompTIA's mid-level expansion play, CAISP ($999 all-in) has the strongest hands-on labs and fullest OWASP LLM Top 10 coverage, GAIPS ($999/$9K) is the SANS gold-standard defender cert with CyberLive exams, and AAISM ($459+) is governance-layer but requires CISM or CISSP first. Under $400 → SecAI+. Want to actually hack and fix → CAISP. Company paying → GAIPS.

The Security Certification Landscape — Where Software Developers Should Start

6+ new AI security certifications launched in just 18 months (2025–2026), while the classic trio — Security+ ($404), AWS Security Specialty ($300), CISSP ($749) — remains foundational. For AI platform developers who aren't security specialists, a three-phase path works best: Security+ → AWS Security → SecAI+ or CAISP → CISSP, totaling $1,800–$2,650.

Security+ → AWS Security → CISSP: ROI Breakdown of Three Classic Security Certifications

Security+ ($404) takes 2–3 months, 50–65% self-study pass rate, +$10K–$20K salary premium. AWS Security Specialty ($300) takes 2–3 months, 45–55% pass rate, highest ROI at 40–67x. CISSP ($749) takes 3–6 months, 50–60% pass rate, +$25K–$35K premium but requires 5 years of experience. Five-year total cost of ownership: Security+ $554, AWS Security $300 (no annual fee), CISSP $1,374.

Taiwan's Security Certification Ecosystem — Regulations, Training, Resources, and Exam Logistics

Taiwan's Cybersecurity Act 2.0 passed and FSC's three-tier system is in place, but neither mandates specific certifications — real demand comes from the job market. Taiwan has 7+ training providers: Uuu has the widest catalog (only one offering both SecAI+ and COASP classroom courses), WUSON is the CISSP legend (monthly cohorts sold out through mid-2027), and DEVCORE exclusively brings OffSec factory instructors. Test centers in Taipei and Kaohsiung; most certs also support online proctoring from home.

Reading Stanford CS329Z Week 1: Stop Tuning Only the Model — State-of-the-Art Is Systems Engineering

The Week 1 anchor reading for CS329Z is Zaharia et al.'s Compound AI Systems: the best results increasingly come from multi-component systems, and even the biggest model is just one part. The post leaves three design questions and three hard challenges — which happen to be exactly what HW1 asks you to answer by building.

Marker: Datalab's Open-Source Pipeline Parser — Faster, CPU-Ready, More Accurate

Marker (Datalab open source, Apache 2.0 license, v2.0.0 released 2026-07-20, 39.5k stars) is a pipeline-style document parsing standard library that outputs Markdown and JSON, supports optional LLM boost (`--use_llm`, default `gemini-3.5-flash`), custom formatting logic, table/formula/inline-math/link/reference/code formatting, image extraction and preservation, header/footer removal, and runs on pure CPU, GPU, or MPS (`Apple Silicon`). Unlike [Docling](/posts/tech/2026-09-06-docling-document-parsing) (structured JSON core, dedicated XML exports, pure MIT), Marker centers on Markdown/JSON under `Apache 2.0` with a separate model-weight license (`AI Pubs Open Rail-M`, $5M commercial threshold); unlike [MinerU](/posts/tech/2026-09-05-mineru-ocr-doc-parsing) (custom agreement with MAU/revenue thresholds + attribution obligations), Marker offers a simpler licensing story (`Apache 2.0` code) but requires accepting a separate model-weight license for weights. The series framework ([three-layer model](/posts/ai/2026-08-06-document-parsing-three-layers)) positions all three (`MinerU`, `Docling`, `Marker`) as pipeline-based parsing-layer options with distinct licensing, output, and speed trade-offs.

Training Small LLMs From Scratch in the Chinese Community: Corpus, Tokenizers, and Three Open-Source Projects

A close look at three open-source projects training LLMs from scratch in the Chinese community — baby-llama2-chinese (218M, 63.4B tokens), ChatLM-mini-Chinese (0.2B T5, 10.23M dialogues), and Steel-LLM (1.12B, 1T tokens, 8 months) — comparing corpus strategy, tokenizer decisions, and community ecosystem. Honest evaluation included: baby-llama2 scored a bottom-ranking 21 in MiniMind's side-by-side test; ChatLM has the strongest knowledge (62) but weak coding.

Learning from Mature Coding Agents (39): Two Rules for Flicker-Free Tool Result Display in TUIs

Four coding agent TUIs all follow the same two rules: tools maintain constant height (one-line summary, never jumping from 0 to N lines) and never auto-collapse (only user-initiated expand/collapse). looplane violates both.

How to Spend Every Parameter: OpenELM's Layer-wise Scaling and MiniCPM's Three-stage Unfreezing

OpenELM uses layer-wise scaling to shift parameters toward layers near the output; with 1.08B parameters and 1.5T tokens it beats OLMo 1.2B (+2.36% on the LLM360 average) despite OLMo training on 3T tokens. MiniCPM trains multimodal small models from scratch with a three-stage unfreezing recipe (Resampler first, vision encoder next, everything unfrozen last); MiniCPM-V 4.5 reaches sub-30B SOTA on VideoMME with only 8B parameters, and 4-bit quantization squeezes fp16's 16–17GB memory footprint down to about 5GB for phones.

Karpathy's nano lineage: nanoGPT, llm.c, and nanochat

A tour of Karpathy's three teaching repos: nanoGPT (2022, ~300 lines each to reproduce GPT-2 124M), llm.c (pure C/CUDA training), and nanochat (2025-10, one speedrun.sh from tokenizer to WebUI). $100 and 4 hours on 8×H100 buys a chatty model; GPT-2-grade capability is now down to about 2 hours and $48.

The framework for training from scratch: LitGPT's no-abstraction rewrites, pretrain flow, and the TinyLlama track record

LitGPT (Lightning AI, ~13,600 stars, Apache 2.0) rewrites 20+ mainstream LLMs — Llama 3, Qwen2.5, Phi 4 — from scratch as single-file, no-abstraction implementations, with a full pretrain / finetune / evaluate / serve CLI. TinyLlama (1.1B parameters, 3T tokens) was trained on this codebase. This post breaks down how it differs from MiniMind, how to actually use it, and where it stops.

MiniMind: Train an LLM From Scratch for $0.40

MiniMind is an open-source project for training LLMs from scratch: a 64M Dense model and a 198M-A64M MoE model that run the entire chain — Pretrain → SFT → LoRA → DPO → PPO/GRPO/CISPO → Agentic RL — in ~2 hours on a single RTX 3090 at roughly 3 RMB (~$0.40). Every core algorithm is implemented natively in PyTorch with no high-level wrappers.

National-Team LLM Training: Apertus' Compliance Route and LLM-jp's Japanese Ecosystem

Switzerland's Apertus (8B/70B, 15T tokens, 1,000+ languages) filters opt-outs and personal data before training to satisfy the EU AI Act; Japan's LLM-jp consortium shipped LLM-jp-4 (12T tokens) in April 2026, claiming wins over GPT-4o and Qwen3-8B on standard benchmarks. Both prove that from-scratch training outside the English sphere is a data-governance problem, not a technical one — plus a note on RWKV-7 as the non-Transformer alternative.

How fully transparent LLMs are built: OLMo 3's model flow and LLM360 K2's 360-degree openness

"Open-source LLM" is a spectrum: weights-only (Llama), weights plus data (most fully open projects), or data order, intermediate checkpoints, and training logs all released (LLM360 K2, OLMo 3's model flow). This piece unpacks the two projects that pushed transparency furthest: OLMo 3 shipped the first fully open 32B thinking model in November 2025, and K2 is the first 65B-class model whose checkpoints even include optimizer states.

Running MiniMind on RunPod: From Zero to a Chatting Model

The hands-on installment of the series: rent an RTX 3090 on RunPod (Secure Cloud $0.5/hr, Community Cloud $0.22/hr), follow the MiniMind README through pretrain (~1.21h) + SFT (~1.10h), spend roughly $0.55–1.50 USD total, and chat with your own 64M model trained from scratch in the terminal.

Training an LLM from Scratch: A Project Spectrum from $0.4 to 65B

Open-source projects have pushed the cost of training an LLM from scratch absurdly low: MiniMind runs the full PreTrain-to-RL pipeline for about $0.4 (2 hours on a single RTX 3090), while at the other end OLMo 3 and LLM360 K2 publish everything — data, code, and stage-by-stage checkpoints of 65B models. This series walks the whole project spectrum from $0.4 to 65B in 11 articles, and flags where the map is biased.

When to Train an LLM From Scratch: A Decision Tree and the Full Cost Ladder

Training from scratch only makes sense in three cases: you want to learn how training works, you have 10B+ clean tokens no open model has seen, or you need a fully transparent training process for research. Otherwise fine-tuning or RAG is almost always cheaper. This post collapses the series' main routes into one cost ladder and a decision tree.

YuLan-Mini: Squeezing a 2.4B Flagship-Scale Small Model Out of 1.08T Tokens

YuLan-Mini is a 2.4B open-source model from Renmin University's AI Box lab, trained on 48 A800 GPUs with only 1.08T tokens — scoring 37.8 on MATH-500 and 64.0 on HumanEval, beating Qwen2/Qwen2.5 peer models trained on 7T–18T tokens at math and code. What's public isn't a slogan: per-phase data mixes, pre-annealing optimizer states, and even W&B logs of the ablation studies.

Docling: IBM's Open-Source, MIT-Licensed Document Parsing Standard Library Built on Structured JSON

Docling (IBM Research Zurich, now governed by the Linux Foundation AI & Data, MIT license, v2.100.0 released 2026-06-09) is an open-source document parsing standard library with structured JSON (DoclingDocument) as its core output. It supports PDF, DOCX, PPTX, XLSX, HTML, EPUB, Apple Pages, video (MP4/AVI/MOV with ASR transcription and keyframes), audio (WAV/MP3), email (EML/MSG), ODF, and XBRL financial reports, with a swappable-stage pipeline parser (pure CPU or GPU-accelerated), VlmPipeline option (GraniteDocling 258M VLM), MCP server and API server (docling-serve), and native integrations with LangChain, LlamaIndex, Crew AI, and Haystack.

Inkling: From an OpenAI Exodus Team to a 975B Open Flagship, and Tinker's Fine-Tuning Bet

Thinking Machines Lab (founded 2025 by Mira Murati, $2B seed at a $12B valuation) released Inkling in July 2026 under Apache 2.0 (975B total / 41B active params, 1M context, native multimodality, controllable thinking effort) plus a smaller Inkling-Small (276B / 12B), paired with the Tinker fine-tuning platform—turning customizability itself into the product.

tech

AI-Native Agent 2026:從 Harness Engineering 到 Skill Engineering 的變革

2026 年 AI Native Agent 時代的核心變化:從 Harness Engineering 到 Skill Engineering,工程師角色從執行者轉向評判者,三個階段演進與五大關鍵技能。

MinerU:從 PDF 到 Markdown 的開源文件解析引擎,為 RAG 與 LLM 訓練而生

MinerU(OpenDataLab / opendatalab/MinerU)是一個開源文件解析引擎,把 PDF、圖像、DOCX、PPTX、XLSX 轉成結構化 Markdown 與 JSON,內建公式識別、表格提取與 109 語言 OCR,起源於 InternLM 前訓練階段,適合 RAG 前處理與知識庫建構。

Agentic Parsing: Letting Agents Decide How to Parse Documents

Traditional document parsing runs a fixed pipeline regardless of input, but contracts, financial reports, and technical manuals each need different strategies. Agentic Parsing lets LLM agents observe a document and dynamically choose tools — AgenticOCR parses only the regions that matter (70%+ visual token savings), and ParseBench shows even the best method scores only 84.9% across 2,000 enterprise pages. No silver bullet.

ColPali: Skip OCR, Retrieve Documents Directly from Images

ColPali renders each PDF page as an image, generates patch-level multi-vector embeddings with a vision-language model, and retrieves via MaxSim late interaction. On table-heavy financial PDFs, recall jumps from 62% to 84% — no OCR, no chunking. The tradeoff: ~100× storage, GPU required, no BM25.

Hierarchical Chunking + Auto-Merge: Small Chunks Search Well, Big Chunks Read Well

Small chunks give precise embeddings but lack context; big chunks have complete context but diluted embeddings. Hierarchical Chunking builds multi-level indexes (2048→512→128 tokens) with an Auto-Merge algorithm: leaf nodes match precisely, and when hit density exceeds a threshold the parent node is returned to the LLM instead. HiChunk shows a 12.7% evidence recall improvement; LlamaIndex and Haystack have it built in.

aideep-dive

LLM Dev Workflow Landscape: When Verification Becomes the Bottleneck

AI boosted task output by 34%, but code review time surged 441% and measured delivery actually slowed 19%. A four-round research survey maps the current landscape: deterministic guardrails (hooks) vs probabilistic ones (prompts), clean-context review, self-improving feedback loops, specification-driven development, AI test quality crisis (high coverage but median 53% mutation score), and the Replit agent fabricating test results.

Multi-hop Retrieval: When Answers Are Scattered Across Documents

Standard RAG retrieves one set of documents per query, but real questions often need reasoning across 2-4 documents. IRCoT pioneered interleaved retrieval-reasoning, PAR²-RAG beats IRCoT by 23.5% accuracy on four benchmarks, and CompactRAG compresses LLM calls down to just two.

aideep-dive

Self-RAG: Teaching the Model to Decide When to Retrieve

Self-RAG trains four reflection tokens (Retrieve / IsREL / IsSUP / IsUSE) into an LLM, letting it decide on-the-fly whether to retrieve, whether results are relevant, and whether its own output is grounded. ICLR 2024 Oral (top 1%), the 7B model beats ChatGPT and Llama2-chat + RAG on multiple QA benchmarks. The catch: it requires fine-tuning—no API-only models.

Table Serialization: How Format Choice Shapes RAG Retrieval for Tabular Data

Markdown-KV format achieves 60.7% LLM comprehension accuracy vs 44.3% for CSV — a 16-point gap from format alone. But retrieval and comprehension have different optimal formats: metadata prepend + row-wise key-value is the current best combination for table RAG.

learningguide

How to Practice Workplace English Speaking: Shadowing, Scenario Drills, AI Apps, and Meeting Phrases

30 minutes a day for 12 weeks — combine shadowing, scenario practice, and AI apps to go from 'I understand but can't speak' to holding your own in meetings.

techdebug

Claude Code Cloud Routines: When outcomes Pushes to a Feature Branch Instead of main

Cloud routine outcomes config creates a feature branch, so the agent commits and pushes there instead of main. Fix: remove outcomes + add explicit git checkout main in skills. Also hit a list API pagination bug (cursor never advances) along the way.

aidebug

Claude Code Routine Connector Keeps Asking for Permission: The Hidden created_via Field

Routines created by Claude itself (created_via: meta_mcp) prompt for connector approval on every call. User-created ones (created_via: http_api) don't. Fix: recreate the routine via the RemoteTrigger API.

tech

Codex 架構總覽:Rust Monorepo、Bazel 建構、跨平台沙箱

Codex 以 Bazel 管理 140+ Rust crate,核心分為 core/tui/exec-server/protocol 四大塊;沙箱用 codex_sandboxing 統一 macOS Seatbelt、Linux Landlock/bwrap、Windows 沙箱三平台介面;exec-server 以 JSON-RPC + Noise Relay 實現遠端執行。

tech

Codex ThreadManager:核心協調者的生命週期、分叉語義與 Subagent 圖譜

ThreadManager 持有 Arc<ThreadManagerState> 統管所有 thread,spawn_thread() 統一處理新建/恢復/分叉/子代理四種啟動路徑;ForkSnapshot 定義 TruncateBeforeNthUserMessage/Interrupted 兩種語義;AgentControl 透過 Weak<ThreadManagerState> 避免循環引用;agent_graph_store 追蹤 ThreadSpawnEdgeStatus::Open/Closed。

tech

Codex Turn 狀態機:TurnContext、StepActivation、Context Manager 與壓縮觸發

TurnContext 在 turn 初始化時捕獲所有設定(模型、審批、token budget),後續 step 透過 StepContext 讀取快照;StepActivation 驗證設定變更不違反 legacy 安全約束;ContextManager 用 Arc<Vec> + 版本號實現 Copy-on-Write 歷史共享;壓縮觸發條件為 token_remaining < threshold,支援 remote v1/v2、local、model fallback 四條路徑。

OMP agent loop: Why two while loops? What the outer 'stopped but woken by steering' layer actually does

omp's runLoopBody uses a double while loop: the inner loop drives the core model call → tool execution rhythm; the outer loop, when the agent would stop, drains queued steering / follow-up / asides to decide whether to run another turn. This design solves delivery timing for 'user typing while model streams' and 'background tasks quietly queueing messages'.

OMP append-only context: Why sync conversation by byte-stable prefix? How Anthropic/DeepSeek KV cache gets protected

omp uses StablePrefix to freeze system prompt + tool specs, AppendOnlyLog for append-only messages, and digest-based longestStablePrefix algorithm. When prune/shake/steering rewrite history, only the tail after the divergence point is resent. This maximizes Anthropic/DeepSeek prompt cache hit rate, fixing the old issue where every turn forced ~40k token re-prefill on llama.cpp (issue #3406).

OMP three-layer approval & fail-closed: why undeclared custom tools become exec, and what yolo still blocks

omp's resolveApproval resolves in three layers: tool declaration → user override → mode tier. Undeclared or malformed approvals default to exec (fail-closed). Tool declares tier + optional policy/override/reason/policyKey; user overrides via tools.approval.<tool>; mode (always-ask/write/yolo) sets auto-allow tier ceiling. Iron laws: tool-side deny and user-side deny can never be crossed by mode; yolo ignores override: true but still honors policy: deny|allow|prompt. bash tokenizes approval: allow must cover whole line, deny/prompt match per segment. Same tool switches read/write via policyKey. checkpoint/rewind are paired sisters. subagent runs headless yolo; parent task is the only auth boundary.

OMP bash tokenized approval: why allow must cover the whole line while deny/prompt match per segment — the design cost of bash.patterns glob

omp's bash approval engine tokenizes commands into segments via a shared shell tokenizer (split on `;`, `&&`, `||`, `|`, `&`, newline, subshell). deny/prompt rules match glob against each segment individually — any hit triggers. allow rules require whole-line match AND no shell control syntax, preventing `cd x && rm -rf /` from slipping through. CRITICAL_BASH_PATTERNS hardcodes 45 dangerous command regexes (`rm -rf /`, `chmod -R 777 /`, `curl | bash`, `kill -9 1`, etc.) that fire before user patterns and cannot be disabled. Design tradeoff: allow is strict for safety, deny/prompt permissive for catch-all.

OMP Internal Design #13: collab-web & wire protocol — How Multi-User Sessions Serialize, Guest Permission Boundaries, and Composer Interrupt Broadcast

omp collab uses a hub topology (host authoritative, guests never peer) with AES-256-GCM sealed JSON frames, 4-byte envelope, and snapshot-chunk sharding. Guest permissions are bound to a 16-byte write token in the link, verified by the host via timing-safe comparison.

OMP Internals #10: hooks/skills/MCP/marketplace — The Four Extensibility Surfaces

OMP unifies hooks, skills, MCP, and marketplace into a cohesive extensibility layer: hooks merge into extension runner's event bus, skills use description for semantic triggering, MCP adopts a 250ms fast-start gate with deferred fallback, marketplace supports dual scopes with Claude Code catalog compatibility.

OMP four compaction strategies: context-full / snapcompact / branch summary / shake — what each solves and how they switch

omp doesn't have just one compaction: context-full uses LLM summarization with iterative windows and budget halving retries; snapcompact skips the LLM entirely, rendering history as dense PNG bitmaps for vision models to read — solving no-API-key, low-latency, vision-model-cheaper scenarios; branch summary summarizes the abandoned branch during `/tree` navigation so file ops aren't lost; shake mechanically replaces tool-result text and large fenced/XML blocks with placeholders — an emergency hatch when summary is too heavy and prune isn't enough. Four strategies, distinct failure modes, orchestrated by session maintenance or manual triggers.

OMP hashline edit & noop-loop-guard: Why file edits must be hash-anchored, and how 182/205 byte-identical no-op retries were tamed

OMP replaces traditional line-number patches with hashline: 4-hex content hash + N* syntactic block locators eliminate whitespace drift. noop-loop-guard tracks per-session, per-canonical-path, per-input-hash consecutive no-ops; after 3 (NOOP_HARD_LIMIT) it throws a ToolError so the agent loop sees a tool failure — breaking the model's 182/205 retry loop from issue #2081 where soft hints were completely ignored.

OMP KDL Rule Tree & 60+ Provider Routing: Why Provider Policy Lives in KDL, Not TypeScript

OMP expresses routing, compat, thinking, quota, and pricing policies for 60+ providers entirely in KDL (taxonomy/classes/providers/runtime), compiles to rules.json, and resolves at runtime via a typed engine. Three reasons: layered ownership, compile-time validation, and mechanical priority resolution—all three would scatter, become unverifiable, or rely on human judgment if hardcoded in TS.

OMP Internals (14): metaharness & Benchmark Infrastructure

OMP's metaharness isn't just a benchmark runner—it's experiment-grade infrastructure for hardware-isolated microVMs, unified auth gateway, and baseline-controlled experimental design.

OMP Internals (8): Provider Quirks & Compat Layer — Why KDL Rules, Not TypeScript

OMP encodes all provider-specific wire behaviors as KDL rules, compiled to JSON and applied by a pure cascade resolver — avoiding TS if-else sprawl, enabling static conflict detection at CI, and making behavior versionable.

OMP Internals 12: Rust Native Crates & FFI Contracts

OMP sinks performance-sensitive, correctness-critical, and determinism-requiring subsystems (grep, AST, PTY, isolation, file walking) into 6 independent Rust crates, then exposes a unified N-API surface via pi-natives; bindings are generated by napi-rs with gen-enums.ts patching runtime enums and explicit ESM exports.

OMP session persistence, fork & tree: entry model, parent chains, /tree navigation, how navigateTree cuts leaf back to old branches

omp stores sessions as append-only JSONL with id/parentId forming a tree; a mutable leaf pointer selects the active path. /tree uses TreeSelectorComponent for navigation; navigateTree() switches leaf with optional branch summary, handles checkpoint/rewind, supports Claude/Codex foreign session import. Fork duplicates the session file inheriting providerPromptCacheKey. Resume/switchSession uses captureState/rollback for atomic transitions. Three-layer storage (indexed/history/artifact) separates concerns.

OMP streaming internals: the event stream is an agent control plane, not a token stream

OMP's Agent stream does more than print tokens: it separates agent, turn, message, and tool-execution events. Understand the contract to render text deltas, tool progress, and errors without mistaking a partial message for committed conversation state.

OMP Internals Deep-Dive (11): TUI Differential Rendering & Composer

OMP builds a custom TUI engine (packages/tui + packages/coding-agent/src/tui) rather than using React/Ink. Core designs: differential rendering (repaint only changed lines via reference equality), explicit history contract, composer multi-modal input & keybindings system. This article analyzes architecture decisions and implementation details from source code.

OMP vs looplane retrospective: what to borrow, what's over-engineered, what not to touch yet

14 posts in, back to the big picture: append-only context, compaction strategy split, three-layer approval, KDL rule tree, session tree/fork are the five most borrowable; snapcompact, metaharness self-built infra are over-engineered; looplane shouldn't self-build full provider catalog, full TUI, or full collab yet. Philosophy diff: omp = batteries-included in-process, looplane = minimal + external runtime.

pi-mono Deep Dive 13: Agent Harness, Skills, System Prompt Assembly — Building Agent Behavior from Scratch

AgentHarness Core Class, System Prompt Dynamic Assembly Flow, Skills Loading & Formatting, Prompt Templates System, How Harness Decides Tool Availability, Result Handling, Telemetry Schema Registration, Default Harness Construction, Extension Harness Extension.

pi-mono Deep Dive 4: Agent Loop — Double-Loop & Event Flow, From Steering to Follow-up Complete Timeline

Heart of pi-agent-core: agentLoop() → runLoop() double while(true). Inner loop handles tool calls + steering messages; Outer loop handles follow-up + prepareNextTurn (compaction, model switch). Enter = steering (inject after current tool), Alt+Enter = follow-up (inject after agent stops). streamAssistantResponse() partial message updates, tool call parsing, parallel/sequential execution, before/after hooks.

pi-mono Deep Dive 1: pi from a CLI User's Perspective — Install, Modes, Session, Model Switch & Message Interjection

Treat pi as a black box first: 4 run modes, session tree persistence, mid-conversation model switching, Enter vs Alt+Enter message interjection, /tree branch navigation. Builds intuition for the architecture parts that follow.

pi-mono Deep Dive 12: Compaction Deep Dive — Strategy, Token Estimation, Branch Summary, Structured Compaction

Complete Compaction Mechanism: shouldCompact Trigger Conditions (Token Ratio, Message Count), estimateTokens Calculation (Char/Word Approximation), findCutPoint Finding Cut Point (Retain Recent N Turns), generateSummary Generating Summary (LLM Call), prepareCompaction Preparing Context, Branch Summary Generation, Structured Compaction (Extension Custom via fromHook), CompactionEntry Details, fromHook Mechanism, Compaction Settings.

pi-mono Deep Dive 15: Containerization, Sandbox, Permission Model — Gondolin, Docker, OpenShell, Security Boundaries

Why Pi has no built-in permission system, Gondolin Extension (micro-VM), Docker mode, OpenShell policy-controlled sandbox, permission model philosophy, three containerization patterns, security boundary comparison, micro-VM vs container vs process isolation.

pi-mono Deep Dive 7: Extension System — Hooks, Custom Tools, UI Components, Lifecycle Complete Mechanism

Complete Extension system analysis: Extension interface definition, onLoad/onUnload lifecycle, four major Hooks (onAgentStart/onBeforeToolCall/onAfterToolCall/onTurnEnd), five extension points (tools/commands/keybindings/ui/settings), ExtensionRunner load order and dependency resolution, ExtensionAPI capabilities, Dynamic Border, Widget, Dialog, Selector UI components, Extension inter-communication, hot reload mechanism, official example Extensions.

pi-mono Deep Dive 9: Model Catalog, Provider Factory, OAuth & Credential Sync — From Auto-Generation to Cross-Device Sync

pi-ai Model Catalog auto-generation flow, Provider Factory registration with Lazy Loading, OAuth 2.0 + PKCE flow implementation, Credential Store (Keychain/Libsecret/Credential Manager/Encrypted File Fallback), Credential Sync cross-device sync mechanism, Model Scope Diagnostics, ModelResolver parsing logic, CredentialSynchronizationOperation state machine.

pi-mono Deep Dive 2: Monorepo Architecture & Core Abstractions — How 7 Packages Divide Work & Why Dependencies Flow One Way

From user-visible features into architecture: 7 npm packages with clear boundaries, one-way dependency flow, why pi-tui/pi-telemetry have zero deps, how pi-ai encapsulates provider details, lockstep versioning avoiding diamond deps. Builds an 'outside-in' mental model.

pi-mono Deep Dive 3: pi-ai — Unifying 15+ LLM Providers, From Lazy Loading to Auto-Generated Model Catalog

pi-ai is pi-mono's anti-corruption layer: upper layers only see Message/Tool/Context/streamFunction; 15+ providers implement details underneath. This post dissects: unified interface design, Provider Factory Registry, Lazy Loading for tree-shaking, Model Catalog auto-generation, OAuth/API Key unification, Credential Sync, Thinking/Reasoning parameter standardization.

pi-mono Deep Dive 16: Release Pipeline — Lockstep Versioning, Binary Build, Trusted Publishing From Code to npm

Full release flow: Lockstep versioning (all packages same version), CHANGELOG, local smoke, release script, Bun+Node binary build, npm-shrinkwrap, GitHub Actions OIDC trusted publishing, R2 release marker, pi.dev/api/latest-version, announcement verification.

pi-mono Deep Dive 10: Remote Session — Client/Server, JSON-RPC 2.0, WebSocket, Reconnection

pi-protocol JSON-RPC 2.0 Definition, pi-client Connection Management & Exponential Backoff Reconnection, pi-server Session Registry, WebSocket Transport, Heartbeat Mechanism, Session Snapshot, Remote Session Handle, RPC Mode Architecture, Streaming Event Transport, Remote Steering/Follow-up Message Interjection.

pi-mono Deep Dive Series: From Zero to Understanding This Minimal Coding Agent's Complete Architecture

This 17-part series takes you from CLI user perspective through pi-mono's Agent Loop, Session Tree, Tool System, Extension System, TUI Architecture, Remote Session, Telemetry, Compaction, and Release process. Ideal for developers wanting to self-host agents, research agent architecture, or contribute to pi.

pi-mono Deep Dive 5: Session Tree — Append-only JSONL, Branching Without History Mutation, Compaction Logic Full Analysis

SessionManager core: JSONL append-only storage, id/parentId tree formation, branch() moves leaf pointer without mutating history, buildSessionContext() handles compaction entries, createBranchedSession() forks to new file. Complete Entry types: message, thinking_level_change, model_change, compaction, branch_summary, custom, custom_message, label, session_info. Migration v1→v2→v3 details.

pi-mono Deep Dive 11: Telemetry — Vendor-neutral Contracts, Schema Definition, Conformance Tests

pi-telemetry Core: TelemetrySchema Defines Span/Event/Attribute, defineTelemetrySchema Creates TypedSpanStarter, InMemoryTelemetryContext/NOOP_TELEMETRY_CONTEXT Zero-overhead Implementations, Conformance Tests Verify Adapter Correctness, AI/Harness Telemetry Schema Complete Definitions, Attribute Type System, Why Not Use OpenTelemetry Directly.

pi-mono Deep Dive 14: Testing, Quality Gates, Supply-chain Hardening — Faux Provider, Browser Smoke, Biome, tsgo, Shrinkwrap, Trusted Publishing

Testing strategy: Faux Provider (no API key e2e), Vitest unit, Browser Smoke (real browser), Biome lint/format, tsgo type check, Pinned Deps, Shrinkwrap, Install Lock, npm Trusted Publishing, CI pipeline.

pi-mono Deep Dive 6: Tool System — 8 Core Tools, Factory Pattern, Parallel/Sequential Execution, Before/After Hooks Interception

Complete analysis of pi-coding-agent's 8 core tools: ToolDefinition (for LLM) vs AgentTool (execution logic), createToolDefinition/createTool Factory, executionMode determines parallel/sequential, beforeToolCall/afterToolCall interception chain, withFileMutationQueue serializes file writes, truncateHead/Line/Tail output truncation, read/write/edit/bash/grep/find/ls/powershell implementation details.

pi-mono Deep Dive 8: TUI Architecture — Differential Rendering, Component Tree, Layout Engine, CSI 2026 Synchronized Output

Complete pi-tui core analysis: Virtual DOM Diff for flicker-free rendering, Component lifecycle, Layout Engine (Flex-like), CSI 2026 Synchronized Output avoiding partial frame tearing, Keybindings Manager, Alt Screen, Bracketed Paste, Kitty/iTerm2 Image Protocol, built-in components (Markdown, Editor, Selector, Diff, Border, Loader, etc.).

Learning Agent Design from Mature Coding Agents (4): Approval Grading and the Audit Trail

Looplane now grades effects as read/modify/modify_execute/execute and fails closed on unclassified tools. Native MCP tools default to execute unless trusted read-only metadata lowers them. Approval events still land in events.jsonl first and grants can scope to one change set or backend; general command rules and universal sandbox coupling remain unfinished.

Learning Design from Mature Coding Agents (37): Code Mode — Compiling Tool Calls into Batches of Executable Code

looplane now ships a bounded tool-program DSL: read-only programs support list/read/search/diff, repeat, and if_contains; modify/check transactions receive whole-transaction approval and roll back touched paths on failure. This is not arbitrary JavaScript/Python code mode, and transaction execution is not parallel.

Learning Design from Mature Coding Agents (26): Context Compression and Compaction — From Gap to Auditable Baseline

Mature-agent compaction must handle triggers, complete-turn cut points, and recovery. looplane now has an 85% high-watermark, automatic compaction, a deterministic native-loop fallback summary, persisted checkpoints, and workspace-context reinjection. Cross-runtime fallback, model-quality summaries, and live-provider long-session validation remain open.

Learning Agent Design from Mature Coding Agents (27): Cross-Session Memory — From Explicit Remembering to Semantic Recall

omp and claude-code provide cross-session memory while the other references mostly rely on instruction files. looplane now has an explicit remember/list/inject baseline: typed JSONL memories enter prompts across sessions, but retrieval is scope-and-recency only, with no semantic ranking, deduplication, forget command, or automatic extraction.

Learning Design from Mature Coding Agents (28): Dangerous Command Interception and Shell Escalation — Between Allowlists and Always Ask

All five projects combine allow/ask/deny decisions, compound-command inspection, and fail-closed behavior. looplane now has a deny-first classifier, critical floor, shell segmentation, timeout-deny, configured allow/deny rules, and visible policy reasons. Broader syntax coverage and live interactive validation remain open.

Learning Agent Design from Mature Coding Agents: Series Overview — Reading Five Codebases to Build My Own

I'm building my own Python coding agent called looplane. This series dissects the source code of five mature projects — pi, oh-my-pi, opencode, codex, and claude-code — topic by topic, while also comparing them with Looplane's current TUI, external CLI runtimes, local gateway, usage/OTel/session tooling, and Cloudflare slice. Every post follows a fixed five-part structure: design problem → how five projects do it → looplane's choice → academic grounding → improvement roadmap, with evidence cited at file#symbol level.

Hooks, Skills, Plugins: The Three-Layer Extension System of Mature Coding Agents

Hooks govern control flow, skills inject knowledge, and plugins package both. looplane now has opt-in deny-only project hooks, a bounded SKILL.md loader, exact enabled_skills selection, plugin manifests/install/list, and external-runtime projection. Input rewriting, full lifecycle coverage, remote registries, and a mature marketplace remain open.

Learning Design from Mature Coding Agents (36): LSP Integration — Pushing Compiler Diagnostics into Agent Context

looplane can now inject repository diagnostics and open-file state into the next model turn, push typed IDE context over WebSocket, package a VS Code bridge, and supervise long-lived LSP subprocesses through ManagedLspServer. Language-specific initialize/didOpen/didChange adapters and live-editor validation remain open.

Learning Design from Mature Coding Agents (30): MCP Integration — the Standard Socket for Tool Ecosystems

An MCP client must handle transports, tool refresh, approvals, and credential boundaries together. looplane now supports allowlisted stdio, Streamable HTTP/SSE, tools/resources/prompts, tools/list_changed, OAuth metadata/PKCE, and a 0600 credential store. A real authorization-server E2E and MCP-specific confirmation UX remain open.

Learning Design from Mature Coding Agents (35): Model Catalogs and Per-Role Routing — looplane's Role Aliases and Reviewer Lane

looplane now has static ModelRole/ModelRoute candidates, opt-in aliases such as --model @cheap, cross-provider fallback, and a no-tool reviewer lane that runs after verification. Role inheritance/override rules and automatic summarizer, parser, or scout routing remain open.

Learning Design from Mature Coding Agents (6): The ModelProvider Abstraction — Why Wrapping an SDK Is Not Enough

Wrapping an SDK directly buys you three walls within months: usage fields that don't agree, error semantics tied to SDK exception types, and tool-call formats that change per provider. All five reference projects separate 'wire protocol' from 'provider identity' as independent dimensions. Looplane goes further with pydantic canonical contracts (Message/ToolCall/Usage/ModelTurn) plus six protocol adapters, forces the OpenAI SDK's built-in retries to zero, and routes every failure through a classified ProviderErrorKind before any retry policy sees it. Its provider table is deliberately copied from pi's packages/ai — lineage, not coincidence.

Learning from Mature Coding Agents (29): OS-Level Sandboxing

An OS sandbox is the kernel boundary beyond path policy. looplane now ships a fail-closed CommandSandbox: sandbox-exec on macOS, Landlock plus seccomp on Linux, and exit 126 when containment cannot be proven. Coverage still focuses on verification commands, and external CI confirmation remains open.

Prompt Version Control: Changing One Word Can Drop an Eval from 5/5 to 0/5

Looplane's prompt is now `m3-exact-edit-v4`: the version persists into artifacts; core/tool/interaction/runtime/instructions/skills/workspace/memory are composed as stable or dynamic sections; and positive/negative examples cover replace_text, unified diffs, and direct replies. Unit tests pin the structure, while live-eval coverage still needs expansion.

Learning Design from Mature Coding Agents (7): Provider Retry Policy — From One 5xx to Bounded Retry and Fallback

Intermittent NVIDIA NIM 500s exposed Looplane's early gap: classified errors with no retry consumer. SDK retries are now disabled; the harness gives each candidate up to five attempts with jittered exponential backoff and capped Retry-After handling, then can move to an explicitly configured fallback model. Both model.retry and model.fallback enter the event log.

Learning Design from Mature Coding Agents (33): Session Recording and Replay — From Event Logs to Safe Forks

looplane now connects events.jsonl to a deterministic reducer, CLI timeline, canonical JSON, SDK replay, and safe event-point forks. Forking never replays prior tools or model calls; provider/live-runtime validation, redaction, and richer replay hooks remain open.

Learning Design from Mature Coding Agents (32): Subagents and Worktree Isolation — Teaching the Main Loop to Delegate

Mature subagents need roles, bounded fan-out, narrowed permissions, and a result contract. looplane now has native named-role schedules, parallel fan-out, child allowed_paths constrained by the parent, unsafe execution disabled by default, and parent-approved transaction proposals. Persistent background lifecycles, recursion trees, and automatic worktree merging remain open.

Learning from Mature Coding Agents (34): Telemetry and Cost Tracking — You Count Tokens, Then What?

looplane now has CostBreakdown, an explicitly estimated static GPT-5-family price table, per-lane usage/cost, and OTel cost fields. Unknown models still show tokens without invented dollars; broader pricing coverage, authoritative external-CLI bills, and live billing reconciliation remain open.

Learning Design from Mature Coding Agents (18): Toolset Design Philosophy — Drawing the Tool Surface Boundary

Looplane's core surface has grown from seven tools to nine with a read-only `tool_program` and rollback-capable `tool_transaction`; search prefers ripgrep and arbitrary shell remains absent. Native MCP tools join only from allowlisted servers and default to execute approval without trusted read-only metadata.

Learning from Mature Coding Agents (3): Workspace Isolation and Path Policy

Looplane's disposable clone and SafePathPolicy protect the source repo. `--sandbox-checks` can now wrap verification commands with macOS sandbox-exec, Linux bubblewrap, or Landlock, while Cloudflare provides a separate bounded Sandbox slice. Network policy, external-runtime coverage, and production hardening are not yet consistent across those backends.

Learning Design from Mature Coding Agents (38): Agent as a Service — Wrapping Your Loop in Something Other Programs Can Call

looplane now has a Cloudflare Durable Object run resource with async creation, status/cancel/artifacts, live NDJSON, and Last-Event-ID SSE; remote approvals use a separate short-lived capability. Python also provides an attach client and a stateful conversation WebSocket. Production deployment, cross-runtime parity, and full multi-tenant hardening remain unverified.

Why Ask AI Could Not List Its Course Maps: A Catalog Retrieval and Convergence Incident

The first observation of 'What course articles do you have?' was contaminated by an old cache entry. A real cache miss retrieved all four university maps but spent 51.169 seconds across three Writer and Critic passes; after catalog-specific retrieval and review fixes, one uncached production observation passed q21 in 26.821 seconds.

How Ask AI Finds Posts: Planner, Hybrid Retrieval, and Retry

Ask AI first extracts intent, complexity, and 1–4 search terms. It then routes across metadata, BM25, Vectorize, and RRF; a retry adds Critic gaps and disables the first-pass-only BM25 short circuit.

How Ask AI Indexes Posts: Chunks, D1 FTS5, and Vectorize

Ask AI indexing runs in two production stages: source-hash changes update D1, post chunks, and FTS5 first; embedding checkpoints and a delete queue then let Vectorize catch up asynchronously. The two stores do not share one transaction.

How a Question Moves Through Ask AI: UI, API, Agents, and Source Cards

Ask AI splits one question across the UI, `/api/chat`, Planner, Research, Writer, Validation, Critic, and Related stages. Answer text, displayed sources, and related-reading cards come from separate paths with separate gates.

Evaluating Ask AI Retrieval: Golden Contracts, Fixtures, Live Runs, and Evidence Boundaries

Ask AI keeps golden contracts, offline fixtures, live SSE output, and production observations separate. A passing fixture proves harness reproducibility; public sources can measure expected-source recall, but they do not expose hidden ranked chunks or establish model-graded faithfulness.

Debugging Ask AI in Production: SSE, Traces, Cache, Checkpoints, and Shadow Runs

One Ask AI request leaves five different evidence surfaces: public SSE, Langfuse traces, D1 logs, semantic cache, and a hidden shadow run. They expose different data, and no single surface reconstructs the complete retrieval context.

When Ask AI May Show Sources: Validation, Critic Review, Degradation, and the Source Gate

Ask AI finding a post does not mean the UI should display it as a source. An answer must pass deterministic Markdown and URL validation, then the Critic's relevance, intent, and grounding checks; if either gate fails, source cards are withheld.

How Ask AI Turns Evidence into an Answer: Writer Context and Citation Contracts

Writer sees the first 8 candidates for a factual query or 12 for a recommendation by default. Citations must use an exact `source_url` from that set, and weak or empty retrieval triggers an instruction to abstain rather than fill gaps from model knowledge.

How to Use Cloudflare Agent Memory: Keep Agent Memory Separate from RAG Documents

Agent Memory is a Cloudflare private beta service for letting agents remember users, teams, projects, and task context across conversations. It fits facts, events, instructions, and tasks; RAG documents, product data, files, and audit logs should still live in AI Search, Vectorize, D1, or R2.

How to Use Cloudflare Agents: Durable Runtime, Tools, and Real-Time Connections

Cloudflare Agents turns an agent session into a durable runtime: each agent instance has stable identity, local SQLite, WebSockets, scheduled work, recoverable execution, and tools. It is not just a chat example; it composes Workers, Durable Objects, AI models, Browser, Sandbox, AI Search, and MCP into a deployable agent app.

Where to Store Cloudflare AI App Data: D1, R2, and Durable Objects

An AI app should not put conversations, artifacts, memory, retrieval documents, locks, and eval traces into one store. D1 fits queryable product data, R2 fits large files and artifacts, Durable Objects fit named coordination and per-session state, and Agent Memory / AI Search / Vectorize handle memory and retrieval.

How to Use Cloudflare AI Gateway: Logging, Caching, Rate Limits, and Fallbacks

AI Gateway is the control plane for AI calls: one layer for logs, analytics, cache, rate limits, retry/fallback, BYOK, and Unified Billing. In Workers, use env.AI.run(..., { gateway }); with external SDKs, change the baseURL or provider-native endpoint.

Cloudflare AI Stack Guide: Building AI, RAG, and Agents on Workers

The Cloudflare AI Stack series covers the infrastructure around AI apps: where models run, how gateway control works, how RAG is built, how agents keep running, how memory is governed, and how browser, sandbox, secrets, data, and observability fit into a product.

How to Use Cloudflare Secrets Store: Worker Secret Reuse and AI Gateway BYOK

Secrets Store is Cloudflare's open beta account-level secret store, currently integrated with Workers and AI Gateway. It fits provider API keys, BYOK keys, and secrets reused across Workers; per-Worker secrets still work, but the governance scope is different.

How to Use Cloudflare Vectorize: Taking Control of RAG Retrieval

Vectorize is Cloudflare's vector database. AI Search is the right starting point for a managed RAG pipeline; Vectorize is the better fit when you need control over chunking, embeddings, metadata filters, hybrid retrieval, reindexing, and fallback behavior.

Looplane remote execution on Cloudflare: Worker, Sandbox, Capability DO, and durable RunSession

Looplane's old synchronous M6 path completed one real deployed coding run. It has since grown into an asynchronous control plane with RunSession, SSE, approvals, cancellation, and artifacts, but that newer path has not been live-revalidated. Audience-separated HMAC capabilities enter the Sandbox while provider credentials stay in the Worker; this is not production-traffic or SLO proof.

Looplane's disposable workspace and run bundle: why the source repository stays untouched

Looplane clones an exact Git commit into a detached-HEAD workspace inside the run directory before a runtime edits or verifies code. The source repository, execution workspace, and run artifacts therefore have distinct boundaries. This provides source isolation and an audit bundle, but it is not an OS sandbox.

Looplane's ExternalCodingRunner: why Codex and Claude Code CLI are external runtimes, not ModelProviders

`ExternalCodingRunner` is Looplane's second runtime lane. The external coding CLI owns its model loop and credentials; Looplane hands off a task and disposable clone, then treats the returned patch as untrusted input and reruns path audit, verification, and the source invariant. This is a capability-bounded handoff, not another `ModelProvider`.

Looplane's ModelProvider multi-gateway: multiple protocols, one canonical contract

Looplane collapses OpenAI-compatible, Responses, Anthropic, Gemini, Workers AI, scripted, and experimental Codex OAuth adapters into one `ModelProvider` contract. The Codex OAuth transport reads SSE but still reduces it inside the adapter into one canonical `ModelTurn`; AgentRunner does not consume token deltas.

Looplane's provider-neutral native loop: from one model turn to a verified terminal state

Looplane's native lane is controlled by AgentRunner: prepare a workspace, request a model turn, execute tool calls, append observations, and enter verification only when the model stops calling tools. Step, wall-time, repetition, token, and cancellation guards can terminate the run independently of the model. Protocol translation belongs to the next article.

Looplane's state-first event journaling: recovering between manifest commits and JSONL appends

Looplane maintains append-only `events.jsonl` and atomically replaced `session.json`, reconciling sequences before crash recovery. The same event contract now supports deterministic replay, canonical JSON replay, fork seeds at a selected sequence, and new workspaces without replaying old side effects; ambiguous `tool.started` or `verification.started` states still hard-fail.

Looplane's tool isolation: path allowlists, strict argv, process groups, and credential-free subprocesses

This article follows one Looplane tool call through its mechanical execution boundary: `SafePathPolicy` for paths and symlink escape, fixed argv with `shell=False`, a sanitized subprocess environment, read-version hashes plus atomic replace for writes, and process-group cleanup at timeout. Permission policy, OS containment, and tool programs are reserved for later articles.

Looplane's TUI and CLI: how a run becomes visible in the terminal

Looplane's TUI and plain CLI are two interfaces over the same runtime paths. The CLI selects a presentation mode from TTY state and flags, runners emit events, and the TUI projects those events into thinking, tool, approval, verification, and terminal states. The screen distinguishes native and external runtimes without treating UI entry points as proof of backend maturity.

How to Use Cloudflare Browser Run: Headless Chrome from Workers

Browser Run gives Workers access to Cloudflare-managed headless Chrome. Quick Actions fit one-shot tasks such as screenshots, PDFs, HTML, JSON, and crawls; Browser Sessions fit Puppeteer, Playwright, CDP, and Stagehand automation where you need full control.

Cloudflare Cache Rules: What to Cache and What Must Stay Dynamic

Cloudflare Cache Rules are zone-level cache policy: request expressions decide what is eligible for cache, how Edge TTL and Browser TTL behave, what dimensions enter the cache key, and how stale content, ETags, and purge interact. Use them for CDN cache policy; use the Worker Cache API for programmatic caching.

How to Use Cloudflare Containers: When Workers Need a Full Linux Runtime

Cloudflare Containers let a Workers app call on-demand serverless containers for workloads that need a full filesystem, a specific runtime, existing container images, or more CPU, memory, and disk. They do not replace Workers; Workers still handle entry, routing, and platform bindings while containers run the heavy runtime-specific work.

Cloudflare Edge Platform Guide: Running Websites and Apps on Cloudflare

The Cloudflare Edge Platform series answers one product question: how do Workers, D1, KV, R2, Durable Objects, Queues, Workflows, Cache, Images, Email, Turnstile, Observability, Browser Run, and Containers help you run a website or app cheaply and reliably?

Cloudflare Edge Platform Production Checklist: Custom Domains, Maintenance Pages, and Workers Limits

Before a Cloudflare app goes live, do not stop at a successful deploy. Check Custom Domains, Routes, www/root redirects, maintenance pages, CPU/memory/subrequest limits, log sampling, and fallback paths. This appendix turns the Edge Platform series into a production checklist.

How to Use Cloudflare Email Service: Sending, Routing, and Product Notifications from Workers

Cloudflare Email Service connects transactional email, magic links, notifications, and inbound routing to Workers. Arbitrary outbound sending currently requires Workers Paid; inbound routing is available on Free and Paid, with DNS, quota, message-size, bounce, and anti-spam limits still shaping the design.

How to Use Cloudflare Hyperdrive: Connecting Workers to Existing Postgres / MySQL

Hyperdrive solves the latency and connection-pooling problem when Workers connect to existing Postgres / MySQL databases. It uses edge connection setup, database-near pooling, and read query caching so a regional database works better with global Workers.

How to Use Cloudflare Images: Variants, Format Conversion, and Delivery Pipelines

Cloudflare Images has two paths: transform images stored in R2/S3/origin at the edge, or store images in Images and deliver named variants. The first is priced by unique transformations; the second also involves stored and delivered images.

How to Use Cloudflare Observability: Workers Logs, Traces, and Analytics Engine

Workers Observability is for debugging and request tracing; Workers Analytics Engine is for high-cardinality product events and custom metrics; GraphQL Analytics API is for querying existing Cloudflare product data. Keeping those roles separate prevents logs from becoming a database and keeps billing, monitoring, and product analytics from blending together.

How to Use Cloudflare Smart Shield: Reducing Origin Load

Smart Shield is Cloudflare's origin protection bundle: Smart Tiered Cache, connection reuse, Argo Smart Routing, Regional Tiered Cache, Cache Reserve, Health Checks, and Dedicated CDN Egress IPs reduce requests and connections reaching your origin.

How to Use Cloudflare Turnstile: Protect Forms and Public APIs Without Classic CAPTCHA

Turnstile is Cloudflare's CAPTCHA alternative: the client widget generates a token, and the server must validate it with the Siteverify API. Tokens expire after 300 seconds and are single-use; a widget without server validation is incomplete.

Cloudflare Workflows: Durable Multi-Step Execution on Workers

Cloudflare Workflows turns multi-step Workers processes into durable steps: each step can retry, sleep, wait for events, and register rollbacks, while instances can be inspected, paused, resumed, or terminated. Queues fit single-step background work; Workflows fit long processes that must remember progress.

techdeep-dive

How to Choose an Execution Environment and Sandbox: From Namespace and gVisor to Firecracker, E2B, and Lambda MicroVMs

A sandbox is not a single package but a spectrum—Namespace, cgroups, seccomp, gVisor, and Firecracker stacked by trust boundary; local OS sandboxes bound blast radius, cloud microVMs bound multi-tenancy, and the choice hinges on trust and ops cost.

A map of Looplane: how one coding-agent task crosses workspaces, runtimes, tools, and events

Looplane turns a coding-agent task into inspectable boundaries: native side effects cross Looplane tools, permissions, and sandboxing, while external runtimes retain their own loops and tools before returning a patch for Looplane audit. This article maps the planned 20-part series.

Looplane context pressure, compaction, and workspace reinjection

Near 85% context pressure, Looplane has two distinct paths: the native loop can apply one bounded deterministic history fallback, while a conversation runtime with native compaction can compact after a completed turn. Both paths re-anchor the next request with workspace context.

Looplane IDE/LSP Context: Diagnostics, Open Files, and the VS Code Bridge

Looplane normalizes up to 200 diagnostics and 32 visible files into bounded, repository-local, untrusted context. Its VS Code and managed-LSP paths supply signals rather than completion, rename, code actions, or full IDE RPC.

Looplane local OS sandboxes: fail-closed execution on macOS, bubblewrap, and Landlock

Looplane can wrap configured local commands and verification in macOS sandbox-exec, Linux bubblewrap, or Landlock/seccomp. A required unavailable backend stops with exit 126 instead of running bare, but external CLIs, MCP/LSP processes, and the entire Looplane process are outside this boundary.

Looplane model roles, fallback, cache hints, and estimated cost

Looplane uses a static model-role catalog and retries or falls back only after retryable provider errors. Cache data is a provider hint plus trace, while cost is a static-table estimate; neither is live routing intelligence or a bill.

Looplane Native MCP: transport, authorization, and approval boundaries

Looplane loads project MCP servers only through an explicit allowlist, projects stdio or Streamable HTTP capabilities into the existing ToolExecutor, and preserves hooks, approvals, timeouts, and cleanup.

Looplane permission layering: how dangerous commands become allow, ask, or deny

Looplane applies a non-bypassable critical floor, evaluates user, organization, and project denies before any allows, and keeps execute operations policy-gated even in dangerous mode. This decides authority; it is not an OS sandbox.

Looplane prompts, instruction precedence, and explicit memory: what the model actually sees

Looplane resolves user and root-to-leaf project instructions before rendering named prompt sections for runtime, skills, workspace state, and the latest 20 explicit memories. The pipeline is traceable and reloadable, but it is not semantic memory and repository text does not become system authority.

Embedding Looplane: SDK, ConversationController, and the WebSocket Boundary

Looplane exposes bounded-run and conversation contracts through a typed 0.x SDK facade. WebSocket attach wraps one prebuilt, controller-owned runtime session rather than providing conversation-ID resume or multi-client routing.

Looplane Skills, Blocking Hooks, and Plugin Packages

Looplane treats skills as bounded repository-local guidance, hooks as opt-in host commands that can only deny, and local plugin manifests as packages for skills and hooks; their authority is deliberately different.

Looplane Subagent Scheduling and Parent-owned Transactions

Looplane normalizes each subagent dispatch into dependency waves of at most four nodes, runs read-only children concurrently in isolated workspaces, then makes the parent repeat hooks, approval, and transaction execution for any modification.

Looplane tool programs, transactions, and safe concurrency

Looplane parallelizes calls only when they are read-only, concurrency-safe, and classified as READ. Tool programs provide bounded read-only repeat and branching, while transactions snapshot and restore possible workspace-file changes; external side effects are not rolled back.

Harvard CS50 AI Week 1: Knowledge — Propositional Logic, Model Checking, Inference Rules & Knowledge Representation

Week 1 shifts to knowledge representation: propositional logic syntax, model checking, Modus Ponens/Resolution inference, CNF conversion. Projects: Knights (logic puzzles) and Minesweeper (probabilistic inference).

Harvard CS50 AI Week 2: Uncertainty — Probability, Bayesian Networks, Markov Models & Genetic Inference

Week 2 shifts from deterministic to probabilistic: Bayes rule, Bayesian nets with D-separation, Markov chains, PageRank random walks. Projects: Heredity (genotype inference) and PageRank (web ranking).

MIT 6.7960 L03: Optimization Overview — SGD, Adam, LR Schedules & Scaling Rules

From SGD to Adam: pick the right optimizer and scale LR with batch size using scaling rules

Harvard CS50 AI Week 3: Optimization — Local Search, Simulated Annealing, CSP & Crossword Generation

Week 3 tackles optimization: hill climbing, simulated annealing escaping local optima, CSP framework with AC-3 arc consistency, backtracking with MRV/degree heuristics. Project Crossword builds a crossword puzzle generator.

MIT 6.7960 L04: Regularization in Practice — Weight Decay, Dropout, Batch Norm & Label Smoothing

Regularization isn't just anti-overfitting — mechanisms & combo strategies for WD, Dropout, BN, Label Smoothing

Harvard CS50 AI Week 4: Learning — Supervised Learning, k-NN, SVM, Reinforcement Learning Q-learning & Nim

Week 4 enters ML: supervised classification (k-NN, SVM, Perceptron), model evaluation, RL basics (MDP, Q-learning, ε-greedy). Projects: Shopping (purchase prediction with k-NN) and Nim (learning to play via Q-learning).

MIT 6.7960 PS1 Walkthrough: From NumPy MLP to PyTorch Autograd Backprop

Hand-write NumPy MLP + backprop → verify with PyTorch Autograd, fully reproducing OCW HW1 core concepts

Harvard CS50 AI Week 5: Neural Networks — Backpropagation, TensorFlow/Keras, CNN & Traffic Sign Classification

Week 5 enters deep learning: perceptron to multi-layer nets, backprop chain rule, loss functions, optimizers, TensorFlow/Keras modeling, CNN conv/pool. Project Traffic trains CNN to classify traffic signs.

MIT 6.7960 L05: CNN Architectures — From Convolution Kernels to Translation Equivariance

Lec 4 core: why CNN is the natural choice for grid data — convolution, translation equivariance, pooling, and classic architectures in one go

Harvard CS50 AI Week 6: Language — N-gram Language Models, TF-IDF QA, Parser & Attention

Week 6 processes natural language: N-gram conditional probability & smoothing, CFG syntax parsing with CYK, TF-IDF vector retrieval, attention mechanism & Transformer basics. Projects: Parser (syntactic generation) and Questions (TF-IDF QA system).

MIT 6.7960 L06: Modern CNN Architectures — ResNet, EfficientNet, ConvNeXt

ResNet's skip connections solve degradation, enabling 100+ layer nets; EfficientNet compound scales depth/width/resolution; ConvNeXt absorbs Transformer design to reclaim CV crown.

Harvard CS50 AI Synthesis (1): From Search to Language — The Complete Arc of Seven Weeks

Synthesis 1: Tracing how seven weeks form a deliberate knowledge arc from symbolic search to language models, revealing the design philosophy from classical AI to modern ML.

MIT 6.7960 L07: Scaling Rules for Optimization — Spectral View, Feature Learning, Hyperparameter Transfer

Optimization is not an isolated numerical problem: view SGD spectrally, the magnitude of weight updates determines feature learning; Maximal Update Parameterization transfers LR/init across width, and the critical batch size sets the marginal return of trading compute for convergence.

Harvard CS50 AI Synthesis (2): Project Portfolio — All 12 Projects Compared, Difficulty Tiered & Skill Mapped

Synthesis 2: Complete comparison of 12 projects — core algorithms, LOC estimates, difficulty tiers, check50 acceptance criteria, transferable skills. With difficulty grading and learning sequence advice.

MIT 6.7960 L08: Transformers — Tokens, Attention, Positional Codes, and How They Relate to MLPs/CNNs/GNNs

A Transformer is not an architecture from nowhere: tokens discretize data, attention does soft aggregation, positional codes restore order. Seen next to MLPs/CNNs/GNNs, all of them are special cases of 'weighted aggregation over neighbors'.

Harvard CS50 AI Wrap-up: What's Timeless, What's Changed, and Where to Go Next

Series finale: Retrospecting timeless core from 7 weeks/12 projects, gaps in 2020/2023 recordings vs 2026 reality, free OCW route completeness, and forward roadmap (Transformers, LLM fine-tuning, RAG, Agents, Evaluation).

MIT 6.7960 L09: Hacker's Guide to Deep Learning — Practical Know-How to Make Nets Actually Obey

Training neural nets is closer to engineering than magic: look at the data, overfit a mini-batch to prove capacity exists, then regularize back the generalization; learning rate is always the highest-leverage knob.

MIT 6.7960 L10: Memory and Sequence Modeling — RNNs, LSTMs, and Vanishing/Exploding Gradients

An RNN compresses the past into a hidden state, but recurrence makes gradients multiply over time — they either vanish or explode; LSTM decouples 'memory' from 'update' via input/forget/output gates so long-range information flows stably. Attention later replaced it because it reaches any history in O(1).

MIT 6.7960 L11: Representation Learning (Reconstruction-Based) — Autoencoders, VQ, Self-Supervision

Representation learning compresses raw data into a 'useful' vector: autoencoders force a meaningful latent space via reconstruction, VQ discretizes it into a codebook, and self-supervision turns 'mask-and-reconstruct' into free supervision.

MIT 6.7960 L12: Representation Learning (Similarity-Based) — Metric Learning, Contrastive, InfoNCE

Similarity-based representation learning does not reconstruct input; it directly shapes latent geometry: pull same-class representations together, push different ones apart. InfoNCE turns this into 'spot the positive among negatives', and alignment / uniformity give it interpretable metrics.

MIT 6.7960 L13: Theory of Representation — Inductive Biases, Gaussian Processes, and the NN–GP Correspondence

Take a net to infinite width and its random-init output becomes a Gaussian process (NN–GP); its training dynamics freeze into the Neural Tangent Kernel (NTK). This theory analyzes nets and, in reverse, guides us to design the 'right inductive bias'.

MIT 6.7960 L14: Generative Models Basics — Density/Energy Models, GANs, Autoregressive, Diffusion

Generative models learn the data distribution p(x). Density models model probability directly, energy models use an unnormalized potential + sampler, GANs let a discriminator force realistic samples, autoregressive predicts the next token step by step, and diffusion dodges tricky maximum-likelihood via 'add noise then learn to denoise'.

MIT 6.7960 Approximation Theory — Universal Approximation, Barron's Theorem, and Why Depth Matters

A single hidden layer can in principle approximate any continuous function (universal approximation), but width can blow up exponentially with dimension; Barron's theorem lets error decay as 1/sqrt(n) independent of dimension for a specific function class; and depth yields exponential width savings on compositional functions — that is the real reason deep beats shallow.

MIT 6.7960 Graph Neural Networks (GNN) — Message Passing, Permutation Equivariance, and the Expressiveness Ceiling

A GNN is essentially 'an MLP with local message passing on a graph' — it generalizes CNN's fixed-grid neighborhood to arbitrary topology. It must satisfy permutation equivariance/invariance. In theory, a first-order GNN's expressiveness is bounded by the Weisfeiler–Lehman graph isomorphism test: some structures it can never tell apart, which is exactly the gap GIN, positional encodings, and subgraph tricks later fill.

MIT 6.7960 L15: Variational Autoencoders (VAE) — ELBO, Reparameterization Trick, and Latent Representations

The core of VAE is ELBO + reparameterization: log p(x) is replaced with E_q[log p(x|z)] − KL(q(z|x)‖p(z)); the encoder outputs μ/σ and z = μ + σ⊙ε (ε ~ N(0,1)) makes sampling differentiable. Training = reconstruction + KL in tension, which gives rise to β-VAE, posterior collapse, VQ-VAE, and related fixes.

MIT 6.7960 L16: Conditional Generative Models — cGAN, cVAE, and Classifier-Free Guidance

The key to conditional generation is 'feed y into the model': cGAN concatenates y into G/D; cVAE passes y to both encoder and decoder; in diffusion, Classifier Guidance uses gradients from an external classifier to push samples toward a class, while Classifier-Free Guidance trains conditional + unconditional together and linearly combines them at inference — the latter is the standard weapon behind Stable Diffusion and Imagen.

MIT 6.7960 L01: Course Introduction — A Map of Deep Learning, Why Depth Works, and Your First Training Loop

Lecture 1 is the 6.7960 opener: deep learning took off because data + compute + algorithms matured together; the course threads from architectures (CNN/GNN/Transformer) through training, representation, generation, transfer, scaling, and LLMs; ends with a ~30-line PyTorch training loop to confirm your environment works.

MIT 6.7960 L17: Out-of-Distribution Generalization — Distribution Shift, Spurious Correlations, and Three Practical Remedies

OOD failure is not a bug, it's the i.i.d. assumption breaking: covariate shift (image style changes), label shift (class proportions change), concept shift (a word's meaning changes) each need different responses; the most common cause is the model latching onto spurious correlations (using grass as a cue for cows); IRM and domain randomization try to fix this in training data structure, test-time adaptation fixes it at inference.

MIT 6.7960 L18: Transfer Learning — Pretraining, Feature Extraction, and Fine-Tuning Strategies

Transfer learning's core insight is 'features learned on big data are good general-purpose representations': freeze the backbone and train only a linear head when downstream data is tiny; full fine-tune when data is plentiful; reach for LoRA / adapter when compute is tight. SimCLR and MAE removed the need for upstream labels and pushed downstream quality another notch.

aiguide

Should You Rent a GPU to Learn Model Training? GPUtw.ai, LoRA, Jupyter, and the First Experiment

GPUtw.ai makes sense as a short-rental GPU learning tool: start with Jupyter, Ollama, or ComfyUI, then try LoRA/QLoRA on a small model. It is not a large foundation-model training platform, and the first run should verify deployment, billing, and data retention with a small budget.

How AI Agent Search Infrastructure Is Changing: Keenable, Independent Indexes, and NEEDLE

Keenable.ai positions itself as search infrastructure for AI agents: a 100B+ document index, Search/Fetch APIs, MCP/CLI entry points, 100K free monthly requests, and keyless public endpoints. It is worth tracking, but the 100B+ index, latency, and quality claims are still mostly company-provided; NEEDLE is open, but needs external reruns and human review.

aideep-dive

How screenshot-to-code Converts Screenshots to Code: Agent Loop, Asset Extraction, Visual Verification

screenshot-to-code is not a one-shot screenshot-to-HTML tool. Its core is a 30-step Agent Loop with 7 tools — extracting real assets from screenshots, self-verifying with Playwright, and running 4 models in parallel so users pick the best output. 74,500+ GitHub stars, MIT License.

aideep-dive

How Agents Accumulate Team Judgment: Warp's Skill Feedback Loop

Warp's self-improving agent pattern is not about dumping every mistake into a prompt. A base skill does the work, humans leave feedback in GitHub or Slack, an improver skill turns repeated signals into a small diff, and humans review the PR before the next run inherits it.

TinyFish: Free Search and Fetch Infrastructure for AI Agents

TinyFish provides four web APIs for AI agents: Search, Fetch, Agent, and Browser. Search and Fetch are permanently priced at $0 with no credit card requirement, making them a practical default layer for RAG and document retrieval.

How Does A/B Testing Turn a Product Change Into an Estimable Effect?

A/B testing turns a product change into an estimate with uncertainty. A useful report covers effect size, confidence, guardrails, randomization, and launch risk.

Why Not Run Many t-Tests? What Is ANOVA Protecting?

ANOVA first checks whether three or more group means differ overall, so you do not inflate false-positive risk by running many pairwise t tests.

Why Do Large-Sample Approximations Work, and When Do They Fail?

Large-sample normal approximation describes the behavior of estimators, not raw data. It is useful, but dependence, boundaries, and distribution shift can make it unreliable.

How Does Bayesian Inference Connect Prior, Data, and Posterior?

Bayesian inference updates uncertainty about an unknown parameter by combining prior belief with the likelihood from observed data, producing a posterior distribution.

What Do Bias, Variance, and Consistency Check in Point Estimation?

Bias checks whether an estimator is centered correctly, variance checks sampling fluctuation, MSE combines both, and consistency asks whether the estimator approaches truth as sample size grows.

When the Formula Distribution Is Unknown, How Does Bootstrap Estimate Uncertainty?

Bootstrap estimates uncertainty by resampling from the observed sample with replacement, rebuilding many sample-like datasets, and watching the statistic fluctuate.

Causal Inference Basics: Why Prediction Accuracy Does Not Mean Real Effect

Causal inference separates prediction from effect. A model can predict who will buy without proving that an intervention will make them buy.

How Do You Tell Goodness-of-Fit From Independence in Chi-Square Problems?

Chi-square tests compare observed counts with expected counts. First decide whether the problem is goodness-of-fit for one categorical variable or independence for two categorical variables.

When Should Bernoulli, Binomial, Normal, and Poisson Appear?

Distributions are names for data-generating situations, not formula cards. Learn when Bernoulli, Binomial, Poisson, and Normal distributions fit a problem.

How Do You Write Confidence Intervals Without Only Memorizing Bounds?

A confidence interval puts a point estimate back inside sampling fluctuation. Computing bounds is only the first step; you also need to explain standard error, critical values, and coverage.

When You See a Dataset, What Statistics Should You Check First?

Data type determines the statistical tools you can use. Start with categorical, numeric, count, and time-ordered data, then choose summaries that fit the question.

How Does the Delta Method Estimate Uncertainty for F1 and Ratio Metrics?

The delta method transfers uncertainty through a smooth function: the local derivative expands or shrinks the estimator's original standard error.

What Makes an Estimator Good: Bias, Variance, or MSE?

An estimator is a rule for using samples to infer a population parameter. To judge whether it is good, look at bias, variance, and MSE together.

When a Mixed Problem Appears, How Do You Pick the Tool in 30 Seconds?

At the final review stage, train problem recognition: identify data type, unknown quantity, and decision goal before choosing a formula and writing a contextual conclusion.

What Do Expectation and Variance Mean in Exams and Model Evaluation?

Expectation describes long-run center; variance describes fluctuation. This post computes E[X], E[X^2], and Var(X), then connects them to average loss and model stability.

How Does Experimental Design Make Results Interpretable Rather Than Merely Correlated?

Experimental design decides whether a result can be interpreted. Randomization, control, blocking, replication, blinding, and pre-specified outcomes give inference a usable foundation.

How Does Fisher Information Tell You Whether a Parameter Is Stable?

Fisher information uses likelihood curvature to measure how well the data locate a parameter; larger information usually means a smaller standard error for the MLE.

Confidence Intervals Are More Than t-Tables: What Is the General Construction?

A confidence interval is built by defining the target estimate, describing its sampling error, and choosing a rule that turns uncertainty into a range.

How Does a GLM Choose Distributions and Link Functions by Data Type?

A generalized linear model starts from the response type, chooses a suitable distribution, and uses a link function to connect the mean to a linear predictor.

From H0 to p-Values, What Decision Is a Hypothesis Test Making?

A hypothesis test is a decision process under uncertainty: write H0/H1, choose alpha, compute a test statistic and p-value, then decide whether the data is strong enough to challenge H0.

How Do Estimation, Testing, Likelihood, and Bayes Fit on One Inference Map?

The inference map starts with the question type: point estimate, uncertainty interval, decision test, likelihood model comparison, Bayesian update, or resampling.

How Does the Likelihood Ratio Test Compare Nested Models?

The likelihood-ratio test compares the log likelihood of a restricted model with a full model; the usual chi-square reference only makes sense under nested-model and approximation conditions.

When OLS Assumptions Fail, How Can the Regression Line Still Be Used?

OLS is a useful baseline, but coefficient interpretation, inference, prediction, and diagnosis depend on assumptions about linearity, errors, independence, and variance.

How Does Logistic Regression Move From Probability to Thresholds and Error Costs?

Logistic regression estimates probabilities first. Classification decisions come later, when thresholds turn those probabilities into actions under real error costs.

Why Should Classification Start With Log Odds?

Logistic regression connects a linear score to a probability between 0 and 1. Understanding odds, log odds, and odds ratios prevents wrong coefficient interpretations.

Why Does MAP Turn Priors Into Regularization?

MAP maximizes the posterior. After taking logs, the prior becomes a penalty term, which connects Bayesian estimation to L1, L2, and regularized ML objectives.

How Do Matching and Weighting Make Observational Data More Experiment-Like?

Matching and weighting do not turn observational data into a true experiment. They try to make treatment and control comparable on observed variables.

Why Does MLE Ask Which Parameter Most Likely Generated the Data?

MLE fixes the observed data and compares which parameter values make that data most plausible; log likelihood turns products into sums and connects directly to negative log loss.

Why Does the Method of Moments Match Sample Moments to Population Moments?

Method of Moments matches sample moments to theoretical population moments, then solves for parameters. It is not always the most efficient method, but it builds the first intuition for parameter estimation.

Missing Data Is Not Just Blank Cells: How Does It Distort Statistics and Models?

Missing data can change representativeness, bias estimates, and mislead ML systems. The first question is why the data are missing.

How Do You Write an ML/AI Evaluation Report That Is More Than a Leaderboard Score?

A useful ML/AI evaluation report turns statistical evidence into a decision: ship, stage, roll back, or run more experiments.

What Do Residuals, Outliers, and Leverage Reveal About Model Failure?

Model diagnostics turn fitted errors into evidence: residual patterns, outliers, leverage, and influential points reveal how a model fails.

How Does Multivariate Analysis Organize Features That Move Together?

Multivariate analysis looks at features together. Covariance, correlation, and PCA reveal shared directions that univariate summaries miss.

What Kind of Optimal Test Is the Neyman-Pearson View About?

The Neyman-Pearson view treats a test as a decision rule: under a fixed Type I error rate alpha, choose the rejection region with the highest power.

What Assumptions Do Nonparametric Methods Relax, and What Do They Cost?

Nonparametric methods are not assumption-free. They relax fixed distributional forms, often gaining flexibility while paying in efficiency, interpretation, or overfitting risk.

How Should You Analyze NTU IM 114-115 Statistics Papers Without Memorizing Answers?

Past papers train question-analysis discipline, not fortune-telling. Each problem should return to data type, unknown quantity, statistical tool, calculation path, and contextual conclusion.

How Do You Avoid Missing Cells in Joint Distribution and PMF Transformations?

Joint PMF problems require listing every cell. Marginalization, conditional probability, and variable transformations are all sums or regroupings of the original cells.

Conditional Probability, Independence, and Bayes: What Viewpoint Is the Problem Switching?

Probability problems are often hard because the viewpoint changes. Define events first, then distinguish conditioning, independence, mutual exclusivity, and Bayes' rule.

How Do Samples, Statistics, and Sampling Distributions Differ?

A sample is the data, a statistic is a function of the sample, and a sampling distribution is the distribution of that statistic under repeated sampling.

How Do PMF, PDF, and CDF Turn Probability Into Computation?

Random variables turn uncertain outcomes into numbers. PMF, PDF, and CDF then let you compute discrete probabilities, continuous interval probabilities, thresholds, and model-score distributions.

How Should coef, SE, t, F, and R-Squared Be Read Together?

A regression table is not a p-value list: coef, SE, t, F, and R-squared answer effect size, uncertainty, single-coefficient tests, overall model signal, and in-sample explanation.

Why Do Ridge, Lasso, and Weight Decay Make Models More Stable?

Regularization adds a preference against extreme parameters. Ridge, Lasso, and weight decay trade some training fit for a model that generalizes more reliably.

How Can Statistics and ML Evaluation Be Rerun to Reach the Same Conclusion?

A reproducible workflow preserves the evidence chain from data to conclusion. Results need data versions, code, seeds, environment, metrics, and raw outputs.

Why Can a Sample Say Something About a Population or Model?

Sampling makes sample statistics fluctuate, and standard error describes that fluctuation. This post separates SD, SE, sampling distributions, and CLT, then connects them to benchmark uncertainty.

How Do Sampling Distributions Become Exam-Ready Reasoning?

A sampling distribution describes how a statistic fluctuates under repeated sampling. Means, proportions, and variances each connect to common distributions used in intervals and tests.

After 53 Posts, How Do You Connect Statistics to ML, Causality, and Mathematical Statistics?

The series does not finish all of statistics. It gives beginners a working map for exams, ML/AI evaluation, causality, Bayesian thinking, time series, and mathematical statistics.

How Does One Regression Line Become Prediction, Interpretation, and Error?

Simple linear regression uses one X to describe the average change in Y. Slope, intercept, residuals, and squared error form the smallest supervised learning model.

How Does Monte Carlo Use Repeated Simulation to Answer Hard Statistical Questions?

Monte Carlo repeats a data-generating process many times so sampling variation, power, coverage, and evaluation instability become visible.

Where Should You Start Statistics If You Need Exams and ML/AI?

Do not start statistics exam prep by memorizing formulas. Start with the sequence of data, probability, sampling, inference, regression, then connect those ideas to model evaluation, A/B testing, and uncertainty in ML/AI.

Why Should Time-Series Data Not Be Randomly Split?

Time-series data have order. Random splits can leak future information into training and make forecasting or monitoring results look better than they are.

Which Test Fits a Two-Group Mean or Proportion Difference?

Two-group comparisons start by classifying the outcome and the design: numeric or binary, independent or paired. That choice determines the standard error, test statistic, and conclusion.

How Does Variable Selection Avoid Memorizing the Training Data?

Variable selection is not only about choosing predictors. It is about avoiding noisy training-set wins that do not generalize.

Statistics Is Not Formula Memorization: What Is It Deciding?

The core of statistics is judgment: describe data, estimate unknowns, compare differences, inspect associations, and make decisions under uncertainty.

How to Use Cloudflare AI Search: Data Sources, Hybrid Retrieval, and Workers Bindings

Formerly AutoRAG, the managed search primitive: drop files into built-in storage or attach R2 and websites, auto-index with Markdown conversion plus vector and BM25, retrieve with hybrid, RRF, and reranking, and query from Workers via namespace or instance bindings, REST, or MCP.

techdeep-dive

What Is GPUtw.ai? Taiwan GPU Cloud, Short-Rental Compute, and Researcher Workflows

GPUtw.ai is a Taiwan-based short-rental GPU cloud. Its main value is not maximum scale, but Taiwan data centers, prepaid credits, Jupyter/ComfyUI/Ollama/vLLM templates, Vault storage, and team billing. Public information is enough for a service introduction, not enough for procurement or production endorsement.

Why Did 'I Want a Beginner AI Course' Return Zero Results? Debugging Chinese Tokenization and a Broken RAG Data Path

Entering '我想找入門的ai課程' in Ask AI showed zero searched posts and triggered a refusal, while Related Reading recommended exactly the right article. The first fix addressed Chinese tokenization, the LIKE fallback, and the Vectorize data path. A second pass added short Han-and-number tokens, post metadata retrieval, and a rule that exposes sources only after both Validation and Critic pass.

Harvard CS181 HW0: Do These 4 Problems First — They Tell You What to Patch

HW0 checks CS181 prerequisites in four problems — y=Xw solvability, optimizing an objective, reasoning about randomness, and OLS in Python. The problem that slows you down most is the gap to patch before HW1.

Harvard CS181 HW1: Ice Core Regression — kNN, Kernels, Basis Functions, and Regularization

HW1 uses an 800,000-year ice-core temperature dataset across four problems: kNN and kernel regression, a geometric proof of least squares, basis-function regression, and a probabilistic derivation of ridge and LASSO, ending with a coordinate-descent LASSO implementation.

Harvard CS181 Machine Learning: Your 2026 Roadmap Through 7 Homeworks (With a 4-Year Comparison)

CS181 2026 is A3 with hw0–6 as the weekly clock (no public recordings); 2025 adds a practical, 2024 has two midterms, 2023 was taught by Weiwei Pan. Start with HW0, then follow hw1→hw6.

Harvard CS50 AI Week 0: Search — From DFS, BFS, A* to Minimax and Alpha-Beta Pruning

Week 0 opens with search algorithms: BFS for shortest paths, Minimax for adversarial play, Alpha-Beta for pruning. Two projects: Degrees (BFS) and Tic-Tac-Toe (Minimax).

techdeep-dive

Choosing a Mobile Client for Claude Code: Moshi's SSH Terminal, moshi-hook, and Pricing

Moshi is an iOS/Android terminal app (plus a free Moshi Desktop web UI) that connects over SSH/Mosh straight to your own machine to drive Claude Code, Codex, and other coding agents. The free tier is a complete terminal; Pro ($7.99/mo and up) unlocks Mosh's connection resilience, deep tmux integration, and the diff viewer.

ADE Workspace Showdown: ADE vs Superset vs Herdr vs Orca

Four philosophies of ADE workspaces: arul28/ADE's Brain+Lane, Superset's 100-agent IDE, Herdr's Rust-native runtime, and Kadro/Orca's pane and Fleet angles — compared in one table with a decision tree for when to pick a workspace over an Omnigent-style control plane.

Governance Deep Dive: Policies, Omnibox, Spend Controls and Credential Brokering

Omnigent moves governance off prompts into a Server-side Policy engine: Python functions returning allow/deny/ask, a three-layer stack with cost budgets and tool caps, plus Omnibox OS-native isolation via bwrap/seatbelt and egress credential brokering — compared with five peer governance stacks.

How to Read 2026 Coding Benchmarks: SWE-bench, Terminal-Bench, DeepSWE, Aider Explained

The same model can score 20 points apart on different harnesses, 32% of SWE-bench Pro verifier judgments were found to be wrong, and DeepSWE's 113 tasks make most models score zero. This guide decodes six major coding benchmarks — what they test, which are easy to game, and which ones you should care about.

Harvard CS50 AI Guide: Seven Weeks, Twelve Projects, and How to Follow a Course Filmed in 2020

CS50 AI's OpenCourseWare edition publishes seven weeks of lectures, slides, notes, and twelve Python projects with autograder feedback, plus a free CS50 Certificate if you score at least 70% on every project. The catch: weeks 0–5 still use the Spring 2020 recordings; only Week 6 (Language) was re-recorded, in 2023.

LLM API Routing: Direct, Aggregator, or Cloud — A Price Comparison

The same model can cost 2-5× more depending on the channel. Direct API is simplest, aggregators (OpenRouter) are most flexible, cloud platforms (Bedrock/Vertex) suit enterprises. This post compares actual August 2026 prices across six channels with a decision tree.

Same Name, Different Layer: meta-harness, ACP, HarnessAgent and Flue

meta-harness means two things: Databricks' control plane and Stanford's outer-loop optimizer. This post uses a four-layer model (MCP/ACP/Runtime/meta-harness) to place Omnigent, Zed ACP, Vercel HarnessAgent and Cloudflare Flue.

Reading MIT 6.7960: One Course, Two Official Editions — Complete the OCW 2024 Package, Read the 2025 Decks for What's New

MIT 6.7960 Deep Learning (Fall 2025) publishes all 21 lecture decks as public Dropbox PDFs, and most required readings map to free textbook chapters; but the five problem sets are released only through Gradescope, and solutions plus recordings live behind Canvas login. This guide covers how the three instructors split the course, a topic map of all 21 lectures, textbook-based substitutes for lectures, and where outside self-learners realistically stop.

Why MoE Wins: The Architecture Behind Every 2026 Frontier Model

Nearly every frontier open-source model in 2026 is MoE: Ornith 35B activates only 3B to beat 31B dense models, MiniMax M3 uses 456B total but 45.9B active to hit SWE-bench Pro 59%, DeepSeek V4 runs 1.6T total with 49B active. This post explains why MoE dominates coding and agentic benchmarks using four case studies.

One Multi-Agent Task, Four Implementations: Omnigent YAML vs LangGraph vs CrewAI vs Goose

The same Polly task — parallel git worktrees plus cross-vendor review — implemented four ways: Omnigent YAML governs at the Server layer, LangGraph controls flow with a StateGraph, CrewAI assembles roles quickly, and Goose ships a desktop Recipe, compared on tokens, latency, and maintainability.

aideep-dive

OLMo: The Only Language Model Family That Open-Sources Its Training Data

Allen AI's OLMo is the only language model family that fully publishes weights, training data (Dolma, 9.3T tokens), training code, all intermediate checkpoints, and evaluation tools. OLMo 3's 32B Think model hits 96.1% on MATH — and you can use OlmoTrace to trace any output back to the exact training data that produced it.

Managing Multiple Agents Together: Omnigent's Meta-Harness, Policies, and Cross-Device Sessions

Databricks' open-source Omnigent wraps Claude Code, Codex, Cursor, Pi and custom agents in a Runner/Server + Omnibox sandbox, adding three-layer Policies and shareable persisted Sessions so you can swap models and harnesses with one-line changes — 9.3k stars, still alpha.

Open-Source AI Licensing Guide: What MIT, Apache 2.0, and Llama License Actually Allow

'Open-source' in AI doesn't mean what it means in software. MIT and Apache 2.0 let you do almost anything; the Llama License requires a separate deal above 700M MAU; old Gemma terms let Google change rules unilaterally (Gemma 4 switched to Apache 2.0). This guide maps what you can and can't do by license type.

Self-Hosting Open-Source LLMs: Framework Choice, Hardware Math, and When It Beats APIs

Open-source models now match closed-source on coding benchmarks, but self-hosting isn't just picking a model — vLLM handles high-concurrency production serving, SGLang is 29% faster on prefix-heavy workloads, Ollama is the local dev default, and llama.cpp runs on the least hardware. A100 cloud rentals run ~$1.4-2.2/hr; self-hosting breaks even at roughly 100M tokens/month.

Three RL Post-Training Playbooks: How Ornith, Nous Research, and MiniMax Built Dark Horse Models

Three non-big-lab teams used different RL post-training strategies to produce benchmark dark horses in 2026: Ornith's self-improvement loop (GRPO), Nous Research's DataForge + Atropos execution-reward RL, and MiniMax's massive-scale RL across 200K real environments. Different strengths, but one shared proof point: post-training RL matters more than pretraining scale.

Tokens, Context Windows, and Inference vs Training: Three Things to Know Before Using AI Models

Models don't read words — they read tokens. A Chinese character is typically 1-2 tokens; an English word is 1-3. The context window is the token limit per request. Inference is using a model; training is teaching one. What you do every day is inference.

Embeddings: How Models Turn Words Into Computable Vectors

Models don't understand text — they only understand numbers. Embeddings map each token to a vector of several hundred dimensions, where semantically similar words end up close together in vector space. This is the shared foundation behind search, RAG, and classification.

How to Read a Model's Report Card: Benchmarks, Arena Elo, and the Traps Behind the Numbers

Benchmark scores in model releases have three common traps: cherry-picking (only showing wins), contamination (test data leaking into training), and saturation (when everyone scores 90%+, the benchmark stops being useful). The most manipulation-resistant signal is Chatbot Arena's Elo ranking — real humans, blind voting, uncontrolled questions.

Fine-tuning vs RAG: When to Teach the Model vs When to Look Things Up

Data changes often and you need citations → RAG. Need consistent style or want to run on a small device → fine-tuning. In practice, many production systems use both: fine-tune a small model that speaks your domain language, then use RAG to supply up-to-date facts.

How Models Improve Themselves: Gradient Descent and the Training Loop

A model uses loss to know how wrong it is and gradients to know which direction to adjust. Gradient descent repeats three things: compute loss, compute gradients, update parameters. The learning rate controls step size — too large and you overshoot, too small and training takes forever.

How a Model Knows It's Wrong: Loss Functions and Cross-Entropy

Every time a model predicts the next token, it assigns a probability to every candidate word. A loss function measures how far that probability distribution is from the correct answer — the further off, the higher the loss, the more the model knows it got it wrong. Cross-entropy is the standard formula; perplexity is its human-readable translation.

Quantization & Inference Optimization: Running a 70B Model on Your Laptop

A 70B model needs ~140GB VRAM in FP16, but 4-bit quantization shrinks it to ~35GB. With llama.cpp's partial CPU offloading, it can run on consumer hardware. GGUF naming conventions (Q4_K_M, Q5_K_S) tell you the precision-size tradeoff. KV cache is why long conversations slow down.

Scaling Laws: How Big Should a Model Be, and Why Bigger Isn't Always Better

Scaling laws show that loss decreases predictably with more parameters, data, and compute — following power-law relationships. The Chinchilla paper's key finding: most models were too large and undertrained. Given the same compute budget, training a smaller model on more data produces better results. This reshaped the entire industry's training strategy.

Understanding AI Models: 18 Articles from Tokens to Self-Hosting

You don't need to become a researcher to understand AI models systematically. This series starts from what you can see (tokens, context windows) and works up to self-hosting open-source models — 18 articles covering everything you need to choose models, read benchmarks, and estimate costs.

Tokenization: The BPE Algorithm, and Why Chinese Costs More Than English

Models charge by tokens, not characters. The BPE algorithm starts from individual bytes and repeatedly merges the most frequent adjacent pair to build a vocabulary. English 'understanding' might be 1-2 tokens, but Chinese '理解' could take 2-3 — same meaning, higher cost.

Pre-training, SFT, RLHF: Three Stages That Turn a Text Predictor into a Useful Assistant

Every LLM goes through three training stages: pre-training reads the internet to learn language, SFT uses example conversations to learn the format, and RLHF uses human preferences to learn what a good answer looks like. The gap between a base model and a chat model is what the last two stages do.

Transformers and Attention: How Models Decide Which Words to Look At

The core of the Transformer is self-attention: for each token, the model computes how relevant every other token is, then takes a weighted sum. This lets the model reach across distance to figure out that 'it' refers to 'cat' not 'mat' — and is the foundation for how it handles long documents.

Ahead of AI: How a Scholar Built 200K Subscribers by Publishing Monthly, Not Daily

Computational biology PhD turned UW-Madison professor Sebastian Raschka launched Ahead of AI on Substack in 2022, publishing monthly deep dives into LLM papers and architectures. Four years later: 200K+ subscribers, zero sponsorships, and a book-newsletter flywheel that proves low frequency and high depth can win in a crowded AI newsletter market.

ByteByteGo: From a Self-Published Book to a Million-Subscriber System Design Empire

Former Twitter/Apple/Zynga engineer Alex Xu self-published System Design Interview in 2020 and hit the Amazon bestseller list. In 2022, he and ex-Discord engineer Sahn Lam launched a Substack newsletter that crossed 26K subscribers in month one and hit one million in two and a half years. From book to newsletter, YouTube, and paid platform, ByteByteGo reached $3.5M ARR in 2024 with a 26-person team — all fully bootstrapped.

Daily Dose of Data Science: From a Cancelled Master's to 200K Newsletter Subscribers

Former Mastercard AI engineer Avi Chawla turned a cancelled US master's admission into a daily Substack newsletter — 150-word visual posts on data science. 10K subscribers in 5 months, income exceeding his full-time job, 200K+ subscribers and a paid course platform four years later.

Dense Discovery: 403 Issues of Deliberately Staying Small

Berlin-born, Melbourne-based designer Kai Brach spun his indie print magazine Offscreen into Dense Discovery, a weekly curated newsletter. Eight years, 403 issues, 36,000 subscribers, 63% open rate — deliberately not scaling, sustained by a single $649+ sponsor slot per issue and a Friends membership program. Proof that 'enough' is a viable business model.

Lenny's Newsletter: How an Ex-Airbnb PM Built the Most Influential Product Management Newsletter

Former Airbnb product lead Lenny Rachitsky left in 2019, wrote a viral Medium post, moved to Substack, and built a 1.2M-subscriber newsletter empire through obsessive quality (50+ revision cycles per post), a 40K-member Slack community, a top-ranked podcast, and an annual summit — all without hiring a single full-time employee.

Morning Brew: From a Michigan Dorm Room to a $75M Media Empire

Alex Lieberman started a PDF called Market Corner in his Michigan dorm, rewriting Wall Street Journal-style news in a casual, conversational tone. Five years later, Morning Brew hit 4 million subscribers and $13M in revenue, then sold to Insider Inc. for $75M. Their referral program accounted for up to 75% of new signups at its peak — one of the most successful growth engines in newsletter history.

Not Boring: When Writing Itself Becomes the Deal Flow

Former investment banker and startup VP Packy McCormick turned a COVID-era social club pivot into Not Boring, a long-form business analysis newsletter. In two years he hit 100K subscribers and $1M in sponsorship revenue, then extended the flywheel into three venture funds totaling $68M+ across 200+ investments — proving that writing can literally be deal flow.

One-Person Media Company: Ten Newsletter Cases and Four Revenue Playbooks

From Stratechery proving in 2014 that one person can make a living writing analysis to TLDR hitting $10M+ ARR in 2024 — ten newsletter cases distilled into four revenue playbooks, three content models, and one universal rule: format choice determines the ceiling.

The Pragmatic Engineer: From Uber Payments Lead to Substack's #1 Tech Newsletter

After six years managing Uber's payments infrastructure and witnessing pandemic layoffs, Gergely Orosz launched The Pragmatic Engineer on Substack. It hit #1 in tech within four months, crossed one million subscribers in three and a half years, and generates $1.5M+ annually — entirely from reader subscriptions, with zero ads or sponsors.

Stratechery: The Paid Newsletter Pioneer Who Proved the Model from a Taipei Apartment

Ben Thompson launched Stratechery full-time from his Taipei apartment in 2014 with a three-tier subscription model. Twelve years later, he has 40,000+ paid subscribers, $5M+ annual revenue, and created Aggregation Theory — the most influential business framework in tech analysis since Clayton Christensen. Substack's seed-round pitch was literally 'Stratechery-in-a-box.'

The Hustle: From a Hot Dog Stand to a $27M SaaS Acquisition

Sam Parr went from selling hot dogs in Nashville to building a 1.5M-subscriber business newsletter, then sold it to HubSpot for eight figures. The buyer didn't want the content — they wanted the mailing list as a SaaS lead funnel.

TLDR: How a Basement Side Project Became an Eight-Figure Ad-Only Newsletter Empire

Dan Ni — Yale math-econ grad, former Jane Street quant trader — left Wall Street after a rare medical condition, and launched TLDR from his parents' basement in Missouri with $50/day in Reddit ads. Eight years later: 13 verticals, 7.2 million subscribers, 22 fully remote employees, and eight-figure annual revenue — all from advertising, not a cent from readers.

FLUX: The Image Model Family Built by Stable Diffusion's Original Team, from 12B to a Self-Flow World Model

FLUX is Black Forest Labs' image-model family. The Stable Diffusion team launched it in August 2024 with a 12B rectified-flow transformer. Two years later it spans klein 4B ($0.014 and the only current Apache-2.0 model) / 9B, pro ($0.03), flex ($0.05), max ($0.07 with live web grounding), and open-weight 32B dev. FLUX 3 extends Self-Flow to video, synchronized audio, and robot actions. This guide covers the FLUX.1-to-FLUX 3 evolution, three-tier licensing, and model selection.

Speech and Audio Models: Four Years from Whisper Rewriting Open ASR to ElevenLabs Consolidating Voice APIs

Speech models split into two lines. Whisper led ASR: its MIT-licensed 1.55B model drove transcription cost toward zero in September 2022; after v2, v3, and turbo cut the decoder from 32 layers to four, OpenAI moved to closed gpt-4o-transcribe. In TTS, ElevenLabs grew to an $11B valuation and $500M ARR, while Kokoro (82M, Apache-2.0) and Chatterbox preserved self-hosting. Speech-to-speech Realtime APIs are now rewriting live conversation.

Video Generation Model Families: Sora Exits after a Two-year Arms Race Dominated by Veo 3.1, Kling 3.0, and Gen-4.5

Sora's February 2024 preview shocked the industry, but the landscape reversed in two and a half years: OpenAI closed the consumer Sora app in April 2026 and scheduled its API for retirement on September 24; Veo 3.1 became the narrative default with native audio and Flow; Kling 3.0 became a unified multimodal model with $240M annualized revenue; and Runway Gen-4.5 briefly led Artificial Analysis in a November 2025 snapshot while defending the professional market through enterprise workflows. This guide compares four families by generation, specifications, pricing, and use case.

When Search Returns Only 10 Results: Fixing CJK Recall in Cloudflare D1 FTS5 Hybrid Search

Querying “認證” returned only ~10 hits while 149 files (509 occurrences) matched; 41 posts and 76 chunks were found via LIKE, but D1 chunks_fts had 0 rows and unicode61/trigram both returned 0 for 2-char CJK terms. Fix: LIKE fallback with char-level OR first, then trigram migration + pnpm sync, then pagination beyond the hard limit of 12.

MiniMax: The Chat App Company That Built a Coding Model to Rival Frontier Labs

MiniMax started as a consumer chat app company, then M2.5 scored 80.2% on SWE-bench Verified at 1/10-1/20 the cost of Claude Opus; M3 (456B total / 45.9B active) became the first open-weight model to clear 59% on SWE-bench Pro, with 1M context powered by their novel Sparse Attention mechanism.

Nous Research: From Research Collective to Open-Source AI Ecosystem Rebel

Nous Research doesn't pretrain — they fine-tune and do RL. Hermes 4 scores 96.3% on MATH-500, NousCoder-14B improves Qwen3-14B's coding ability by 7% using only 24K training samples. But the real moat is Hermes Agent: 236K GitHub stars, #19 globally, 3,000 contributors.

Ornith: The Open-Source Coding Dark Horse Built on Self-Improvement RL

DeepReinforce's Ornith 1.5 family, trained with self-improvement RL: the 397B flagship scores 86.0 on SWE-bench Verified, matching Claude Opus 4.8; the 35B-A3B activates only 3B parameters per token yet leads every coding benchmark in its class; the 9B runs on phones. MIT-licensed, fully open-source.

Managing Multiple Claude Code Sessions: Agent View, Dispatch, State Monitoring, and Cross-Session Messaging

`claude agents` gives you one screen listing every background session, with Needs input, Ready for review, Working, Completed, and related states managed in one table. Press Space to peek, Enter to attach. Combined with cross-session messaging (ListAgents / SendMessage, v2.1.224+), sessions can also message each other; same-machine delivery uses a local socket and never touches Anthropic servers.

The .claude Directory, Explained: settings, rules, skills, and auto memory

Claude Code splits its configuration across the project `.claude/` folder and your home directory — 20+ file locations. Only two mental models matter: settings merge across layers and are enforced; CLAUDE.md and rules concatenate into context as guidance. And every file is committed, gitignored, or Claude-written.

How Claude Code Reviews Your PRs: Multi-Agent Analysis, REVIEW.md, and ultrareview

After GitHub PR review is configured, a fleet of agents reviews PRs according to the repo's trigger mode — 20 minutes on average, about $15–25 per review, with findings posted as inline comments on the offending lines. For larger changes, /code-review ultra launches a cloud deep review that reports independently verified bugs in 5–10 minutes at roughly $5–25 per run; Pro/Max plans include 3 free runs.

Managing Claude Code Costs: Token Tracking, Model Choice, Effort, and Team Analytics

Claude Code costs accumulate with context size: enterprise deployments average ~$13 per developer per active day and $150–250 per month. This post covers /usage and /insights tracking, six token-saving tactics, and a systematic answer to 'which model should I use': provider-dependent model aliases, effort levels, fast mode ($10/$50 per MTok for Opus 5/4.8), and the advisor tool.

Claude Code Config Not Taking Effect: Diagnosing with /context, /doctor, /mcp, and Error References

When CLAUDE.md rules are ignored, hooks never fire, or an MCP server shows no tools, the file usually didn't load, loaded from an unexpected location, or got overridden. This guide covers what the diagnostic entries (/context, /memory, /skills, /doctor, /mcp, and more) actually show, safe-mode bisection, and a table of six high-frequency error messages with fixes.

How to standardize Claude Code dev environments: devcontainer.json, CI consistency, and team rollout

Add Anthropic's official Dev Container Feature (`ghcr.io/anthropics/devcontainer-features/claude-code:1.0`) to `.devcontainer/devcontainer.json` and three steps—write the config, rebuild the container, run `claude` to sign in—put every teammate's Claude Code behind the same container definition. The same definition feeds GitHub Codespaces and CI; a five-step rollout gets the whole team there.

How Claude Code Orchestrates Subagents at Scale: Dynamic Workflows, ultracode, and Rerunnable Scripts

Dynamic workflows let Claude write multi-agent orchestration as a JavaScript script that a runtime executes in the background — up to 1,000 agents per run, savable as a /<name> command. This piece covers trigger methods, the save-and-rerun flow, three fit scenarios (codebase audit, large migration, cross-checked research), and where they differ from Agent Teams.

How Claude Code Works: The Agentic Loop, Built-in Tools, and Two Safety Rails

Claude Code runs an agentic loop — gather context, take action, verify results — until the task is done. This entry to the series breaks down its five tool categories, the model/harness split, and the two safety rails: checkpoints and permission modes.

How to Choose Claude Code's Multi-Agent Options: Subagents, Agent View, Agent Teams, Dynamic Workflows

The official docs split Claude Code's parallel work into 4 approaches: subagents delegate inside one session, agent view lets you supervise background sessions yourself, agent teams coordinate workers through a lead, and dynamic workflows run scripted fleets of subagents with cross-checks; file collisions are always handled by worktrees. Includes a translated comparison table and a three-question decision guide.

Claude Code in the Cloud: on the web, --cloud/--teleport, and Steering from Mobile

Claude Code on the web runs tasks in cloud environments, Anthropic-managed VMs by default or self-hosted environments when routed there: authorize GitHub, dispatch from browser or mobile, start cloud sessions with --cloud, and pull them back local with --teleport. Research preview on Pro/Max/Team; no separate compute charge, but rate limits are shared.

How Much Autonomy to Give Claude Code: Permission Modes, the Auto Mode Classifier, and Allow/Deny Rules

Claude Code ships six permission modes; day to day you cycle Manual, Accept edits, Plan, and Auto with Shift+Tab. On Pro/Max/Team plans, eligible interactive terminal and VS Code sessions start in auto mode by default, with a background classifier reviewing most actions and blocking force pushes, `curl | bash`, production deploys, and more by default. This post covers the four-mode spectrum, permission rule syntax, and organization-level trust config.

How prompt caching shapes Claude Code's speed and bill: prefix matching, invalidation triggers, and hit rate

Claude Code's prompt caching works by exact prefix matching: a cache read bills at roughly 10% of the standard input rate, but switching models, changing effort, enabling fast mode, toggling MCP servers, or denying an entire tool forces the next turn to reprocess everything. The TTL defaults to five minutes; the main conversation and a few helper requests on a subscription get one hour.

How to manage Claude Code sessions: --continue, --resume, /branch, and JSONL transcripts

Claude Code writes every session line by line to a JSONL file under ~/.claude/projects/, kept for 30 days by default. This post breaks down --continue vs --resume, session naming rules, /branch fork semantics, and transcript export and cleanup settings.

Claude Code install and login troubleshooting: PATH, install sources, proxy, OAuth callback

Work through install and login failures in five steps: verify PATH on your OS, confirm there is only one installation, test downloads.claude.ai for a 200, recover failed OAuth callbacks by pasting the login code or using claude auth login, then finish with claude doctor.

Claude Code Runtime Troubleshooting Guide: CPU/Memory, Session Hangs, Auto-Compact Thrashing, Tables, Search Failures

Five classes of runtime fixes: diagnose high memory with /compact plus /heapdump; recover a hung session with Ctrl+C and claude --resume; write large tables to files instead of forcing terminal output; escape autocompact thrashing by reading files in chunks or running /compact with a focus; fix broken search by installing system ripgrep and setting USE_BUILTIN_RIPGREP=0.

Agentic / Reasoning RAG: From Search-R1's RL Multi-Turn Search to Deep Research and MCP's Reasoning × Retrieval Paradigm

In 2025 RAG stopped being 'retrieve once, generate once.' Search-R1 trains models to search autonomously in multiple turns with RL, REX-RAG/AlignRAG add policy and alignment branches, OpenAI Deep Research productizes the loop, and MCP generalizes retrieval into unified tool invocation. This post unpacks the design philosophy, trade-offs against ten generations, and when to adopt the new paradigm.

aideep-dive

Apple Opens Free Private Cloud Compute Access: AFM 3 and What Developers Need to Know

Apple is giving App Store Small Business Program developers free access to AFM 3 models on Private Cloud Compute if their apps have fewer than two million first-time downloads. The five-model family includes the sparse 20B-parameter AFM 3 Core Advanced, which activates only 1–4B parameters on-device, and AFM 3 Cloud Pro on Google Cloud NVIDIA GPUs, refined with outputs from Gemini.

aideep-dive

BytePlus ModelArk Coding Plan: ByteDance's AI Coding Subscription

BytePlus ModelArk Coding Plan offers Lite ($10/month) and Pro ($50/month) subscriptions covering models such as DeepSeek-V4, GLM-5.2, and Seed-2.0 in tools including Claude Code and Cursor. Lite includes about 24,000 requests per month; Pro includes five times as many.

Learning Agent Design from Mature Coding Agents (2): The Shape of the Agent Loop — Event Streams, Checkpoints, Resume

pi's loop is a double while-loop wrapped in an EventStream; claude-code's source openly says stop_reason is unreliable and uses tool_use blocks observed during streaming as the sole continue signal; codex models a turn as a cancellable SessionTask and records sessions with a dedicated rollout crate. looplane chose an ordering — manifest first, JSONL second — that turns Ctrl-C into verified resumption instead of a rerun. All evidence cited at file#symbol level.

Learning from Mature Coding Agents (13): CLI Ergonomics — Make New Tools Feel Already Familiar

Mature coding-agent CLIs have converged on the same conventions: positional prompt, -p means print, exec is headless, resume is a first-class command, -C changes directory; looplane inherits this vocabulary directly, driving learning cost close to zero.

Learning Design from Mature Coding Agents (10): Edit Tool Trade-offs — unified diff, exact edit, hashline, and whole-file

LLMs break unified diffs on bookkeeping: wrong hunk counts, hallucinated context lines. The five reference projects split into two camps — simplify the diff grammar (Codex drops line numbers), or drop diffs entirely (Claude Code/Pi/OpenCode exact replace); OMP goes further by binding read state into the format via hash anchors. looplane took the minimal-intervention path: keep the guarded apply_patch, add a zero-fuzzy replace_text, and its qwen3:4b eval went from stable failure to 5/5.

Learning Agent Design from Mature Coding Agents (9): External CLIs as a Backend — Where Does the Security Boundary Go?

Every mature coding agent ships a machine interface: codex has `exec --json` plus a full app-server JSON-RPC protocol, claude-code has `-p` with stream-json, and pi/opencode/omp each expose a JSON event stream. Wrapping these CLIs as your backend is the fastest path to subscription-backed coding — but they own their agent loop, their login, and their permission model. looplane's answer: let the external CLI fully own its loop while looplane holds only three things — an isolated working copy, patch audit, and final verification. One runtime never impersonates another.

Learning Design from Mature Coding Agents (22): The Gateway Pattern — Turning Any Provider into an OpenAI-Compatible Endpoint

The ecosystem treats /v1/chat/completions as the lingua franca, but your providers don't all speak it. The five reference projects split into three camps: pi and OpenCode make the client speak every dialect natively so no gateway is needed; OMP builds a real protocol translator (foreign wire → neutral context → provider adapter, no raw passthrough); Codex and Claude Code run proxies that translate nothing and exist purely to force traffic through a controllable path. Looplane copies OMP's boundary but narrows it to one wire in, one out: strictly parse OpenAI Chat into a canonical contract, then dispatch to any ModelProvider — and along the way hit a cross-event-loop client-close bug whose lesson is that provider lifecycles belong to the ASGI lifespan, not the signal handler.

Learning Design from Mature Coding Agents (21): Headless Mode and CI Usage — When Nobody Can Click Approve

The biggest problem when an agent enters CI is approval: no TTY, nobody to click approve. The five reference projects converge on two strategies — delegate permission decisions to the calling program (claude-code's control protocol), or replace approval semantics entirely (codex defaults to Never plus sandboxing, opencode auto-rejects). looplane keeps one AgentRunner loop and injects a different ApprovalPolicy: headless uses HeadlessApprovalPolicy, which never reads stdin so it cannot hang the pipeline, and denies EXECUTE by default — fail closed.

Learning from Mature Coding Agents (14): Onboarding Design — Provider-Aware Init and Instant Verification

A blank config file drives people away; a bad credential discovered too late drives them away faster. All five mature agents treat setup as a first-class state, and looplane adds the step most of them skip: verify the key right after saving it.

Learning Design from Mature Coding Agents (20): The Run Artifacts Contract—What Makes a Run Auditable After It Ends?

After an agent run finishes, 'the model said it's done' is not evidence. Codex splits traces into a manifest + JSONL + payloads bundle, omp mirrors on-disk files into SQLite, pi indexes native session files with runs.jsonl. Looplane picked the strictest option: six fixed files per run, the run is incomplete if any is missing, and patch review reads changes.patch—not anyone's verbal claim.

Learning from Mature Coding Agents (16): Runtime Abstraction and Capability Handshake

Five external CLIs expose five different machine interfaces: JSONL event streams, JSON-RPC handshake, HTTP API, ACP, stream-json. The right way to support them is not one interface that pretends they're identical — it's a narrow runtime boundary plus an honest capability matrix. Availability means installed, not authenticated; protocol drift fails closed.

Learning from Mature Coding Agents (11): Sandboxes and Remote Execution — Deploying on Cloudflare Sandbox

A local sandbox limits the blast radius of an agent on your machine; a cloud sandbox is about moving code safely onto someone else's machine. All five mature projects solve the first problem; only looplane actually deployed the second. Lessons from production: mocks can't catch SSE framing, green CI can't catch a stale wheel, and cleanup paths deserve timeouts just as much as success paths.

Learning Agent Design from Mature Coding Agents (19): Session Persistence and Crash Recovery — Rescuing State After the Agent Dies

All five agents store sessions as append-only JSONL plus some form of single-writer protection, but crash recovery lives in the details: pi repairs torn tails, codex reopens and retries after write failures, and looplane picked a 'manifest first' ordering that reduces the only crash window to one repairable slot. This post dissects each project's write ordering and fail-closed conditions, all cited at file#symbol level.

Learning from Mature Coding Agents (12): Can Small Models Code? — Capability Boundaries and Eval Discipline

Small models don't fail at reasoning first — they fail at format stability: tool-call JSON, diff hunk arithmetic, and context budgets all break. The mature harnesses build evals on real model behavior (pi's model-backed evals, OMP calibrating benchmarks from real session logs, Codex even relaxing its parser for weaker models). looplane picks the narrowest but hardest path: one fixture, five real Ollama runs, a manifest declaring exactly which files and patch fragments count as success — and M2's failure kept verbatim as evidence. Never pass mock off as E2E; never spin partial success into full passes.

Learning Design from Mature Coding Agents (17): Startup Performance and Engineering Discipline — It Was Never the Language

A CLI tool pays its startup cost on every invocation, and performance optimization without a baseline means no regression protection. codex uses daemon reuse and skill snapshot caches; claude-code splits its entrypoint into dynamic imports plus a built-in startup profiler; opencode and omp each maintain lazy-loading discipline; pi does none of it and leans on Bun being fast. looplane is Python — slow by birth — so it applies the full discipline: lazy imports, single-flight disk cache, background controller prewarming, and hyperfine paired benchmarks wired to a CI gate that fails on >10% regression.

Learning Design from Mature Coding Agents (8): The Right Way and the Wrong Way to Use Subscriptions — OAuth and Credential Boundaries

The five reference projects split into three camps on subscription auth. Codex and Claude Code implement OAuth only for their own official clients and store tokens in the OS keyring. pi and OMP directly reuse Claude Code's client ID to implement Pro/Max OAuth — technically feasible, but Anthropic's docs explicitly bar third parties from offering claude.ai login without approval. OpenCode removed its bundled Pro/Max plugins entirely, the cleanest policy precedent in the ecosystem. Looplane's rules: own your grant, never scrape another CLI's credentials, accept third-party OAuth only when the provider clearly supports it, and never copy or forward credentials.

Learning Agent Design from Mature Coding Agents (24): Testing a Moving Agent — fake-CLI Contracts, Recorded Streams, TUI Pilot

An agent's two dependencies — the LLM and external CLIs — are both non-deterministic, but mature projects separate 'the moving parts' from 'the shape of the boundary': codex fakes the Responses API with wiremock plus a scripted SSE server and pins its TUI with insta snapshots; opencode built a VCR-style http-recorder package; pi splits model-backed evals from unit tests into two vitest configs; omp wraps its edit benchmark itself in unit tests. looplane stacks four layers against external CLIs: unit tests, fake-CLI contract tests, recorded-stream integration proofs, and Textual pilot TUI tests. The methodology in one line: record real non-deterministic output, then make deterministic assertions about it.

Learning Design from Mature Coding Agents (15): From Full-Screen TUI to Semantic Transcript

Mature coding agent TUIs never print the event stream directly — they build a typed projection layer first and update it in place. looplane took three steps (full-screen composition, runtime-first dual modes, removing the Ask/Agent split) before two old constraints — non-streaming output and resume-without-replay — were truly lifted.

Learning Agent Design from Mature Coding Agents (5): The Verification Gate — Changed Files Isn't Success, Verified Is

None of the five reference projects enforces 'all declared verification commands pass' at the harness level: pi leaves verification to the model, OpenCode and Codex put it in the system prompt, Claude Code uses a separate adversarial verifier subagent but as a soft contract, and only OMP's cleanse actually runs checks from harness code. looplane takes the hardest path: if files changed, every declared verification command must pass before terminal_reason=verified; with no changes, checks don't rerun (no_changes). Whether to verify is decided by code, not by the model.

Why Python: The Cost and Compensation of Language Choice for Coding Agents

None of the five mature coding agents use Python — pi/opencode/claude-code run on TypeScript, codex rewrote TS into Rust, omp bolted ~80k lines of Rust native crates onto its hot path. looplane still chose Python; the costs are startup performance and packaging, compensated by lazy imports, uv, and Cloudflare Sandbox.

Which Graph RAG to Choose: GraphRAG v3.1.2 vs LightRAG vs HippoRAG 2 — Design, Cost, and Selection

Same 'knowledge graph + retrieval' label, three different bets: Microsoft GraphRAG v3.1.2 pays indexing cost for global summarization, LightRAG cuts cost with dual-level retrieval and incremental updates, HippoRAG 2 turns RAG into growing associative memory via PPR — this guide splits the trade-offs by component with four query modes, indexing pipelines, and a selection matrix.

Late Chunking vs Contextual Retrieval: Encode-First Zero-Cost Context vs LLM-Prefix Precision and Cost

Anthropic Contextual Retrieval uses an LLM to prefix each chunk with 50-100 tokens, cutting failure rate from 5.7% to 1.9% with rerank at ~$1.02/1M tokens; Late Chunking encodes the full 32K-window document first then mean-pools by chunk boundaries for zero extra LLM cost — the trade-off is window, latency, and update shape.

aiguide

The Complete Unsloth Guide: Fine-Tune and Run LLMs Locally, Faster

Unsloth is the fastest, most VRAM-efficient local LLM fine-tuning tool — 2× training speed and 70% less VRAM. In 2026 it added a Desktop app that bundles inference, training, image/video generation, web search, and agent integration into a complete local AI workstation.

Apple Foundation Models: Privacy-first Ecosystem AI with a 20B Sparse Model on Phones

Apple Foundation Models (AFM) is Apple's closed-ecosystem AI family. It evolved from a 3B dense model with LoRA adapters in 2024 into five models in 2026. AFM 3 Core Advanced runs a 20B IFP sparse architecture on phones while activating only 1–4B parameters; Cloud Pro runs on Google Cloud NVIDIA GPUs and is refined through Gemini distillation. There is no public API price or third-party benchmark, and access is limited to Apple's Foundation Models framework.

techguide

Choosing Mac Remote Desktop: Tailscale, Jump Desktop, RustDesk, and Built-in Screen Sharing

No budget needed: Tailscale plus built-in Screen Sharing is free and the most reliable stack for daily remote work on Mac — upgrade to Jump Desktop only if you want it smoother. This guide compares four options and fixes lid-close and sleep pitfalls in 5 minutes.

Self-Hosted Inference Overview: When Running Your Own Models Makes Sense

The key question in self-hosted inference isn't how fast the engine is — it's your GPU utilization. A fully saturated A100 costs ~$0.70 per million output tokens; at 10% utilization that becomes $7, more than most cloud APIs. This overview maps seven tools across three layers to help you decide which layer you need.

TensorRT-LLM: The Compile-for-Performance NVIDIA-Only LLM Inference Engine

TensorRT-LLM is NVIDIA's open-source LLM inference library (Apache 2.0). It offline-compiles model weights and compute graphs into optimized TensorRT engines, then serves them with custom CUDA kernels, in-flight batching, and multi-dimensional parallelism. The cost: NVIDIA GPUs only, compilation takes tens of minutes, and switching models or quantization means rebuilding.

TGI: HuggingFace's LLM Inference Server, and Why It Entered Maintenance Mode

Text Generation Inference (TGI) is HuggingFace's own LLM inference server, built in Rust and Python. It pioneered continuous batching and Flash Attention in open-source inference engines. The GitHub repository was archived on March 21, 2026, and HuggingFace recommends migrating to vLLM or SGLang. TGI still matters: it defined the architectural baseline that successor engines inherited, and many HuggingFace Inference Endpoints still run it.

2021 AI Conference Guide: Computer Vision

2021 was the year Transformers decisively entered computer vision: Swin Transformer won the ICCV Best Paper award, DINO showed that a self-supervised ViT could learn object segmentation without labels, and NeRF grew from one paper into an entire subfield. CVPR and ICCV both moved fully online because of the pandemic, yet the work published that year shaped architectural choices across computer vision for years to come.

2021 AI Conference Guide: Machine Learning

2021 was the year diffusion models surpassed GANs, self-supervised learning made theoretical breakthroughs, and reinforcement learning confronted weaknesses in its evaluation methodology. NeurIPS received a then-record 9,122 submissions, ICLR’s Score-Based Generative Modeling paper became a theoretical foundation for the diffusion ecosystem, and ICML delivered substantial work on optimization theory and the dynamics of self-supervised learning.

2021 AI Conference Guide: Natural Language Processing

2021 marked NLP’s shift from fine-tuning an entire model to adapting only a small fraction of its parameters. Prefix-Tuning at ACL, LoRA on arXiv, and Prompt Tuning at EMNLP all appeared that year; ACL Rolling Review launched; and the Findings track established itself as a second publication channel.

What AI Conferences Published in 2021: Transformers Spread, Self-Supervised Learning, and the Start of Diffusion

2021 was a dividing line for major AI conferences. Transformers spread from NLP throughout computer vision and time-series research, self-supervised learning became the most common cross-conference theme, and a diffusion model won an ICLR Outstanding Paper award before anyone realized it would displace GANs. Meanwhile, GNNs and federated learning reached historic peaks in paper volume before beginning to decline.

2022 AI Conference Guide: Computer Vision

2022 marked computer vision’s turn from recognition toward generation. Latent Diffusion Models appeared at CVPR and led to Stable Diffusion; NeRF research jumped from 25 papers in 2021 to more than 50 at CVPR alone; ConvNeXt mounted a compelling counterattack for CNNs; and ECCV in Tel Aviv set a record with 157 oral papers.

2022 AI Conference Guide: Machine Learning

2022 was the year diffusion models took center stage, Chinchilla scaling laws rewrote large-model training, and Chain-of-Thought turned reasoning into an ability that prompts could elicit. NeurIPS passed 10,000 submissions; three of its 13 Outstanding Papers directly concerned diffusion; and Chinchilla and data pruning both challenged the belief that bigger was always better. On the eve of ChatGPT’s release, every required piece fell into place at that year’s conferences.

2022 AI Conference Guide: Natural Language Processing

2022 marked NLP’s shift from demonstrating model capabilities toward aligning and controlling them. InstructGPT brought RLHF into the mainstream, Chain-of-Thought showed that prompts could unlock reasoning, and Flan 2022 matured instruction-tuning methodology. ACL and NAACL adopted ARR as their sole review path, exposing infrastructure and reviewer-load problems. ChatGPT launched at year-end and rewrote the rules of NLP research.

What AI Conferences Published in 2022: The Diffusion Boom, Chain-of-Thought, and the Eve of ChatGPT

2022 was a turning point at major AI conferences. Diffusion models moved from emerging to mainstream, with two NeurIPS Outstanding Papers; Chinchilla rewrote scaling laws; Chain-of-Thought showed that large models could reason; and InstructGPT used RLHF to teach language models to follow instructions. When ChatGPT launched at year-end, these academic topics instantly became global news.

A Guide to the Top AI Conferences of 2023: Computer Vision

In 2023, computer vision moved from seeing images to understanding, generating, and controlling them. Segment Anything turned segmentation into a general zero-shot capability, ControlNet made diffusion models precisely controllable, and 3D Gaussian Splatting challenged NeRF with real-time rendering. CVPR received more than 9,000 submissions and ICCV more than 8,000 as both conferences returned to in-person events.

A Guide to the Top AI Conferences of 2023: Machine Learning

In 2023, LLMs took over the machine-learning conference agenda. NeurIPS received more than 12,000 submissions; both Outstanding Papers addressed large models, while runner-up DPO became a practical alternative to RLHF within two years. DreamFusion opened the text-to-3D field, ICML spotlighted LLM watermarking and learning-rate adaptation, and the Mamba preprint emerged as the first serious architectural challenger to the Transformer.

A Guide to the Top AI Conferences of 2023: Natural Language Processing

2023 was the first full academic year after ChatGPT, and LLMs rewrote the NLP conference agenda. ACL's Best Papers examined humor understanding and the propagation of political bias; an EMNLP Best Paper explained in-context learning through information flow; and the HackAPrompt competition paper also won an EMNLP Best Paper award, signaling that security research had entered the mainstream. The year's largest shift was from asking how to make models more accurate to asking how we can tell when a model is misleading us.

What Topics Dominated the Top AI Conferences of 2023? The Year LLMs Rewrote the Research Agenda

2023 was the first year in which LLMs comprehensively rewrote the AI research agenda. DPO received a NeurIPS Outstanding Paper Runner-Up award, ReAct became an ICLR Oral, and hallucination grew from a marginal term into a major track at every conference. Meanwhile, 3D Gaussian Splatting swept through computer vision after its SIGGRAPH debut, Mamba emerged at the end of the year to challenge the Transformer attention monopoly, and publication volume for traditional NLP pipelines began a clear decline.

2024 AI Conference Review: Computer Vision

In 2024, 3D Gaussian Splatting took over 3D reconstruction, video generation moved from research toward products, and vision-language models spread into specialized domains. CVPR received a record 11,500-plus submissions; its Best Papers were Google Research's Generative Image Dynamics and the UCSD/Google collaboration Rich Human Feedback for Text-to-Image Generation. ECCV gave its Best Paper award to Columbia's Minimalist Vision with Freeform Pixels, an unconventional return to the physics of optics.

2024 AI Conference Review: Machine Learning

ML conference submissions exploded in 2024: NeurIPS received a record 15,671 papers, while ICML and ICLR passed 9,000 and 7,000. Research shifted from training ever-larger models toward spending inference compute more intelligently, making test-time compute scaling the year's defining new direction. VAR beat diffusion with next-scale image prediction, Rectified Flow became the theoretical foundation for Stable Diffusion 3, and ICLR gave its inaugural Test of Time Award to the original VAE paper.

2024 AI Conference Review: Natural Language Processing

NLP conferences redefined themselves under LLM dominance in 2024. ACL made open science its annual theme, and four of its seven Best Papers probed fundamental limits of language models. EMNLP turned toward multilingual and cross-cultural work, with Best Papers spanning speech representations and gradient interpretability. ACL and EMNLP received more than 10,000 submissions combined, but the deeper anxiety was what remains of NLP when LLMs can perform nearly every traditional NLP task.

What Top AI Conferences Accepted in 2024: The Year of Agents and the Scaling Debate

The defining conference keywords of 2024 were agents, alignment, multimodal LLMs, and inference-time compute. The LLM share at five major conferences doubled again after its sharp 2023 rise; agent-related terms grew 4.3 times; and diffusion models graduated from an emerging topic to a second generative-AI pillar alongside LLMs. Traditional task-oriented NLP continued to contract, while GANs almost disappeared from top venues.

2025 AI Conference Review: Computer Vision

2025 was a two-conference year for computer vision, with CVPR and ICCV both taking place. CVPR received a record 13,008 submissions; Best Paper VGGT turned 3D reconstruction from iterative optimization into feed-forward inference. ICCV's Marr Prize went to BrickGPT, which generates brick structures from text that can actually be assembled. 3D Gaussian Splatting displaced NeRF, video generation moved toward products, and flow models began replacing diffusion, completing several paradigm shifts in one year.

2025 AI Conference Review: Machine Learning

ML conferences broke every submission record in 2025 and pushed peer review to its limit. NeurIPS received 21,575 papers and used more than 20,000 reviewers; ICML passed 12,000 for the first time, and ICLR reached 11,565. Reasoning and agents were the strongest trends. One NeurIPS runner-up, the conference's only perfect-score paper, challenged whether RLVR creates new reasoning ability. Awards for Alibaba Qwen's Gated Attention and a mechanistic theory of neural scaling laws showed a community moving from scaling at all costs toward understanding why scaling works.

2025 AI Conference Review: Natural Language Processing

NLP conference submissions nearly doubled in 2025: ACL received 8,360 papers and EMNLP 8,174. China-based first authors exceeded 51% at ACL, and DeepSeek's Native Sparse Attention won Best Paper. The deeper story was an identity crisis: an ACL president said 'ACL is not an AI conference,' a quantitative study asked 'Has ACL Lost Its Crown?', and EMNLP faced questions about what still distinguished it from ACL or NAACL.

What Top AI Conferences Accepted in 2025: The Agent Breakout and Reasoning Revolution

The two strongest signals at AI conferences in 2025 were reasoning papers jumping from 47 to 216, a 4.6-fold rise, and agent-related terms exceeding 150 papers with 4.3–11-fold growth. Diffusion moved from breakout topic to infrastructure; RAG became a mainstream enterprise architecture with unusual coverage across all five conferences; state-space models and world models began tracing the early 2020–2021 path of Vision Transformers. Pure prompt-engineering papers encountered reviewer fatigue.

Submitting to Top AI Conferences as an Independent Researcher: A Reality Check

Publishing at a top conference as an independent researcher is possible, but the numbers are harsh: single-author papers have fallen to a single-digit share, the average author count has risen from 3 to 5, and the top 20 institutions account for 35-50% of authorships. Andreas Madsen spent eight months working without pay, earned an ICLR Spotlight, and still ended up returning for a PhD. This article examines real cases, evidence of review bias, and viable paths for researchers without a large lab behind them.

What Happens to a Top-Conference Paper from Submission to Publication

An AI conference paper passes through anonymized submission, format screening, reviewer bidding and assignment, independent scores from 3-4 reviewers, an author rebuttal, AC/SAC/PC decisions, and camera-ready revision—a process lasting about 4-5 months. ACL-family conferences add ARR's rolling-review model, in which review comes before the author commits the paper to a venue.

Main Track, Findings, and D&B Track: Three Publication Routes at Top AI Conferences

A paper submitted to a major conference can follow three very different routes: the Main Track is the highest-threshold formal publication, Findings is the ACL family's companion venue for solid work that misses the main program, and NeurIPS created the D&B Track specifically for datasets and evaluation methodology. Their review standards, prestige, and career signals differ enough that understanding the route matters before writing the paper.

Who Submits to Top AI Conferences: Labs, Companies, and the Global Map

The institutional map of top AI conferences is being rapidly redrawn. Industry labs dominate frontier model development—nearly 90% of notable models came from industry in 2024—but academia remains the largest source of highly cited research. Chinese universities went from challengers to nearly half of NeurIPS paper volume in five years, while OpenAI and Anthropic have nearly vanished from conference author lists. The decoupling of publication volume from research capability is the defining signal.

aideep-dive

How Marin Trains 535B: Scaling Ladder, MoE Expert Parallel, Harrier Data and Live W&B

Stanford Marin pre-registers a paloma macro-loss of 2.04 with a 5-rung Scaling Ladder at 1% cost, then trains 535B-A23B on 11×GB200 in public with live W&B telemetry — 847 training buckets already show the most teachable frontier run.

techguide

AI Model Evaluation Sources: How to Judge Whether a Model Is Actually Good

You cannot take model vendors' self-reported scores at face value. This guide covers the most important independent evaluation platforms, domain benchmarks, adoption indicators, and official sources in 2026: what each measures, how to read it, where it is biased, and which figures matter for different use cases.

Claude——From AI Safety Lab to SWE-bench Champion, the Strongest Closed-Source Agent Choice

Claude is Anthropic's closed-source LLM family, known for Constitutional AI training, agent capabilities, and coding performance. In July 2026, Opus 5 scored 96% on SWE-bench Verified to claim the coding crown, while Fable 5 led general capability at 83% on LiveBench. Four tiers (Fable / Opus / Sonnet / Haiku) span $1–$10, making this the only family in the series with zero open weights.

Cohere — The RAG-Native Outlier: How Command, Embed, Rerank, and Aya Fit Together

Cohere is the only family that ships generation, retrieval, reranking, and multilingual as distinct products. Command A runs 256K context on two GPUs at 111B, Embed v4 does mixed image-text retrieval, Rerank v4 handles 32K semi-structured data, and Aya covers 101 languages — a four-piece stack built for RAG. This post breaks down each pillar's positioning, licensing, and selection guide.

DeepSeek: From an MoE Lab to OpenRouter's Most-used Open Model

DeepSeek used MLA and MoE innovations to drive inference costs to an industry low. V4 Flash activates only 13B parameters while approaching frontier-model quality and ranks first by OpenRouter usage. This guide traces V1 through V4, the R1 reasoning branch, and how to choose each version.

Gemini——Google's Native Multimodal Flagship: 1M Context and Scientific Reasoning Champion

Gemini is Google DeepMind's native multimodal LLM family, famed for a 1M-token context window and native video/speech input plus scientific reasoning. 3.1 Pro tops GPQA Diamond 94.1% and ARC-AGI-2 77.1% to claim science-reasoning dual crowns, at $2/$12—1/6 of Claude. 3.7 Flash delivers near-Pro agent capability for $0.75/$3.75.

GLM——From a Tsinghua Lab to a 744B Open-Source Flagship, and GLM-5.3's Cybersecurity Surge

GLM is Zhipu AI (Z.ai)'s open LLM family from Tsinghua's KEG Lab. GLM-5.3 (2026/08) lifts coding +50% over the previous generation, hits 84.5% on CyberGym ahead of Anthropic Mythos 5 and OpenAI GPT-5.6 Sol, and scores 60 on the Artificial Analysis Intelligence Index tied with Kimi K3 for open-source #1. The only frontier open model trained entirely on Huawei Ascend.

GPT——Closed API for Revenue, Open GPT-OSS for Ecosystem: the Unified Routing Platform Behind the World's Largest AI Service

GPT is OpenAI's LLM family, from 117M parameters in 2018 to the three-tier GPT-5.6 Sol/Terra/Luna lineup in 2026, serving 1B+ users and 2M enterprise customers. GPT-5.6 Sol leads LiveBench 81.1%, Terminal-Bench 2.1 88.8%, and Artificial Analysis Coding Agent Index 80 across multiple agentic benchmarks, while OpenAI's first open-weight model GPT-OSS ships under Apache 2.0.

Grok — From a 314B Open-Source Bet to Grok 4.6/Build/Imagine, xAI's Distribution-Driven Catch-Up

Grok is xAI's LLM family: founded July 2023, opened with a 314B MoE under Apache 2.0 in March 2024, and two and a half years later spans Grok 4.6 (500K, $2/$6, four reasoning levels), Grok 4 Fast (2M), Imagine for image/video, and Grok Build for terminal coding — its moat is distribution (X / grok.com / Tesla / Bedrock), not single-model supremacy. This post traces Grok 1→4.6, sub-line positioning, pricing, and licensing traps.

Kimi——From a 200K Long-Context Tool to a 2.8T Open-Source Frontier, and K3's Architectural Leap

Kimi is Moonshot AI's LLM family, born from ultra-long context. Kimi K3 (2026/07) is the world's first open 3T-class model—2.8T params, 104B active, 1M context, scoring 60 on the Artificial Analysis Intelligence Index tied with GLM-5.3 for open-source #1. Its Kimi Delta Attention brings a 2.5× scaling efficiency gain.

Llama——From Open-Source Experiment to the Most Deployed Open LLM, and Meta's Closed-Source Pivot

Llama is Meta's open-source LLM family, with the largest enterprise deployment footprint and the most mature ecosystem. Llama 4 Scout (10M context) and Maverick (17B active / 400B total MoE) are the current open multimodal benchmarks, but Meta pivoted to closed-source Muse Spark in April 2026—Llama 4 is likely the last major open Llama, and its license is not truly open (Llama 4 Community License, separate license required above 700M MAU).

Mistral——Europe's Open AI Challenger: Smaller Models and European Sovereignty as a Different Bet

Mistral is Europe's most successful AI startup, cutting through the market with a 'smaller, faster, cheaper' strategy and European data-sovereignty positioning. Mistral Large 3 is Europe's strongest commercial LLM, Small 4 is the 24B efficiency king, and Medium 3.5 is the open Modified-MIT model optimized for agentic coding. Its moat is not technical scale but the 'European compliance' card.

Qwen: Open Weights at Every Size from 0.8B to 2.4T — How HuggingFace's Download Champion Runs a Two-Track Play

Qwen is the most-downloaded model family on HuggingFace, spanning sizes from 0.8B to 2.4T. In August 2026, Alibaba open-sourced a Max-tier flagship for the first time (Qwen3.8-2.4T-A95B) — but swapped the customary Apache 2.0 license for custom terms. Meanwhile the other new release, Qwen3.8-27B, runs native vision on laptop-class hardware and is the only one shipping under Apache 2.0. This post traces the family from 2023 through generation 3.8, explains how the open line and the commercial line split apart, and helps you pick the right model at each tier.

AI Model Landscape: The 2026 Map You Need

In 2026, AI models span seven major categories and more than 20 subcategories. This introduction to the AI Model Families series maps use cases to models and models to families, with current rankings and selection advice for each use case.

Antigravity CLI: Google Replaces a 100K-Star Open-Source Tool with a Closed-Source Go Binary

At Google I/O 2026, Antigravity CLI (agy) replaced Apache 2.0 Gemini CLI with a closed-source Go binary. Technical upgrades — multi-agent orchestration, native sandbox, millisecond startup — but free tier cut 98%, open-to-closed source, 28-day transition window. Community reaction was sharp.

Grok Build: xAI's Rust Coding Agent That Uploaded Your Repo Before Going Open Source

Grok Build is xAI's Rust coding agent — 845K LOC, 8 parallel sub-agents, Arena Mode. May 2026 beta, July open-sourced (Apache 2.0) — but the direct trigger for open-sourcing was a privacy incident: it silently uploaded entire repos (including SSH keys, .env files) to Google Cloud Storage at a 27,800x traffic ratio. The exfiltration code remains in the binary, disabled only by a server-side flag.

Muse Code: Meta's First Coding Agent, Trading Training Rights for a 20x Discount

In August 2026, Meta Superintelligence Labs released Muse Code beta. Closed-source static binary, Muse Spark 1.2 model, parallel persistent sub-agents with worktree isolation. The biggest controversy is pricing: Standard at $1.25/$4.25 per M tokens, or Contributor at $0.10/$0.20 — 20x cheaper, but your code enters Meta's training pipeline.

How to Pick a Self-Hosted Inference Server: From Ollama to Xinference, Six Tools and Their Trade-Offs

Self-hosted inference servers fall into three layers: execution engine (llama.cpp), serving engine (vLLM, SGLang), and model management platform (Ollama, Xinference, Triton). Picking the right layer matters more than picking the right tool — ask where your bottleneck is before deciding where to add complexity.

Xinference: One Platform to Manage LLM, Embedding, Speech, and Image Models

Xinference wraps vLLM, SGLang, llama.cpp, Transformers, and MLX under a single management layer, using a Web UI and OpenAI-compatible API to manage LLMs, embedding, rerank, speech, and image models — suited for self-hosted deployments that need multiple model types to coexist. But the management layer's parsing logic also creates a larger attack surface than pure serving engines (CVE-2026-61539 is a case study).

What Is an AI 'Top Conference': Why CCF, CORE and h5-index Disagree

There's no official certificate for being an 'AI top conference.' It's a community consensus built from four independent signals — CCF-A, CORE-A*, a high Google Scholar h5-index, and a low acceptance rate — and those four signals frequently disagree. ICLR being completely absent from CCF's list is a live example.

techdeep-dive

Agent Platform Deep Dive (8) — Context/Memory and Cloudflare Deployment: Seamless Migration from Local Development to Production

Agent Platform uses a Cloudflare-first architecture: local `npm run dev` runs Node-based simulations, while production maps to Workers + Workers Assets + D1 + KV + R2 + Vectorize + Queues + Workflows + Durable Objects + Workers AI. The Runtime interfaces stay the same (InMemory → Cloudflare implementations), so upper layers migrate without noticing. Deployment requires only `wrangler login` → create resources → fill in IDs → `wrangler secret put` → `wrangler deploy`. CI/CD watches the main branch and runs typecheck + build + dry-run + migration + deploy.

techdeep-dive

Agent Platform Deep Dive (VII)—Evaluation & Quality Gates: Comprehensive Evaluation, Regression Prevention, and an Immune System for Skill Releases

Evaluation is Agent Platform's quality immune system: instead of collecting statistics only after a run, it enforces checks throughout Pre-run, In-run, and Post-run execution. Seven eval categories cover Flow → Step → Skill → Artifact → Evidence → Policy → Regression. A Skill release must pass five gates—Trigger, Functional, Policy, Regression, and Human Review—and any failure blocks it. The Learning Loop moves from Run signals through Proposal, Human Review, Sandbox Eval, Quality Gate, and Publish, under one strict rule: agents propose, humans review, and eval gates decide whether a change can ship.

techdeep-dive

Agent Platform Deep Dive (Part 2) — Flow Runtime: Versioned Flows, Checkpoints, and Resume/Retry Mechanisms

Flow Runtime is the heart of Agent Platform: a Flow becomes immutable when published, each Run is bound to a specific version and preset, Steps move through a DAG according to edge conditions, every boundary saves a checkpoint, and resume/retry-step preserves the complete trace history.

techdeep-dive

Agent Platform Deep Dive (Part 6) — Observability, Evidence, and Artifacts: Structured Traces, Claim-to-Source Lineage, and Versioned Outputs

Observability is a first-class capability, not logging added after the fact: a structured trace connects FlowRun→StepRun→SkillInvocation→ProviderCall→ToolInvocation→GuardResult→EvidenceItem→ArtifactVersion. The Evidence Store traces every claim back to its source, excerpt, citation, confidence, and conflicts. Artifact versioning supports approve/reject/regenerate without deleting history. Context Snapshots allocate token budgets by category and record automatic compression when a block exceeds its budget. Procedural, episodic, and semantic memory can be written only through proposals reviewed by a human.

techdeep-dive

Agent Platform: An In-Depth Look at an Open-Source AI Workflow Control Plane (Part 1)—Architecture and Positioning

Agent Platform turns AI agents from a blank chat window into a structured workflow platform whose behavior can be defined, versioned, observed, verified, and improved. Its built-in Deep Research seed flow demonstrates the complete feedback loop.

techdeep-dive

Agent Platform Deep Dive (Part 5) — Policy Engine: Runtime Guards, Budget Control, Human Approval, and Loop Protection

The Policy Engine acts as the Agent Platform's constitution and enforcement layer: policies are versioned and bound to flows and presets; four guard layers enforce rules at step boundaries; budgets cap cost, tokens, runtime, iterations, and tool calls; external writes require human approval; loop detection trips circuit breakers; and escalation records provide an auditable trail. Rules are configuration-driven, so adding one means changing JSON rather than hard-coded logic.

techdeep-dive

Agent Platform Deep Dive (Part 4) — Provider Router & MCP: Multi-Provider Routing, Fallback Chains, and an OpenAI-Compatible Proxy

The Provider Router is Agent Platform's model and tool gateway: it unifies 30+ providers, MCP tool discovery, step-local permission control, fallback chains with RRF fusion, and an OpenAI-compatible Proxy that existing SDKs can use without code changes. It is configuration-driven rather than hard-coded, with provider-health-aware routing.

techdeep-dive

Agent Platform Deep Dive (3) — Skill System: Versioned Capability Packages, Explicit Binding, and the Learning Loop

A Skill is a versioned, installable, and auditable capability package. Its dual-file architecture separates metadata from instructions, explicit binding replaces model-driven routing, and every invocation is recorded. The Learning Loop turns run signals into proposals, sandbox evaluations, human review, and publication while enforcing the principle: agents propose, humans review, and evals serve as the gate.

techdeep-dive

Groundlane Series Part 1: Why AI Agents Need a Controlled Web Access Layer

Groundlane is an open-source TypeScript remote MCP server (v0.1.0) giving AI agents web_search, web_fetch, and web_extract through a single stable contract, with auth, provider routing, and resource limits kept at the operator boundary.

techdeep-dive

Groundlane Series Part 2: Actual Calls, Response Structures, and Error Boundaries for the Three MCP Tools

Hands-on parameter choices and response structures for web_search (ten adapters, RRF merge, dual-provider default), web_fetch (format/render strategies, finalUrl provenance), and web_extract (CSS selector determinism, no implicit LLM step), with verifiable error boundaries.

techdeep-dive

Groundlane Series Part 3: Comparing with Traditional Approaches — WebFetch, stealth_fetch, puppeteer, and requests

A four-dimension comparison (determinism, replaceability, identity boundary, operational cost) between Groundlane's controlled remote MCP contract and traditional local approaches (WebFetch, stealth_fetch, puppeteer, requests), with verifiable scenario recommendations.

techdeep-dive

Groundlane Series Part 4: In-Site Application — Verified Workflow with Existing MCP Tools and Usage-Mode Rules

Based on the in-site groundlane skill (mcp__groundlane__*) and usage-modes rules, this part describes reproducible steps for reference verification and digest data collection — without assuming unimplemented features or using deprecated stealth_fetch.

techdeep-dive

Groundlane Series Part 5: Pitfalls and Best Practices — timeout, selector, render mode, version-change risk, and security boundaries

A reproducible checklist of the most common operational pitfalls: truncated results from fixed caps, selector errors tied to DOM stability, render-mode cost/determinism tradeoffs, version-change verification (v0.1.0 preview), and the non-negotiable security boundaries (URL policy, auth, concurrency, budget).

techdeep-dive

Building a Taiwan Stock Research Agent (Part 1): Why Taiwan Needs Its Own Research Agent

US-stock LLM agents have attracted nearly 100,000 GitHub stars, yet no Taiwan-stock project has even passed 10. I consolidated three side projects into a Taiwan-stock research agent where every conclusion must first survive a backtest; this article explains why.

techdeep-dive

Building a Taiwan Stock Research Agent (Part 2): LangGraph Parallel Architecture—Five Analysts Working at Once

Five analysts fan out in parallel within one superstep, so latency is max rather than sum; backtesting and reflection stand before synthesis, restricting the LLM to explaining evidence that already exists.

techdeep-dive

Building a Taiwan Stock Research Agent (Part 3): Tiered LLMs and a Degradation Chain—API, Local CLI, and Dictionary Fallbacks

Only two roles call an LLM; every other analyst remains fully programmatic. Each call follows an Anthropic API → local Claude CLI → rules-based degradation chain, and cost accounting trusts only provider-reported values—unknown cost is never treated as $0.

techdeep-dive

Building a Taiwan Stock Research Agent (Part 4): Backtest Accountability—Why Backtests Lie

This project has one core rule: every LLM conclusion must first pass a historical backtest of the same signals. When expectancy is negative, synthesis cannot issue an optimistic verdict. Each of the four traps that make backtests lie has a programmatic countermeasure.

techdeep-dive

Building a Taiwan Stock Research Agent (Part 5): Walk-Forward Evaluation, Run Cards, and an Honest 50% Baseline

I do not measure whether the agent ‘feels accurate.’ I freeze parameters in walk-forward OOS tests, record a hash of every input in run cards, and keep the honest 5/10 = 50% golden-eval baseline so the agent has to admit that it is not accurate yet.

techdeep-dive

Building a Taiwan Stock Research Agent (Part 6): Making Every Number in an LLM Report Auditable

Numbers are the easiest part of an LLM report to hallucinate. I therefore put every trusted number into a SHA-256-addressed evidence manifest and let the LLM cite only {{fact.id}} placeholders. If it writes a bare number, the entire output is discarded and replaced with a deterministic template.

techdeep-dive

Building a Taiwan Stock Research Agent (Part 7): The Copilot Loop—Plan Contracts, Verifiable Sources, and Human Review

A research request first becomes a ResearchPlan that requires human approval. External documents must be fetched in full, and verbatim quotes must be verified before they can enter a report. Quant review is always append-only, and free-text feedback never flows back into a prompt. This is the complete M5 Copilot loop.

techdeep-dive

Building a Taiwan Stock Research Agent (Part 8): The Boundary Between Research and Paper Orders—Content-Addressed Execution Contracts

Three frozen Pydantic contracts weld the boundary between a research artifact and order-placement authority shut: content addressing, eight hard gates, and paper-only execution, while the agent never touches credentials.

techdeep-dive

Building a Taiwan Stock Research Agent (Part 9): Deployment Boundaries—from Docker to a Public API on Cloudflare Containers

The full deployment path for a Python agent, from local uv run to Docker to a public API on Cloudflare Containers: the Worker enforces authentication, the Container runs FastAPI, secrets never enter the image, and the service sleeps automatically after 10 idle minutes—the right way for a side project to save money.

Building an Academic Search Pipeline: The Roles of arXiv, OpenAlex, Crossref, Semantic Scholar, and PubMed

An academic-search pipeline cannot simply concatenate five APIs: use arXiv or PubMed for domain discovery, align OpenAlex and Semantic Scholar records through DOI, PMID, and arXiv IDs, then use Crossref and PubMed relationships to check the version of record, corrections, and retractions.

aideep-dive

AG2: Organizing Multi-Agent Collaboration with Conversations and GroupChat

AG2 continues AutoGen's ConversableAgent model: agents collaborate through messages, while GroupChatManager selects the next speaker by round robin, manual choice, randomness, or an LLM.

aideep-dive

Choosing an Agent Framework in 2026: LangGraph, CrewAI, MAF, AG2, Mastra, Pydantic AI, and DSPy

These seven tools are not one product category: LangGraph, MAF, and Mastra emphasize durable workflows; CrewAI and AG2 emphasize multi-agent collaboration; Pydantic AI emphasizes typed Python agents; DSPy optimizes AI programs against data and metrics. Choose the control model first.

Writing Search Queries for Agents: Keywords, Semantic Descriptions, Decomposition, and Rewriting

An agent should not send the user's sentence unchanged to every search service. Classify the need as exact lookup, keyword, semantic, or fielded search; move source, date, language, and field constraints into native provider parameters; then rewrite according to zero-result, overbroad, stale, or source-mismatch symptoms.

aideep-dive

Amazon Bedrock Deep Dive: Putting Model APIs Inside the AWS Governance Boundary

Amazon Bedrock is more than a reseller for multiple model APIs. It brings model invocation, IAM, Regions, Knowledge Bases, Guardrails, and CloudWatch into one AWS control plane. It fits teams already on AWS that value governance overhead more than the lowest token price.

aideep-dive

Arize Phoenix: Turning Traces into Datasets, Experiments, and Evaluators

Phoenix is an MIT-licensed open-source LLM observability and evaluation platform. It collects traces with OpenTelemetry and OpenInference, turns production failures into versioned datasets, compares prompt, model, or RAG changes in experiments, then writes code, human, and LLM evaluator scores back as annotations. It is not Arize AX, and self-hosting defaults require security work.

Giving an Agent Access to Logged-In Websites: Sessions, Permissions, and Automation Boundaries

Authenticated browser state is not a convenience setting; it is a credential that can impersonate its owner. Use a dedicated low-privilege account and isolated profile, separate reading from reversible writes and high-risk transactions, and leave MFA plus final submission to a human.

aideep-dive

Baseten: The Model Inference Lifecycle from Truss Packaging to Autoscaling

Baseten puts custom-model packaging, GPU deployment, inference engines, autoscaling, and release workflows on one platform. Its value is not another OpenAI API, but retaining runtime control while operating less GPU orchestration.

aideep-dive

Braintrust: Closing the LLM Evaluation Loop from Datasets Back to Production

Braintrust connects versioned datasets, immutable experiments, scorers, and production traces into one evaluation loop. Its value is not another score but the ability to turn production failures into offline tests. The company announced an $80 million Series B in February 2026; its customer list is company-reported.

aiguide

Brave Search API Complete Guide: An Independent Search Index for Agents

Brave Search API exposes five endpoint families—Web, News, Images, Videos, and LLM Context—backed by Brave's own Web index and ranking models. Its core search is not merely a Google SERP wrapper.

aideep-dive

Browserbase: Turning Agent Browsers into Operable Infrastructure

Browserbase combines remote Chromium, persistent Contexts, proxies, and a Session Inspector in one control plane. It operates browser fleets; it does not decide an agent's next action. As of August 2026, the company reports more than 35 million monthly browser sessions and over 10,000 customers.

aideep-dive

Cartesia Deep Dive: From Sonic Streaming TTS to a Real-Time Voice Agent Pipeline

Cartesia's core is Sonic real-time TTS, Ink STT, and streaming inference. Although it offers the Line voice-agent platform in 2026, buyers must still separate the model layer from telephony orchestration and design consent, retention, and fallback for cloned voices.

aideep-dive

Cerebras Inference: Know the Bottleneck Before Putting Wafer-Scale Speed in an Agent Loop

Cerebras can dramatically accelerate generation on supported models, but agent latency still depends on prefill, tool I/O, model quality, and platform compatibility.

aideep-dive

Chroma Vector Database: From Local RAG to Distributed Retrieval

Chroma manages embeddings, documents, and metadata through collections; it embeds into Python locally, uses HNSW on a single node, and separates compute from storage with object storage, SSD caches, and SPANN in distributed deployments.

aideep-dive

Claude Code Startup Playbook: Five Operating Principles from Anthropic's Guide

Anthropic interviewed 15 startups and distilled five Claude Code operating principles: everyone ships, automate the tedium, trust but verify, build for rebuilding, prototype to productionize. ClickHouse shipped 30% more features, Clay automated 100% of bug triage, Artemis Security hit 6,000+ PRs per week.

aideep-dive

Cloudflare Kitesurf: An Agent Browser That Is Not Chromium—and What It Trades for Scale

Kitesurf is a non-Chromium browser backend in Browser Run that remains in beta. It trades pixel compatibility, persistent authenticated sessions, WebGL, and full anti-bot behavior for low CPU and memory through Workers isolates, Rust/Wasm, and stateless components.

Cloudflare Sandboxes Deep Dive: How Workers, Durable Objects, and Containers Form an Agent Runtime

Cloudflare Sandboxes uses a Worker as the entry point, a named Durable Object as the control plane, and a Container inside an isolated VM as the execution plane. It fits Cloudflare-native fleets of ephemeral Linux workspaces, but persistence, security boundaries, and three layers of billing remain your responsibility.

Completing CMU 07-280: What You Know, What Is Missing, and What Comes Next

Finishing 07-280 means more than reading 24 guides: produce a search engine, supervised-model comparison, CNN/GPT-2 experiments, and a small RL-plus-MCTS system before choosing 07-380, 10-301, or a specialist course.

Reading CMU 07-280: Why Search, GPT-2, and AlphaZero Belong in One Course

07-280 is CMU's new Spring 2026 AI+ML core: 24 lectures and 12 main assignments move from heuristic search and CSPs to AlexNet, GPT-2, and AlphaZero. Its public material supports self-study, but complete recordings, Canvas checkpoints, Gradescope, and staff feedback remain unavailable.

CMU 07-280 Lecture 1: The Shared Problem Behind AI, ML, and Representation Learning

Lecture 1 uses an alien autoencoder, the scope of AI and ML, and AI history to establish the course's coordinate system: an intelligent system turns inputs into representations and decisions under uncertainty.

CMU 07-280 Lecture 2: Heuristic Search from UCS and Greedy to A*

Lecture 2 decomposes search into a problem, frontier, and priority: UCS uses paid cost, Greedy uses estimated remaining cost, and A* combines them as `f=g+h`; tree and graph search require different optimality conditions.

CMU 07-280 Lecture 3: Minimax, Alpha-Beta, and Expectimax

Lecture 3 turns a single path into a contingent plan: minimax faces an optimal opponent, alpha-beta skips branches without changing the root value, and expectimax replaces worst-case choice with probability.

CMU 07-280 Lecture 4: CSPs, AC-3, and Search Order

Lecture 4 exposes structure through variables, domains, and constraints, then upgrades DFS with backtracking, forward checking, AC-3, MRV, and LCV; the goal is to prove failure earlier.

CMU 07-280 Lecture 5: Defining Machine Learning with Loss, Risk, and ERM

Lecture 5 formulates machine learning through `X → Y`, loss, risk, and empirical risk minimization: a training set only gives average observed loss, while the real objective remains generalization over an unknown distribution.

CMU 07-280 Lecture 6: How Decision Trees Split Data with Mutual Information

Lecture 6 recursively grows a tree from decision stumps, measures label uncertainty with entropy, and selects splits by `I(Y;W)=H(Y)-H(Y|W)`; this is computationally practical greedy ERM, not a global optimal-tree guarantee.

CMU 07-280 Lecture 7: Linear Regression and the Normal Equation

Lecture 7 applies ERM to linear functions and squared loss, moves from a one-dimensional slope to `argmin ||y-Xθ||²`, and derives the normal equation when `XᵀX` is invertible.

CMU 07-280 Lecture 8: Gradient Descent, SGD, and Learning Rate

Lecture 8 moves from a one-dimensional parabola to vector gradients and compares batch GD, SGD, and mini-batches; the learning rate determines whether updates converge, oscillate, or diverge.

CMU 07-280 Lecture 9: Logistic Regression as Probability Estimation

Lecture 9 models P(y=1|x) with a sigmoid instead of directly predicting 0 or 1, learns parameters with cross-entropy and convex optimization, and extends naturally to softmax regression.

CMU 07-280 Lecture 10: Trading Expressiveness for Stability with Features and Regularization

Lecture 10 uses φ(x) to let linear models express nonlinear functions, then controls the resulting overfitting with train/validation/test separation, L1/L2 regularization, and model selection.

CMU 07-280 Lecture 11: Building a Neural Network from Logistic Regression

Lecture 11 expands a logistic unit into a multilayer network: linear layers produce z, activations produce a, and multiple neurons jointly learn a feature transform trained through a final loss.

CMU 07-280 Lecture 12: How Backpropagation Reuses the Chain Rule

Lecture 12 treats a network as a computation graph: the forward pass stores intermediates, the backward pass propagates upstream gradients, and local linear, activation, and softmax rules compute every parameter gradient efficiently.

CMU 07-280 Lecture 13: From Reward Hacking to Auditable AI Scientists

Lecture 13 separates alignment into specification, distribution shift, oversight, and corrigibility, then uses benchmark selection, leakage, and post-hoc selection experiments to show why a final paper cannot audit an autonomous research workflow.

CMU 07-280 Lecture 14: Encoding Image Structure with Convolutional Networks

Lecture 14 replaces dense image models with local connectivity and parameter sharing, moving from convolution, stride, padding, and pooling to AlexNet, GPU data parallelism, ResNet skip connections, and BatchNorm.

CMU 07-280 Lecture 15: Separating Pretraining, Transfer Learning, and Fine-Tuning

Lecture 15 splits a pretrained model into representation g and task head h: freeze g and train only the head, or fine-tune some or all parameters at a smaller learning rate depending on data volume and source-target distance.

CMU 07-280 Lecture 16: Unifying Logistic and Linear Regression with Maximum Likelihood

Lecture 16 starts from likelihood p(D|θ), uses i.i.d. to factor the joint probability and logs to turn products into sums; Bernoulli MLE yields sample proportions, conditional Bernoulli yields logistic cross-entropy, and Gaussian noise yields squared error.

CMU 07-280 Lecture 17: From Tokenization to N-gram Language Models

Lecture 17 first decides how text becomes tokens, then uses N-grams to turn sequence probability into conditional probabilities estimated from corpus counts. Tokenization is the first design decision about what a model can see.

CMU 07-280 Lecture 18: How N-grams Train, Sample, and Fail

Lecture 18 truncates the chain rule with an N-gram Markov assumption, estimates probabilities from corpus counts, and contrasts greedy, categorical, and temperature sampling. The real bottlenecks are zero probability for unseen contexts and a fixed window.

CMU 07-280 Lecture 19: Turning Next-token Prediction into Geometry

Lecture 19 builds a minimal next-token model from two embedding matrices, dot-product similarity, softmax, and cross-entropy. Shared vector parameters replace the isolated count cells of an N-gram table.

CMU 07-280 Lecture 20: From Position Encoding to Causal Self-Attention

Lecture 20 expands one-token embeddings into sequences, adds positional information, derives Q/K/V scaled dot-product attention and causal masking, and assembles multi-head blocks into a GPT-2 skeleton.

CMU 07-280 Lecture 21: How Bellman Equations Solve Markov Decision Processes

Lecture 21 formulates stochastic sequential decisions as an MDP with known dynamics, defines value and Q-values through Bellman backups, and solves for an optimal policy with value or policy iteration.

CMU 07-280 Lecture 22: Q-learning When Dynamics Are Unknown

Lecture 22 keeps the MDP structure but removes known transitions and rewards. TD learning updates value from one sample, and Q-learning uses an off-policy target to learn optimal action values directly.

CMU 07-280 Lecture 23: From Approximate Q-learning to DQN

Lecture 23 replaces a huge Q-table with Qθ(s,a): first derive a gradient update for linear features from squared TD error, then add replay data and a fixed target network to form DQN.

CMU 07-280 Lecture 24: How Monte Carlo Tree Search Connects to AlphaZero

Spring 2026 Lecture 24 is MCTS, not Fall 2026 LLM post-training. It allocates simulations through selection, expansion, rollout, backup, and UCB, then connects policy/value heads and self-play to AlphaZero.

CMU 07-280 Stage Review I: From Search Problems to Supervised Learning

Lectures 1–12 form one decision pipeline: define states, moves, and objectives, then use heuristics, losses, regularization, and backpropagation to control an otherwise intractable search space.

CMU 07-280 Stage Review II: Building AlexNet and GPT-2 as Working Systems

Stage II uses HW8 and HW11 to test whether representation, computation graphs, training, transfer, and generation actually connect, rather than treating CNNs and Transformers as diagrams to memorize.

CMU 07-280 Stage Review III: From MDPs and Q-learning to AlphaZero

Stage III connects value, policy, bootstrapping, function approximation, and MCTS into AlphaZero: a network supplies priors and estimates, search improves decisions, and self-play creates the next training set.

CMU 11-785 Lecture 1: Introduction

Spring 2026 Lecture 1 focuses on neurons, perceptrons, connectionism, and the problem framing of deep learning. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 2: Neural Nets as Universal Approximators

Spring 2026 Lecture 22 focuses on latent variables, the ELBO, the KL term, and the reparameterization trick. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 3: Training I: Learning and Empirical Risk Minimization

Spring 2026 Lecture 3 focuses on data distributions, hypotheses, losses, empirical risk, and their roles in generalization. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 4: Training II: Gradient Descent

Spring 2026 Lecture 4 focuses on gradients, learning rates, parameter updates, and the training of a linear neuron. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 5: Training III: Backpropagation

Spring 2026 Lecture 5 focuses on computational graphs, the chain rule, local derivatives, and gradient reuse. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 6: Training IV: Convergence, Loss Surfaces, and Momentum

Spring 2026 Lecture 6 focuses on non-convex loss surfaces, curvature, saddle points, and momentum's accumulated direction. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 7: Training V: SGD and Second-order Methods

Spring 2026 Lecture 7 focuses on the tradeoffs among full-batch, mini-batch, stochastic gradients, and second-order information. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 8: Training VI: Optimizers and Regularization

Spring 2026 Lecture 8 focuses on AdaGrad, Adam, regularization, BatchNorm, Dropout, and loss selection. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 9: CNNs I

Spring 2026 Lecture 9 focuses on local connectivity, weight sharing, convolution kernels, and feature maps. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 10: CNNs II

Spring 2026 Lecture 10 focuses on stride, padding, receptive fields, and multi-channel convolution. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 11: CNNs III

Spring 2026 Lecture 11 focuses on stacked convolutional architectures, feature hierarchies, and design tradeoffs. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 12: CNNs IV

Spring 2026 Lecture 12 focuses on CNN training, architecture selection, and the end-to-end assembly of a vision model. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 13: RNNs I

Spring 2026 Lecture 13 focuses on sequence state, temporal unrolling, parameter sharing, and recurrent computation. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 14: RNNs II

Spring 2026 Lecture 14 focuses on backpropagation through time, gradient stability, and LSTM-style gated memory. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 15: Seq2Seq and Connectionist Temporal Classification

Spring 2026 Lecture 15 focuses on variable-length input/output, unknown alignment, and the CTC objective. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 16: CTC Blanks and Beam Search

Spring 2026 Lecture 16 focuses on blanks, collapse rules, prefix probabilities, and approximate decoding. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 17: Language Models and Translation

Spring 2026 Lecture 17 focuses on autoregressive factorization, conditional language models, and translation decoding. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 18: Attention and Transformers

Spring 2026 Lecture 18 focuses on queries, keys, values, scaled dot-product attention, and the Transformer block. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 19: Transformers and Newer Architectures

Spring 2026 Lecture 19 focuses on encoder/decoder structures, masks, residual paths, and architecture variants. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 20: Large Language Models

Spring 2026 Lecture 20 focuses on scaled autoregressive models, training stages, inference, and capability boundaries. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 21: Representations and Autoencoders

Spring 2026 Lecture 21 focuses on bottleneck representations, reconstruction objectives, dimensionality reduction, and representation quality. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 22: Variational Autoencoders

Spring 2026 Lecture 22 focuses on latent variables, the ELBO, the KL term, and the reparameterization trick. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 23: Diffusion Models

Spring 2026 Lecture 23 focuses on forward noising, reverse denoising, score or noise prediction, and sampling. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 24: Generative Adversarial Networks

Spring 2026 Lecture 24 focuses on the generator, discriminator, minimax objective, and training instability. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 25: Graph Neural Networks

Spring 2026 Lecture 25 focuses on message passing, aggregation, node representations, and permutation symmetry. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 26: Reinforcement Learning

Spring 2026 Lecture 26 focuses on states, actions, rewards, returns, values, and policy learning. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 27: Hopfield Networks

Spring 2026 Lecture 27 focuses on associative memory, energy functions, fixed points, and pattern retrieval. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

CMU 11-785 Lecture 28: Boltzmann Machines

Spring 2026 Lecture 28 focuses on energy-based probability models, stochastic units, the partition function, and learning difficulty. This guide follows the official slides and recording and adds a small self-check that does not depend on the enrolled-course grader.

A Complete Guide to CMU 11-785: 28 Public Lectures, but an Incomplete Assignment Chain

CMU 11-785 Spring 2026 publishes official slides and YouTube recordings for all 28 content lectures, plus extensive bootcamps and recitations. Its HW1–HW4 specifications, starters, and evaluation still depend on Autolab, Piazza, and Kaggle.

aideep-dive

Python Coding Agent M11: Why an Exec Loop Cannot Reproduce the Claude Code Conversation Experience

A Claude Code- or Codex-style TUI depends on long-lived sessions, typed transcripts, and approval at tool boundaries—not a screen full of color.

aideep-dive

Cognee Complete Guide: Turning Documents into Graph Memory for Agents

Cognee is a data-to-memory pipeline: a relational store preserves sources and provenance, a vector store finds semantically similar content, and a graph store represents entity relationships, exposed through remember, recall, improve, and forget.

CS124 Week 1 Introduction and Setup: Turning Language Problems into Computable Components

CS124 Winter 2026 opens by mapping a ten-week path from tokenization and classification to retrieval, speech, networks, and LLMs, while PA0 establishes the Jupyter environment used throughout the quarter.

CS124 Week 10 PageRank and Social Networks: From Anchor Text and Centrality to the Course Wrap-Up

Week 10 models the Web with anchor text, PageRank, and centrality; post-training, multilinguality, and speech belong only to a public final-deck outline labeled 2025, not the 2026 live narration.

CS124 Week 2 Words, Tokens, Edit Distance, and N-grams: Decide What the Model Sees First

Week 2 builds three layers: a token vocabulary with BPE, sequence comparison with dynamic-programming edit distance, and probability approximation with n-grams; PA1 turns regex and BPE into executable work.

CS124 Week 3 Logistic Regression and Text Classification: From Features to Probability and Loss

Week 3 connects text features, sigmoid probabilities, cross-entropy loss, and gradient descent, producing a classifier whose feature contributions remain inspectable.

CS124 Week 4 Information Retrieval: The Indexing and Ranking Layer Beneath RAG

Week 4 builds candidates with an inverted index, ranks them with tf-idf and cosine similarity, and then connects retrieved evidence to generation; PA3 exposes RAG's inspectable retrieval half.

CS124 Week 5 Embeddings and Social NLP: Context Vectors and the Public-Evidence Boundary

Week 5's public materials support the distributional hypothesis, word embeddings, and cosine similarity; the paired Social NLP lecture is unrecorded and restricted, so concrete audit methods are labeled as author extensions.

CS124 Week 6 Neural Networks and LLMs: From Units and Backpropagation to Decoder-Only Models

Week 6 uses public neural-network slides for weighted sums, nonlinearities, loss, and backpropagation, then a public LLM/Transformer deck labeled 2025 for decoder-only architecture without treating it as the 2026 live transcript.

CS124 Week 7 Transformers and Speech Processing: Causal Attention, Generation, and an Unrecorded Lecture

Week 7's public path is PA6a: implement causal self-attention, train a small Shakespeare Transformer, sample text, and compute perplexity; the live speech lecture remains an explicit source gap.

CS124 Week 8 Speech and the PA7/Git Lab: Auditing Information Loss in a TTS-to-STT Pipeline

Week 8 sends text through TTS and back through STT, requiring error classification, formatting-loss analysis, and accent stress tests, while Lab 4 prepares Git collaboration for the team agent project.

CS124 Week 9 Collaborative Filtering and LLM Agents: From Movie Similarity to Search and Memory Tools

Week 9 builds movie recommendations with item-item collaborative filtering, then packages recommendation, web search, databases, and memory as agent tools under API-budget and team constraints.

CS224N Lecture 3: Matrix Calculus and Backpropagation

Lecture 3 decomposes neural-network training into computation graphs, local derivatives, and the chain rule: the forward pass computes a result; backprop accumulates gradients from the output so every parameter knows how to move.

CS224N Lecture 11: Why LLM Benchmarks Expire

Lecture 11 divides evaluation into what to test, how to measure it, and when the result stops being trustworthy. Benchmarks saturate or leak, prompts change scores, and an LLM judge remains a biased model.

CS224N Lecture 9: Prompting, LoRA, and Parameter-Efficient Adaptation

Lecture 9 compares prompting, pruning, LoRA, prompt tuning, and adapters. Each asks the same question: how many parameters must change, and how much task-specific state must be stored, to adapt a large pretrained model?

CS224N Lecture 6: Turn a Final Project into a Testable Question

Lecture 6 completes the Transformer picture with encoders, decoders, and cross-attention, then breaks the final project into formats, assessment, research topics, and data. A viable topic needs one explicit baseline and metric.

CS224N Lecture 1: Four Paradigm Shifts in NLP

Winter 2026 Lecture 1 divides NLP into four eras: early exploration, symbolic systems, statistical machine learning, and deep/self-supervised learning. The point is not the dates but how each era redefined the language problem.

CS224N Lecture 15: Reading Agentic Interpretability Without Public Slides

Lecture 15 is Been Kim's interpretability guest session, but the Winter 2026 site publishes no slides or agenda. This article does not invent lecture content; it maps the five official readings across concept discovery, agentic investigation, and new vocabulary.

CS224N Lecture 17: An Official Reading Map for Multimodality

Lecture 17 is Luke Zettlemoyer's multimodality guest session, but the site publishes no slides or agenda. Its official readings establish three routes: visual reasoning workspaces, early-fusion token models, and text autoregression with image diffusion.

CS224N Lecture 19: How Small Models Can Move Beyond Brute-Force Scaling

The final lecture frames Open Questions in NLP 2026 as smart scaling: prolonged RL, Prismatic synthetic data, RL as pretraining, and open collaboration seek reasoning gains beyond adding parameters.

CS224N Lecture 8: From Instruction Tuning and RLHF to DPO

Lecture 8 explains how instruction tuning, preference data, and RLHF turn a pretrained model into an assistant, then derives DPO from winner–loser pairs. Every step converts human judgment into signal—and imports its biases.

CS224N Lecture 7: Pretraining, Subwords, and In-Context Learning

Lecture 7 decomposes pretraining into scalable data, subword tokenization, three model objectives, and in-context learning. A general self-supervised objective yields reusable representations; downstream signals specify their use.

CS224N Lecture 10: Six Components of RAG and Language Agents

Lecture 10 moves from question answering and RAG into language agents, then decomposes them into reasoning and planning, memory, tools, data, and evaluation. An agent is an inspectable loop between a model and external state.

CS224N Lecture 12: Decoding, DeepSeek-R1, and Reasoning Training

Lecture 12 shows that output policy is not a detail: greedy, beam, and sampling produce different text. It then moves from R1-Zero/R1 into PPO, GRPO, and DAPO, asking when longer reasoning actually helps.

CS224N Lecture 13: Speculative Decoding and Test-Time Scaling

Lecture 13 moves from inference efficiency to inference capability: speculative decoding drafts with a small model and verifies with a large one; on-policy distillation addresses drift; long context and test-time scaling spend inference resources.

CS224N Lecture 4: Language Models, RNNs, and Vanishing Gradients

Lecture 4 defines a language model as a next-word probability distribution, then uses an RNN to compress an arbitrarily long prefix. It also exposes recurrence's central cost: information and gradients travel one time step at a time.

CS224N Lecture 16: Hallucination, Creativity, Work, and Alignment

Lecture 16 divides NLP's social impact into four questions: why models hallucinate, why AI-assisted creativity may homogenize output, how work is reorganized, and why value alignment cannot be reduced to one reward.

CS224N Lecture 18: Material-Gap Record for Tinker and LoRA Without Regret

Lecture 18 is a John Schulman guest session. The official page gives only the title Tinker and LoRA Without Regret, date, and speaker—no slides, agenda, or readings—so this article records confirmed facts and unknowns only.

CS224N Lecture 14: How Tokenization Creates Multilingual Cost Gaps

Lecture 14 moves from word, character/byte, and subword segmentation to BPE failures and cross-lingual fairness. A tokenizer determines sequence length, compute cost, and the units a model sees; it is not neutral preprocessing.

CS224N Lecture 5: From Recurrence to the Transformer

Lecture 5 moves from the long-range and sequential bottlenecks of RNNs to self-attention and the Transformer. It shortens information paths and enables parallel computation, at the price of quadratic attention and separately encoded position.

CS224N Lecture 2: How word2vec Turns Meaning into Vectors

Lecture 2 moves from word2vec's prediction task, objective, and gradients to count-based vectors and evaluation. Meaning becomes a high-dimensional position learned from context, not a label retrieved from a dictionary.

Stanford CS224V Lecture 10: How SPINACH Explores Wikidata and Builds SPARQL

SPINACH does not guess complete SPARQL in one shot. It searches entities and properties, inspects Wikidata entries and examples, executes small queries, and composes a final query under explicit action and stopping rules.

Stanford CS224V Lecture 12: CHURRO Makes Multilingual Historical Documents Searchable

CHURRO represents full-page text, layout, and metadata in HDML, unifies multilingual historical data for a page-level VLM, and connects extraction to HistoryGenie for searchable, conversational archives.

Stanford CS224V Lecture 14: Scaling Language Models When Data Is the Bottleneck

The final lecture is not a complete LLM-training tutorial. It studies data efficiency under fixed data and abundant compute, revisiting epochs, batches, ensembles, self-training, and conditions for synthetic continued pretraining.

Stanford CS224V Lecture 5: WikiChat's Seven-Stage Defense Against Hallucination

The [WikiChat paper](https://aclanthology.org/2023.findings-emnlp.157/) expands RAG into query formulation, retrieval, filtering, generation, claim extraction, renewed retrieval and verification, and removal of unsupported content—and evaluates retrieval separately from factuality.

Stanford CS224V Lecture 1: Turning Hallucinating LLMs into Dependable Assistants

Fall 2025 opens with computational thinking: reliability comes from decomposing retrieval, formal representation, verification, and generation into testable algorithms, not from one heroic prompt.

Stanford CS224V Lecture 2: STORM, Co-STORM, and Knowledge Curation

STORM uses perspective-guided questions, simulated interviews, and outlines to broaden research; Co-STORM keeps a person in the loop so discovering unknown questions and co-editing become part of the system.

Stanford CS224V Lecture 8: SLIDERS Turns Long-Document Sets into Queryable Tables

SLIDERS induces a question-specific schema, applies semantic chunking and contextualized extraction, reconciles duplicate rows, and answers with SUQL instead of feeding every long document directly to one model.

Stanford CS224V Lecture 13: ReactGenie Gives Voice and Native GUIs Shared State

ReactGenie annotates React components to expose data, actions, and views, parses composite voice commands into a DSL, and renders native graphical output against shared UI context.

Stanford CS224V Lecture 11: Translate Trial Criteria into SMT Instead of Asking an LLM to Decide

The lecture parses patient records and trial criteria into SMT, retrieves candidates through a weaker propositional projection, and runs a solver on the reduced set. Reasoning is inspectable, but NL-to-SMT remains the main error boundary.

Stanford CS224V Lecture 9: Why Automated Qualitative Coding Still Needs Expert Review

Automated qualitative coding defines event types and arguments in a codebook, then separates document classification, structured extraction, and entity linking. Constrained JSON fixes form, not expert judgment.

Stanford CS224V Lecture 6: Why Database Agents Begin with Semantic Parsing

Reliable database agents map language to executable queries, resolve schemas and enumerated values, and evaluate execution separately from answer generation; hybrid questions additionally require explicit source routing.

Stanford CS224V Lecture 7: SUQL Unifies SQL and Free-Text Retrieval

SUQL adds answer and summary functions over text to SQL. A semantic parser emits one hybrid query, while an optimizing compiler applies predicate pushdown, top-k pruning, and lazy evaluation.

Stanford CS224V Lecture 4: Task-Agent Evaluation Beyond Human-Like Answers

CS224V splits task-agent evaluation into state updates and complete interaction: isolate the semantic parser, then test task completion, grounded queries, and valid actions with real users.

Stanford CS224V Lecture 3: Building Task-Oriented Agents with Genie Worksheets

Genie Worksheets declare task capability as a form-like specification. A contextual semantic parser updates formal dialogue state while the runtime controls queries, actions, and responses.

CS336 Lecture 3: Transformers Have Many Variants but Few Stable Defaults

Lecture 3 does not turn its survey of modern LLMs into a single best recipe. It finds a conservative consensus—pre-norm, RMSNorm, no biases, SwiGLU, and RoPE—plus a small set of deviations justified by inference cost or stability.

CS336 Lecture 4: Attention Has Alternatives, and MoE Does Not Scale for Free

Lecture 4 studies two kinds of sparsity: linear/recurrent attention reduces sequence-length cost, while MoE activates only part of a model for each token. Both turn saved FLOPs into routing, balancing, communication, and kernel problems.

CS336 Lecture 14: Filtering, Deduplication, and Mixing Turn Raw Web Data into Training Data

Lecture 14 moves raw documents through language, quality, and safety filtering; exact and near deduplication; and source mixing. Each stage reshapes model behavior, while synthetic instruction and agent trajectories extend the pipeline into executable environments.

CS336 Lecture 13: Data Does Not Fall from the Sky, and Every Source Has Access and License Costs

Lecture 13 traces training sources through Common Crawl, Wikipedia, GitHub, arXiv, books, and open datasets. Technically accessible is not the same as licensed, and raw data is not training data; provenance must precede cleaning and mixing.

CS336 Lecture 12: There Is No Single True LLM Score, Only Different Games

Lecture 12 moves from perplexity to exams, chat preferences, agents, reasoning, and safety. Every benchmark changes the capability definition, scaffold, judge, and contamination risk, so evaluation must first say whether it compares a method, model, or complete system.

CS336 Lecture 5: GPUs Win by Moving Data Less, Not by Making Each Thread Fast

Lecture 5 explains GPUs through SMs, warps, and the memory hierarchy, then unifies common optimization under low precision, fusion, recomputation, coalescing, and tiling. FlashAttention combines those principles for attention.

CS336 Lecture 10: LLM Inference Is About Reading Weights and KV Cache Less Often

Lecture 10 separates prefill from decode: prefill parallelizes and is often compute-bound, while decode is sequential and commonly bandwidth-bound. GQA/MLA, quantization, speculative decoding, continuous batching, and PagedAttention reshape that cost.

CS336 Lecture 6: Benchmark and Profile Before Writing a Triton Kernel

Lecture 6 turns GPU principles into kernels: benchmark scaling across shapes, profile actual calls and time, then implement GeLU, softmax, reductions, and tiled matrix multiplication in Triton. Speed begins with measuring correctly.

CS336 Lecture 17: Multimodal Models Turn Images into Tokens, Then Reconcile Semantics with Detail

Lecture 17 organizes CLIP/SigLIP, LLaVA, Qwen-VL, and Chameleon into three paths: contrastive encoders learn semantics, vision-encoder/projector/LM stacks provide understanding, and discrete image tokens enable generation. Resolution, token budgets, and modality balance constrain them all.

CS336 Lecture 1: From Bytes to a Tokenizer—and What Deserves to Scale

CS336's first lecture does not treat building a language model from scratch as reenacting every old technique. It separates mechanics, mindset, and intuitions, then uses BPE to show how raw bytes become trainable tokens.

CS336 Lecture 7: Build Data, Tensor, and Pipeline Parallelism from Collectives

Lecture 7 starts below FSDP APIs, building a communication language from broadcast, all-reduce, all-gather, reduce-scatter, and all-to-all before assembling data, tensor, and pipeline parallelism.

CS336 Lecture 8: Align ZeRO, FSDP, and 3D Parallelism with Hardware Topology

Lecture 8 moves from parallel primitives to system design: ZeRO progressively shards optimizer state, gradients, and parameters; TP, PP, SP, and EP split width, depth, sequence, and experts. Their composition must follow topology and dynamic activation memory.

CS336 Lecture 2: Count FLOPs and Memory Before Asking Whether a Model Fits

Lecture 2 reduces model training to tensors, FLOPs, bytes, and time: use einops to track dimensions, arithmetic intensity and roofline analysis to identify bottlenecks, then trade compute for memory with gradient accumulation and activation checkpointing.

CS336 Lecture 16: RLVR Scales Reasoning with Verifiable Rewards, but GRPO Is Not Free PPO

Lecture 16 moves from PPO to GRPO and RLVR. Math, code, and environment outcomes provide scalable rewards and avoid some preference-model overoptimization, but group-normalized advantages introduce difficulty and length bias while rollout infrastructure becomes the dominant cost.

CS336 Lecture 9: Scaling Laws Are Extrapolation Tools, Not Crystal Balls

Lecture 9 begins with log-log linear relationships between data and error, then uses scaling laws to compare architectures, optimizers, batches, and model-data allocations. The Chinchilla dispute shows how fitting methods, observed ranges, and deployment objectives change the answer.

CS336 Lecture 11: Scaling Laws in Practice Must Scale Learning Rate and Batch Too

Lecture 11 reads public recipes from MiniCPM, DeepSeek, Qwen, and Llama 3: hold most architectural ratios fixed, sweep learning rate and batch at small scale, then choose model/data allocation with IsoFLOPs. μP helps, but normalization, optimizers, and weight decay can break transfer.

CS336 Lecture 15: SFT Teaches Imitation; RLHF Begins Direct Preference Optimization

Lecture 15 divides post-training into imitation and optimization. SFT extracts pretrained capabilities from instruction-response data; RLHF uses pairwise feedback to bridge demonstrations and preferences. PPO and DPO both inherit data bias, reward overoptimization, and mode collapse.

aideep-dive

Daytona Agent Sandbox: A Forkable Computer for Every Agent

Daytona treats a sandbox as a long-lived computer that can start, pause, snapshot, and fork. It raised a $24 million Series A in 2026, while a Laude Institute case study reports 37,000 sandboxes in one week. It fits parallel evaluations and coding agents, but its core open-source repository is no longer maintained.

aideep-dive

Deepgram Voice Agent API: From Streaming STT and Turn Detection to TTS

Deepgram combines streaming STT, LLM orchestration, turn detection, barge-in, and streaming TTS over one WebSocket while preserving paths for standalone speech models and bring-your-own LLM or TTS.

aideep-dive

Dify as a Low-Code Agent Platform: From a Working Workflow to a Published AI App

Dify puts models, Knowledge, visual Workflows, Agents, Plugins, and application APIs in one workspace; this guide builds a minimal Workflow that can be tested, published, and called through the API, then explains when an Agent is actually warranted.

aideep-dive

DSPy: Compiling AI Programs with Signatures, Metrics, and Optimizers

DSPy replaces handwritten prompt strings with task Signatures, execution Modules, and Optimizers that compile better instructions and examples against a dataset and metric.

aideep-dive

E2B Agent Sandbox: Put Model-Generated Code in a Resumable microVM

E2B combines Templates, Firecracker microVMs, and process, file, and network APIs into an agent execution layer. Its real selection advantage is preserving memory and processes across pause and resume, not merely providing another code interpreter.

aideep-dive

ElevenLabs ElevenAgents: The Lifecycle from Realtime Speech to Phone Agents

ElevenLabs has expanded from a TTS vendor into the ElevenAgents platform: Scribe Realtime listens, Flash speaks, and the platform connects the LLM, turn-taking, tools, and telephony. The key choice is whether you need a voice model or the whole agent control plane.

aideep-dive

Fireworks AI: From Serverless APIs to Custom Model Deployments

Fireworks AI puts open-weight model evaluation, dedicated GPU deployments, and LoRA customization behind one API surface. Serverless fits low-volume starts, On-demand fits sustained traffic and custom models, while reserved capacity adds enterprise capacity guarantees.

aideep-dive

Flowise Deep Dive: From Assistant, Chatflow, and Agentflow to an EOL Migration Decision

Flowise uses Assistant, Chatflow, and Agentflow to cover simple assistants, single-agent systems, and multi-agent orchestration; however, its repository was archived in August 2026 and official EOL is scheduled for August 31, so new projects should not adopt it without a maintained fork and migration plan.

aideep-dive

Galileo Deep Dive: Experiments, Evaluators, and the Agent Observability Loop

Galileo connects dataset experiments, LLM/code/Luna evaluators, production traces, and runtime guardrails. It fits enterprises that need observability plus intervention, but the old Protect surface is deprecated and evaluator models do not replace human calibration or application security.

The H2 2026 Harness War: Eight Frameworks Rewriting, Three Model Makers Entering, 110+ CLIs — How to Make Sense of It

In August 2026, it's not just five frameworks moving. Beyond OMP 2, Pi v2, Opencode 2, dsh, and Claude Code, three model makers — Google (Antigravity CLI), Meta (Muse Code), and xAI (Grok Build) — are building coding agents directly. Add Amp, Cline 2.0, and the Codex CLI Rust rewrite, and eight-plus frameworks are undergoing architecture-level changes simultaneously. Factor in 110+ total CLI tools, and H2 2026 is a divergence period for harness methodology. This article analyzes four architectural approaches, one shared direction, and one emerging trust crisis.

aideep-dive

Haystack Deep Dive: Testable RAG with Components and Pipelines

Haystack turns indexing, retrieval, generation, and evaluation into replaceable Components connected by directed-multigraph Pipelines; it fits Python teams that want RAG flows to be tested, versioned, and deployed as code.

aideep-dive

Helicone Deep Dive: LLM Gateway, Request Tracing, and Cost Analytics

Helicone is an open-source LLM gateway and observability platform: requests sent through its compatible endpoint automatically capture model, latency, tokens, cost, and custom properties, while managed credits or BYOK enable routing and fallbacks.

aideep-dive

Hugging Face Is More Than a Model Download Site: Hub, Datasets, Spaces, and Inference

Hugging Face Hub is a collaboration layer for versioned models, datasets, and applications. Datasets handles data, Spaces runs demos, while Inference Providers and Endpoints provide managed inference.

aideep-dive

Hyperbrowser Deep Dive: Browser-as-a-Service Infrastructure for Agents

Hyperbrowser packages Chrome sessions, proxies, stealth, profiles, and recordings behind managed Playwright and Puppeteer APIs. It fits agents that need to scale real-browser work quickly, while profile credentials, anti-bot compliance, and proxy bandwidth costs remain application responsibilities.

Jina Reader Guide: Turn Web Pages into Agent-Readable Markdown

Jina Reader turns a known URL into LLM-friendly Markdown; production use still requires explicit rendering, scope, token-budget, validation, and fallback decisions.

aideep-dive

LanceDB Deep Dive: Embedding Vector Search in Arrow Data Workflows

LanceDB stores vectors, metadata, and multimodal source data in the Lance columnar format. Its OSS edition embeds in Python, TypeScript, or Rust processes; distributed Enterprise becomes relevant when the data or service outgrows one machine.

aideep-dive

LangChain v1 Agents: create_agent, Middleware, and the LangGraph Runtime

LangChain v1 provides a high-level agent loop through create_agent, runs it on LangGraph, and treats tools, structured output, and middleware as its extension boundaries.

aideep-dive

LangSmith Deep Dive: From Agent Traces to Offline and Online Evaluation

LangSmith structures LLM applications as projects, traces, runs, and threads, then uses datasets, evaluators, and experiments to turn production failures into offline regression tests. It observes any LLM application and does not require LangChain.

aideep-dive

Letta and MemGPT Complete Guide: Memory Inside a Stateful Agent Runtime

Letta extends MemGPT's operating-system analogy but is not a standalone memory API. The runtime persists agent state, editable in-context blocks, conversation history, and external archival memory, while the model can actively curate memory through tools.

aideep-dive

LiteLLM: From a Python SDK to a Self-Hosted AI Gateway

LiteLLM is not a model provider. It is a Python SDK and self-hosted proxy that normalizes 100+ LLM APIs, then centralizes routing, fallbacks, virtual keys, budgets, and observability at the gateway layer.

aideep-dive

LiveKit Voice Agents: From WebRTC Rooms to Interruptible Voice Pipelines

LiveKit models a voice agent as a server participant in a realtime media room, with AgentSession orchestrating STT, turn detection, LLM, TTS, and interruption. It raised a $100 million Series C at a $1 billion valuation in 2026. It fits products needing WebRTC, multiple client platforms, telephony, and swappable models, but self-hosting the media server does not self-host the entire AI pipeline.

aideep-dive

Mastra: Agents, Workflows, Memory, and Evals in TypeScript

Mastra is a TypeScript agent framework that combines agents, typed workflows, memory, MCP, tracing, and scorers in one Node.js development environment.

Mem0 Complete Guide: Controlled Long-Term Memory for AI Agents

Mem0 sits between an agent and storage: it extracts durable facts from interactions, scopes them by user, agent, or run, and searches them before a later generation. Its appeal is a small API; its risks are extraction errors, stale memories, and authorization boundaries.

aideep-dive

Milvus Vector Database Deep Dive: Segments, Indexes, and Distributed Operations

Milvus separates real-time ingestion, historical queries, index building, and persistence into independently scalable components. It fits large, continuously updated retrieval services, but smaller projects often pay too much operational complexity for that architecture.

MIT 6.S191 Lecture 1: The Minimal Structure of Deep Learning

Lecture 1 of the 2026 course builds the vocabulary shared by the rest of the course: perceptrons, forward propagation, loss, and gradient descent.

MIT 6.S191 Lecture 2: Sequence Modeling: From RNNs to Attention

Lecture 2 of the 2026 course addresses data where order changes meaning—text, audio, and time series—and connects directly to music generation in Lab 1.

MIT 6.S191 Lecture 3: Computer Vision: How Convolution Preserves Spatial Structure

Lecture 3 of the 2026 course moves from image tensors, convolution, and pooling to recognition systems, preparing for MNIST and face detection in Lab 2.

MIT 6.S191 Lecture 4: Generative Modeling: From Latent Spaces to Diffusion

Lecture 4 of the 2026 course separates generative from discriminative tasks, organizes VAE, GAN, and diffusion objectives, and leads into Lab 2’s DB-VAE.

MIT 6.S191 Lecture 5: Reinforcement Learning: Learning from Return Instead of Labels

Lecture 5 of the 2026 course connects agent, environment, state, action, reward, and policy into an interaction loop, introducing credit assignment and exploration.

MIT 6.S191 Lecture 6: New Frontiers: Choosing the Problem Beyond the Model

Lecture 6 of the 2026 course places deep learning in emerging applications and real constraints, emphasizing data, outputs, evaluation, and failure conditions.

MIT 6.S191 Lecture 7: The Three Laws of AI: Safety Through Observability and Evaluation

Lecture 7 of the 2026 course starts from Asimov’s literary laws and examines modern safety protocols through traces, test data, and continuous evaluation.

MIT 6.S191 Lecture 8: AI for Science: Putting Domain Structure into Learning

Lecture 8 of the 2026 course uses the scientific-discovery loop to show how simulators, AI emulators, and experiments cooperate instead of reducing science to generic prediction.

MIT 6.S191 Lecture 9: Massively Parallel Training: Memory and Communication Set the Boundary

Lecture 9 of the 2026 course starts with GPU memory pressure and moves through checkpointing, offloading, ZeRO, FSDP, and multiple forms of parallelism.

MIT 6.S191 Lab 1: Generate Music with PyTorch and an LSTM

In the 2026 lab, students cover tensors, autograd, and modules before turning ABC notation into character sequences for LSTM music generation.

MIT 6.S191 Lab 2: From MNIST to Facial Debiasing with a DB-VAE

In the 2026 lab, part 1 classifies MNIST with dense and convolutional networks; Part 2 learns a facial latent distribution with a DB-VAE and changes training sampling.

MIT 6.S191 Lab 3: LoRA Fine-Tuning and LLM-as-a-Judge Evaluation

In the 2026 lab, students build chat templates and generation with LFM2-1.2B, adapt style through LoRA, and combine OpenRouter with Opik for a judge workflow.

aideep-dive

n8n Deep Dive: From Triggers and AI Agents to Human Review and Operations

n8n is automation-first: a webhook, schedule, or application event starts a workflow, then an AI Agent may choose tools inside it; production still requires deliberate memory, approvals, credentials, execution data, and scaling architecture.

aideep-dive

OpenRouter: One API Key for Multi-Model, Multi-Provider LLM Routing

OpenRouter exposes many models and inference endpoints through an OpenAI-compatible API, with provider ordering, failover, BYOK, and zero-data-retention controls in one routing policy.

OpenViking: Agent Memory as a Virtual Filesystem

Volcano Engine's open-source OpenViking stores agent memory, knowledge, and skills as a viking:// virtual filesystem — browsable with ls, tree, and find. Three-tier loading (L0/L1/L2) averages just 550 tokens per retrieval, boosting LoCoMo memory accuracy from 24–57% to 80–83%.

aideep-dive

Parallel Web Systems: Search, Extraction, and Deep Research for Agents

Parallel Web Systems separates Search, Extract, and Task APIs into web-access layers with different latency and cost profiles, while Basis maps citations, excerpts, and confidence to output fields.

aideep-dive

Patronus AI Deep Dive: From Evaluators and Experiments to Production Monitoring

Patronus AI treats evaluators as reusable scoring units, then applies them to offline experiments and production traces. It suits teams that want managed hallucination, safety, and multimodal evaluators, but judge scores cannot replace human labels, deterministic tests, or real security validation.

aideep-dive

pgvector Deep Dive: Bringing Vector Search Back into PostgreSQL

pgvector is a PostgreSQL extension, not a standalone vector database. It adds exact and approximate vector search to the same data model, transactions, and operations stack, while leaving index tuning and horizontal scaling as PostgreSQL concerns.

aideep-dive

Portkey: Put LLM Routing, Observability, and Governance Behind One AI Gateway

Portkey sits between applications and model providers: one OpenAI-compatible endpoint adds routing, fallbacks, request logs, budgets, and guardrails, with an open-source gateway available for self-hosting.

Securing Private-Corpus Queries: ACLs, Deletion Propagation, and Freshness

Authorization must take effect before candidate generation, while ACLs, deletion events, and source versions must propagate to every derived index; freshness needs measurable event-time SLOs too.

Private Corpus Search Boundaries: Decide Where Data May Go First

The first private-corpus decision is not which vector database to buy. Define data classes, trust zones, policy enforcement points, and freshness SLAs so every index, model, and observability system receives only the minimum data it is allowed to process.

Private-Corpus Retrieval Eval: Turning a Traditional Chinese Query Set into a Reproducible Benchmark

The repository has a 20-query Traditional Chinese/English golden dataset, but no document-level qrels, retrieval runs, raw latency data, or executable benchmark script. Reporting Recall@k, MRR, or nDCG as measured results would therefore be dishonest; this article defines the contract needed to run them reproducibly.

From Source to Index: Sync and Incremental Updates for Private Corpora

Private-corpus sync is not periodic refetching. It requires stable canonical IDs, source versions plus checksums for change detection, idempotent upserts, and tombstones that propagate deletion through every index.

aideep-dive

Promptfoo Deep Dive: Local-First LLM Evaluation and Red Teaming

Promptfoo combines prompts, providers, test cases, and assertions in YAML to produce repeatable local and CI evaluation matrices, with red teaming against the same targets. It lowers the testing barrier but does not remove output variance, LLM-judge bias, or hosted data-flow concerns.

aideep-dive

Pydantic AI: Building Python Agents with Types, Dependencies, and Validation

Pydantic AI models an agent as Agent[Deps, Output]: dependencies, tool inputs, and final outputs are typed, and model results must pass Pydantic validation.

aideep-dive

R2R Deep Dive: Ingestion, Hybrid Search, and RAG behind an API

R2R packages document ingestion, hybrid search, knowledge graphs, RAG, Agents, and access controls behind a REST API; it fits teams that already own their product frontend and backend and need a retrieval service.

aideep-dive

Choosing a RAG Framework: LlamaIndex, Haystack, RAGFlow, Dify, and R2R Operate at Different Layers

LlamaIndex and Haystack are code-first frameworks; RAGFlow and Dify are managed application platforms; R2R packages retrieval as an API service. Choose how much control your team needs over ingestion, retrieval, and operations before choosing a tool.

aideep-dive

RAGFlow Deep Dive: From Document Parsing and Chunk Review to Cited Answers

RAGFlow puts document parsing, human chunk review, retrieval tests, chat, and citations in one platform; it fits layout-heavy PDFs and tables, but carries more deployment weight and platform state than a Python library.

aideep-dive

Runloop: Devbox Infrastructure Built for Coding Agents

Runloop combines isolated microVMs, reproducible images, disk branching, credential proxies, and evals in one coding-agent platform; an official case study reports more than 10,000 concurrent Devboxes in one workload.

aideep-dive

Sail Research: Trading Latency for Cost in Long-Horizon Agent Inference

Sail Research lets each inference request declare a completion window, scheduling patient background agents on cheaper capacity, while Sailboxes provide persistent long-running execution environments.

From Search Results to Reliable Citations: URL Deduplication, Source Tiers, and Claim-Source Mapping

Reliable citation is not appending URLs to an answer. Separate URLs, content copies, and source independence, then connect atomic claims to quote spans and snapshots through a rerunnable claim-source matrix.

aiguide

SerpAPI Complete Guide: Multiple Engines, Structured SERPs, and Async Queries

SerpAPI primarily manages search-results-page retrieval and parsing: select an engine, receive a structured SERP, then handle location, pagination, asynchronous polling, and validation in your application.

aiguide

Serper Search API Guide: Turn Google Results into Agent-Ready JSON

Serper is a third-party Google SERP API: one POST request returns structured JSON such as organic, knowledgeGraph, and peopleAlsoAsk, but production code still needs optional-field validation, URL checks, retries, and source verification.

aideep-dive

Slack Code: Multiplayer AI Coding and the Agent Control Plane Landscape

Slack Code moves AI coding agents from individual terminals into shared Slack channels where teams can see diffs, previews, and plans in real time. But it solves management's visibility anxiety, not engineers' productivity bottleneck — the real battle is over who becomes the agent control plane.

CS221 Lecture 1: Overview: Defining Intelligence Under Resource Constraints

Lecture 1 of Stanford CS221 Autumn 2025 develops operational representations and algorithmic intuition through Overview: Defining Intelligence Under Resource Constraints.

CS221 Lecture 2: Learning I: From Computation Graphs to Linear Regression

Lecture 2 of Stanford CS221 Autumn 2025 develops operational representations and algorithmic intuition through Learning I: From Computation Graphs to Linear Regression.

CS221 Lecture 3: Learning II: Linear Classification, Features, and Cross-Entropy

Lecture 3 of Stanford CS221 Autumn 2025 develops operational representations and algorithmic intuition through Learning II: Linear Classification, Features, and Cross-Entropy.

CS221 Lecture 4: Learning III: Deep Networks as Composable Computation Graphs

Lecture 4 of Stanford CS221 Autumn 2025 develops operational representations and algorithmic intuition through Learning III: Deep Networks as Composable Computation Graphs.

CS221 Lecture 5: Search I: Define the State Before Choosing the Algorithm

Lecture 5 models search with states, actions, successors, and costs, then uses acyclic dynamic programming to show that an efficient algorithm still solves the wrong problem when state omits information needed by the future.

CS221 Lecture 6: Search II: Priorities in UCS and A*

Lecture 6 of Stanford CS221 Autumn 2025 follows the official material on Search II: Priorities in UCS and A* and makes its assumptions and limits explicit.

CS221 Lecture 7: MDPs I: Putting Uncertainty into State Transitions

Lecture 7 of Stanford CS221 Autumn 2025 follows the official material on MDPs I: Putting Uncertainty into State Transitions and makes its assumptions and limits explicit.

CS221 Lecture 8: MDPs II: Learning Q-Values Without a Transition Model

Lecture 8 of Stanford CS221 Autumn 2025 follows the official material on MDPs II: Learning Q-Values Without a Transition Model and makes its assumptions and limits explicit.

CS221 Lecture 9: MDPs III: Differentiating Expected Return Directly

Lecture 9 moves from tabular RL to function approximation, derives REINFORCE with the log-derivative identity, and connects the derivation to the executable PyTorch implementation.

CS221 Lecture 10: Games I: From Expectimax to Minimax

Lecture 10 extends single-agent search into adversarial game trees: expectimax averages chance outcomes, minimax takes the opponent's worst case, and alpha-beta removes irrelevant branches without changing the answer.

CS221 Lecture 11: Games II: TD Learning, Simultaneous Games, and Nash Equilibria

Lecture 11 first learns game values from experience with temporal-difference updates, then moves from sequential play to simultaneous games described by mixed strategies, minimax guarantees, and Nash equilibria.

CS221 Lecture 12: Bayesian Networks I: From Joint Distributions to Factorization

Lecture 12 builds a joint distribution from random variables and factors, then uses Bayesian-network factorization to express conditional independence and make conditioning and marginalization executable.

CS221 Lecture 13: Bayesian Networks II: Gibbs Sampling and the Markov Blanket

Lecture 13 replaces costly exact inference with Gibbs sampling: resample one variable at a time from a conditional determined by its Markov blanket, then approximate query probabilities with sample frequencies.

CS221 Lecture 14: Bayesian Networks III: From Counts and Smoothing to EM

Lecture 14 moves from maximum-likelihood counts and Laplace smoothing with complete data to EM, which alternates posterior responsibilities for latent variables with parameter updates.

CS221 Lecture 15: Logic I: Models, Entailment, and SAT

Lecture 15 separates propositional syntax from semantics: model checking defines entailment through satisfying assignments, SAT finds witnesses, and inference rules must be judged for both soundness and completeness.

CS221 Lecture 16: Logic II: Quantifiers Beyond Individual Propositions

Lecture 16 compresses knowledge across objects with predicates, quantifiers, and functions, then derives conclusions through substitution, unification, and definite-clause forward inference while exposing termination and completeness limits.

CS221 Lecture 17: Language Models: From Next-Token Prediction to Generation

Lecture 17 defines a language model as a chain-rule factorization of sequence probability, compares n-gram and neural conditional models, and shows how sampling, temperature, and evaluation shape generation.

CS221 Lecture 18: AI & Society: Benefits, Misuse, Accidents, and Institutions

Lecture 18 classifies AI's social effects as benefits, misuse, accidents, and structural harms, then connects fairness audits, research ethics, copyright, and platform terms to accountable institutional choices.

CS221 Lecture 19: AI Supply Chains: Resources, Labor, and Markets Behind Models

Lecture 19 uses the Economics of AI deck to connect compute, data, distribution, and organizational complements to GDP, labor, and ideas-driven growth.

CS221 Lecture 20: Fireside Chat, Conclusion: Turning Twenty Lectures into Modeling Choices

Lecture 20 is Percy Liang's fireside chat on career and research, CS221 and Stanford, and AI's future, with every attribution tied to the official video and editorial synthesis kept separate from auto-caption uncertainty.

Stanford CS224W Lecture 1: Introduction: Why Relational Data Needs Graph Machine Learning

A slide-grounded reconstruction of Fall 2025 Lecture 1, covering Course map and tools, A common language for graph data, Hand-designed features and representation learning while documenting the classroom material unavailable to self-learners.

Stanford CS224W Lecture 2: Node Embeddings: From Random Walks to node2vec

A slide-grounded reconstruction of Fall 2025 Lecture 2, covering Encoder-decoder view, Similarity and the objective, Random walks while documenting the classroom material unavailable to self-learners.

Stanford CS224W Lecture 3: Graph Neural Networks: A First Complete Message-Passing Model

A slide-grounded reconstruction of Fall 2025 Lecture 3, covering From fixed embeddings to deep encoders, The message-passing framework, Aggregation and update while documenting the classroom material unavailable to self-learners.

Stanford CS224W Lecture 4: A General Perspective on GNNs: Turning a Model into Design Components

A slide-grounded reconstruction of Fall 2025 Lecture 4, covering The GNN design space, Message, aggregation, and update, GraphSAGE while documenting the classroom material unavailable to self-learners.

Stanford CS224W Lecture 5: GNN Augmentation and Training: Co-designing Data, Tasks, and Models

A slide-grounded reconstruction of Fall 2025 Lecture 5, covering Graph-data augmentation, Feature and structural augmentation, Supervision and loss while documenting the classroom material unavailable to self-learners.

Stanford CS224W Lecture 6: Theory of GNNs: The WL Test, GIN, and Expressive Limits

A Fall 2025 slide-grounded reconstruction of Lecture 6, covering What distinguishability means, The Weisfeiler–Lehman test, An upper bound for message passing while documenting unavailable classroom material.

Stanford CS224W Lecture 7: Designing Powerful Graph Encoders: Structural and Positional Awareness

A Fall 2025 slide-grounded reconstruction of Lecture 7, covering The perfect-GNN thought experiment, Three levels of standard-GNN failure, Identity-aware encoding while documenting unavailable classroom material.

Stanford CS224W Lecture 8: Graph Transformers: Connecting Attention to Graph Structure

A Fall 2025 slide-grounded reconstruction of Lecture 8, covering Self-attention and message passing, The scope of graph attention, Positional and structural encodings while documenting unavailable classroom material.

Stanford CS224W Lecture 9: Heterogenous Graphs: Adding Node and Relation Types to Message Passing

A Fall 2025 slide-grounded reconstruction of Lecture 9, covering Heterogeneous graph schemas, Relation-specific messages, R-GCN while documenting unavailable classroom material.

Stanford CS224W Lecture 10: Knowledge Graphs: Modeling Relations with TransE, ComplEx, and RotatE

A Fall 2025 slide-grounded reconstruction of Lecture 10, covering Knowledge graphs and completion, Triple scoring, TransE and relation patterns while documenting unavailable classroom material.

Stanford CS224W Lecture 11: GNNs for Recommender Systems: From Collaborative Filtering to LightGCN

A Fall 2025 slide-grounded reconstruction of Lecture 11, covering Graph formulation of recommendation, The matrix-factorization baseline, Message passing in NGCF while documenting the public-material boundary.

Stanford CS224W Lecture 12: Relational Deep Learning: Turning Databases Directly into Prediction Graphs

A Fall 2025 slide-grounded reconstruction of Lecture 12, covering Limits of the tabular pipeline, Mapping relational databases to graphs, Temporal entity graphs while documenting the public-material boundary.

Stanford CS224W Lecture 13: Advanced Architectures in RDL: RelGNN and the Relational Graph Transformer

A Fall 2025 slide-grounded reconstruction of Lecture 13, covering The multi-relational bottleneck, RelGNN composite message passing, Relation-specific aggregation while documenting the public-material boundary.

Stanford CS224W Lecture 14: Advanced Topics in GNNs: In-Context Learning and Uncertainty on Graphs

A Fall 2025 slide-grounded reconstruction of Lecture 14, covering The goal of relational foundation models, Zero-shot relational transfer, PRODIGY's prompt graph while documenting the public-material boundary.

Stanford CS224W Lecture 15: Foundation Models for Knowledge Graphs: New Entities, New Relations, and Double Equivariance

A Fall 2025 slide-grounded reconstruction of Lecture 15, covering Limits of transductive KG embeddings, Entity-inductive link prediction, The relation graph while documenting the public-material boundary.

Stanford CS224W Lecture 16: LLM + GNN: Letting Language Models Read Graphs and Graph Models Read Text

A Fall 2025 slide-grounded reconstruction of Lecture 16, covering Complementary gaps in LLMs and GNNs, Text-attributed graphs, The LLM as predictor or encoder while documenting the public-material boundary.

Stanford CS224W Lecture 17: Agents + Graphs: Retrieval, Planning, and Action in Structured Worlds

A Fall 2025 slide-grounded reconstruction of Lecture 17, covering From graph QA to agents, Multimodal retrieval in STaRK, Tool use and traversal while documenting the public-material boundary.

Stanford CS224W Lecture 18: Deep Generative Models for Graphs: GraphRNN and Goal-Directed Molecular Generation

A Fall 2025 slide-grounded reconstruction of Lecture 18, covering The graph-generation problem and representation, Evaluating generation quality, GraphRNN's autoregressive factorization while documenting the public-material boundary.

Stanford CS224W Lecture 19: Ranking 315K GNN Designs with Anchor Models

The Fall 2025 conclusion studies roughly 315K GNN designs across 32 tasks: run a small set of anchor models, derive task similarity from rankings, and transfer the best designs from similar tasks.

Linear Regression: From LMS to Locally Weighted Regression

Linear regression is more than a best-fit line: Chapter 1 connects squared loss to gradient descent, normal equations, maximum likelihood, and locally weighted regression.

Classification and Logistic Regression: Decision Boundaries and Newton's Method

Chapter 2 derives logistic loss from a sigmoid probability model, then contrasts it with the perceptron and extends it through softmax and Newton's method.

Generalized Linear Models: Unifying Regression and Classification

Chapter 3 uses exponential families, natural parameters, and link functions to place least squares and logistic regression inside one modeling template.

Generative Learning Algorithms: GDA, Naive Bayes, and Smoothing

Chapter 4 models p(x|y) and p(y), using GDA, Naive Bayes, and Laplace smoothing to expose both the power and price of generative classification.

Kernel Methods: Nonlinear Learning Without Explicit Features

Chapter 5 replaces high-dimensional feature inner products with kernels, letting inner-product-based linear algorithms learn nonlinear functions without constructing the features.

Support Vector Machines: Margins, Duality, and SMO

Chapter 6 formalizes classification confidence as geometric margin, then builds an implementable SVM through Lagrange duality, kernels, and SMO.

Deep Learning: Modules, Backpropagation, and Vectorization

Chapter 7 decomposes neural networks into composable modules and uses backpropagation and vectorization to explain how deep models can be trained efficiently.

Generalization: Bias–Variance, Double Descent, and Sample Complexity

Chapter 8 decomposes test MSE into irreducible noise, squared bias, and variance, then uses uniform convergence and VC dimension to explain when training performance transfers to new data. Double descent shows why parameter count is not a universal measure of complexity.

Regularization and Model Selection: Explicit, Implicit, and Cross-Validated

Chapter 9 presents three controls on generalization: explicit complexity penalties, optimizer-induced implicit regularization, and model selection on data excluded from training. MAP estimation then connects a Gaussian prior to an L2 penalty.

Clustering and k-Means: A First Alternating-Optimization Algorithm

Chapter 10 introduces unsupervised learning through k-means: alternating updates make distortion non-increasing and numerically convergent, but do not guarantee a global optimum.

EM Algorithms: From Gaussian Mixtures to VAEs

Chapter 11 starts from soft assignments in Gaussian mixtures, uses Jensen's inequality to construct the ELBO, interprets EM as alternating maximization over a variational distribution and model parameters, and extends the idea to VAEs through approximate posteriors and reparameterization.

Principal Components Analysis: Projection, Reconstruction, and Reduction

Chapter 12 formulates PCA as geometric optimization: maximize projected variance along a unit direction to obtain the leading eigenvector of the covariance matrix. The top k eigenvectors give both maximum retained variance and minimum linear reconstruction error.

Independent Components Analysis: Recovering Independent Sources

Chapter 13 models ICA as x=As: observations are unknown linear mixtures, and the goal is to estimate W=A^{-1} to recover independent, non-Gaussian sources. A Jacobian determinant enters the transformed density and leads to the Bell–Sejnowski likelihood update.

Diffusion Models: Forward Noise, Reverse Generation, and the ELBO

Chapter 14 starts with a fixed Gaussian noising Markov chain and learns to reverse each transition. The ELBO turns reverse-kernel matching into weighted noise prediction, while the continuous-time view explains reverse drift through the score ∇log p_t.

Foundation Models Overview: Linear Probes, Fine-Tuning, and LoRA

Chapter 15 compares linear probing, full fine-tuning, and LoRA—not only by trainable parameter count, but by representation movement, data needs, and memory cost.

Representation Learning: Contrastive Learning, Retrieval, and RAG

Chapter 16 connects representation learning to systems: contrastive objectives shape an embedding space, semantic retrieval finds neighbors in it, and RAG passes retrieved context to a generator.

Large Language Models: Tokenization, Transformers, MoE, and SFT

Chapter 17 runs from next-token loss through Transformers, KV caches, MoE, and SFT, connecting an LLM's objective and architecture to its inference costs.

Reasoning in LLMs: Chain of Thought and Long-Reasoning RLVR

Chapter 18 separates two levers for LLM reasoning: chain of thought adds test-time computation, while verifiable rewards and policy gradients train long-reasoning behavior.

Reinforcement Learning: MDPs, Value Iteration, and Continuous States

Chapter 19 uses Bellman equations to turn long-horizon decisions into one-step updates, moving from value iteration in known MDPs to model learning and continuous-state approximation.

LQR, DDP, and LQG: From Linear Control to Uncertainty

Chapter 20 exploits linear dynamics and quadratic objectives to solve LQR, then uses DDP for local nonlinearity and Kalman filtering with LQG for partially observed state.

Policy Gradient and Its Variants: REINFORCE and PPO

Chapter 21 derives REINFORCE with the log-derivative trick, then uses reward-to-go, baselines, and PPO clipping to control policy-gradient variance and update size.

aideep-dive

Steel Browser: An Open-Source Browser API and the Boundary of Self-Hosting

Steel packages Chromium sessions, CDP, proxies, stealth, and debugging behind an Apache-2.0 browser API. Its public repository has about 7,400 stars and it entered the Stripe Projects developer preview in 2026. Self-hosting fits development and data-control needs; Cloud addresses concurrency, managed proxies, CAPTCHA, recordings, and SLAs.

aideep-dive

Together AI: From Serverless Inference to Dedicated Endpoints and Fine-Tuning

Together AI puts serverless APIs for open-weight models, dedicated GPU endpoints, batch inference, and fine-tuning on one platform, letting teams validate per token before moving to reserved deployment when traffic or customization justifies it.

aidebug

How to Write a Claude Code Skill That Doesn't Eat Context: Entry Point, Thresholds, Cost, Sources

I turned Anthropic's session-cost advice into a global skill, and the first version made the very mistakes it was meant to prevent: a description stuffed with trigger keywords, hard thresholds based on file counts and minutes, and 'protect the main context' conflated with 'spend fewer tokens overall'. Three rounds later the entry point is one page, details live in references, numeric thresholds became four judgment dimensions, and every claim from a draft post was checked against official docs.

aideep-dive

Vapi: Managed Voice-Agent Orchestration and the Safety Boundaries Before Going Live

Vapi connects phone and web audio, STT, LLMs, TTS, tool calls, and call observability in a managed voice runtime. Providers are swappable, but Vapi's realtime orchestration is not portable. In May 2026, the company reported one million developers and announced a $50 million Series B.

aideep-dive

Vercel Sandbox Deep Dive: Putting the Agent Execution Layer Inside the Vercel Ecosystem

Vercel Sandbox isolates untrusted code in Firecracker microVMs and integrates with Fluid compute, Active CPU pricing, and Vercel OIDC. It fits agents already running on Vercel, but network defaults, memory billing, and persistence still require deliberate design.

aideep-dive

Vertex AI Explained: From Model APIs to Gemini Enterprise Agent Platform

Vertex AI is more than the Gemini API: it puts access to 200+ models, training, evaluation, deployment, and governance under one Google Cloud control plane. Since April 2026, its products and roadmap have moved into Gemini Enterprise Agent Platform, while the Vertex AI API, documentation paths, and many resource names remain in active use.

Web Extraction Quality Benchmark: Crawl4AI, Firecrawl, Jina Reader, and Readability

Extraction tools cannot be compared by HTTP 200s. The same 20 URLs must be scored for body text, headings, tables, code, links, metadata, noise, latency, and cost. This article publishes the corpus, adapter contract, and gates, but no winner without a same-version raw run across all four paths.

aideep-dive

Zep Complete Guide: Temporal Knowledge Graphs for Agent Memory

Zep does not merely vectorize chat history. It turns episodes into entities and facts with validity time, allowing new information to invalidate an old relationship without erasing history. Graphiti is the open-source framework; Zep adds managed scale and governance.

careerguide

Beyond Upwork: Seven Remote Work Platforms Worth Bookmarking

From the free We Work Remotely to the elite 3%-acceptance Toptal, seven platforms each serve a different niche — job boards (FlexJobs, WWR, Remote OK), startup talent (Wellfound), community (Remotive), market research (Working Nomads), and premium freelancing (Toptal). All friendlier than competing on price at Upwork.

CS188 Bayes Nets and Ghostbusters: Inference When Ghosts Are Invisible

Lectures 13–18 and Project 4 move from factor operations and variable elimination to exact inference and particle filtering, letting Pacman track invisible ghosts through noisy distance sensors.

Completing CS188: Turn 28 Lectures and Projects P0–P5 into a Portfolio

Lectures 26–28 close with nuclear monitoring, AI safety, and reflection. Independent completion should preserve assumptions, test evidence, and failure analysis for Projects 1–5 instead of reporting only autograder scores.

CS188 CSPs and Multi-Agent Search: Choosing Minimax, Alpha-Beta, and Expectimax

Lectures 5–8 use CSPs to practice variables, constraints, and search order before Project 2 implements minimax, alpha-beta, and expectimax. Their key difference is the assumption made about other agents.

CS188 Decisions and Machine Learning: From VPI and Naive Bayes to Attention

Lectures 19–25 connect rational decisions and VPI to machine learning, while Project 5 uses PyTorch for regression, classification, CNNs, attention, and an optional character-GPT.

CS188 MDPs and Reinforcement Learning: From Value Iteration to Q-Learning

Lectures 9–12 and Project 3 use the same Gridworld to contrast value iteration with a known model, Q-learning from unknown dynamics, and approximate Q-learning that generalizes through features.

CS188 Search and Heuristics: Pacman from DFS and BFS to A*

Lectures 1–4 and Project 1 connect DFS, BFS, UCS, A*, state representation, and heuristic design. The goal is not memorizing algorithms but separating what the frontier, cost, and state each control.

Berkeley CS188 Spring 2026: Learn AI Through Projects P0–P5

CS188 Spring 2026 publishes 28 recordings, 27 lecture slide sets, 11 discussions, and Projects P0–P5. P0 is a Python/autograder tutorial, P1–P4 use Pacman settings, and P5 contains general machine-learning tasks.

Berkeley CS189 Spring 2025 Overview: HW1–7 with Code and Data You Can Run, Plus What Fall 2026 Looks Like

CS189 exists online in several terms. From post 2 onward this series uses Spring 2026 (eecs189.org/sp26, A3) as its base, walking through each lecture block and assignment; Spring 2025 (Shewchuk) is a different classic route with 25 lectures of notes, HW1–7 and past exams, but its official recordings sit behind a bCourses login; Fall 2026 is still in progress and serves only as a reference. This post is the series entry point and full table of contents.

Berkeley CS285 L19–25: Exploration, RL Theory, Multitask Learning, and Open Problems

The final seven lectures move from exploration and theoretical limits through two review lectures to advanced exploration, multitask RL, and unresolved research problems.

Berkeley CS285 Homework and Final Projects: The CPU, GPU, and H100 Boundary

Five assignments move from CPU-friendly imitation learning to H100-based LLM RL and six-hour offline-RL runs; self-learners should use three compute tiers instead of copying the entire enrolled workflow.

Berkeley CS285 L1–4: Imitation Learning, Distribution Shift, and RL Basics

The first four lectures move from behavioral cloning to MDPs; HW1 turns distribution shift into an observable failure through MSE policies, DAgger, and flow matching.

Berkeley CS285 L11–18: From Variational Inference and LLM RL to Offline RL

L11–18 connect control as inference, LLM RL, model-based RL, and offline RL, with HW4 and HW5 providing two compute-intensive implementations.

Berkeley CS285 L5–10: Policy Gradients, Actor-Critic, DQN, and SAC

L5–10 build the deep-RL core through policy- and value-based routes; HW2 is CPU-friendly, while HW3's Atari and HalfCheetah runs can require hours of GPU time.

Berkeley CS285 Spring 2026 Guide: 25 Lectures, Five Assignments, and the Self-Study Boundary

Spring 2026 CS185/285 publishes slides for 25 lectures, nine discussion units, five assignments, and starter code; current recordings require bCourses access, while HW4 defaults to an H100, so this is not a zero-cost open course.

Berkeley CS288 Part 5: Inference-time Compute, Reasoning, and Embodied Agents

Units 15–18 place NLP models inside perception, reasoning, tool, and environment loops; the question shifts from next-token prediction to allocating inference compute and validating multi-step action.

Berkeley CS288 Part 1: From N-grams and Word Representations to Text Classification

The first four units make text countable, representable, and classifiable; A1 then moves from n-grams and perceptrons to an NBOW MLP.

Berkeley CS288 Spring 2026: 18 Slide Units, Three Assignments, and the Limits of Self-Study

CS288 moves from n-grams to RAG, reasoning, and agents through 18 public slide units and three assignments; Berkeley-only recordings make this an A3 materials route, not a public video course.

Berkeley CS288 Part 3: Pre-training, Post-training, Generation, and Evaluation

Units 08–12 turn a base model into an interactive system: pre-training establishes capability, post-training shapes behavior, and generation plus evaluation determine how outputs are used.

Berkeley CS288 Part 4: Turning Retrieval, RAG, and Advanced Architectures into a System

Units 13–14 connect models to external knowledge; A3 requires data collection, QA annotation, indexing, and ablations under CPU and latency constraints.

Berkeley CS288 Part 2: Sequence Models, Seq2Seq, and Transformers

Units 05–07 move from recurrent state to encoder-decoder models, then rewrite the information path with attention and Transformer blocks.

CMU 07-380 Fall 2026 Overview: 26 Lectures from Logic and Planning to Diffusion, HW and Project Not Yet Fully Released

07-380 Fall 2026 is the first offering of CMU's new AI II: 26 lectures from logic, planning and optimization to probabilistic graphs and generative systems. As of the 2026-09-29 course site, the Lec1–9 slides, PR1–6 notes, Rec1–5 (with solutions) and HW1–3 are public, and this site now has 14 lecture-by-lecture guides for them. Slides after Lec10, HW4–7 and the final project are not out yet, so the course as a whole is still A2.

CMU 10-301 HW1: Find ML Foundation Gaps with Mathematics and Python

HW1 is written and programming work: mathematical and CS foundations followed by a majority-vote classifier.

CMU 10-301 HW2: From Information Calculations to a Complete Decision Tree

HW2 moves from hand-calculated entropy and mutual information to an end-to-end tree learner, predictor, and evaluator.

CMU 10-301 HW3: Compare K-NN, Perceptron, and Linear Regression

HW3 is written work: a decision-tree review followed by K-NN, Perceptron, and Linear Regression through inductive bias, errors, and model selection.

CMU 10-301 HW4: Turn Logistic Regression Likelihood into a Classifier

HW4 joins probabilistic interpretation, cross-entropy gradients, and implementation into one traceable training pipeline.

CMU 10-301 HW5: Expose Neural Networks and Backpropagation with NumPy

HW5 avoids automatic differentiation so learners must track forward shapes, caches, and backward gradients themselves.

CMU 10-301 HW6: Learning Theory, MLE/MAP, and Fairness Metrics

HW6 combines generalization, MLE/MAP, probabilistic learning, fairness metrics, and social impact in one written assignment about assumptions and tradeoffs.

CMU 10-301 HW7: Move from Basic Neural Networks to Deep Learning

HW7 builds on HW5 backpropagation to address deep-model architecture and training failures, emphasizing diagnosis over merely adding layers.

CMU 10-301 HW8: From MDPs to Reinforcement-Learning Updates

HW8 connects states, actions, rewards, transitions, and value updates while separating environment dynamics, policy, and estimation error.

CMU 10-301 HW9: Close the Course with Ensembles, k-Means, PCA, and Recommenders

The final written assignment combines ensembles, clustering, representation, and recommendation to test whether you can choose a learning paradigm from problem structure.

CMU 10-301/601 Spring 2026: Learn Machine Learning Through Nine Assignments

Spring 2026 publishes material for 27 lectures and nine homework bundles; outsiders can do the core work but cannot access Panopto, Piazza, Gradescope, or official homework solutions.

learningdeep-dive

CMU's AI Core Redesign: From 15-281 + 10-315 to 07-280 + 07-380

In 2026, CMU recombined its separate general-AI and SCS machine-learning introductions into the 07-280 → 07-380 sequence. This is a redistribution of content and prerequisites, not a pair of simple course renames.

Stanford CS107 Lecture 4: Bitwise Operators, Conversions, and Masks

Lecture 4 first shows that signed/unsigned conversion can preserve bits while changing meaning, that mixed comparisons may surprise, and how sign extension, zero extension, and truncation alter width. It then derives AND, OR, NOT, XOR, and bitmask idioms for testing, setting, clearing, and combining fields.

Stanford CS107 Lecture 3: Integers, Bytes, and Two's Complement

Lecture 3 starts with 32/64-bit address spaces, derives the ranges of unsigned and two's-complement signed integers, inversion-plus-one, and shared addition hardware, then separates unsigned modular arithmetic from C signed overflow and tests the model against four failure cases.

Stanford CS107 Lecture 5: Bit Shifts, Bit Tricks, and GDB

Lecture 5 extends masks to shifts, power-of-two and popcount tricks, then uses an absolute-value example to expose signed intermediate overflow at INT_MIN. Its second half establishes a GDB workflow around breakpoints, execution control, formatted printing, memory examination, and backtraces.

Stanford CS107 Lecture 2: A First C Program, Binary, and Hexadecimal

Lecture 2 puts C back into its Unix history and development environment: headers, main, printf, argc/argv, ssh, emacs, make, and executables. It then derives 8 bits = 1 byte, 256 byte patterns, and reliable conversion among decimal, binary, and hexadecimal.

Stanford CS107 Lecture 1: From the Course Map to the Unix Command Line

Winter 2026 opens by explaining why CS107 goes below programming-language abstractions: from bytes and memory through assembly and heap allocators. It then lays out the 40/10/20/30 grading structure and closes with a first tour of the Unix command line.

Harvard AI/ML Course Guide: Do CS50 AI, CS181, and CS182 Videos Match Their Assignments?

CS50 AI is Harvard's most complete public entry point, but the Summer 2026 course still uses 2020 recordings and assignment assets while the rolling OCW projects have moved to other editions. CS181 Spring 2026 exposes current homework and notes without current recordings; CS182 Fall 2026 has not yet completed an offering.

learningdeep-dive

The Pacman AI Project Lineage: How Berkeley CS188 and CMU 15-281 Restructure the Same Material

CMU 15-281's Search and Games explicitly credits Berkeley's Pacman AI projects. The official course site separately lists a zero-point P0 tutorial and five programming assignments, P1–P5.

Stanford CS103 Lecture 0: From Set Language to Cantor's Diagonal

Starting with elements, subsets, and power sets, this lecture culminates in Cantor's diagonal proof that no set is as large as its own power set.

Stanford CS103 Lecture 1: Building a First Direct Proof from Even and Odd

The even-square and odd-sum examples show how arbitrary choices, assumptions, witnesses, and a want-to-show become a checkable direct proof.

Stanford CS103 Lecture 2: Negation, Contraposition, and Contradiction

This lecture identifies exactly when an implication is false, then turns quantified negation, contraposition, and contradiction into checkable proof tools.

Stanford CS103 Lecture 3: Propositional Logic, Truth Tables, and Equivalence

Propositional logic abstracts English statements into Boolean variables, then uses truth tables to check connectives, translation direction, and equivalences.

Stanford CS103 Lecture 4: Objects, Quantifiers, and Types in First-Order Logic

This lecture extends propositional logic into a language about objects: distinguish constants, predicates, functions, and propositions, then express some and every with existential and universal quantifiers.

Stanford CS103 Lecture 5: First-Order Logic II—Nested Quantifiers, Negation, and Uniqueness

Translate natural language one layer at a time: identify universal and existential forms, then handle quantifier order, negation, restricted quantifiers, and uniqueness.

Stanford CS103 Lecture 6: Functions I, from Definitions to Injection and Surjection Proofs

A function is more than a formula: domain, codomain, totality, and determinism are essential, while the quantifiers defining involutions, injections, and surjections dictate their proofs.

Stanford CS103 Lecture 7: Functions II—Surjections, Assumptions, and Composition

This lecture uses surjections and a proof about birds to separate assuming from proving, then shows that involutions are injective and surjective and carries those ideas into function composition.

Stanford CS103 Lecture 8: Cardinality by Bijections and Cantor's Diagonal Argument

Two sets have equal cardinality when a bijection pairs their elements; Cantor's diagonal set defeats every function from S to its power set by constructing a value it misses.

Stanford CS103 Lecture 9: Graphs, Part I

This lecture moves from the formal definitions of graphs and digraphs to independent sets, vertex covers, and their complement relationship.

Stanford CS103 Lecture 10: Walks, Graph Complements, and the Pigeonhole Principle

Starting with walks, paths, cycles, and components, this lecture proves that a graph or its complement is connected and develops the pigeonhole principle through degrees and monochromatic triangles.

Stanford CS103 Lecture 11: Generalized Pigeonhole, Ramsey Theory, and Average Load

Use the generalized pigeonhole principle to force a monochromatic triangle at a six-person party, then solve a movie-preference puzzle through average load and contradiction.

Stanford CS103 Lecture 12: Induction, Counterfeit Coins, and Invariants

Induction is not a list of checked examples: establish a true starting point, prove that an arbitrary true case transmits truth to the next case, and invoke the induction principle.

Stanford CS103 Lecture 13: Mathematical Induction, Part II

This lecture connects starting from ordinary induction to induction may start later, following the official examples and proof obligations.

Stanford CS103 Lecture 14: Finite Automata, Part I

This lecture connects why begin with a weak computer to from device behavior to a state machine, following the official examples and proof obligations.

Stanford CS103 Lecture 15: Finite Automata, Part II

This lecture connects the dfa definition connects the first half of cs103 to regular means that some dfa exists, following the official examples and proof obligations.

Stanford CS103 Lecture 16: Finite Automata, Part III

This lecture connects the automata ladder measures power with languages to dfa transition tables, following the official examples and proof obligations.

Stanford CS103 Lecture 17: Regular Expressions

This lecture connects from closure properties to a language syntax to regex is mathematics, not one library, following the official examples and proof obligations.

Stanford CS103 Lecture 18: Nonregular Languages

This lecture connects four equivalent descriptions of regularity to the precise finite-memory intuition, following the official examples and proof obligations.

Stanford CS103 Lecture 19: Context-Free Languages

This lecture connects from finite-state limits to recursion to the arithmetic grammar, following the official examples and proof obligations.

Stanford CS103 Lecture 20: Turing Machines, Part I

This lecture connects why the model changes after cfgs to long addition and local access, following the official examples and proof obligations.

Stanford CS103 Lecture 21: Turing Machines, Part II

This lecture connects the sample tm looks back from the end to beyond pairwise marking, following the official examples and proof obligations.

Stanford CS103 Lecture 22: Turing Machines, Part III

This lecture connects a quick quantifier audit for recognizers and deciders to why every decision problem can be represented as a language, following the official examples and proof obligations.

Stanford CS103 Lecture 23: Unsolvable Problems, Part I

This lecture connects returning from r, re, and utm to three self-reference warm-ups, following the official examples and proof obligations.

Stanford CS103 Lecture 24: Unsolvable Problems, Part II

This lecture connects defining and locating halt to why halt is recognizable, following the official examples and proof obligations.

Stanford CS103 Lecture 25: Unsolvable Problems, Part III

This lecture connects the lava diagram's two classification tasks to the deck's operational reading of rice's theorem, following the official examples and proof obligations.

Stanford CS103 Lecture 26: Complexity Theory

This lecture connects decidable does not mean feasible to efficiency requires choosing a resource, following the official examples and proof obligations.

Stanford CS103 Wrap-Up: Four Foundations and Where to Go Next

The final deck reconnects proofs, graphs, automata, and computability, then maps those foundations to Stanford courses that use them.

Stanford CS107 Lecture 15: Reading x86-64 Addressing Modes Without Confusing Addresses and Values

CS107 Lecture 15 decomposes x86-64 mov operands into immediate, register, absolute, indirect, displacement, indexed, and scaled-indexed forms, then unifies pointer dereference and array access with D + R[b] + R[i]×s.

Stanford CS107 Lecture 16: From Subregisters to x86-64 Arithmetic and Logic

CS107 Lecture 16 connects b/w/l/q data widths, subregisters, movs/movz, lea, calling conventions, arithmetic and logic, and shifts through one method: establish operand width before tracing sources, destinations, and real memory accesses.

Stanford CS107 Lecture 18: From Condition Codes to x86-64 Loops

CS107 Lecture 18 connects ZF/SF/CF/OF to cmp, test, signed and unsigned conditional jumps, then reconstructs if statements, loops, dynamic instruction counts, setcc, and cmovcc.

Stanford CS107 Lecture 17: From Multiply and Divide to x86-64 Control Flow

CS107 Lecture 17 completes full-width x86-64 multiplication and division, traces %rip through instruction bytes, and uses direct and indirect jmp to show how execution leaves its default sequential path.

Stanford CS107 Lecture 19: Understanding x86-64 Function Calls and Calling Conventions

CS107 Lecture 19 traces %rsp, push/pop, call/ret, parameters, return values, stack locals, and caller/callee register discipline to build the ABI contract that preserves data and control across functions.

Stanford CS107 Lecture 14: From C to x86-64, Reading Disassembly for the First Time

CS107 Lecture 14 dissects the ten x86-64 instructions for sum_array: addresses and machine bytes appear on the left, AT&T assembly on the right, and the reader's job is to recover C-level effects from opcodes, operands, registers, and control flow—not to write assembly from scratch.

Stanford CS107 Lecture 7: From String Search to Buffer Overflows—Input Validation Is Not Capacity Checking

CS107 Lecture 7 builds pointer-based string scanning with strchr, strstr, and strspn, then shows why valid content can still overflow a buffer: safety requires input rules, destination capacity, termination, and memory-error detection.

Stanford CS107 Lecture 25: Caching, Memory Hierarchy, and Locality

CS107 Lecture 25 builds the essential cache model from a concise deck: memory access costs are nonuniform, smaller and faster layers retain data likely to be reused, and temporal and spatial locality determine whether a program benefits.

Stanford CS107 Lecture 6: A C String Is Not a Type but a Memory Contract

CS107 Lecture 6 reduces C strings to character arrays, a terminator, and an address: every convenience in strlen, strcmp, strcpy, strncpy, and strcat depends on the caller preserving capacity and termination invariants.

Stanford CS107 Lecture 24: Profile with Callgrind, Then Read What GCC Optimized

CS107 Lecture 24 builds a measurement workflow with matrix multiplication and Callgrind, then examines GCC constant folding, common-subexpression elimination, dead-code elimination, strength reduction, code motion, and recursion-to-loop conversion. Optimization starts with bottleneck evidence.

Stanford CS107 Lecture 23: The Allocator Invariants Behind In-Place realloc

CS107 Lecture 23 advances the explicit free list to in-place realloc: split a useful remainder when shrinking, absorb free right neighbors when growing, and allocate-copy-free only as a fallback, while preserving both the physical heap and logical list.

Stanford CS107 Lecture 13: From Comparators to a Fully Generic Bubble Sort

CS107 Lecture 13 upgrades a Boolean callback to a three-way comparator, then combines void *, element width, and const void * callbacks into a fully generic bubble sort before mapping the design to qsort, bsearch, lfind, and lsearch.

Stanford CS107 Lecture 12: Function Pointers Inject Ordering into Generic C

CS107 Lecture 12 first uses char * for byte-wise generic swap and rotate, then uses a function pointer to separate bubble sort's traversal mechanism from its ordering rule: void * abstracts data types, while callbacks abstract behavior.

Stanford CS107 Lecture 11: How void * Gives C Generics Without Pretending Types Still Exist

CS107 Lecture 11 finishes the heap contracts of calloc, strdup, free, and realloc, then turns several typed swap functions into void * plus a byte count: C generics do not preserve an unknown type; they explicitly transfer responsibility for addresses, widths, and interpretation.

Stanford CS107 Lecture 21: A First Heap Allocator and the Tension Between Speed and Space

CS107 Lecture 21 starts with alignment, throughput, and utilization, then uses a bump allocator and an implicit free list to explain metadata, splitting, placement, internal and external fragmentation, and the need to coalesce freed blocks.

Stanford CS107 Lecture 22: Why an Explicit Free List Lives in Two Orders at Once

CS107 Lecture 22 replaces an implicit list with an explicit free list. Searches visit only reusable blocks, but every free block now has both physical neighbors and logical links, so unlinking, coalescing, and reinsertion must preserve both structures.

Stanford CS107 Lecture 8: A Pointer Is Not Magic, but a Copyable Address

CS107 Lecture 8 starts with address-of and dereference, explains why C pointer parameters are still passed by value, and shows how int *, char *, and char ** can modify caller-owned ints, chars, and pointers respectively.

Stanford CS107 Lecture 9: An Array Is Not a Pointer, but They Cooperate in Expressions

CS107 Lecture 9 uses seven C-string rules to separate array objects, pointer variables, and string literals: arrays often convert to first-element pointers in expressions, but storage, assignment, mutability, and sizeof remain different.

Stanford CS107 Lecture 20: After Reverse Engineering, Ask About Privacy and Trust Before Building a Heap Allocator

CS107 Lecture 20 places reverse-engineering capability in an ethical context: privacy has individual and social models, while trust combines reliance with a risk of betrayal. It then reviews process memory and shifts from heap-allocation client to allocator implementer.

Stanford CS107 Lecture 10: Stack vs. Heap Is About Lifetime and Ownership, Not Just Speed

CS107 Lecture 10 moves from sizeof and pointer arithmetic to stack-frame lifetime: returning a local array leaves a dangling pointer; malloc crosses function returns but makes NULL handling, size arithmetic, ownership, free, and leaks the programmer's responsibility.

Stanford CS107 Lecture 26: Wrap-up, Six Systems Questions, and What Comes Next

CS107 Lecture 26 closes ten weeks through six big questions: representation, text, memory, generics, execution, and allocation. It checks the learning goals through the explicit allocator and points toward CS111 and other systems courses.

Stanford CS109 Lecture 1 | What is Probability?: List outcomes first; only then assign probabilities to events.

List outcomes first; only then assign probabilities to events.

Stanford CS109 Lecture 2 | Conditional Probability: A condition restricts the sample space to outcomes still compatible with the evidence.

A condition restricts the sample space to outcomes still compatible with the evidence.

Stanford CS109 Lecture 3 | Bayes Theorem: Bayes’ theorem turns an easier generative direction into the inferential direction we need.

Bayes’ theorem turns an easier generative direction into the inferential direction we need.

Stanford CS109 Lecture 4 | Counting and Combinatorics: Decide whether order matters and repetition is allowed before choosing a formula.

Decide whether order matters and repetition is allowed before choosing a formula.

Stanford CS109 Lecture 5 | Random Variables and Expectation: A random variable maps outcomes to numbers; expectation is a weighted average, not necessarily an attainable value.

A random variable maps outcomes to numbers; expectation is a weighted average, not necessarily an attainable value.

Stanford CS109 Lecture 6 | Moments: Expectation, LOTUS, and linearity

Expectation compresses a distribution into a weighted average; LOTUS handles transformed values, while linearity makes sums tractable even without independence.

Stanford CS109 Lecture 7 | Variance and Poisson: From spread to rare-event counts

Variance describes a random variable's spread; Poisson models counts in a fixed interval and approximates a large-n, small-p binomial.

Stanford CS109 Lecture 8 | Continuous Random Variables: PDFs, CDFs, Uniform, and Exponential

A continuous variable assigns zero probability to a point and area to intervals; CDFs, Uniform, Exponential, and memorylessness build on that distinction.

Stanford CS109 Lecture 9 | Normal Distribution: Standardization, Phi, and continuity correction

Standardization maps Normal variables to Z; Phi, linear transforms, and continuity correction turn intervals and large binomials into computable probabilities.

Stanford CS109 Lecture 10 | Probabilistic Models: Joints, marginals, independence, and Bayes

A joint distribution retains the full relationship among variables; marginals, conditionals, independence, and Bayes extract different answers from it.

Stanford CS109 Lecture 11 | Inference: Prior times likelihood, then normalize

Inference multiplies each hidden-variable prior by an observation likelihood and normalizes; the same loop handles repeated evidence and discretized continuous beliefs.

Stanford CS109 Lecture 12 | General Inference: Bayesian networks, sampling, and rare evidence

A Bayesian network factorizes a huge joint through conditional independence; ancestral sampling generates joint samples, and rejection sampling filters them into a conditional.

Stanford CS109 Lecture 13 | Multinomial: Category counts, bag of words, and log probability

The Multinomial extends two-category Binomial counts to many categories; the same PMF models documents as word counts for Bayesian authorship with log-scores.

Stanford CS109 Lecture 14 | Beta: Turn an unknown probability into an updatable random variable

A Beta distribution represents full belief about an unknown success rate; success/failure data updates two parameters for posteriors, smoothing, and Thompson-sampling decisions.

Stanford CS109 Lecture 15 | Adding Random Variables and the Central Limit Theorem

A few independent sums have closed forms; general IID sums become approximately Normal under the CLT, with continuity correction for discrete sums.

Stanford CS109 Lecture 16 | Bootstrapping: Sampling statistics, error bars, and p-values

The bootstrap treats a sample histogram as a population proxy, resampling with replacement to approximate a statistic's sampling distribution, error bar, or null p-value.

Stanford CS109 Lecture 17 | Algorithmic Analysis: Conditional expectation, indicators, and recursion

Expected cost in randomized code can be conditioned on the first random choice; counting problems become indicator sums, often avoiding the full distribution entirely.

Stanford CS109 Lecture 18 | Information Theory: Surprise, entropy, information gain, and KL

Surprise turns rare events into bits; entropy is expected surprise, information gain selects uncertainty-reducing questions, and KL measures excess cost from a model distribution.

Stanford CS109 Lecture 19 | Maximum Likelihood Estimation: Hold data fixed and optimize the parameter

MLE fixes observed data and optimizes parameters; log-likelihood turns products into sums, but a maximum can also lie on a boundary.

Stanford CS109 Lecture 20 | Logistic Regression: Derive the gradient from Bernoulli likelihood

Logistic regression turns a linear score into a Bernoulli probability with sigmoid; the gradient xⱼ(y-ŷ) follows directly from the log-likelihood chain rule.

Stanford CS109 Lecture 21 | Comparing Classifiers: Beyond accuracy to calibration, error costs, and fairness

Classifier comparison requires held-out data, baselines, calibration, precision/recall, and an explicit fairness criterion—not accuracy alone.

Stanford CS109 Lecture 22 | Deep Learning: Derive backpropagation with the chain rule

A neural network stacks logistic units; a forward pass computes probabilities, while backpropagation reuses output error to obtain every gradient.

Stanford CS111 Lecture 1: Welcome to CS111!

Lecture 1 follows shared I/O cards in the 1940s, batch processing, multiprogramming, and personal computers to explain how OS responsibilities accumulated as hardware costs and user needs changed.

Stanford CS111 Lecture 2: Threads, Processes, and Dispatching

Lecture 2 defines shared and private process/thread state, then uses fork, execvp, waitpid, and thread creation to show how the kernel creates execution units.

Stanford CS111 Lecture 3: Threads, Processes, and Dispatching, Continued

Lecture 3 follows running, blocked, and ready transitions to show how PCBs, context save/restore, and the dispatcher complete one CPU-control handoff.

Stanford CS111 Lecture 4: Concurrency

Lecture 4 defeats each Too Much Milk attempt with an explicit schedule, deriving race condition, atomicity, critical section, and synchronization requirements from concrete interleavings.

Stanford CS111 Lecture 5: Mutexes, Condition Variables, and Mesa Semantics

Lecture 5 uses an eight-slot circular Pipe to prove that a mutex supplies exclusion, while a condition variable atomically releases the lock and blocks when a predicate is false; under Mesa semantics, wait must return to a while loop that rechecks the predicate.

Stanford CS111 Lecture 6: Implementing Locks

Lecture 6 evolves a one-core interrupt-masking lock through multicore version 5, tracking guard, lock, and wait-queue state to prevent races and lost wakeups.

Stanford CS111 Lecture 7: Deadlock Conditions and Global Lock Ordering

Lecture 7 extracts four necessary deadlock conditions from request/ownership graphs, then compares detection, prevention, and lock ranking; breaking circular wait is common in practice, but every module must obey one global order.

Stanford CS111 Lecture 8: FIFO, Round Robin, Priorities, and Multicore Scheduling

Lecture 8 moves from FIFO and round robin through the unimplementable SRPT ideal to adaptive priority queues and the multicore conflict among queue contention, core affinity, and work conservation.

Stanford CS111 Lecture 9: Linkers and Dynamic Linking

Lecture 9 follows source through assembly, object, executable, and process, explaining the linker's three passes and how a dynamic loader resolves shared-library addresses through a jump table at startup.

Stanford CS111 Lecture 10: Dynamic Storage Management

Lecture 10 moves from predictable LIFO stacks to heap free lists, first/best fit, and slabs, then compares reference counting with mark-and-sweep across dangling pointers, leaks, cycles, and fragmentation.

Stanford CS111 Lecture 11: Dynamic Storage Management, Continued

Lecture 11's official PDF is byte-identical to Lecture 10; this article preserves that artifact gap and focuses on reachability, dangling pointers, leaks, reference-count cycles, and mark/compact garbage collection.

Stanford CS111 Lecture 12: Trust and Operating Systems

Lecture 12 defines trust as voluntary vulnerability, separates over-trust from untrustworthiness, and applies assumption, inference, and substitution to the Linux TCB, the xz attack, and AI-code policy.

Stanford CS111 Lecture 13: Virtual Memory

Lecture 13 starts from the failures of single-tasking and load-time relocation, uses an MMU with base/bound to create isolated virtual and physical address spaces and traps, then introduces segmentation to escape one contiguous region.

Stanford CS111 Lecture 14: Virtual Memory, Continued

Lecture 14's official PDF is byte-identical to Lecture 13; this article records the gap and focuses on how multiple base/bound/protection entries enable growth, sharing, and compaction while retaining fixed-count, fragmentation, and rigid-layout limits.

Stanford CS111 Lecture 15: Paging

Lecture 15 uses fixed pages to remove inter-process external fragmentation, then connects x86-64's four-level walk, sharing and aliasing, and the TLB to trade-offs among translation speed, sparse tables, context switches, and page size.

Stanford CS111 Lecture 16: Page Faults, Demand Fetching, and Prefetch

Demand paging loads pages only when needed; present bits, precise exceptions, and restartable instructions let the kernel safely fill them from executables, zero-fill, or backing store.

Stanford CS111 Lecture 17: From Page Faults to Clock—Who Leaves When Memory Is Full?

Lecture 17 separates demand paging into fetching and replacement: MIN cannot know the future, exact LRU is too expensive, and Clock uses reference/dirty bits to find a page old enough to evict; when active working sets exceed RAM, even a 1% fault rate can cause an approximately 1,000-fold slowdown.

Stanford CS111 Lecture 18: Disk Geometry, Interrupts, and DMA

A disk hides mechanical seek and rotation behind a linear block API; modern I/O then uses memory-mapped registers, DMA queues, and interrupts so the CPU mainly issues commands and receives completions.

Stanford CS111 Lecture 19: File Abstractions, Allocation, and FAT

A file system maps durable byte collections onto disk blocks; contiguous, linked, and FAT allocation trade locality, growth, random access, and metadata cost.

Stanford CS111 Lecture 20: Multilevel Inodes, Index Walks, and Disk Scheduling

The 4.3BSD inode uses direct, single-indirect, and double-indirect tiers so lookup depth scales with file size; FIFO, SPTF, SCAN, and CSCAN then trade seek cost, fairness, and wait time.

Stanford CS111 Lecture 21: Block Cache, Free Bitmaps, and Delayed Allocation

Block cache retains hot indexes, bitmap slack preserves placement choices, and fragments plus delayed allocation trade later, better information for locality.

Stanford CS111 Lecture 22: Directory Lookup, Hard Links, and Symbolic Links

Directories map text names to file-system-local inode numbers; hard links share inode identity and reference counts, while symlinks store paths and permit cross-filesystem references with loops and dangling targets.

Stanford CS111 Lecture 23: From fsck and Ordered Writes to Write-Ahead Logging

A single file-system operation updates several blocks, but a crash can occur between any two writes; this lecture compares how fsck, ordered writes, and write-ahead logging trade recovery time, performance, durability, and consistency.

Stanford CS111 Lecture 24: Journaling, Transactions, and Checkpoints

Lecture 24 continues from the WAL entry point into transactions, idempotent replay, and checkpoints, showing why consistency is not durability and why a journal does not replace fsync or backups.

Stanford CS111 Lecture 25: Truth, Trust, and Technology—How Algorithms, Generative AI, and Deepfakes Reshape Trust

Lecture 25 separates assumption, inference, and substitution as ways to establish trust, then examines how social recommendations, generative AI, and synthetic media amplify over-trust; the response is preserved provenance, independent validation, and coordinated responsibility.

Stanford CS111 Lecture 26: Flash Translation Layers, Garbage Collection, and Wear Leveling

Flash programs pages but erases whole units; an FTL hides the asymmetry with out-of-place mapping, then manages amplification through garbage collection, temperature segregation, wear leveling, and TRIM.

Stanford CS111 Lecture 27: Trap-and-Emulate, Virtual I/O, and Nested Page Tables

A VM expands the process interface into a machine interface; the hypervisor directly executes ordinary instructions, traps privileged operations, and virtualizes interrupts, I/O, and two-stage address translation.

Stanford CS111 Lecture 28: Four Ideas Connecting Concurrency, Memory, and Storage

Lecture 28 reduces the semester to concurrency, memory, and storage, then uses four ideas—virtualization, atomicity, locality, and layering—to explain how operating systems manage shared resources.

Ably: Global Realtime Messaging with Channels, Presence, and Recovery

Ably manages global realtime connections, channels, presence, and short-window recovery; applications still own idempotency, durable business state, token capabilities, and offline resynchronization.

Delegated Authorization for AI Agents: Do Not Hand User Tokens to the Model

An agent should execute one task with a short-lived, audience- and permission-restricted credential while preserving user and agent identities, execution-time authorization, confirmation, and audit lineage.

AI Agent Sandbox Escapes and Permission Boundaries: A Container Is Not the Whole Boundary

Agent execution must constrain kernels, filesystems, processes, networks, credentials, and tool authorization; sandbox escape is only one path, and an overpowered API token is often more direct.

Linode and Akamai Cloud: Connecting Classic VPS Compute to Edge and Managed Kubernetes

Linode is now the compute foundation of Akamai Cloud Computing; VMs, LKE, storage, and databases retain regional and network boundaries, so the Akamai brand does not imply complete integration.

techdeep-dive

Algolia Site Search Deep Dive: Hosted Indexing, Ranking, and InstantSearch

Algolia packages indexing, search-as-you-type, facets, and UI components as a hosted service; it ships quickly, but data synchronization, relevance, and usage costs remain your responsibility.

Apache Kafka: A Replayable Event Log, Not Merely a Message Queue

Kafka is a distributed log ordered by partition and retained by policy. Consumer groups divide work through offsets, while exactly-once processing holds only inside boundaries covered by Kafka transactions.

Apache Pulsar: An Event Platform Separating Stateless Brokers from BookKeeper Storage

Pulsar separates serving from storage: stateless brokers handle connections while BookKeeper stores ledgers and subscription cursors. Elasticity and multi-tenancy come with more operational components.

Appwrite: A Self-Hostable BaaS for Auth, Databases, Storage, Functions, and Realtime

Appwrite combines Auth, TablesDB, Storage, Functions, Realtime, and Messaging behind consistent APIs; Cloud and self-hosted products resemble each other but have different operational ownership.

techdeep-dive

assistant-ui Explained: Runtime and Primitives for Backend-Portable Agent Chat

assistant-ui separates Agent Chat into headless React primitives, a conversation runtime, and backend adapters, so the UI does not have to bind directly to one model SDK's message state.

AWS App Runner: The Shortest AWS Path from Source or Container to a Web Service

App Runner packages build, deployment, TLS, load balancing, and autoscaling as a web service, trading orchestration control for a simpler platform with explicit VPC, instance, and health boundaries.

AWS Fargate: No EC2 Management Does Not Mean No ECS Management

Fargate removes container-host operations, but task definitions, ECS services, VPCs, IAM, scaling, deployment, and observability remain your system.

AWS Lambda: Understand Events, Retries, and Concurrency Before Choosing Functions

Lambda fits short-lived, event-driven, bursty work; its design center is invocation, retries, idempotency, and downstream capacity—not merely smaller containers.

AWS SQS and SNS: One Stores Work, the Other Fans Out Events

SQS is a pull-based durable queue and SNS is a push-based topic. Reliable fan-out commonly connects one SNS topic to multiple SQS queues rather than sharing one queue across services.

Azure App Service: The App Service Plan Matters More Than the Container

App Service manages web runtimes, TLS, deployment, and scaling, while capacity, cost, and isolation live in the shared App Service Plan rather than one app.

Azure Container Apps: Serverless Containers with Revisions, KEDA, and Environments

Azure Container Apps hides Kubernetes while exposing HTTP/TCP ingress, revisions, KEDA scaling, jobs, and Dapr; replica concurrency, event idempotency, VNet design, and identity remain yours.

Better Auth: A TypeScript Authentication Framework, Not Application Authorization

Better Auth unifies login, sessions, providers, and plugins; applications still own resource authorization, revocation latency, and policy for agent actions.

techdeep-dive

Better Auth: Should Authentication Live Inside Your TypeScript App?

Better Auth trades a managed identity platform for an in-app library and your own database; that gives you control, but migrations, security updates, and incident response become your responsibility.

techdeep-dive

Bright Data Deep Dive: From Proxies and Web Unlocker to Browser API and Datasets

Bright Data splits web data access into four layers: proxies preserve control, Web Unlocker returns unblocked content, Browser API hosts interactive browsers, and Web Scraper APIs or Datasets deliver structured data.

CapRover: Self-Hosted PaaS with Docker Swarm, Nginx, and Persistent Apps

CapRover wraps Docker Swarm, Nginx, and captain-definition in a simpler PaaS; stateless apps scale, while local persistent apps remain pinned to one node.

techdeep-dive

Clerk Authentication Platform: From UI Components and Session Tokens to Organization Authorization

Clerk's real value is an integrated identity lifecycle, not a sign-in box; resource authorization, tenant isolation, and business-data consistency remain your application's responsibility.

Cloudflare Durable Objects: The Stateful Coordination Layer for Workers and WebSockets

Durable Objects map a name or ID to a globally unique, single-threaded actor with private SQLite storage. They fit per-room, per-user, per-tenant, and per-run coordination boundaries; the real design question is where the object key belongs.

Cloudflare Queues: Move Work off the Request Path into Retryable Batches

Cloudflare Queues is the message queue beside Workers: producers enqueue slow work, consumers process it with batching, ack/retry, delays, and DLQs. It is good for single-step background jobs; durable multi-step state belongs in Workflows.

CodeQL: Extracting a Codebase into a Database and Querying Data Flow

CodeQL builds a code database with language extractors, then queries syntax, types, calls, control flow, and data flow; its depth depends on models and carries extraction and query-maintenance costs.

Convex: Building Reactive TypeScript Backends with Queries, Mutations, and Actions

Convex combines typed backend functions, a transactional document database, and reactive query subscriptions; correctness depends on separating deterministic mutations from side-effecting actions.

Coolify: Control Planes, Docker Servers, and the Self-Hosted PaaS Boundary

Coolify controls Docker, proxies, and resources on your servers over SSH; deployment gets easier, but OS, security, capacity, data backup, and recovery remain yours.

techdeep-dive

CopilotKit Explained: Bring Agent State, Tools, and Human Approval into React

CopilotKit is more than a chat box. Its React components, AG-UI events, shared state, and interrupt flows connect an agent's execution to an existing product interface.

CoreWeave: An AI Cloud Built from Kubernetes, GPU Fabric, and Storage

CoreWeave is more than rented GPUs: it combines Kubernetes, GPU networking, storage, and inference into AI infrastructure, while platform engineering and capacity governance remain yours.

Crusoe Cloud: From GPU VMs and Managed Kubernetes to Managed AI

Crusoe offers both Infrastructure Cloud and Managed AI: GPU VMs and clusters provide control, while serverless and dedicated inference provide higher abstractions with different responsibilities.

DeepSeek Harness (dsh): A Coding Agent Framework That Takes Everything-is-a-Plugin All the Way

DeepSeek Harness (dsh) is DeepSeek's official open-source coding agent framework, released as a v0.1 developer preview on 2026-08-13, accumulating 184,000+ stars in 9 days. Its core is the Cordis plugin kernel — model adapters, tools, agent loop, and UI are all swappable plugins. Four runtime modes, with the ability to use Claude Code and Codex as sub-agents. Web UI first, no native CLI.

Deno Deploy: Deno 2, Revisions, Timelines, and Global TypeScript Serverless

The new Deno Deploy runs application revisions on Deno 2; understand it through timelines, contexts, databases, and telemetry rather than Deploy Classic assumptions.

DigitalOcean App Platform: Managing PaaS Topology with Components and App Specs

DigitalOcean App Platform composes Services, Workers, Jobs, Static Sites, and Functions into an App, with an App Spec as the reviewable deployment contract.

DigitalOcean: A Simplified Cloud from Droplets to App Platform

DigitalOcean covers common product architectures with Droplets, DOKS, Managed Databases, and App Platform; simplicity comes from a smaller surface, not from eliminating OS, network, backup, or HA design.

Django: Python Web Applications with ORM, Admin, Auth, and Long-Term Evolution

Django combines data models, migrations, authentication, admin, forms, and security defaults into one system; using it well still requires understanding QuerySets, middleware, async boundaries, and production settings.

Dokku: Turning One Docker Host into a Git-Push PaaS

Dokku combines a Git receiver, buildpacks or Dockerfiles, process models, Nginx, and plugins for a single-host Heroku workflow; simplicity comes from narrow orchestration scope.

Dokploy: Self-Hosted PaaS with Applications, Compose, Remote Servers, and Swarm

Dokploy supports single-container Applications and Compose or Stack, while treating one host, independent remote servers, and a Swarm cluster as distinct topologies.

DuckDB: An OLAP Query Engine Inside Your Process

DuckDB is an in-process columnar OLAP database for Parquet, CSV, and DataFrames, not a conventional multi-user OLTP server for web applications.

techdeep-dive

Elasticsearch and OpenSearch: Choosing a Lucene-Based Site Search Engine

Elasticsearch and OpenSearch both build text analysis, BM25, aggregations, and vector search on Lucene, but licensing, governance, hybrid-search APIs, and managed ecosystems have followed separate paths since the 2021 fork.

Elysia: Bun-First APIs with Runtime Schemas and End-to-End Types

Elysia connects runtime schemas, TypeScript inference, OpenAPI, and Eden clients into one contract pipeline, but Bun-first performance, plugin scope, and cross-runtime compatibility still require separate verification.

Fastify: Plugin Encapsulation, JSON Schema, and Efficient Node.js APIs

Fastify is more than benchmarks: plugin scopes, hooks, decorators, and compiled JSON Schema build composable Node.js APIs with explicit request and response contracts.

Firebase: The BaaS Boundary of Auth, Firestore, Functions, and Security Rules

Firebase moves quickly because client SDKs directly access managed Auth, Firestore, and Storage; the real backend contract lives in data models, Security Rules, Functions, and cost limits.

Fly.io: The Real Boundaries of Machines, Fly Proxy, and Multi-Region Deployment

Fly.io places fast-starting Machines in chosen regions and connects them through Fly Proxy and private 6PN; cross-region state consistency remains your hard problem.

gitleaks: Secret Detection across Working Trees, Git History, and CI Diffs

gitleaks scans files or Git patches with rules, regexes, entropy, and allowlists; after a finding, revoke and rotate first rather than merely deleting a file or rewriting history.

Google Cloud Model Armor: Runtime Filters around Prompts, Responses, and Agent Tools

Model Armor can inspect prompt injection, jailbreaks, sensitive data, malicious URLs, and unsafe content at runtime; it is a probabilistic detector, not an authorization or sandbox boundary.

Google Cloud Run: Services, Jobs, and Worker Pools Have Different Container Lifecycles

Cloud Run is more than an HTTP container platform: choose request-serving, run-to-completion, or always-on pull work first, then design concurrency, identity, and scaling.

Google Kubernetes Engine: GKE Autopilot Reduces Node Operations, Not Kubernetes

GKE manages the Kubernetes control plane and Autopilot manages most node infrastructure, while workloads, policy, networking, upgrade compatibility, and cost governance remain yours.

GraphQL and Code Generator: Type Safety Comes from Schema Plus Operations

A GraphQL schema defines available capabilities; GraphQL Code Generator combines it with actual queries, mutations, and fragments to emit precise results, variables, and typed documents.

gRPC and Connect: One Protobuf Contract across Services and Browsers

gRPC generates cross-language stubs from Protobuf services. A Connect server can support gRPC, gRPC-Web, and the Connect protocol together, avoiding a translating proxy for browsers.

Hatchet: One Engine for Task Queues, DAGs, and Durable Tasks

Hatchet unifies regular tasks, DAGs, and durable tasks behind a Postgres-backed control plane. Durable tasks checkpoint at waits and child tasks, then replay deterministic orchestration code on recovery.

Hetzner Cloud: Cheap VMs Mean Owning the Entire Operating System

Hetzner Cloud offers lean IaaS through servers, networks, volumes, load balancers, and firewalls; its price advantage is real only after patches, HA, backups, egress, and on-call are counted.

Hono RPC: Infer Fetch Client Types from Route Implementations

Hono RPC exports a route's `typeof AppType` so `hc` can infer inputs, response bodies, and status codes. The shared artifact is a TypeScript type, not an independent wire schema.

Inngest: Turn Serverless Functions into Recoverable Workflows with Steps

Inngest makes steps the persistence boundary for ordinary TypeScript, Python, and Go functions. Recovery re-executes the function while memoized steps avoid repeating completed side effects.

Kamal: Docker Deployment Without a Resident Control Plane

Kamal deploys immutable images from an operator over SSH and switches traffic through kamal-proxy; it is not a scheduler and does not operate hosts or data.

Koyeb: A Global Serverless Model of Apps, Services, and Instances

Koyeb groups Services in Apps, runs revisions as Instances in selected regions, and integrates global routing, autoscaling, private discovery, and CPU or GPU compute.

Kubernetes: Container Orchestration from Pods and Controllers to Declarative Reconciliation

Kubernetes is fundamentally API objects, controllers, and reconciliation loops—not YAML; it manages workload lifecycles without solving application state, data consistency, or organizational governance.

Kysely: A Type-Safe TypeScript Query Builder That Keeps SQL Visible

Kysely derives query results from database types while preserving SQL and escape hatches, but teams must still keep migrations, the live schema, and generated types aligned.

Lambda Cloud: GPU Compute from One VM to Multi-Node AI Clusters

Lambda Cloud provides on-demand GPU VMs and 1-Click Clusters; it offers direct AI compute environments rather than automatically solving training, serving, and MLOps.

Liveblocks: Collaboration Backends with Presence, Storage, Comments, and Notifications

Liveblocks packages rooms, presence, conflict-free storage, comments, and notifications as managed product primitives; adoption still requires alignment with existing auth, canonical databases, and data lifecycle.

MongoDB: Document Flexibility Moves the Cost to Data Boundaries

MongoDB fits aggregate-oriented data and evolving documents; the hard decisions are embedding, transaction boundaries, indexes, and shard keys.

MySQL: A Mature Relational Database, Not Merely the Popular Default

MySQL offers a mature ecosystem, InnoDB transactions, and predictable operations, but teams must still own indexes, isolation, replication lag, and migrations.

NATS and JetStream: Choose Ephemeral Messaging or Durable Events First

Core NATS is storage-free, at-most-once pub/sub. JetStream adds streams, consumers, acknowledgments, retention, and replay. They share subjects but expose fundamentally different reliability contracts.

Nebius AI Cloud: A Full Platform for GPU Clusters, Managed Kubernetes, and Serverless AI

Nebius combines GPU VMs and clusters, Kubernetes, Slurm, storage, and Serverless AI; choose the responsibility layer before comparing hardware and price.

Neon vs Turso: Do Not Call Both Serverless Databases Managed Postgres

Neon is serverless PostgreSQL; Turso Cloud currently follows the libSQL and SQLite-compatible path. Their compatibility boundaries are fundamentally different.

NestJS: A Node.js Architecture Framework of Modules, Dependency Injection, and Request Lifecycles

NestJS is valuable not for decorators alone, but for Modules, Providers, DI, and a predictable pipeline across HTTP, GraphQL, WebSocket, and microservice architectures.

Netlify: A Web Platform for Atomic Deploys, Functions, and Edge Functions

Netlify centers on atomic deploys and previews, then adds Functions, Edge Functions, Blobs, and Database; each runtime and state layer has distinct consistency and limits.

ngrok: From Localhost Tunnels to Controlled Global Ingress

ngrok is an agent-initiated reverse proxy and ingress, not a VPN for the whole machine; public endpoints still need explicit authentication, traffic policy, and data boundaries.

Nhost: A BaaS of PostgreSQL, Hasura GraphQL, Auth, and Storage

Nhost uses PostgreSQL as the source of truth, Hasura to generate GraphQL, and connects Auth claims, role permissions, Storage, and Functions into one platform.

OMP 2 (Oh My Pi 2): From Pi Fork to Full Rust Rewrite as an Independent Coding Harness

OMP 2 is no longer a Pi fork. The entire codebase has been rewritten from scratch in Rust, with ~41 crates covering a custom bash engine, GPU-accelerated GUI, embedded CPython 3.14t, gRPC transport, and Kokoro-82M TTS. Currently in pre-release with no stable version yet.

openapi-typescript: Turn OpenAPI into Runtime-Free Types and a Fetch Client

openapi-typescript converts OpenAPI 3.0 and 3.1 into pure TypeScript types; openapi-fetch then infers methods, literal paths, parameters, and response unions from that schema.

Opencode 2: The Cost of Swapping Bun for Node, Tauri for Electron, and Rebuilding the Entire API

Opencode 2 is a major rewrite led by Anomaly (Dax Raad). Runtime migrated from Bun to Node.js (memory issues), desktop from Tauri to Electron (WebKit perf and Node integration), v1 API intentionally incompatible. New: multi-tab parallel sessions, persistent backend service, HTTP API + SDK. Currently beta, stable estimated ~September 2026. ~200K stars.

OpenStack: Not a Virtualization UI, but Cloud Services You Operate Continuously

OpenStack combines Keystone, Nova, Neutron, Glance, Placement, Cinder, and other services into multi-tenant IaaS; adopting it means staffing a cloud platform team, not finishing an installation.

Oracle Cloud Infrastructure: Start with Compartments, VCNs, and Fault Domains

OCI is a full hyperscale cloud; architecture starts with tenancy and compartment IAM, region/AD/fault domains, and VCNs before selecting Compute, OKE, databases, and storage.

oRPC: Put End-to-End Type Safety and OpenAPI on the Same Path

oRPC supports implementation-first and contract-first APIs, offers an RPC client, and can expose the same router through OpenAPI 3.1.1 HTTP endpoints.

OVHcloud: Combining Public Cloud, OpenStack, vRack, and Dedicated Servers

OVHcloud combines Public Cloud, OpenStack APIs, Managed Kubernetes, vRack, and dedicated or private cloud; that flexibility also creates more networking and responsibility boundaries.

Oxc and Oxlint: Speed Comes From Reconnecting Parsing, Types, and Rules

Oxlint has grown from a fast ESLint companion into a standalone linter with type-aware rules and JavaScript plugins. Migration depends on rule, framework-file, and plugin compatibility—not only a 50–100x benchmark.

techdeep-dive

Pagefind Explained: Full-Text Search for Astro Without a Search Backend

Pagefind scans static HTML after an Astro build and ships its index with a WebAssembly search runtime; the browser fetches only the index chunks required by a query, so no search server is needed.

PartyKit: Turning Collaboration Rooms into Stateful Edge Servers

PartyKit concentrates WebSocket coordination in room-keyed stateful servers for multiplayer and presence, while durable documents, authorization, hibernation, and platform ownership still need explicit design.

Pi v2: AgentHarness API Goes Stable, Earendil Incorporates — Minimalism Enters Its Next Chapter

Pi v0.84.0 (2026-08-06) promotes the AgentHarness v2 API to stable. Lane-based v4 Session model makes operations durable and interruptible. CBOR replaces JSON, Unix sockets replace HTTP. Earendil Inc. (Armin Ronacher's PBC) behind it has secured initial funding. 95.4K stars, still MIT, still minimal.

PocketBase: A Single-Binary Backend with SQLite, Auth, and Realtime

PocketBase packages SQLite, collections, Auth, file storage, SSE realtime, and an admin UI into a small executable; deployment is easy, but single-host and pre-v1 compatibility limits matter.

Promptfoo Red Team: Turning Prompt Injection, Tool Misuse, and Data Leaks into Regression Tests

Promptfoo plugins generate risk probes, strategies transform attacks, targets execute the system, and graders judge outcomes; useful red teams exercise the full agent application rather than only a foundation model.

Protobuf and Buf: Put Cross-Language Schema Linting, Codegen, and Compatibility in CI

Protobuf defines binary messages with stable field numbers; Buf adds modules, linting, remote plugins, generation, and breaking-change checks to govern those schemas.

Proxmox VE: Combining KVM, LXC, Clusters, Ceph, and Backups On-Premises

Proxmox VE integrates VMs, containers, clusters, HA, storage, and backup; it simplifies virtualization management while hardware, quorum, networks, capacity, and DR remain yours.

Pulumi: Infrastructure as Code in TypeScript, Python, Go, and Other Languages

Pulumi lets general-purpose programs register cloud resources for a deployment engine, providers, and stack state to preview and update; greater language power demands stronger abstraction discipline.

RabbitMQ: Express Routing with Exchanges, Preserve Work with Quorum Queues

RabbitMQ's strength is routing through exchanges, bindings, and queues. For replicated work, default to quorum queues and combine publisher confirms, manual acknowledgments, and idempotent consumers.

Railway: Deploying an Application Topology with Projects, Services, and Environments

Railway is not merely one-click deployment; it puts container services, environments, variables, and private networking into one operable application project.

Self-Hosting Inference with Ray Serve: Python Service Graphs, GPU Scheduling, and Autoscaling

Ray Serve is a distributed serving layer on Ray. Deployments and handles compose Python service graphs, while replicas, CPU/GPU scheduling, autoscaling, and model multiplexing handle orchestration; it complements rather than replaces vLLM or SGLang.

Redis Streams: An Append-Only Log and Consumer Groups for Reliable Redis Messaging

Redis Streams stores replayable entries with `XADD`; consumer groups add a Pending Entries List and `XACK`. It is more durable than Pub/Sub, but it does not automatically become Kafka.

Redpanda: Kafka API Compatibility Still Requires Broker Migration Testing

Redpanda reimplements a Kafka-compatible event log with C++/Seastar, thread-per-core execution, and a Raft group per partition. Client compatibility is broad, but operational and edge semantics still require testing.

Render: A Full PaaS Topology of Web, Private, Worker, and Cron Services

Render's value exceeds turning a repository into a URL: distinct service types model public HTTP, private listeners, queue workers, cron, data stores, and Blueprint infrastructure.

Renovate: Turning Dependency Updates into a Governed Continuous Process

Renovate is more than an update-PR bot: packageRules, grouping, schedules, minimum release age, and automerge policy determine update speed, review noise, and supply-chain exposure.

Replicate: Turn Model Versions into Prediction APIs Instead of Renting GPUs

Replicate abstracts GPUs behind versioned models, predictions, Cog, and deployments; integrators still own version pinning, async workflows, webhook verification, data persistence, and spending limits.

Restate: Put Journals, Durable State, and Service Calls in One Execution Model

Restate journals operations and results, then re-executes handlers while skipping completed work. Virtual Objects and Workflows add keyed state, single-writer semantics, and long-lived coordination.

RunPod: GPU Pods and Serverless Endpoints Are Different Products

RunPod Pods fit interactive and persistent GPU work, while Serverless fits queued or load-balanced inference; choosing incorrectly mixes persistence, cold starts, and retry semantics.

Scaleway: European Cloud Instances, Kapsule, Serverless, and Managed Data

Scaleway now spans compute, Kapsule, serverless, databases, storage, AI, and IAM rather than only low-cost VMs; maturity and integration still require per-region verification.

techdeep-dive

Scrapy Deep Dive: A Self-Hosted Crawler from Engine to Pipeline

Scrapy separates crawling into the Engine, Scheduler, Downloader, Spider, Item Pipeline, and middleware; it fits high-volume, rule-driven HTTP crawling where you need control over scheduling, throttling, retries, and storage.

techdeep-dive

Selenium Deep Dive: Browser Automation from WebDriver Sessions to Grid

Selenium drives real browsers through standardized WebDriver sessions, making it useful for cross-browser workflows, existing test assets, and remote Grid capacity; it can render JavaScript applications, but it does not guarantee bypassing CAPTCHAs or other anti-automation controls.

Semgrep: Encoding Security Policy as Readable, Tested Static Analysis

Semgrep lets teams express SAST policy with source-like patterns and taint rules; rule quality depends on positive and negative tests, framework modeling, and exception lifecycle.

Server-Sent Events: Recoverable One-Way Push over HTTP Event Streams

SSE sends text/event-stream over ordinary HTTP, with browser reconnection and Last-Event-ID; simple transport is not durable unless the server retains events behind that cursor.

Self-Hosting Inference with SGLang: RadixAttention, OpenAI APIs, and Multi-GPU Serving

SGLang is an inference engine for generative models. RadixAttention reuses KV cache across shared prefixes, while OpenAI-compatible APIs, structured output, and multi-GPU parallelism support production LLM serving; it is not a complete product backend.

Sigstore and SLSA: Verifying Who Built an Artifact with Which Process

Sigstore provides identity-bound signing, short-lived certificates, and transparency logs; SLSA describes trustworthy build provenance. They protect only when admission verifies identity, issuer, digest, and build expectations.

Snyk: Connecting SCA, SAST, Container, and IaC Findings to Development

Snyk maps code, open-source, container, and IaC findings to projects, remediation paths, and developer workflows; successful adoption depends on baselines, ownership, and executable policy.

Socket.dev: Blocking Malicious Package Behavior at Dependency-Diff Time

Socket.dev goes beyond CVEs by analyzing install scripts, obfuscation, network and shell access, and ownership changes when packages enter a dependency diff.

Socket.IO: Realtime Events with Rooms, Acknowledgements, and Recovery

Socket.IO is an event protocol and runtime over WebSocket or long-polling, not a native WebSocket compatibility layer; ordering, arrival, recovery, and horizontal scaling need separate designs.

Speakeasy: Manage Multi-Language SDKs with OpenAPI Overlays and Workflows

Speakeasy records OpenAPI, Overlays, targets, and generator versions in `.speakeasy/workflow.yaml`, enabling local or CI generation, compilation, and publishing of multi-language SDKs.

SST v3: Composing Full-Stack Cloud Apps with Components, Link, and Dev Mode

SST v3 composes cloud resources through TypeScript, high-level app components, Pulumi and Terraform providers, and resource linking; it is application-first IaC, not a hosted PaaS.

Stainless: Continuously Generate Publishable Multi-Language SDKs from OpenAPI

Stainless uses OpenAPI plus its configuration to generate multi-language SDKs, docs, CLIs, and MCP servers. Its value is a continuous preview, publishing, and upgrade pipeline rather than one-off codegen.

techdeep-dive

Stytch Deep Dive: From B2C Login and Sessions to B2B Organizations and Authorization

Stytch is an API-first managed identity platform: choose the Consumer or B2B model, converge authentication factors into sessions, then enforce organization and RBAC boundaries on the server.

Teleport: Infrastructure Access with Short-Lived Credentials, RBAC, and Session Audit

Teleport is a protocol-aware infrastructure access platform: its Auth Service signs short-lived credentials, while Proxies and Agents mediate SSH, Kubernetes, database, and app access with audit evidence.

Terraform: The IaC Workflow of Providers, Plans, Applies, and State

Terraform compares provider schemas, configuration, state, and real APIs to create a plan and apply it; controlled state and change workflow matter more than HCL itself.

NVIDIA Triton Inference Server: Multi-Framework Models, Dynamic Batching, and Pipelines

Triton Inference Server serves TensorRT, ONNX, PyTorch, and other models through consistent HTTP and gRPC APIs. Its defining tools are the model repository, dynamic batching, instance groups, and ensembles—not LLM-specific KV-cache scheduling.

tRPC: Connect Client and Server Contracts through TypeScript Inference

tRPC lets a client reference the server router type without a separate schema or code generation. Version 11 also has an official, alpha-stage OpenAPI 3.1 generator.

ts-rest: A TypeScript Contract-First API that Preserves REST Semantics

ts-rest describes methods, paths, status codes, and schemas in a shared contract, providing end-to-end server and client types without a code-generation step.

Twingate: Identity-to-Resource Access Instead of Whole Networks

Twingate uses clients, connectors, a controller, and relays to narrow user authorization to specific resources; it is managed ZTNA rather than a general peer-to-peer overlay.

TypeScript 7: A Native Go Rewrite Makes Type Checking Fast, but Migration Goes Beyond tsc

TypeScript 7 rewrites the compiler and language service in Go. The order-of-magnitude gains come from native execution and shared-memory parallelism, while old APIs and compiler options define the migration cost.

techdeep-dive

Typesense Site Search: Full-Text Search You Can Treat as a Product Feature

Typesense is a search server centered on instant, typo-tolerant keyword search. Collection schemas, field weights, facets and sorting are enough to build site search, with a choice between self-hosting and Typesense Cloud; Chinese fields need locale: zh, and domain terminology may require custom segmentation.

Vercel: Frontend Cloud Gains Its Advantage from Framework-Aware Deployments

Vercel binds framework build output, previews, CDN caching, and Functions into deployments; deep integration adds speed while defining runtime, locality, cost, and portability boundaries.

Vite 8 and Rolldown: One Pipeline Replaces the Two-Bundler Split

Vite 8 replaces the esbuild-for-development and Rollup-for-production split with Rolldown. Speed is the visible result; consistent bundler semantics across development and production are the deeper change.

Vitest: Share Vite Configuration Without Mistaking Component Tests for E2E

Vitest's advantage is not merely a familiar Jest-style API. Tests share Vite transforms, aliases, and plugins with the application; Browser Mode adds real-browser confidence but does not replace full E2E testing.

Vultr: Choosing Between Cloud Compute, VKE, and GPUs

Vultr spans VMs, bare metal, GPUs, VKE, databases, and storage; breadth and regional choice help, but product availability alone does not create an integrated architecture.

WebSocket: Duplex Connections Are the Start; Protocol and Backpressure Are the Work

WebSocket supplies a duplex message transport, not auth renewal, schemas, acknowledgements, replay, rooms, or backpressure policy; those layers determine production correctness.

WireGuard: Minimal VPN Tunnels with Cryptokey Routing

WireGuard is a small, explicit layer-3 encrypted tunnel that binds public keys, peers, and AllowedIPs, but it does not supply identity, device management, or a policy control plane.

techdeep-dive

WorkOS Enterprise Auth: Growing from AuthKit to SSO, SCIM, and Audit Logs

WorkOS lets a B2B SaaS add SAML/OIDC, SCIM, and audit exports around one organization identity model; application authorization and data governance remain your responsibility.

Yjs and CRDTs: Concurrent Editing as Commutative, Replayable Document Updates

Yjs uses shared types and commutative, associative, idempotent binary updates to converge concurrent edits; it does not prescribe transport or include authorization, persistence, or domain conflict resolution.

ZeroTier: Virtual Networks with Controllers and Flow Rules

ZeroTier places devices on a managed virtual L2/L3 network, attempts peer-to-peer transport, and uses a controller to publish membership and policy; it resembles software-defined networking more than a single tunnel.

zizmor: Finding Template Injection and Token Risks in GitHub Actions

zizmor performs domain-specific static analysis on workflow and action YAML for template injection, broad permissions, artifact credential leaks, and unpinned uses; it does not analyze called shell scripts.

Zodios: Build an Axios Type-Safe Client from Zod Endpoint Definitions

Zodios uses a central Zod endpoint definition for Axios client types, runtime validation, and aliases, with optional Express and OpenAPI packages.

techdeep-dive

Zyte Deep Dive: From Scrapy Development and Anti-Bot Fetching to Scrapy Cloud

Scrapy owns crawl flow and data models, Zyte API handles fetching, browsers, and anti-bot infrastructure, and Scrapy Cloud adds deployment, scheduling, and output; each layer can be adopted independently.

Agent Plugins 1.0: OpenAI, Google, and AWS Unite to Standardize AI Agent Extensions

Agent Plugins 1.0 is a packaging format that bundles Agent Skills (markdown instructions) and MCP server configs into a single directory, loadable by ChatGPT, Cursor, GitHub Copilot, Kiro, and VS Code. It's not a new protocol — it's the wrapper above protocols. Vercel initiated it, OpenAI/AWS/Microsoft/Cursor co-authored it, and Google joined on launch day. Anthropic isn't on the governance board, but MCP is a core primitive of the spec.

aiguide

AgentQL Complete Guide: Semantic Web Extraction and Playwright Automation

AgentQL replaces brittle CSS and XPath selectors with queries shaped like the data you want: `query_data` returns structured values, while `query_elements` returns interactive Playwright locators. The public Starter plan lists 50 free API calls per month, but its payment, hard-stop, and remote-browser reset rules still need to be verified in Billing.

aiguide

Apify Complete Guide: How Actors, Tasks, Schedules, and Datasets Form a Scraping Platform

Apify is not a single crawler. It packages scraping programs as Actors, saves reusable configurations as Tasks, triggers them with Schedules, and delivers results through Datasets. It fits teams that do not want to operate queues, schedulers, and workers, but Actor fees, compute, proxies, storage, and transfer all draw from the same platform budget.

aiguide

Browser Use Complete Guide: The Agent Loop Behind Browser Automation

Browser Use combines browser state, model decisions, and actions such as click, type, and extract into a repeatable loop. The open-source package favors custom tools and execution control; Cloud manages browsers, profiles, proxies, and concurrent work.

aiguide

changedetection.io Complete Guide: Selectors, Notifications, and Browser Steps

changedetection.io is a web-change signal layer: narrow the monitored content, suppress noise, and notify downstream systems only when a meaningful change occurs. It is neither a search API nor a crawler replacement.

Composio: Who Holds Every User's Token When Your Agent Connects a Hundred SaaS Apps

This site covers MCP thoroughly but has never written about the layer underneath it: when your agent acts for ten thousand end users reading their own Gmail, whose database holds those refresh tokens, who rotates them, who revokes them. Composio is currently the most complete answer — MIT-licensed SDKs, a commercial hosted execution and OAuth layer. It claims 1,000+ toolkits; the managed-auth page actually lists 121 with a Composio OAuth app and 96 that require your own credentials. New pricing effective 2026-08-15: 100K free tool calls, $29/mo Pro. This post takes the authorization model down to an operational level and draws the line between wiring up MCP servers yourself and buying an integration platform.

Seven Answers to a Full Context Window, and No Consensus

Chroma's controlled study shows that even when it fits, a full context degrades performance. Coding agent vendors have landed on seven different responses: compact, hand off, prune, defer loading, isolate, train it into the model, or change the unit of work. Amp removed /compact outright, Atlassian argues summarization should be a last resort, and Cursor's A/B test measured a 46.9% token reduction. The three real disagreements come down to what each team is measuring.

aiguide

Crawl4AI Complete Guide: From Markdown Crawling to Structured Extraction

Crawl4AI handles retrieval after a URL is known: use JsonCssExtractionStrategy for stable DOMs, and switch to LLMExtractionStrategy only when extraction needs semantic judgment or must tolerate irregular layouts.

CrewAI: Organizing Multi-Agent Collaboration Through Role-Playing

CrewAI (GitHub 57.4k stars, MIT, PyPI 11.6M weekly downloads) defines agents by role, goal, and backstory, then groups them into crews for collaboration. Unlike LangGraph's graph-first and MAF's workflow-first approach, CrewAI is team-first — you don't draw nodes and edges, you describe who's on the team and what each person does. It fully removed its LangChain dependency in late 2024 and is now a standalone framework. The commercial side splits into the open-source package and AMP, a managed platform adding visual building, deployment, tracing, and compliance.

Exa: Neural Search Built for Agents, Not People

Exa turns every indexed web page into an embedding and retrieves by vector similarity instead of keyword matching. Official pricing as checked on 2026-08-21: $7 / 1k requests for /search (first 10 results included), $1 / 1k pages for /contents, $12–15 / 1k for the deep tiers, with $20 in free credits for new accounts. This blog's CLAUDE.md puts Exa first among cloud fetch tools, 16 of its 38 skills reference it directly, and only four existing posts mention it in passing — with zero dedicated posts. This is that post.

aiguide

Firecrawl Complete Guide: Choosing Scrape, Crawl, Map, and Structured Extraction

Firecrawl puts single-page scraping, site discovery, whole-site crawling, and JSON extraction behind one API. Cloud removes browser, proxy, and worker operations; self-hosting gives infrastructure control, but not the complete Cloud feature set.

aideep-dive

Choosing Free Search, Scraping, and Browser APIs: Recurring Quotas, Trials, and Self-Hosting

Free access is not one model: recurring allowances, balance top-ups, rate-limited access, one-time credits, and self-hosting have different steady-state costs.

Linkup Search API Guide: From standard and deep to Structured Output

Linkup separates search depth from response shape: start most agent queries with standard + searchResults, move to deep only for multi-step browsing, and treat the monthly $20 as a balance refill rather than a new $20 grant.

LlamaIndex Is Not a RAG Framework Anymore, and Old Tutorials Won't Tell You

LlamaIndex (51,775 GitHub stars, MIT, verified 2026-08-21) has moved its center of gravity from indexing to Workflows: the standalone llama-index-workflows package pulls 2.81M weekly PyPI downloads, more than the 1.97M of the llama-index umbrella package itself. This post covers the core abstractions, the trade-off against hand-rolling a pipeline, and a hands-on test of its defaults on Traditional Chinese text — at the same chunk_size=1024, English fits 4,645 characters and Traditional Chinese only 1,332. Plus one fact you need before choosing: the TypeScript port is archived and unmaintained.

aideep-dive

Meilisearch Complete Guide: Indexing, Chinese Search, and Tenant Security

Meilisearch turns application data into fast, typo-tolerant full-text search; the hard parts are index settings, asynchronous tasks, Chinese tokenization, access filters, and tested recovery.

Microsoft Agent Framework: After the Merge, Who Does the Name AutoGen Point To?

Microsoft merged Semantic Kernel and its own AutoGen into Microsoft Agent Framework, which hit 1.0 GA on 2026-04-02 for .NET and Python (Go is still public preview). The absorbed autogen-agentchat has not shipped since 2025-09-30. But AG2, the fork on the original authors' side, never merged — it shipped 1.0.2 six days ago, and `pip install autogen` gets you AG2, not Microsoft. This post covers MAF's abstractions, the migration clock, and how to read the tangle of names.

MIT 6.S191 Guide: Nine Lectures and Three Labs Are Public, but the Full Path Still Uses Three External Services

MIT 6.S191's 2026 edition publishes nine lecture videos, slides, three software labs, and solutions, making it an A3 self-study course. The supplied path still depends on Google/Colab, Comet, and OpenRouter for Lab 3, while unaffiliated learners do not receive MIT credit, project feedback, or API credits.

Modal: The Layer Your Inference Engine Runs On — and When the Premium Isn't Worth It

Modal is a per-second-billed serverless GPU platform that also treats agent sandboxes as a first-class primitive (company-reported: over 1 billion sandboxes launched, more than a third of revenue). The selection question isn't how convenient it is — it's your GPU utilization. Verified 2026-08-21: Modal's A100 80GB works out to $2.50/hr against RunPod's $1.59/hr for the same card, so above 64% utilization renting your own is cheaper. But on the same day, H100 SXM is $3.95/hr on Modal against $3.99 on Lambda — on that card the premium is gone.

aideep-dive

Qdrant Complete Guide: Collections, Hybrid Search, and Self-Hosted Operations

Qdrant is not just a place to store embeddings: define the vector schema, index frequently filtered payload fields, then add dense+sparse queries, tenant boundaries, snapshots, and monitoring to build an operable retrieval service.

aiguide

Scrapling Complete Guide: From Adaptive Selectors to Concurrent Spiders

Scrapling puts HTTP, Playwright browsers, CSS/XPath extraction, and a Spider API behind one Python interface. Adaptive selectors save element properties and relocate a target by similarity after a layout change, but the output still needs validation.

aiguide

SearXNG Complete Guide: Engine Tuning, JSON API, and Self-Hosted Operations

SearXNG is a metasearch engine, not a crawler, and it does not own a web-wide index. Based on the official 2026.8.20 documentation, this guide covers Compose installation, settings.yml, engine selection, the JSON API, and empty-result diagnosis.

Build Your Own Search Backend: SearXNG + Crawl4AI, From Zero to Claude Code

Three components, one job each: SearXNG finds, Crawl4AI reads, and a thin layer of your own code glues them together. Three defaults will stop you cold — `formats` only emits HTML, `secret_key` ships as the literal string `ultrasecretkey`, and turning on the limiter blocks your own code. Also, the official install path changed in March 2026, so most tutorials online point at a repository that is now archived.

Tavily and Exa Can't Be Self-Hosted: How to Build Your Own

Tavily and Exa are cloud-only APIs and can't be self-hosted. What you can assemble instead is SearXNG (269 upstream engines, 82 on by default) plus Crawl4AI (78.8k stars, Apache-2.0), and the ready-made Tavily-compatible wrappers are all still double-digit-star solo projects you should not depend on. But SearXNG has no index of its own, and running it from a datacenter IP gets you empty results — those two facts decide whether self-hosting is worth it.

Stanford CS124: Numbered 100, Four Prerequisites Written Into the Catalog, and Not Offered at All Next Year

CS124 is the first course in Stanford's NLP branch. Its textbook is Jurafsky's own Speech and Language Processing, free online, and all nine assignment repos are public. But a banner sits on the course homepage: it will not be taught at all in AY 2026–27. And the chapter numbers the syllabus points at no longer match the August 2026 textbook.

Stanford CS221: The AI Intro Course Whose Prerequisites Field Reads CS103, CS106B, CS109, CS161

CS221 lays AI out along one axis, and reflex models — deep learning — sit in the lowest slot, with states, variables and logic above them. When Percy Liang took over in Autumn 2025 he replaced the slides with runnable Python and wrote 'Cut constraint satisfaction problems :(' into the source of the first lecture — yet ExploreCourses and Stanford Online both still advertise constraint satisfaction as a course topic. The project has gone from 20% of the grade in 2019 to extra credit only.

Stanford CS224N: Open the 2019 Syllabus and Transformers Are Still Lecture 14

CS224N has kept every course website since 2000 online. In Winter 2019, Transformers were lecture 14, taught by a guest. In Winter 2026 they are lecture 5, and every lecture after that assumes you already know them. The machine translation assignment is gone; assignment 3 now has you code a decoder-only Transformer from scratch, with pytest suites that run on your laptop.

Stanford CS224U: The Course Site Stopped in Spring 2023, but You Can Clone the Whole Thing

CS224U's teaching material isn't a slide deck — it's an Apache-2.0 GitHub repo holding the lecture notebooks, all three assignments, and the grading document for the final project. But the on-campus course has skipped three straight academic years since Spring 2023, and ExploreCourses briefly put it back on the books for Spring 2026-27, then dropped that section again by 29 September 2026. The official description still lists relation extraction and semantic parsing; the 2023 syllabus covers neither. And the data-loading cell in the first assignment breaks in a fresh environment today, on a Hugging Face compatibility change.

Stanford CS224V: Renamed to Agentic AI in 2026, but What It Teaches Is Formal Methods Against Hallucination

CS224V only became Agentic AI in the 2026–2027 catalog, and the rename changed nothing underneath: the course still translates natural language into formal semantics and constrains agents with SMT solvers and knowledge graphs instead of wiring frameworks together. Seven of the eleven mandatory readings come out of the instructor's own lab. Every slide deck is public, and the course site says outright that they are deliberately incomplete.

Stanford CS224W: Every Assignment Runs in Colab, but the Biggest Slice of the Grade Is Closed to Self-Learners

All six CS224W Colabs download and run today, and the first one needs only NetworkX — no PyG install at all. But the exam is 35% of the grade, the largest single piece, and it's an in-person closed-book sitting. The public recordings stop at 2021 and cover none of the current syllabus's second half: graph transformers, relational deep learning, LLM+GNN.

Stanford CS228: The Prerequisites Are One Sentence About Probability and Algorithms — But the Course Hasn't Run in Two Years

CS228's official prerequisite is a single line — 'basic probability theory and algorithm design and analysis' — with no named course. But ExploreCourses shows it was last offered in Winter 2024, and the next slot, Winter 2027, still has a blank instructor field. What a self-learner can actually get is cs228-notes: 16 chapters, complete, last touched in June 2025.

Stanford CS229: Notes Rewritten Every Year, Public Problem Sets Frozen at 2020, and an Official Self-Test From 2008

The three things you need to self-study CS229 run on three different clocks. The lecture notes are 278 pages and were recompiled in August 2026. The newest problem sets you can download are from summer 2020. The self-assessment Stanford Online tells you to attempt before enrolling is a PDF created in 2008. Seventeen lectures from spring 2026 are public, and the last three are mislabeled.

aideep-dive

Stanford CS25 V6: A Course Called Transformers United Whose First Two Talks Weren't About Transformers

CS25 is Stanford's 1-unit seminar where attendance is the only homework and anyone can audit. Of the nine talks in the Spring 2026 season, the three worth your time are Albert Gu on the inductive biases of SSMs vs Transformers, Charles Frye on serving inference across thousands of GPUs, and Victoria Lin on what native multimodality still hasn't solved.

Stanford CS329Z: No Frameworks, Just One Chat-Completion Call — Grow an Agent Harness from Scratch

CS329Z is a new three-unit agent engineering course debuting at Stanford in Autumn 2026. Its first homework bans every agent framework: one chat-completion call plus code you write yourself, grown on a real corporate email archive from a RAG pipeline into an agent harness with tools, a terminal, memory and a human in the loop. DSPy is still in the lectures, but no longer in the homework. The course site lives in a public GitHub repo, and its commit log records every syllabus revision: three assignments cut to two, peer review grown into a fifth of the grade, and the project topic changed from fixed to open.

Stanford CS336: The Lectures Are Runnable Python, and From Assignment 2 On You Pay for the GPUs

Of the seventeen regular CS336 lectures, only nine are executable Python programs; the other eight are PDF slide decks — and the split falls exactly along the two instructors. Assignment 1's handout carries eight 'Low-Resource Tips' for finishing it on a laptop. Assignments 2 through 5 carry none. The course page lists the hourly price of a B200; the handouts list how many B200 hours each problem needs.

aiguide

Tavily Search API Complete Guide: Search, Extract, Map, and Crawl

Tavily exposes Search, Extract, Map, and Crawl through one web API for agents. The free plan includes 1,000 credits per month; basic, fast, and ultra-fast Search cost 1 credit each, while advanced costs 2.

vLLM: The Default Choice for Self-Hosted Inference — and When It's Over-Engineering

vLLM is the de facto standard for self-hosted LLM inference (89,470 GitHub stars, verified 2026-08-21), built on managing the KV cache the way an OS manages paged memory. But the selection question isn't how fast it is — it's your GPU utilization. Using Red Hat's measured 793 output tokens/second, a fully saturated A100 costs roughly $0.70 per million output tokens; at 10% utilization that becomes $7, more than most cloud APIs.

How to Evaluate Agent Search Quality: Building a Web Retrieval Benchmark

A web retrieval benchmark must evaluate complete tasks, not HTTP 200s: 30 fixed cases across five failure strata and three live channels, measuring answers, citations, freshness, latency, cost, and unnecessary escalation. This article delivers the harness and gates, but no fabricated ranking while the three live channels remain unconfigured.

A Complete Web Retrieval Route for AI Agents: When to Use Search, Fetch, Crawlers, and Browsers

An agent should not open a browser for every web task: route first to Search or Fetch, then escalate on explicit signals such as status codes, weak content, JavaScript shells, authentication, or challenge pages, with retry, budget, cache, deduplication, and provenance constraints at every step.

Berkeley AI/ML Course Guide: From CS61A to CS288, What Can You Actually Study Online?

Berkeley has no standalone undergraduate AI degree. A workable path builds on the CS BA or EECS BS foundation, enters through either CS188's broad AI curriculum or CS189's mathematical machine learning curriculum, then branches into deep learning, NLP, vision, or reinforcement learning. Many 2025–2026 courses are A3, but the newest class, the newest stable URL, and the best self-study edition are not always the same.

learningdeep-dive

CMU's AI Degrees: The First U.S. AI Bachelor's Turned 'What Should AI Students Learn?' into Graduation Requirements

Stanford has no AI degree; AI is a track inside CS. CMU launched the first U.S. B.S. in Artificial Intelligence in 2018, divided AI into four clusters, required one course from each, and made ethics a graduation requirement. At the master's level, MSAII sits not in CS but in the Language Technologies Institute; 84 of its 195 units cover an innovation process ending in a fundable capstone. Two official-page conflicts emerged during verification: whether the AI Core has two or three courses, and whether MSAII totals 192 or 195 units.

CMU AI/ML Course Guide: The New 07-280 Core and a Public Self-Study Route

CMU's current BSAI now runs through 07-280 and 07-380 before branching into an NLP/vision core and four AI clusters. 07-380 debuted in Fall 2026, and the same semester launched a graduate-level 11-768 AI Agents course. The residual Spring 2026 materials for 07-280 and the complete 10-301/601 site already support self-study; retired 15-281 remains a useful legacy route.

learningdeep-dive

The Conference as a Content Factory: AI Engineer's Structural Advantage

AI Engineer reached 600,000 YouTube subscribers in under three years not because it mastered video production, but because it barely needs to produce videos at all: recordings from eight conferences a year create an inexhaustible supply of YouTube material. The real constraint on content creation is structure, not skill.

A Global Map of AI and CS Courses: Which Ones Can You Actually Study in Public?

This map audits AI and CS courses at Stanford, CMU, MIT, UC Berkeley, Harvard, and National Taiwan University (NTU) in 2025–2026 using four access labels: A0 for a visible catalog entry, A1 for a public syllabus, A2 for partial materials, and A3 for a self-study-ready package. A course site or YouTube playlist can exist without giving outsiders access to the current videos, assignments, or starter code.

MIT AI/ML Course Guide: Course 6-4 Is a Real AI Degree, but Its Public Materials Span Three Eras

MIT has offered Course 6-4, a formal BS in Artificial Intelligence and Decision Making, since 2022. For an outside learner, however, the current degree requirements, the 2025–2026 course sites, and the best OCW editions rarely line up. A workable route follows 6-4's programming, algorithms, linear algebra, and probability foundation, then selects among 6.S191, 6.3900, 6.4110, 6.7960, vision, and robotics according to what is actually public.

Stanford CS103: A Math Course Whose First Assignment Is Installing a C++ Compiler

CS103 teaches you how to write proofs, then teaches you what can't be proven — but the part nobody mentions is that it ships C++ programming assignments, starting with PS0: install Qt Creator. Its real asset is a shelf of homegrown 'Guide to X' handouts and a Proofwriting Checklist that graders actually deduct points against, all public. Solutions and practice exams sit behind Stanford login, and the Honor Code page explains why.

Stanford CS107: The Same Course Weights Assignments at 40% One Quarter and 20% the Next

CS107 runs from Unix and C all the way to x86-64 and writing your own malloc, across seven assignments. But line up four archived syllabi and the course stops looking like one course: assignments are worth 40% in three quarters and 20% in Summer 2026, where in-class quizzes take 40%. The resubmission policy exists only in the quarters Cain taught; Troccoli's quarter has none. The one assignment that accepts no late days is the final heap allocator. And what blocks a self-learner isn't the autograder — it's that every starter repo lives on AFS.

Stanford CS109: A Probability Course That Turned "How to Read This Lecture With an LLM" Into Official Coursework

Every lecture in CS109's Summer 2026 offering ships with an official LLM Learning Guide — six concepts, a Learn prompt and a Test me prompt for each, written week by week across the quarter for a total of 23 PDFs. The same course's honor code Rule 4 forbids asking an LLM to solve your homework, and 65% of the grade sits in proctored exam rooms. Those two facts are halves of one design.

Stanford CS111: Nine Assignments Build an Operating System, and the Exams Don't Test Them

CS111's nine assignments run from lambdas to crash recovery in a journaling file system. Reading the site page by page turns up three things the syllabus blurb never mentions: assignment 3 is the point of no return, because assignment 4 compiles your assignment 3 code; a whole block of the final exam asks for definitions of ethics terms, and the public practice sheet ships with answers; and pasting your own code into an AI tool to ask about it is written down, in plain words, as an Honor Code violation.

Stanford CS161: The Algorithms Course That Lists Writing Clearly as Its Third Learning Goal

The first slide of CS161 names three goals: design, analysis, communication. The third one is why handwritten homework scores zero and why solutions have to read like a memo to a colleague. Of the eight problem sets, HW2 is the wall. The lecture notebooks exist to show that timing runs can't tell you which algorithm is faster. And the summer offering is a completely different course wearing the same number.

Stanford CS161 Lecture 1: Why Algorithm Analysis Starts with Karatsuba Multiplication

Splitting two n-digit integers in half still creates four recursive products and leaves the runtime at n². Karatsuba reconstructs the cross term with (a+b)(c+d)-ac-bd, cuts the branching factor to three, and reaches roughly n^1.585.

Stanford CS161 Lecture 2: From an InsertionSort Proof to MergeSort's n log n

Lecture 2 turns 'fast' into a worst-case bound that can be proved. A loop invariant establishes InsertionSort's correctness while its worst case is n²; a recursion invariant and O(n) work per level give MergeSort O(n log n).

Stanford CS161 Lecture 3: Reading a Recursion Tree Through the Master Theorem

For T(n)=aT(n/b)+O(n^d), the central comparison is branching growth a versus per-problem shrinkage b^d. Equality makes every level equally heavy, a<b^d makes the root dominate, and a>b^d makes the leaves dominate; outside the template, use substitution.

Stanford CS161 Lecture 4: How Median of Medians Guarantees Linear-Time Selection

Selection does not require sorting. Median of medians groups elements by five, selects the median of the group medians as a pivot, and guarantees that the larger recursive side has at most 7n/10+5 elements; substitution proves O(n) worst-case time.

Stanford CS161 Lecture 5: Proving Randomized QuickSort's Expected Time

Randomized QuickSort has O(n log n) expected time on every fixed input but Θ(n²) worst-case time. The valid proof does not substitute expected subproblem sizes into a recurrence; it computes the probability that each pair is compared.

Stanford CS161 Lecture 6: Sorting Lower Bounds and Linear-Time Radix Sort

The Ω(n log n) lower bound applies to comparison sorting. When integer keys can index buckets directly, stable Counting Sort can power Radix Sort and achieve O(n) under conditions such as M≤n^c.

Stanford CS161 Lecture 7: Binary Search Trees, Red-Black Trees, and the Source of Worst-Case O(log n)

Ordinary BST operations cost O(h) and can degrade to O(n); five red-black invariants cap the height at 2 log₂(n+1), giving search, insertion, and deletion worst-case O(log n) bounds.

Stanford CS161 Lecture 8: Hashing, Collisions, and What Expected O(1) Actually Guarantees

A universal hash family only needs to keep the collision probability of every distinct key pair at most 1/n; that makes the expected bucket size below 2, yielding expected O(1), not per-operation worst-case O(1).

Stanford CS161 Lecture 9: Graph Representations, DFS, BFS, and Proofs About Search Order

DFS and BFS both scan an adjacency-list graph in O(n+m); DFS finish times produce a topological order for a DAG, while BFS layers equal exact unweighted shortest-path distances.

Stanford CS161 Lecture 10: Why Two DFS Passes Find Strongly Connected Components

Contracting each SCC always produces a DAG; first-pass DFS finish times order those components, and a second pass on the transposed orientation discovers exactly one SCC per DFS tree in O(n+m).

Stanford CS161 Lecture 11: Dijkstra, Bellman-Ford, and Two Orders of Relaxation

Dijkstra finalizes the minimum estimate and relies on nonnegative weights; Bellman-Ford repeatedly relaxes every edge, spending O(nm) to support negative edges and detect a negative cycle reachable from the source.

Stanford CS161 Lecture 12: Dynamic Programming with Bellman–Ford and Floyd–Warshall

Dynamic programming starts by defining subproblems, derives a recurrence from optimal substructure, and evaluates states in dependency order; Bellman–Ford layers by edge count, while Floyd–Warshall layers by allowed intermediate vertices.

Stanford CS161 Lecture 13: Designing Dynamic Programs for LCS, Knapsack, and Independent Set

Lecture 13 turns dynamic programming into five steps: choose a state, derive transitions, fill the table, reconstruct a solution, and then improve the implementation. LCS takes O(mn), both knapsack variants take O(nW) pseudo-polynomial time, and maximum-weight independent set on a tree takes O(|V|).

Stanford CS161 Lecture 14: When a Greedy Algorithm Turns Local Choices into a Global Optimum

A greedy algorithm is not merely 'pick what looks best.' It keeps one choice at each step and needs an exchange argument proving that the choice preserves an optimum. Lecture 14 develops that proof pattern through activity selection, weighted completion time, and Huffman coding.

Stanford CS161 Lecture 15: Proving Prim and Kruskal with the Cut Property

The heart of MST algorithms is an invariant: the selected edges remain contained in some MST. The cut property proves that every step of Prim and Kruskal is safe.

Stanford CS161 Lecture 16: Ford–Fulkerson, Residual Networks, and Max-Flow Min-Cut

Ford–Fulkerson augments through a residual network. When no path remains, residual reachability yields a cut equal to the flow, certifying max flow, min cut, and their equality.

Stanford CS161 Lecture 17: Gale–Shapley and Revocable Greedy Choices

Deferred Acceptance permits tentative choices to be revoked. Monotone proposals prove O(n²) termination and stability, with an outcome favoring the proposing side.

Stanford CS161 Lecture 18: From the Algorithmic Toolbox to LP, Coding, and ML

The finale recaps the CS161 toolbox and points toward LP duality, Reed–Solomon coding, and ML-assisted algorithms. Officially, this lecture has slides but no notes.

Reading Guide: Pick What Most People Use — the Other Five Criteria Are Tie-Breakers

The primary criterion has not changed: it is still adoption — and AI makes it matter more, not less, because more users means more training data means higher agent accuracy. The five criteria this series collects (machine-readable docs, types, whether the source is in your repo, data shape, machine-callability) are for breaking ties when adoption is comparable, or for costing out what picking the less popular option will charge you.

AI SDK Message Parts: The Data Skeleton of a Conversation UI

The AI SDK splits an AI message into a parts array — text, reasoning, source-url, tool-* — each an independent typed fragment (introduced in v5, unchanged since). That data structure dictates how modern AI conversation UIs are written: render by switching on part.type, handing each fragment to its component. This post unpacks the design logic of the parts model, useChat's streaming behavior, and how it became the foundation for component libraries like AI Elements.

Drizzle ORM: A SQL-First Database Access Layer for TypeScript

Drizzle ORM is a SQL-first TypeScript ORM — its query builder reads like SQL, so queries written by agents are auditable in diffs. Zero dependencies, ~7.4 KB gzipped, native support for edge databases like Cloudflare D1, Neon, and Turso. Still at version 0.45.2 with no 1.0, yet weekly downloads have reached 16.9 million — surpassing Prisma's 13.8 million.

llms.txt: The Copy of Your Docs Written for Machines

llms.txt is a convention proposed by Jeremy Howard on 2024-09-03 (the spec is now at v2): a Markdown index at your site root written for LLMs. Hand-tested across six frontend docs sites: TanStack, shadcn, Zustand, AI SDK, and Next.js all ship it; React Router is the lone 404. The companion llms-full.txt (full-text version) is live at Anthropic, Cloudflare, and others. This post covers the spec, who uses it, and why it has started to influence library selection.

shadcn Registries and MCP: The Third Way to Distribute Components

Component distribution used to offer two roads: npm packages (black-box dependencies) or manual copy-paste. The shadcn registry standardizes a third — components described as JSON with embedded source and dependencies, installed by CLI straight into your repo as your own code. Anyone can host a registry (AI Elements is one), and the official MCP server lets AI agents browse and install components directly.

Supabase: A Platform Built Entirely on PostgreSQL

Supabase isn't just an open-source Firebase alternative — its core design builds Auth, Storage, and Realtime entirely on PostgreSQL schemas and WAL. The result: everything is queryable with SQL, pgvector works out of the box, and AI agents can operate the entire platform by writing SQL. 108k GitHub stars, Apache 2.0, free tier with 500 MB database.

Tailscale: Your Agent Lives at Home, You're Not Dialing Home

Self-hosting an agent that runs 24/7 means opening something on your own network that must be reachable from outside and must never sit on the public internet. This post takes apart what each Tailscale mechanism actually solves: the tailnet for reachability, subnet routers for private resources, tags plus ACLs for the permission boundary, and seconds-fast policy propagation plus Tailnet Lock for revocation. Pricing checked 2026-08: Personal is free, up to 6 users, unlimited user devices, 50 tagged resources included.

TanStack Router: Making Routes Compile-Time Verifiable

TanStack Router (1.0 in December 2023, ~20M weekly downloads) makes paths, params, and search params compile-time inferred: navigating to a nonexistent route is a type error, not a runtime 404. This post unpacks its three core designs — type safety, first-class search params, and Query-integrated loaders — and why AI agents writing code amplifies their value.

Temporal: Write the Process as Code, and It Finishes Even After a Crash

Temporal is a durable execution platform (Server 1.31.2, Python SDK temporalio 1.31.0, MIT, verified 2026-08). What separates it from BullMQ / Celery isn't scale but the guarantee: a queue guarantees a message gets consumed, Temporal guarantees a multi-call process runs to completion. The price is that Workflow code must be deterministic — and LLM calls are inherently non-deterministic. This post covers how to resolve that tension and when the constraint isn't worth it.

Trigger.dev: Durable Tasks via Process Snapshots, No Determinism Required

Trigger.dev is an Apache 2.0 durable task platform (v4.5.12, checked 2026-08) that uses CRIU to snapshot entire Node.js processes for pause and resume. Unlike Temporal's replay model, it never re-executes your orchestration code and imposes no determinism constraint — LLM calls go directly in the task. The tradeoff: snapshots can't preserve TCP connections (you reconnect manually), and checkpointing is cloud-only — self-hosted deployments don't get it.

WebMCP: Letting a Web Page Hand Its Own Functions to an Agent

WebMCP lets a page register its own functions as agent-callable tools via document.modelContext.registerTool(), replacing the agent's guess-the-button DOM scraping. Chrome opened an origin trial in 149 and estimates stable in 157; Edge followed in 150. But WebKit has formally opposed it ('an agent acting on a user's behalf is, in effect, assistive technology... the site should not single it out for different treatment') and Mozilla filed neutral. This post covers both APIs, where the security gates sit, and whether to invest now with one and a half engines behind it.

techdeep-dive

Testing Five zh-TW Terminology Linters: One Is Usable, One's --fix Turns 只是 Into 隻是

On 109 posts that genuinely contain Mainland vocabulary, zhtw-mcp scored 29.4% precision at 85.2% recall; twlint scored 21.2% / 82.5% — but 195 of twlint's error-level findings are legitimate Traditional characters misread as Simplified (干→幹 60 times, 只→隻 15), so --fix corrupts the text. The most useful result was not a winner: zhtw-mcp found five Mainland terms my own list had missed, and correctly declined to flag one it wrongly included (審計). These tools calibrate your wordlist; they do not replace it.

Zod: From Form Validation to TypeScript's Universal Contract

Zod's 224M weekly downloads (checked August 2026) put it far beyond 'form validation library': API boundaries, environment variables, route search params, LLM tool schemas and structured output all run on the same schemas. The core mechanism is one definition, two payoffs — runtime validation and static types derived from a single source. Zod 4 (on npm July 2025) is faster, slimmer, and easier on tsc.

Behavioral & Ethics Interview Guide: AI Ethics, Teamwork, and Impact Narratives

Behavioral interviews aren't about improvisation — they're about a pre-prepared story library. AI Engineer behavioral interviews have unique focus areas: AI ethics (bias, fairness, privacy), technical decision impact narratives (why you chose this model/architecture), and experience driving ML projects across teams. Strategy: build 8-10 STAR stories, practice each until you can deliver it in under 2 minutes.

Coding Interview Guide: Strategies for ML-Flavored Programming Problems

AI Engineer coding interviews aren't identical to SWE — beyond LeetCode medium, you'll face ML-flavored problems (implementing a tokenizer, writing a batch inference pipeline, handling sparse matrices). Strategy: practice LeetCode medium to 70% pass rate, then spend remaining time on numpy/pandas operations, data processing pipelines, and ML-related programming problems.

Deep Learning Interview Guide: Core Intuitions from CNN to Transformer

Deep learning interviews don't ask you to derive backpropagation — they test whether you can explain the design intuition behind architectures. High-frequency topics: CNN's locality and translation invariance, why the evolution from RNN to Transformer was necessary, self-attention computation and complexity, BatchNorm vs LayerNorm use cases, and common training tricks (learning rate scheduling, gradient clipping, mixed precision).

LLM Application Design Interview Guide: From RAG to Agent Architecture

LLM Application Design is the hottest new interview topic in 2025-2026. Key focus areas: RAG pipeline chunking/retrieval/reranking design, agent tool-use and planning loops, context window management strategies, guardrails and safety design, and LLM application evaluation methods. Interviewers especially value whether you've hit real-world pitfalls.

ML Fundamentals Interview Guide: From Bias-Variance to Evaluation Metrics

ML fundamentals interviews don't test formula memorization — they test whether you can explain concepts intuitively and hold up under follow-up questions. High-frequency topics: the practical meaning of bias-variance tradeoff, the selection logic for L1/L2 regularization, why cross-entropy beats MSE for classification, SGD vs. Adam tradeoffs, and how precision/recall priorities differ by scenario.

ML System Design Interview Guide: From Requirements to Production Architecture

The core of ML System Design interviews isn't choosing the model — it's how to turn a business objective into a system that's deployable, monitorable, and iterable. Interviewers want to see if you can: translate business goals into ML objectives, design data pipelines and feature stores, choose reasonable serving strategies, and plan monitoring and A/B testing.

MLOps & Deployment Interview Guide: From CI/CD to Model Monitoring

MLOps interviews test whether you have experience pushing models to production. Key topics: ML pipeline CI/CD (how it differs from software CI/CD), model registry and version management, A/B testing design and pitfalls, inference scaling strategies (horizontal scaling, model compression, caching), and production monitoring and alerting design.

NLP & LLM Interview Guide: From Tokenization to RLHF

The dividing line in LLM interviews is whether you've actually used these things. High-frequency topics: BPE tokenization logic and multilingual challenges, pretraining objectives (CLM vs MLM), three levels of fine-tuning (full/LoRA/prompt tuning), RLHF workflow and failure modes, prompting as engineering practice, and the difficulty of LLM evaluation with current methods.

AI Engineer Interview Overview: From Company Types to Preparation Strategy

AI Engineer interviews go beyond ML — big tech emphasizes system design and coding, startups look for end-to-end delivery, and AI-native companies test LLM engineering depth. Strategy: identify your target company types first, then allocate prep time across six dimensions (ML fundamentals, system design, LLM applications, coding, paper reading, and behavioral).

Paper Reading Interview Guide: How to Read, Discuss, and a Must-Read List

Paper reading interviews don't test whether you've read that specific paper — they test whether you can quickly understand a new method and identify its limitations. AI-native companies (Anthropic, OpenAI) particularly favor this format. Strategy: practice reading a paper in 30 minutes and verbally stating contribution + limitation, build your own must-read list, and practice summarizing each paper in three sentences.

Stanford CS329A: A Course on Self-Improvement That Says Out Loud What It Can't Improve

CS329A is built around the generation–verification gap: models can produce the right answer but can't tell which one it is. The conclusion the course draws about itself matters more — today's methods make models more consistent, not smarter. Nine lectures are public, out of twenty.

A Reading Guide to Stanford's CS Courses: Ordered by Prerequisites, from CS106A to CS336

Stanford CS rests on CS103, CS107, CS109, CS111, and CS161; CS221 names three of those plus CS106B as preparation. This guide combines official prerequisites with an explicitly editorial reading order and marks public-material and offering risks.

In the AI Era, Taste Is an Amplifier

AI pushes execution cost toward zero. People with good taste create more value; people with poor taste create more garbage. The difference is not whether you can use AI, but whether your mind contains something worth amplifying before you use it. This series documents my attempt to sharpen judgment systematically.

AI Product Design Interview Guide: From Human-in-the-Loop to Trust Building

AI Product Design is the hottest new interview topic in 2025-2026. Core areas: when to use AI (not every problem needs it), human-in-the-loop design patterns (when to let humans intervene), trust building (how to make users believe AI output), AI product challenges (hallucination, latency, cost), and AI product evaluation metrics.

Behavioral & Leadership Interview Guide: Influence, Conflict Resolution, and Vision

Product Builder behavioral interviews differ from SWE — they don't just test teamwork, they specifically test how you drive things without formal authority. Core skills: influence narratives (how to convince engineers to build your feature), conflict resolution (disagreements with designers/engineers/stakeholders), vision expression (how to make someone understand your product direction in 30 seconds), and failure stories (learning from failure without deflecting blame).

Execution Interview Guide: From Roadmap to Cross-Team Collaboration

Execution interviews test whether you can turn ideas into deliverables. Core skills: roadmap planning (how to prioritize with limited resources), priority defense (why A before B), cross-team collaboration (how to drive engineering and design), stakeholder management (how to handle conflicts), and the ability to track progress with data.

Growth & Experimentation Interview Guide: From Growth Loops to Experiment Design

Growth interviews don't test whether you can growth hack — they test whether you have systematic growth thinking. Core skills: growth loop design (the acquisition → activation → retention → referral flywheel), experiment design (the full hypothesis → metric → experiment → analysis process), retention strategy (finding the aha moment, designing habit loops), and using data to decide what's worth continued investment.

Metrics & Analytics Interview Guide: From North Star to Experiment Design

Metrics interviews test whether you can make decisions with numbers, not how much statistics you know. Core skills: north star metric selection logic (why this one and not that one), metric tree decomposition (finding actionable levers), funnel analysis (which step's drop-off is most worth fixing), A/B testing design and pitfalls, and judgment when facing counterintuitive data.

Product Builder Interview Overview: From PM to Builder Mindset

A Product Builder isn't a traditional PM — you need to build from 0 to 1, not just write PRDs. Interviews test the intersection of product intuition, metrics thinking, technical understanding, and execution ability. Prep strategy: first figure out whether your target company wants a PM or a Builder, then allocate time across nine dimensions.

Product Design Interview Guide: From Problem to Solution

Product Design interviews don't test whether you can draw wireframes — they test how you go from problem to solution. Core skills: MVP scope judgment (what to build and what not to), trade-off analysis (speed vs completeness, generic vs custom), communication of design decisions (why A instead of B), and iterative thinking.

Product Sense Interview Guide: From User Insight to Feature Prioritization

Product Sense interviews don't test how many features you can think of — they test whether you can find the problem truly worth solving within a vague requirement. Core skills: user segmentation thinking, problem reframing (turning 'add a feature' into 'what problem are we solving'), structured reasoning for feature prioritization, and the ability to hold or revise your judgment under follow-up questions.

Strategy Interview Guide: From Market Positioning to Competitive Moats

Strategy interviews don't test whether you can recite frameworks — they test whether you can make judgments with incomplete information. Core skills: market sizing (the practical use of TAM/SAM/SOM, not rote numbers), competitive moat analysis (network effects, switching costs, brand), go/no-go decisions for new markets, and using elimination rather than addition for strategic trade-offs.

Technical PM Interview Guide: From API Design to Architecture Understanding

Technical PM interviews don't require you to write production code, but you need to be able to read trade-offs. Core skills: API design fundamentals (RESTful, versioning, error handling), high-level system architecture understanding (microservices, database selection, caching), collaboration patterns with engineers (RFC process, technical spec review), and making product decisions under technical constraints.

Choosing Among the Three AWS AI Certifications: Current Guidance After MLA-C02 Beta Opens

AIF-C01, MLA-C02, and AIP-C01 are not a difficulty ladder but three job-function slices: AIF tests AI business judgment, MLA tests whether you can put traditional ML, foundation models, and agentic workflows into production, and AIP tests whether you can integrate foundation models into a GenAI system. The MLA-C02 English beta is open for registration; general-availability dates and non-English versions remain unannounced.

Choosing Among the Four Claude Certifications: First Check Whether You Can Even Register

Before comparing exam objectives there is one fact that outranks all of them: registration for the Claude certifications is open only to organizations in the Claude Partner Network — individuals cannot sign up. For those who clear that gate, four things actually decide the answer. CCAO-F ($99) does not count toward partner tier eligibility while the other three do. Claude Code is 20% of CCAR-F but only 3.1% of CCDV-F — the architect exam tests the tool far more heavily than the developer exam. CCAR-P is 28% non-technical (governance 14% plus stakeholder communication 14%), which nothing else in this series is. And CCAO-F is the cheapest but only 14% Prompting; Output Evaluation at 21% is its real spine. All four are valid 12 months, with retakes at 14 / 30 / 90 days and 4 attempts per rolling 12 months.

Which Microsoft AI Certification: The Forks Between AI-103, AI-500, AB-620, and AB-100

Across Microsoft's four AI/agent certifications, only AI-103 → AI-500 is an official ladder; everything else is positioning. Three forks decide it: whether you write Python (AI-103/AI-500 vs AB-620), whether you build or judge (AB-100 vs the rest), and whether you can actually start today — AI-500's four official learning paths currently 404, AB-620 has no practice test, AB-100 has a free one. For readers who prefer Chinese there is a fourth fork: AI-103 and AB-620 offer Traditional Chinese, AI-500 and AB-100 are English only. All four cost $165, expire after one year, and renew free but only inside a six-month window.

Choosing Among NVIDIA's Four: Two Can't Be Registered For, Training Is All Paid, and the Docs Contradict Themselves

NVIDIA's generative AI line has four exams: NCA-GENL and NCA-GENM ($125 each, associate), NCP-GENL and NCP-AAI ($200 each, professional). Three decision inputs no other vendor forces on you. One: both professional exams still show 'Coming soon' next to Register, so any near-term plan is down to the two associates. Two: NVIDIA is the only vendor in this series whose official prep courses are all paid — real cost is exam fee plus courses, and the self-paced totals are $390 (NCA-GENL), $210 for only three of five courses (NCA-GENM), and $1,620 list price across NCP-GENL's five. Three: the official documents disagree with themselves — NCP-AAI's weights total 98% on the web page and 92% in the PDF, and two cells of NCP-GENL's web table carry misplaced text, one of it about OpenUSD. Lock-in also varies sharply: NCP-AAI is 7% NVIDIA-specific, NCP-GENL is 31% GPU and model-compression work.

Aider: The Oldest Terminal AI Pair Programmer, and Where Its Maintenance Stands

Aider is a terminal AI pair programmer dating back to 2023 (Python, Apache-2.0, ~48.3k stars), designed against the grain of today's autonomous agents: you control context by hand with /add, every edit becomes its own atomic git commit, and architect/editor mode splits planning from editing across two models. But note the maintenance cadence: the latest PyPI release is 0.86.2 from 2026-02, the last commit was 2026-05, and the site still recommends Claude 3.7 Sonnet and o1.

Amp: The Coding Agent That Defines Itself by What It Deletes

Amp spun out of Sourcegraph in December 2025 as Amp Frontier Corporation, and its npm package moved from @sourcegraph/amp to @ampcode/cli. Its defining trait is deletion: the editor extension, Amp Tab, TODO lists, Fork, custom commands, and public threads have all been removed. Monthly subscriptions only arrived on 2026-07-18 (Megawatt $20, Gigawatt $200); before that it was pay-as-you-go only. The current focus is orbs — remote machines that keep working after you close your laptop.

GitHub Copilot CLI: An Agent That Runs on GitHub the Platform

Copilot CLI went GA on 2026-02-25 and is included in every Copilot plan, Free included. Its differentiator isn't the agent — it's the GitHub integration: a built-in GitHub MCP server that works on issues and PRs, org policies inherited automatically, and an `&` prefix that hands work to the cloud coding agent. Billing runs on GitHub AI Credits (1 credit = $0.01): Pro $10/mo includes $15, Pro+ $39 includes $70, Max $100 includes $200.

omp (Oh My Pi): The Fork That Inverts Pi's Minimalism

omp is a fork of Pi, but it is not just a plugin layer stacked on top: it adds roughly 80,000 lines of Rust, pulling grep, shell, AST, and PTY in-process. Built-in tools go from Pi's 7 to 31, plus 14 LSP ops, 28 DAP ops, and 60+ providers. One codebase, two opposite bets.

Choosing a React Stack in the AI Era: From the TanStack Trio to the Full Map

TanStack Router (19.7M weekly downloads) + Query (55.8M) + Zustand (44.5M) as the core, with Vite, react-hook-form + Zod (224M), Tailwind + shadcn, and Vitest + Playwright — the current default stack for serious SPAs. The AI era adds three new selection criteria: does the docs site ship llms.txt (all of TanStack does; React Router doesn't), can type safety act as an agent guardrail, and does the source code live in your repo where an agent can read it.

AI Elements: Vercel's ChatGPT-Style Interface as Copy-In shadcn Blocks

AI Elements is Vercel's React component library for the AI SDK ecosystem — the registry currently holds 48 components covering Conversation, Reasoning, Sources, Tool, and the rest of the AI-interface vocabulary. It follows the shadcn model: npx ai-elements@latest copies the source into your project, fully editable, mapped one-to-one onto useChat's message parts.

aideep-dive

AI Agents Generating Slides: Letting the Model See Its Own Layout

The 2026 consensus for agent-built slide decks: outline-first, separate content from construction, then render to images and let a fresh-eyes subagent do visual QA. Anthropic's and OpenAI's official slides skills both converged on PptxGenJS plus a visual verification loop, and the research line (PPTAgent → PreGenie → DeepPresenter) points the same way. But two later corrections matter: PresentBench shows the widely cited PPTEval scores too generously, and SeaSlides argues the model should not write free-form HTML/SVG at all.

AI Governance Frameworks vs. Exam Objectives: EU AI Act, NIST AI RMF, ISO/IEC 42001 — and Why No Certification Names Them

Governance carries more weight on these exams than most engineers expect — CCAR-P is 14% governance plus 14% stakeholder work (28% non-technical), CCAO-F is 15%, AB-100's deploy-and-govern block is 40–45%, AIF-C01 is 14% responsible AI plus 14% security/compliance/governance. But across all fifteen official exam guides in this series, not one names the EU AI Act, the NIST AI RMF, or ISO/IEC 42001; the only regulations any of them names are CCAR-P's GDPR, HIPAA, and FedRAMP. So the use of this post isn't memorizing frameworks for an exam — it's using the three frameworks as a skeleton to file six certifications' scattered governance objectives. The three split cleanly: the EU AI Act is law (fully applicable 2026-08-02, high-risk duties pushed to 2027-12-02 by the AI Omnibus), the NIST AI RMF is voluntary (GOVERN/MAP/MEASURE/MANAGE, and 1.0 is currently being revised), and ISO/IEC 42001 is a certifiable management system standard whose clauses sit behind a paywall — so this post uses only what ISO's own public page states.

AWS AI Practitioner (AIF-C01): v1.1 Turned It Into an Agentic AI Exam

The AIF-C01 exam guide moved to v1.1 on April 30, 2026, adding seven objectives — MCP, multi-agent patterns, context engineering, token-based pricing, and hallucination detection all became testable, with Bedrock AgentCore, Kiro, and Strands Agents joining in-scope services. Almost every summary online describes the older version. This guide builds on the official five-domain weighting, integrating AWS's four-step prep method and field-tested advice from 10+ candidates who passed. $100, 90 minutes, 65 questions (50 scored), pass at 700, valid 3 years — the only exam in this series offered in Traditional Chinese.

AWS GenAI Developer Professional (AIP-C01) Prep Guide: Exam Scope, Resources, and a Four-Step Plan

AIP-C01 tests integrating foundation models into production AWS applications: RAG, agents, security, cost, and evaluation. This guide follows the five official domains with a prerequisite self-check, four-step preparation plan, resource choices, and hands-on checkpoints, plus a ten-week schedule, an AI study prompt, and exam-day guidance. Start with a diagnostic, then close gaps through one RAG-and-agent project.

Claude Certified Architect Foundations Exam Complete Guide

A complete study guide for Claude's official architect certification (CCAR-F): five domains weighted 27/18/20/20/15, four scenarios drawn from six, common anti-patterns, and hands-on preparation. Official specs are 60 items / 120 minutes / $125 / 12-month validity / 720 to pass; registration is limited to Claude Partner Network members, and on-time renewal is free and non-proctored.

Claude Certified Architect Professional (CCAR-P): 28% of It Isn't Technical

CCAR-P is the most expensive and most senior of Anthropic's four exams ($175, 63 items, 120 minutes). Integration is the heaviest domain at 19%, but what really separates it from everything else in this series is the other two: Governance, Safety & Risk Management at 14% and Stakeholder Communication & Lifecycle Management at 14% — 28% combined on compliance, risk, discovery interviews, and delivery lifecycle rather than code. The guide names GDPR, HIPAA, and FedRAMP, and its Intended Audience explicitly excludes entry-level developers and anyone doing 'prompt writing without broader system design responsibility.'

Claude Certified Associate (CCAO-F): The Heaviest Domain Is Knowing When Claude Is Wrong

CCAO-F is the cheapest of Anthropic's four exams ($99, 60 items, 120 minutes), aimed at people who work with Claude rather than build against it. The heaviest of its seven domains is Output Evaluation and Validation at 21% — spotting hallucinations, deciding when human review is required, and adapting outputs — with Governance, Risk, and Responsible Use at another 15%. Anthropic states plainly that it is not for developers building against APIs or designing agentic systems. One easily missed limitation: this credential does not count toward Claude Partner Network tier eligibility, while the other three do.

Claude Certified Developer (CCDV-F): A Third of It Is Ordinary Software Engineering

CCDV-F is the engineer's exam among Anthropic's four certifications. The official blueprint has eight domains, and the heaviest — Applications and Integration at 33.1% — is led by Claude Application Design (8.6%) and Software Engineering Foundations (7.4%), meaning a third of the exam is API mechanics and ordinary software engineering. The counterintuitive part: Claude Code is only 3.1% and Eval only 2.6%, while the sibling Architect exam gives Claude Code 20%. Official specs: $125, 53 items, 120 minutes, pass at 720, valid 12 months, registration limited to Claude Partner Network organizations.

Cost, Latency, and Availability Across Six Exams: One Topic Tested From Three Altitudes

Google PMLE, AWS AIF-C01 and AIP-C01, Microsoft AI-103 and AI-500, and NVIDIA NCP-GENL all test how to make a GenAI application fast, cheap, and reliable — and they form a three-rung ladder: AIF-C01 asks whether you know cost scales with tokens, AIP-C01 and the two Microsoft exams ask whether you can instrument and control it, NCP-GENL asks whether you can change the model and the hardware. Three different altitudes. NVIDIA works at the kernel and quantization layer (Model Optimization 17% + GPU Acceleration 14% = 31%, the heaviest single cost/latency block in the whole series), AWS and Microsoft at the application layer (three caching tiers, token caps, chargeback), and Google at the MLOps layer (CPU/GPU/TPU evaluation, data vs model parallelism, scaling serving backends by throughput). The shared core is eight levers, but each lever becomes a different question at each altitude. This post deliberately carries no prices and no hardware specs — that is the part of this topic that rots fastest.

Preparing for Google PMLE After the Exam Guide Rewrite

Google's Professional ML Engineer exam guide was rewritten in 2026: Vertex AI is renamed Gemini Enterprise Agent Platform throughout, so older study material no longer matches the product names in the questions. This guide uses the official six-section weighting as its skeleton, listing what each section tests, which official materials cover it, and what to build — plus a study schedule whose reasoning is spelled out. Official specs: $200, two hours, 50–60 multiple-choice and multiple-select questions, two-year validity, 3+ years of industry experience recommended including 1+ year on Google Cloud.

The Research Side of Hermes Agent: Batch-Running Thousands of Prompts Into Training Data

`batch_runner.py` runs thousands of prompts in parallel into ShareGPT-format tool-calling trajectories, lets each prompt name its own container image, and resumes by matching prompt content rather than index. Two quality filters run before you see the data: samples with zero reasoning are discarded, and entries calling hallucinated tool names are dropped at merge time. This is why a research lab builds a personal agent — the agent is the data pipeline.

The Hermes Agent Gateway and Scheduler: An Unattended Agent's Biggest Risk Is Spending Your Money

One gateway process fronts 30-plus chat platforms and denies every user not on an allowlist or paired by DM. The scheduler adds two unusual guards: pre-dispatch validation marks a misconfigured job `blocked_config` without making a single LLM call, and the model drift guard makes unpinned jobs fail closed when the global model changes — protecting you from an hourly job quietly following you onto a paid model.

Installing and Upgrading Hermes Agent: Check the Support Tier Before You Pick a Path

Hermes install paths come in three support tiers: macOS (Apple Silicon), Windows 10/11, Linux/WSL2, and Docker are Tier 1; Termux and Nix are Tier 2; pip, brew, AUR, and Intel Macs are explicitly unsupported — fixes for those won't be merged. On the upgrade side, `hermes update` snapshots state first, then compiles nine critical files after the pull and hard-resets the checkout if any fail to parse.

Hermes Agent: Nous Research's Self-Improving Agent, and Its Real Relationship With OpenClaw

Hermes Agent is Nous Research's MIT-licensed agent framework, built around a learning loop: it writes its own skills, curates its memory, and searches past sessions with FTS5. It ships `hermes claw migrate` to move you off OpenClaw — but OpenClaw was not replaced, and both projects are still moving. This is the series opener: what it is, how it differs, and when not to pick it.

Memory and Skills in Hermes Agent: A System That Rewrites Itself, and Where You Can Intervene

Hermes memory has hard caps: 2,200 characters for MEMORY.md and 1,375 for USER.md, and an over-limit write returns an error instead of auto-compacting, forcing the agent to make room itself. Skills are maintained by a curator that runs every 7 days after 2 hours of idle, marks skills stale at 30 days and archives at 90 — but never deletes. The switches actually worth flipping are `memory.write_approval` and `skills.write_approval`, which stage the background self-improvement writes for review.

Migrating From OpenClaw to Hermes Agent: What Moves, What Doesn't, and the Archive Directory

`hermes claw migrate` imports persona, memory, skills from four locations, model and provider config, platform tokens, and the approval allowlist — but secrets are never imported silently, and even `--preset full` requires an explicit `--migrate-secrets`. What can't move (cron jobs, plugins, hooks, the multi-agent list, deep channel config) isn't discarded but parked in `~/.hermes/migration/openclaw/<timestamp>/archive/` for manual work. Coming from Claude Code or Codex is a different command: `hermes import-agent`.

Hermes Agent Model Providers: The Subscription Billing Trap, and Why Fallback Fires Only Once

Hermes supports 40+ providers, and the consumer-subscription OAuth paths are where billing surprises live: Anthropic OAuth only spends Claude Max extra-usage credits, and Claude Pro can't use it at all. Auxiliary tasks default to `provider: auto`, meaning your expensive main model does compression and vision grunt work. The fallback chain is a one-shot switch per session, not continuous retry.

The Hermes Agent Security Model: There's a Floor Below --yolo That You Can't Remove

Approvals default to smart mode: an auxiliary model waves through low-risk commands, auto-denies genuinely dangerous ones, and escalates the uncertain cases to you. Neither `--yolo` nor `approvals.mode: off` can disable the hardline blocklist (`rm -rf /`, fork bombs, `dd` to a physical disk), and `approvals.deny` is its user-editable counterpart, evaluated before yolo. Upstream is explicit that the threat model is an honest-but-wrong agent, not an adversarial process.

Hermes Agent's Seven Terminal Backends: Moving to a Sandbox Turns Off Dangerous-Command Approval

Hermes can run commands on seven backends: local, ssh, docker, singularity, modal, daytona, and vercel_sandbox. The decisive trade-off isn't performance, it's approval — local and ssh run dangerous-command checks, the other five skip them entirely because the container is treated as the boundary. Also, Docker defaults to one long-lived container shared across sessions, not a fresh environment per conversation.

The Nous Tool Gateway: One Subscription Instead of Four Accounts, at the Cost of Concentrating Your Tool Supply Chain

The Tool Gateway routes four tool categories — web search (Firecrawl), image generation (nine FAL models), TTS (OpenAI), and cloud browser (Browser Use) — through Nous infrastructure, replacing four signups with one OAuth. It's per-tool rather than all-or-nothing, and `use_gateway: true` overrides any direct key in your `.env` — the precedence rule people most often get wrong.

The Hermes Agent Tool Layer: What Happens When 3,300 MCP Tools Won't Fit in Context

Attach enough MCP servers and the tool schemas alone eat your context — upstream's extreme example is Cloudflare's ~3,300 tools, whose names alone run about 32K tokens. Hermes answers with Tool Search: MCP and non-core plugin tools collapse into three bridge tools and schemas load on demand, while core tools never defer. Separately, plugins are disabled by default and only run when named in `plugins.enabled`.

Microsoft AB-100: The Architect Exam — Don't Prepare From the Blurb on Its Own Page

AB-100 is the architect tier of Microsoft's agent line, weighted 25-30 / 25-30 / 40-45 with deployment and governance heaviest. Its outline runs on verbs like design, recommend, and propose — it tests judgment, not configuration. Three things to know first: the scope blurb on the official exam page is wrong (it is information-protection and DLP boilerplate, which I verified verbatim), so prepare from the study guide instead; it has a free practice assessment, the only one of Microsoft's three agent credentials that does; and the 15 associate certifications it lists are described as usable, not required. Official specs: $165, English only, pass at 700, one-year validity.

Microsoft AB-620: The Low-Code Agent Track on Copilot Studio

AB-620 is the low-code branch of Microsoft's agent certification line — it tests agent flows, adaptive cards, computer use, MCP tools, A2A, and Fabric data agents in Copilot Studio, not Python. The three skill areas weigh 30-35 / 40-45 / 20-25, with integration the heaviest. Official specs: $165, 120 minutes, pass at 700, one-year validity, and 13 languages including Traditional Chinese — the only localized exam of Microsoft's three agent credentials. It is generally available, but the practice assessment is not out yet.

Microsoft AI-103: After the Foundry Rename, Every Older Azure AI Study Guide Is Void

AI-103 replaces AI-102, retired June 30, 2026, and the objectives were rewritten around Microsoft Foundry — prompt flow, Azure AI Studio, Azure OpenAI Service, and Azure AI Agent Service appear nowhere in them. The five skill areas weigh 25-30 / 30-35 / 10-15 / 10-15 / 10-15, with generative AI and agents the largest. Official specs: $165, 120 minutes, pass at 700, offered in Traditional Chinese, may include interactive components — and it is valid for only one year, though renewal is a free, open-book, unproctored online assessment.

Microsoft AI-500: The Objectives Are Published, the Training Isn't

AI-500 is a rare thing among the major clouds — an expert-level certification dedicated to multi-agent systems, weighted 15-20 / 30-35 / 20-25 / 20-25, naming Agent Framework, LangGraph, Hugging Face Transformers, MCP servers on Azure Functions / Logic Apps / API Management, A2A, Key Vault, and the AI Red Teaming Agent. Three constraints come first, though: it is still in beta (scores wait for rescoring), it requires AI-103 before you can take it, and the official training is not live — the four learning paths listed on the exam page all return 404 today, and the instructor-led course opens 2026-09-30.

Multi-Agent Architecture Across Five Exams: The Shared Core and What Doesn't Transfer

Microsoft AI-500, AB-620, AB-100, NVIDIA NCP-AAI, and Claude CCAR-F all test multi-agent architecture, and they overlap on seven things: orchestration topologies, A2A and MCP, per-agent identity boundaries, three-layer memory, observability and agent replay, human-in-the-loop, and guardrails at four intervention points. But four vendors use four vocabularies for the same ideas, and each exam has objectives that don't transfer — Microsoft names four context-window failure modes nobody else names, 7% of NVIDIA's is locked to NeMo and NIM, and Claude tests SDK-level details like stop_reason. One correction along the way: Google PMLE's wall-to-wall 'Agent Platform' is a Vertex AI rename, not a multi-agent domain.

aideep-dive

Multimodal Models, First Half of 2026: Native Fusion vs. Bolted-On Vision, and Why Leaderboards Contradict Each Other

Pure image understanding has flattened out — four frontier models all clear 80% on MMMU-Pro within 3 points of each other. The real differentiation is video, long-document OCR, and realtime speech, each with a different leader. But the most useful lesson from assembling these rankings is that two credible sources named different Video-MME leaders more than 10 points apart — and that July and August each turned the field over again.

NVIDIA NCA-GENL: The Name Says LLM, Half the Blueprint Is Classical ML

NCA-GENL is usually what a job posting means by 'NVIDIA Generative AI / LLM certification.' But the official blueprint diverges sharply from the name — Core Machine Learning and AI Knowledge 30%, Software Development 24%, Experimentation 22%, Data Analysis 14%, Trustworthy AI 10% — with LLM and RAG content scattered at bullet level rather than forming a domain, alongside spaCy, NumPy, Keras, and cross validation. The other thing to know first: NVIDIA's official preparation courses all cost money ($30–$500), making it the only vendor in this series without a free official learning path. Official specs: $125, 1 hour, 50–60 items, two-year validity, English only, pass/fail with no score reported.

NVIDIA NCA-GENM: The Multimodal One, With Two Required Courses Only Sold as $500 Workshops

NCA-GENM matches NCA-GENL on price, length, and level but not on emphasis: Experimentation rises to 25% (the heaviest), Core ML drops from 30% to 20%, and two new areas appear — Multimodal Data 15% and Performance Optimization 10%. The content covers U-Net, CLIP, diffusion models, multimodal loss functions, attention maps, and NVIDIA's Riva / NeMo / Triton / ACE SDKs. Watch the cost structure: two of the five recommended courses exist only as $500 workshops with no self-paced option, so a self-study path cannot cover the official set. Official specs: $125, 1 hour, 50–60 items, two-year validity, English only.

NVIDIA NCP-AAI: Registration Isn't Open, and the Official Weights Contradict Each Other

NCP-AAI is NVIDIA's professional-level agentic AI credential — $200, 120 minutes, 60–70 items, two-year validity. Two things come first: registration is not open (the Register button carries a 'Coming soon' label), and NVIDIA's own web page and PDF study guide disagree on the weights — Deployment and Scaling is 13% on the page and 5% in the PDF, Run/Monitor/Maintain is 5% on the page and 7% in the PDF, and the two versions total 98% and 92% respectively. Both are nvidia.com. This guide treats that as a range and an uncertainty rather than picking one.

NVIDIA NCP-GENL: 31% Is GPU and Model Optimization, and Two Cells of the Official Table Are Broken

NCP-GENL is NVIDIA's professional-level LLM credential — $200, 120 minutes, 60–70 items. What separates it from every other GenAI exam is where the weight sits: Model Optimization 17% plus GPU Acceleration 14% is 31% on quantization, distillation, pruning, distributed parallelism, and CUDA profiling — not on calling APIs. Two things first: the Register button says Coming soon, so you cannot sit it yet; and two description cells in the official weight table are corrupted — Fine-Tuning is described with OpenUSD data-interchange text and Model Optimization with deployment text. I verified both verbatim; the correct descriptions are in the official PDF.

How Ten Certifications Actually Test Prompting: Exam Framing vs. Practice

Most people assume GenAI certifications are built around prompt writing. CCAO-F gives Prompting 14% while Output Evaluation gets 21%; CCDV-F gives Prompt and Context Engineering 11.0%. What actually gets tested is structured output, injection-resistant prompting, dynamic context injection, context compression and caching, prompt lifecycle governance, and proving a prompt change helped — closer to context engineering and software engineering than to writing craft. None of the ten asks you to write a prompt on the spot; they are all multiple choice, so explaining why beats having a feel for it. The single most useful line comes from CCAR-F: when business logic must be guaranteed, 'change the prompt first' is usually the wrong answer.

RAG and Retrieval Evaluation Across Four Exams — and One Everyone Assumes Tests It, Which Doesn't

Four certifications genuinely test RAG and retrieval evaluation: AWS AIF-C01 (chapters 2 and 3 total 52%, covering RAG, vector stores, and FM evaluation metrics), AWS AIP-C01 (11 of the 27 skill points in its 31% Domain 1 sit in vector storage and RAG), NVIDIA NCP-AAI (Knowledge Integration 10% plus Evaluation and Tuning 13%), and Microsoft AI-500 ('multi-agent RAG architecture' inside its 30–35% Develop area). Google PMLE contributes exactly one LLM-as-a-judge objective, and Claude CCDV-F — the developer certification people most readily assume covers RAG — has no retrieval objective across its eight domains, with Eval at just 2.6%. Includes a same-vendor foundational-vs-professional comparison, a four-vendor terminology map, non-transferable objectives, and a practice project.

aideep-dive

Nine Self-Hosted Personal Agents, One Security Question: Where Does the Execution Boundary Go?

OpenClaw has 386k stars to Hermes Agent's 232k, yet Hermes passed it on OpenRouter daily tokens back on 2026-05-10 (224B vs 186B). The nine self-hosted agents that appeared this year aren't nine competitors — they're nine incompatible answers to one question. CVE-2026-44112 broke OpenClaw's own sandbox, and in the Meta alignment director's inbox incident there was no attacker at all: context compaction ate the safety instruction.

Cloudflare Workers AI Model Picking Guide: By Use Case, Price, and Context

The Workers AI catalog currently holds 84 models. For general chat pick glm-4.7-flash ($0.06 / $0.40 per M, 131K context), for vision pick gemma-4-26b-a4b-it ($0.10 / $0.30, 256K), for cheap high-volume steps pick granite-4.0-h-micro ($0.017 / $0.112), and for embeddings pick qwen3-embedding-0.6b or bge-m3 (both $0.012 per M). This post is updated on a schedule.

travelguide

Do Frequent Japan/Korea Travellers Need a Foreign Currency Account? Breaking Down the 1.5% Card Fee and Real Exchange Costs

The 1.5% overseas card fee is 1% network + 0.5% issuer, and dual-currency cards pay it too. Bank of Taiwan quotes KRW only as cash with a 15.7% round-trip spread, versus 2.45% for JPY spot — so a JPY account is worth it, a KRW one isn't (and mostly can't be opened).

aideep-dive

The 45 Rules of microsoft/AI-Engineering-Coach: An Opinion About Agentic Engineering, Written as Executable Thresholds

A VS Code extension open-sourced by Microsoft employees that reads your local Claude Code / Codex / OpenCode session logs. The real payload is 45 Markdown rules: prompts under 30 characters, sending the next message within 15 seconds of receiving 20 lines of AI code, instruction files over 4,000 bytes — turning 'context engineering' into numbers you can argue with.

CS146S Week 4: What Goes in CLAUDE.md, What Hooks Should Block, Where Subagents Cut

The course lists four techniques for directing agents: instruction files, hooks, commands, subagents. The instruction file is the only one loaded in full every startup, making it config rather than memory; hooks cover what instructions can't, because a rule can be ignored and a hook cannot; commands are the only one a human triggers. The course also marks just one and a half of seven task steps as human work.

CS146S Week 1: A Coding Agent Is, Underneath, a While Loop

Week 1 of CS146S is 'build Claude Code in 200 lines' plus a dissection of production system prompts. The agent loop really is that small. The course slides close with four things Claude does underneath, one of them being `<system-reminder>` tags scattered everywhere to stop the model drifting — which appears in no official documentation.

CS146S Week 5: Express Scores 28, CockroachDB Scores 74 — Agent Readiness Is Measurable

Factory breaks 'can an agent work in this repo' into eight pillars and five levels, and published real scores: CockroachDB L4 (74%), FastAPI L3 (53%), Express L2 (28%). The thesis is that agent readiness approximates the density of deterministic validation loops — linters, type checkers, tests are reward signals for agents.

CS146S Week 7: o3 Found a Linux Kernel Zero-Day at a 1:50 Signal-to-Noise Ratio

The course measured AI SAST false positive rates at 50–100%, against 50%+ for traditional SAST — the genuinely new problem is nondeterminism: run the same prompt twice, get different results, and you can never answer "am I done scanning?" The course lists five agent attack vectors, one of which, intent breaking, attacks the agent's plan itself.

CS146S Week 3: An Agent Skill Is a Folder — the Hard Part Is Two Lines of Description

The Agent Skills spec fits in a sentence: a directory containing a SKILL.md. The real design is three levels of progressive disclosure — only name and description load at startup, the body loads on a match, bundled files load on demand. This site's own repo carries 35 skills and 7,893 lines of SKILL.md, and startup still costs only those 35 metadata pairs.

CS146S Week 6: To Make AI Review Useful, Google Deleted 17 Rules First

Google deployed AutoCommenter to tens of thousands of engineers and published the whole tuning process: suppressing 17 'technically correct but low-value' rules raised the useful ratio from 54% to 66%, with 80% set as the bar for the next rollout stage. Final comment-resolution rate landed around 40%. The bottleneck in AI code review was never detection — it's volume.

CS146S Week 9: One Person Wiring Up MCP Is Fine; Three Hundred Need a Gate

How an individual connects tools is a preference; how an organization does it is governance — who can touch what data, where keys live, whose budget it lands on. Anthropic's published record of ten internal teams contains a good indicator: security engineering accounts for 50% of all custom slash commands in the entire monorepo. Adoption doesn't spread evenly; it takes off first in teams that already build their own tools.

CS146S Week 8: Once Agents Run in the Cloud, the Bottleneck Moves from Waiting to Reviewing

Background agents replace 'you watch it run' with 'it finishes and opens a PR.' Every vendor's design converges on the same parts: an isolated environment, external triggers (issues, Slack, Linear), and a PR as the output. The genuinely new problem is that you become the bottleneck — five agents finish at once, five diffs queue for you, and none of them know the others exist.

CS146S Week 2: Context Engineering, RePPIT, and MCP's 98.7% Cut

Fall 2026 compresses a full week of prompting into one bullet here and adds RePPIT (Research, Propose, Plan, Implement, Test) and MCP. Two RePPIT rules are worth stealing outright: always ask for exactly two proposals, and never let the instance that wrote the code review it. On the MCP side, Anthropic measured turning tools into code calls dropping 150,000 tokens to 2,000.

Stanford CS146S, Two Syllabi Side by Side: What Changed in a Year

Stanford CS146S's Fall 2026 syllabus compresses prompting from a full week into a single bullet, drops the terminal and UI-generation weeks, and adds Agent Skills, Agent-Ready Codebases, Background Agents, and AI-Native Team. Grading moved too: the final project fell from 80% to 50%, with 30% now on open source contributions. This series reads all ten weeks.

CS146S Week 10: The Software Factory Isn't Automation — It's Handing Over the Feedback Loop

The final session is 'self-running, self-improving software systems.' The parts all appeared in the previous nine weeks: deterministic validation loops, skills that can be written back, background agents, centralized governance. One easily missed proportion from the slides — coding is 30% of engineering time, and running it in production is the other 70%.

Adversarial Robustness and Generative Models: It's Not Nonlinearity, It's Linearity

Researchers initially assumed neural networks are easy to fool because they're nonlinear. That was wrong — Goodfellow's 2014 paper argues the primary cause is their linear nature, and high dimensionality lets every tiny perturbation compound. The second half covers generative models: GANs' three pathologies, and why diffusion sidesteps two of them by adding noise and learning to remove it.

Agents, Prompts, and RAG: What's Left After the Lecture Is the Hard Part

A BCG experiment found a jagged frontier: inside it, AI substantially improved consultants' work; outside it, AI made results worse — and people fell asleep at the wheel. The lecture also takes a strong position: avoid fine-tuning wherever possible, because by the time you're done tuning, the next model already beats your fine-tuned version.

AI Project Strategy: Three Hours in a Spreadsheet Buys Back Weeks

Andrew Ng demonstrates error analysis on a deep researcher: columns are the pipeline stages, rows are 10 to 100 queries, you only look at the ones that went badly, and you mark each cell where something broke. The percentages don't have to sum to 100%. He says it takes three or four hours and saves weeks of going the wrong direction — and the fraction of people who actually do it is far below 100%.

Deep Reinforcement Learning: Putting RLHF Back Inside the RL Frame

The third reason Go can't be learned with supervision is the interesting one: the ground truth itself is ill-defined — the strongest human doesn't play their best moves every day, and even their best move isn't optimal. The last 20 minutes map RLHF fully back onto RL: the agent is the model being fine-tuned, the action is the next token, an episode is one full generation, and the reward is extremely sparse.

Full Cycle of a DL Project: You Get Two Days to Collect Data

Andrew Ng walks a face-recognition door system through the entire project lifecycle, and the whole lecture has one thesis: speed. He gives teams a two-day deadline, on the reasoning that 'time spent preparing data should be commensurate with the time it takes to train the model once.' It closes on a line: my job is to build something that actually works, and that is not the same as building something that works on the test set.

Supervised, Self-Supervised & Weakly Supervised Learning: From Comparing Pixels to Comparing Meaning

CS230's second lecture derives embeddings through three case studies: day/night classification teaches you to use humans as a proxy for choosing resolution, trigger-word detection teaches you to manufacture a million training examples in three hours, and face verification walks you through designing your first loss function. The final step — from supervised triplets to self-supervised pairs — is why modern models can consume billions of unlabeled images.

What's Going On Inside My Model? Where You Look First When It Regresses

Ask a model what a goose looks like to it and it draws a whole flock — because the labeled data tagged a flock as 'goose,' so it thinks the flock is the label. This lecture gives seven ways to open a CNN up, then says honestly: applied to transformers, even the frontier of this research only explains two layers.

Introduction to Deep Learning: The Two Moments Prompting Stops Being Enough

CS230's first lecture is a course overview, but Andrew Ng spends most of it on three things: why scaling works, when prompting stops being enough, and why he thinks 'don't learn to code' is one of the worst pieces of career advice ever given.

Scanned PDF Benchmark: How Did 10 Parsers Handle Graduate Entrance Exams?

I tested 10 open-source PDF parsing tools on four scanned NTU graduate entrance exams. VLM-based tools—Firecrawl, MinerU 3.4, and Marker v2—overwhelmingly beat conventional OCR on formulas and code, but installation was the real barrier: MinerU's old package name creates dependency hell, Marker's first model download takes 10 minutes, and PaddleOCR needs a separate engine. In practice, use RapidOCR for screening and MinerU or Firecrawl for close inspection.

Career Advice in AI: The 10x Engineer Who Failed 300 Applications

An elite-level engineer applied to over 300 jobs and failed every one, because the interview guides told him to stand his ground and have a backbone, and he read that as being hard-nosed. The interviewers' read: this is the 10x engineer, and I don't want him anywhere near my team. This lecture also gives the best criterion anyone has offered for vibe coding — technical debt.

aideep-dive

X Open-Sources Its Algorithm: Publishing the Weights Made It Clearer Where the Black Box Is

The 2026-08-13 release added 363,246 lines and published the For You ranking weights for the first time: favorite 0.5, reply 5.0, report −234.0. But the weights are constants — the P(action) that actually decides order comes from a 2560-dim, 8-layer transformer.

Context and Memory: Where Agents Actually Fail

Chroma tested 18 frontier models and all of them degrade as input grows — as a cliff, not a slope. Memory failures are usually retrieval failures in disguise. And the real cost of KV cache is bandwidth, not storage: every generated token reads the whole cache.

Security: Prompt Injection Can Only Be Contained in the Harness

In November 2025 three frontier labs jointly broke all 12 previously proposed prompt-injection defenses. EchoLeak's payload passed Microsoft's own dedicated classifier. So the goal is not blocking every attack — it is surviving the ones that land, and that is harness work.

Drawing the Lines: Agent, Workflow, RAG, and MCP

The line between workflow and agent is who decides the steps — the developer at design time, or the model at run time. By that definition most LLM systems in production today are workflows. Plus a usable test for choosing between RAG and an agent.

Launch Is Where the Work Starts: Enterprise Agent Cases Read Sideways

Salesforce's number from 20,000 deployments: 90% of the work on an agent happens after launch, the reverse of traditional software. Stripe merges 1,300 PRs a week with no human-written code, and credits the environment rather than the model.

The Protocol Layer: MCP, A2A, ACP, Skills

MCP governs agent-to-tool, A2A governs agent-to-agent, Skills govern reusable knowledge. The test is whether the data changes: if it changes between calls you need MCP; if it's stable enough to write down, a skill file is simpler and has no runtime that can fail on its own.

The Model Is a Component, the Harness Is the System

Microsoft, OpenAI, Salesforce, Stripe and three others independently say the same thing: reliability comes from the engineering around the model. And 'give the deterministic parts back to code' has been shipped as a product four separate times — Agent Script, Procedures, runtime, blueprints.

Three Shapes of RAG and the Evaluator Paradox

Standard RAG gives a wrong answer when it retrieves the wrong chunk, and nothing in the system will notice. Agentic RAG adds a self-check, at the cost of the evaluator paradox: the ceiling on self-correction is whatever the evaluating LLM can judge about relevance.

Drones Have Closed Taiwanese Airports 23 Times, and Airports May Only 'Enforce Against'

Three Legislative Yuan budget evaluations log 23 drone-caused airport closures in Taiwan between FY2019 and August 2023, and every row's response column repeats one sentence: on notification, go to the scene with the Aviation Police and investigate. Article 99-13(6) of the Civil Aviation Act gives an airport the verb 'enforce against', not 'stop or remove' — yet the FY2024 report says every airport has already procured handheld jammers.

Power Plants and Fabs Have No Row in the Civil Aviation Act's Authority Table

Paragraphs 3 and 4 of Article 99-13 of the Civil Aviation Act say 'government agencies, schools or legal persons', but paragraph 7's proviso — the only sentence letting a site operator 'take appropriate measures to stop or remove' a drone — says only 'government agencies', and Taipower, CPC and the fabs are legal persons. Their remaining path is to have a municipality announce a restricted zone and 'enforce against' violators, which drops the penalty from an NT$300,000 minimum to an NT$300,000 maximum, while importing a jammer requires 'critical infrastructure provider' status that you learn you hold by receiving a letter.

Police May Lawfully Fly a Drone, but That Is Not Authority to Film

Eight municipal police departments now run drone units, and the authority to fly is clear — Article 99-16(2) of the Civil Aviation Act exempts agencies performing statutory duties. But that is a flight-safety exemption: a Taipei City legal opinion says it 'may be unable to serve as the legal basis for conducting administrative investigation by drone', and the special-compulsory-measures chapter passed in July 2024 authorises GPS, IMSI-catchers and private-space imaging with no aerial-investigation provision at all.

Who May Bring a Drone Down? Taiwan's Law Authorises the Outcome and No Means of Achieving It

'Drone' appears in 19 of Taiwan's central statutes and regulations; 'counter-drone' appears in exactly one, and that one is a procurement act. Six settings are authorised to 'stop or remove' an intruding drone and none of them says by what means — Article 99-13 of the Civil Aviation Act says only 'take appropriate measures', while the 31-item police instrument list includes mortars but no jammer, and Article 67(1) of the Telecommunications Management Act has no public-duty exception.

A Flight Controller's Autonomy Has 56 Entries, One About Opportunity

ArduPilot's ModeReason enum is the exhaustive list of on-board autonomous decisions: of 56 values, 43 are the aircraft deciding for itself, 21 of those because it detected something dangerous, and exactly one — SOARING_THERMAL_DETECTED — because it found an opportunity. As for end-to-end, PX4 mainline's mc_nn_control has a 10 KB tensor arena and a Kconfig default of n; ArduPilot mainline has none.

Why Drone Thermal Camera Prices Jump: The Cost Steps Export Control Draws

Thermal camera prices jump in steps because ECCN 6A003.b.4.b splits control on 111,000 focal-plane elements: 384×288 clears it by 408 pixels, 640×512 is three times over, nothing common sits between, and FLIR prices the Boson 640 at 2.04–2.31× the 320. The Note 3.b carve-out then limits not resolution but focal length — you may have a thermal camera, just not a telephoto one.

Production Ramp Evidence Isn't in the Factory, It's in Failed-to-Award Notices

The public evidence of a production ramp isn't in the factory but in the government e-procurement system: the fire service's 88 thermal drone sets were priced centrally, six fire departments' first tenders drew no qualifying bidder, all seven awards I pulled equal the budget to the dollar, and the bidder pool is about six firms. A control group added afterwards — aerial ladder trucks, at 0.54 failed-to-awarded versus the drones' 0.33 — refutes this post's original first conclusion: failed tenders are the norm in Taiwanese fire-service procurement.

There Is No Swarm in a Drone Light Show: 200 Aircraft Share One Integer

Grepping all 9,199 lines of Skybrush's open-source light-show firmware for neighbour, collision and swarm turns up nothing: no aircraft knows another exists, and the entire content two hundred of them synchronise is one GPS time-of-week integer plus a millisecond offset, with collision avoidance settled on the ground in the choreography software. Chapter 7 of Taiwan's drone cybersecurity spec — the swarm chapter — leaves the 'drone' column a dash down all twelve items and tests switches and routers instead.

Taiwan's Drone Security Spec: The Five Items That Test Resilience Are Optional

Reading the whole specification turns up three counter-intuitive things: of the seven mandatory items only three sit on the aircraft and the rest on the ground control station; an unencrypted link still passes, as long as the manual or the box states the reason or the risk — the rule asks for disclosure, not encryption; and the five items that actually test resilience (spoofing, jamming, link loss, known firmware vulnerabilities, app certification) are all in Chapter 8, which is optional.

From 40 Minutes to 6 Hours: It's the Airframe, Not the Battery

The same 5 kg aircraft on the same 2 kg pack needs 512 W to hover for 41 minutes as a multirotor, or 128 W to cruise for 164 minutes and 197 km as a fixed wing — and that range is independent of cruise speed. A VTOL's cost is not the energy burned in transition (2.4%) but that it carries both a rotor system and a wing, worth about a quarter of the range; Taiwan builds something at every rung from 40-minute electric multirotors to a 6-hour piston helicopter.

Why Drones Only Fly 30 to 45 Minutes: One Equation Gives the Answer

Hover power scales with takeoff weight to the 1.5 power while battery energy scales linearly, and that one exponent gap makes added battery hit diminishing returns fast — the optimum battery fraction solves to two-thirds of takeoff weight, independent of rotor diameter, rotor count, efficiency and energy density. What actually caps endurance is payload, not the regulatory weight thresholds: 60 minutes needs a 314 Wh/kg cell, and the best high-rate cell today is 242 Wh/kg.

Frequency Hopping Is Not Encryption, and Taiwan Caps Power by Channel Count

ExpressLRS derives its hop sequence by running the bind phrase through MD5 into a UID and feeding that to a linear congruential generator — fully reproducible, and I ported it to Python and matched the original C bit for bit — while the link itself has no encryption at all, only a 14-bit CRC. Taiwan's LP0002 then turns channel count into a power ceiling: 75 or more hopping channels at 2.4 GHz allows 1 W, fewer allows 0.125 W.

Analogue FPV Video Links Have No Clause to Walk Through in Taiwan

Analogue FPV video in Taiwan has two separate problems: LP0002's 5.8 GHz window is only 125 MHz wide while the market's standard channel table spans 300 MHz, leaving eighteen of the forty classic channels inside the window once transmit width is counted. The harder one is that §4.10.1.1 recognises only frequency-hopping transmitters at 5.8 GHz, and analogue video is fixed-frequency — it doesn't violate a clause so much as fail to match any clause in the document.

The Seven Seconds After GPS Jamming: How It Notices, and Why Detection Is Off

Flying a drone to 28 m in SITL and switching on the simulator's built-in GPS jamming, the flight controller declared EKF failure about seven seconds later, switched to LAND, landed and disarmed — it neither flew away nor crashed. Reading PX4's source then shows that of the 12 gates in EKF2_GPS_CHECK, bit 11, jamming detection, is off by default: the estimator catches jamming on its own, while spoofing can only be reported by the receiver.

PX4 or ArduPilot: The Real Fork Is the Licence

After cloning and building both, three things differ from the stereotype: ArduPilot's EKF3 header credits the derivation to PX4/ecl, so the hardest layer is shared; PX4's last year of commits comes from company domains (380 from Auterion alone) while ArduPilot's comes from personal addresses with one contributor at 37%; and what really decides the choice is not performance but BSD-3 versus GPLv3, and which layer you need to modify.

The Filings Answered What I Assumed Needed an Interview

I had filed "what's the real margin on military-grade-commercial tenders, and how long is the cash cycle" under questions requiring interviews. The published filings answer both, more precisely: gross margin runs 35–39%, normal for hardware; but operating expenses consume it, and operating income has been negative for three straight quarters while reported net income came from non-operating items. The real constraint is inventory — roughly 385 days of it, producing a cash conversion cycle near 377 days. The money in this business isn't stuck in margin, it's stuck in inventory.

The CAA Published the Entire Question Bank: What Four Exam Subjects Reveal About the Regulator

In 2022 Taiwan's Ministry of Transportation issued a press release titled 'Drone Licensing Overhaul! Obscure Questions Removed, Question Banks Fully Published,' putting all 1,420 questions online — 388 general, 588 professional, 324 renewal, 120 simplified renewal. The stated reason was candid: candidates said it was too hard. And the published bank became a policy document in its own right — the meteorology subject imports the manned-aviation syllabus wholesale, and flight principles lean heavily on fixed-wing aerodynamics, while most candidates fly multirotors.

The Drone Chapter Has No Privacy Provision, Only the Criminal Code

Taiwan's Civil Aviation Act drone chapter regulates flight safety, not the people being flown over — and that isn't my commentary, it's the Legislative Yuan's own 2020 research report: the chapter 'contains no specific provision on the privacy management that matters most to the public.' So privacy falls back on Article 315-1 of the Criminal Code and the Personal Data Protection Act, and case law shows that catches part of it: pointing a drone at a hot spring room drew an indictment, raising a phone to a window drew four months. What it doesn't catch is evidence — the report's own words: 'by the time the victim notices it, it may already have vanished.'

After "NT$2M a Year Flying Drones": Agricultural Spraying Has Run a Full Cycle

Agricultural spraying is the one drone application whose unit economics are fully transparent: the billing unit is the fen (about 970 m²), rates run NT$150–300, spraying one fen takes about two minutes, equipment costs NT$300–500k, and even the substitute's price is public (manual spraying, roughly NT$200 per fen). And precisely because anyone can run the arithmetic, everyone did — operators who entered in 2018 sprayed over a thousand hectares a year and cleared over a million; later entrants broke even after two or three years and quit. Licensed operators charge NT$300 per fen; unlicensed ones undercut to NT$150. This is the industry cycle in miniature, compressed into about six years.

Inspection Is Taiwan's Furthest-Along Drone Application — Because It Routed Around BVLOS

I assumed inspection was blocked by beyond-visual-line-of-sight rules the way logistics is. It isn't. Bridge, transmission tower, and high-speed rail viaduct inspection are all running with hard numbers: one bridge went from 8 inspectors, 4 vehicles and 2 days to 5 people, 1 vehicle and half a day with no traffic control at all, at 60% of conventional cost; high-speed rail crews covered at most 700 metres a day on foot and now save 3–5x; a private power plant cut headcount by three quarters and cost by half with no outage required. The reason is that all of these are segmented, fixed-point tasks completable within visual line of sight. What BVLOS actually blocks is continuous long-range routes, not 'inspection' as a category.

Taiwan Already Has 24 Drone Logistics Corridors — It Didn't Take the Wait-for-Regulation Route

I assumed Taiwan had no real drone logistics because it has no BVLOS framework. Wrong: the CAA has approved 10 cases across 24 flight corridors, and the Institute of Transportation has run a six-year PoC → PoS → PoB path since 2020, entering commercial validation in 2025. The way it routes around regulation is the mirror image of inspection — inspection cuts work down to within visual line of sight, logistics gets corridors approved one at a time. And its value isn't being cheaper than a boat: during the typhoon sailing suspensions at Liuqiu, a drone made the crossing in a bit over ten minutes, which is what you have when the boats don't run.

Search and Rescue Drones: The One Application Whose ROI Isn't Money — and the Easiest Budget to Cut

Agricultural spraying runs NT$150–300 per fen, computable to the decimal. Search and rescue has no such number, because a life recovered has no price. That difference determines two things: first, the specification is written by terrain rather than performance — the defining feature of Taiwan's fire agency drones is that they do NOT depend on GPS, because mountain signal drops and they must thread through trees; second, when the legislature moved to cut the budget, the ministry could only point to one man pulled from a flooded river in June. Defending a budget with an anecdote is fragile, and this application has no better weapon.

Why Countering Drones Is Hard: Jamming Is Failing, and Taiwan's Problem Isn't Only Technical

Electronic warfare has one structural limitation: it needs a signal to attack. Fiber-optic control emits no radio, satellite links bypass ground jammers, and AI terminal guidance needs no link at all for the final leg — three methods systematically defeating jamming, which is why defense is shifting toward interception. Taiwan carries an extra layer: a Control Yuan investigation documents the Ministry of National Defense revising its drone response SOP three times in two months, from "may shoot down" to "flare warning only, first shot requires the Minister's authorization" and back to "shoot down with 7.62mm or smaller." That is not a technology problem.

Taking Apart Two TTSB Crash Reports: Neither Was the Operator's Fault

Only four drone occurrences have entered Taiwan's official aviation safety statistics — because the threshold is 'substantial damage to a drone over 25 kg.' A hobbyist crash never enters the count. The two published investigation reports are the same model, the same manufacturer, and the same agency, and both probable causes were hardware failures: main rotor servo electrical failure in one, a fractured tail rotor pitch link in the other. In both, the flight control computer and the operator were explicitly cleared.

How to Read a Drone Spec Sheet: Which Lines Regulation Turned Into Boundaries

The three most important lines on a drone spec sheet are exactly the three lines spec sheets don't print. Taiwan's Drone Cybersecurity Testing Specification defines a product 'series' as units whose flight control, communications, and satellite positioning chip modules are all identical — regulation decides whether two drones are the same drone by those three modules, not by looks or endurance. And the weight field isn't a marketing number either: 250 g, 1 kg, 2 kg, 15 kg, and 25 kg are five separate legal thresholds.

aideep-dive

What AI Certifications Engineers Can Take in 2026

Every AI certification an engineer can register for in 2026, listed one by one: AWS AIP-C01 / MLA-C01 / AIF-C01, Google PMLE, Microsoft AI-103 and AI-500, all twelve NVIDIA exams, Databricks, Snowflake, Oracle's Agentic AI track, IBM watsonx, Salesforce Agentforce, GitHub GH-300, Anthropic's four Claude exams, Taiwan's iPAS AI Application Planner, plus the governance and audit line (IAPP AIGP, ISACA AAISM / AAIA, CertNexus CAIP) — prices, validity, and registration gates all checked against vendor pages. Two things that hit your wallet: Google's PMLE exam guide has renamed Vertex AI to Gemini Enterprise Agent Platform throughout, making pre-mid-2026 study material worthless, and the iPAS intermediate certificate is valid 5 years, not permanently.

anydoc: 14 Office Formats to Markdown, Firecrawl's Rust Answer

Firecrawl's open-source Rust conversion library turns 14 office formats (including legacy .doc / .ppt / .xls) into GFM at a 4.7ms median — 109× faster than Docling under the same timing basis. The trade-off: it does no OCR at all.

The Parsing Layer: When Structure Must Be Inferred — and Licensing Is the Real Selection Axis

Scans and complex layouts leave you no choice but to infer structure with a model. But the technical gap between MinerU, Marker, and Docling is far smaller than the licensing gap — MinerU needs a separate license past $20M monthly revenue, Marker's model weights need payment past a funding threshold, and only Docling is cleanly MIT. Read the LICENSE before the benchmark.

The Three-Layer Ladder of Document Parsing: Pick the Layer Before the Tool

The most common mistake in feeding documents to an LLM isn't picking the wrong tool — it's picking the wrong layer. Structure already in the file goes to the conversion layer (milliseconds); text without structure goes to extraction; only inferred structure needs parsing. anydoc's 4.7ms against Docling's 513.6ms is a 109× gap, and most people jump straight to the most expensive layer.

The Deterministic Extraction Layer: Solve 80% of Your PDFs With No Model At All

Digital-native PDFs already contain readable text — what's missing is structure, and heuristics can recover it. PyMuPDF, pdfplumber, pypdf, and Tika do this with zero GPU and zero inference cost. The biggest selection trap isn't accuracy; it's PyMuPDF's AGPL-3.0 license.

The Drone Industry Job Map: Eleven Roles, and Which Ones a Software Person Can Actually Enter

Put drone jobs back into the five-layer supply chain and the picture clarifies: Layer 2 belongs to mechanical and electrical engineers, while Layer 3 — flight control firmware, sensor fusion, RF, edge AI — is the real entry point for software people, and also exactly the layer Taiwan is short on. But check the scale first: 267 companies and NT$12.9B of 2025 output means far fewer openings than the news volume suggests.

Four Gates into Taiwan's Drone Industry: The Entry Mechanics Public Records Can Tell You

TEDIBOA has grown from 50 founding members to over 260, but since 1 July 2025 you must first be a full member of the Defense Industry Development Association before you can even apply. R&D grants under the Industrial Innovation Platform cap out at 50% of total project cost, three concurrent cases per company, three years maximum — and explicitly ban red-supply-chain components while requiring joint cybersecurity lab testing. This piece covers only what public documents can actually establish.

From Software into Drones: Use the PX4 Architecture Diagram as a Job Map

PX4's own architecture docs say the companion computer runs Linux because 'Linux is a much better platform for general software development than NuttX; there are many more Linux developers.' That sentence is the entry point. The three transition paths differ sharply in friction — CV and MAVLink work on the companion computer transfers almost directly, flight controller firmware means learning RTOS and work-queue constraints, and estimation and control means quaternions and Kalman filters.

Four Ways to Learn Drones in Taiwan: Universities, Licences, Competitions, and Vocational Training

Taiwan has no 'drone department.' The system's core is National Formosa University's Department of Aeronautical Engineering — a NT$90M Ministry of Education training base sited inside the Chiayi drone cluster, the only one of 18 regional bases located in an industrial park, plus a NT$50M lab with a large low-speed wind tunnel. And the 2026 Presidential Cup's three challenge topics (edge AI recognition, high-precision positioning, GNSS-denied autonomous navigation) are almost a mirror of the industry's Layer 3 gap.

Following Taiwan's Drone Defense Money: Three Budgets and a Bill Stuck for Two Months

Taiwan's public drone money splits three ways: an approved NT$44.2B coordination program (R&D grants), a proposed NT$210B defense procurement special statute (stuck in cross-party negotiation), and annual agency budgets (NT$7.2B+ for 2026). The largest was written to run from 1 August 2026 — that date has passed with the bill still unresolved. And the Executive Yuan version buys only three specific items.

The Drone Supply Chain Against a Four-Criteria Framework: Only One of Four Holds

Measuring the drone sector against supply chain chokepoint / structural demand / high switching cost / long-term institutional holding: Taiwan sits at the most substitutable layer, 80% of demand comes from public budgets rather than end-user behavior, and only the certification-driven switching cost genuinely holds. The Army's NT$988M counter-drone contract — failed three times, terminated in full, NT$98.78M performance bond forfeited — is the most expensive lesson in why winning a bid is not revenue.

What Paper Still Earns in a Digital-First Study Stack: Three Places It Works, and One It Doesn't

Once studying went fully digital, paper still holds three places with real evidence: reading (paper over screens at g ≈ −0.21 across 171,055 participants, widening to 0.35–0.48 when scrolling is required), writing while you answer (on screen, harder questions draw less scratch work, not more), and drawing (45% recall against 20% for writing). The one claim most people lean on, that handwritten notes stick better, spans −0.008 to +0.248 across four meta-analyses with no consensus.

Getting a Taiwanese Drone Licence: Tiers, the No-Skipping Rule, Fees, and Timeline

The general licence is written-test only and costs NT$450 total (NT$200 test + NT$250 certificate). The professional basic tier adds a practical test and medical exam for NT$1,900. Professional tiers must be climbed in order — reaching the top bracket realistically takes 18 months to two years. Each advanced group (G1/G2/G3) costs another NT$1,200. Fee figures come from Annex 17 of the regulations; the widely circulated 'NT$500 practical test' is wrong.

Taiwan's Drone Rules in Plain Language: Registration, Licence, Penalties

Register anything 250g or heavier; the registration number expires after 2 years. Individuals only need a licence at 2kg–15kg with navigation equipment. The licence term is now 3 years, not 2, and the student licence age dropped to 14, not 16. This piece covers only the currently effective text of the Remotely Piloted Drone Management Regulations, and flags the three most common outdated claims circulating online.

Four Drone Business Models, and Why Selling Airframes Is the Worst One

Hardware runs 35–55% gross margin under permanent DJI price pressure; autonomy software and DaaS subscriptions run 60–80% and recur. Skydio's software subscriptions were already ~30% of revenue in 2023 at a 38% blended margin; India's Garuda had DaaS at 62% of FY24 revenue with a 351-day cash conversion cycle against a defense-heavy peer's 597. Taiwan is almost entirely concentrated in the lowest-margin, most substitutable cell.

BVLOS in Three Jurisdictions: Taiwan Has No Framework At All

The EU has had a workable path since the end of 2020 — the Specific category grants an operational authorisation based on risk assessment, with U-space Regulation (EU) 2021/664 rolling out on top. The US Part 108 rule was still at OIRA and unpublished as of July 2026. Taiwan doesn't have the framework at all: its regulations offer 'extended visual line of sight' (900m, 400ft, observer required), while true BVLOS runs on per-activity permits valid for three months.

Drone Industry Cycles: How the 2016 Bubble Burst, and What's Different This Time

The 2016 consumer drone bubble left specific wreckage: 3D Robotics stopped making hardware, GoPro recalled all 2,500 Karma units six weeks after launch and cut 15% of staff, Parrot cut 35% of its drone workforce, and Lily Robotics collapsed after taking $34M in pre-orders. In 2023 even Skydio — $570M raised — exited consumer, and three years later it is valued at $4.4 billion. This wave runs on a completely different engine, but three things are exactly the same.

The Drone Industry Map: Components, Regulatory Ceilings, and the Non-Chinese Supply Chain Rebuild

The global drone market is roughly US$69B in 2026 (IDTechEx). China holds about 80% of it (CSIS) and DJI over 70% of multi-rotor. The FCC put every foreign-made drone on its Covered List in December 2025; Taiwan's drone output jumped from NT$5.0B to NT$12.9B in one year, and Q1 2026 exports already beat all of 2025. This piece breaks down the five-layer supply chain, the four demand blocks, and the two ceilings holding back scale.

techdeep-dive

Your Phone Isn't Listening: What FTC Filings and Meta's Own Docs Say About Why the Ads Are So Accurate

Northeastern tested 17,260 Android apps and found zero activating the microphone. In May 2026 the FTC ruled that Cox Media Group — the company that claimed to be listening — collected no voice data at all and was reselling data-broker email lists, settling for $930,000. The real pipelines are off-site event feedback, lookalike spillover, contact-graph uploads, and location brokers.

Taiwan's Drone Supply Chain: Where the 267 Companies Are, and Which Layer They're Stuck On

Of the 267 companies the Ministry of Economic Affairs counted, 164 are in northern Taiwan. But geography is not division of labor — Thunder Tiger's published bill of materials shows motors, batteries, frames, and propellers sourced locally, while flight control, comms/GPS, and camera modules go to US, European, and Japanese partners. Exports were only 23% of 2025's NT$12.9B output, and 88.1% of export value sits in the 2–7 kg weight band per Ministry of Finance statistics.

techdeep-dive

Three Routes to Hand-Drawn SVG Icons: A ~88k Free Library, a Generator That Bends Lucide, and the License Page Nobody Reads

Koboyo claims close to 90,000 free hand-drawn SVG icons (the count oscillates: 92,967 → 87,954 → 90,150), but its sitemap only lists about 17,930 icon pages, and its license page explicitly forbids building an icon library or canvas app with them. There are actually three routes to a hand-drawn look: collect a library, bend existing geometry programmatically (sketchyicons turns every straight run in Lucide into a quadratic Bézier, seeded by icon name for byte-for-byte reproducibility), or generate with AI. This piece compares seven libraries on scale and license, unpacks the algorithms behind sketchyicons and tldraw, and surveys the icon search tools now shipping MCP servers.

AI Makes Things Smooth Exactly Where They Should Be Hard: What Generative AI Does to Learning

The single most-cited meta-analysis on ChatGPT in education (g = 0.867, ~500k views) was retracted by Nature in April 2026. But the positive finding was not overturned — the issue is that it measures performance while the AI is available. Bastani's PNAS RCT measured something else: +48% accuracy during practice with GPT-4, then 17% below never-users once access was removed.

Learning How to Learn: Auditing the Course 4.17 Million People Took — What Holds Up, What's Just a Metaphor

Dunlosky's 2013 review rated 10 study techniques; only self-testing and distributed practice earned 'high utility'. But a 2026 systematic review puts the effect at 0.22–0.46, and Pan & Rickard's transfer meta-analysis finds 'no positive transfer' once publication bias is corrected — making the premise in the framework's own name the piece that tests worst.

aideep-dive

Digital Employees: Reliability Comes From the Harness, Not the Model

"Digital employee" isn't a technology — it's a pricing and accountability unit. Anthropic's Project Vend had Claude actually run three shops, and found the most effective intervention wasn't a smarter model but forcing it to follow procedures. Their words: "we rediscovered that bureaucracy matters." Gartner estimates only ~130 of the thousands of vendors claiming to be agentic actually are.

aideep-dive

The Image-to-Video Landscape: Architecture, Models, and Real Prices in 2026

Every serious image-to-video model in 2026 runs latent diffusion on a DiT backbone, so visual quality is no longer a useful axis for choosing one. The real axes are native audio, self-hostability, and dollars per second. Three widely-repeated errors worth correcting: Sora's app shut down on April 26 and its API goes on September 24; Wan 2.7 is described everywhere as Apache 2.0 open weights but no first-party source has them; Veo 3.1 officially costs $0.40/s, not the $0.75/s that circulates on review sites.

aideep-dive

The 2026 Map of 3D Modeling Tools: AI Generation, Scanning, CAD, or Manual

There are four paths to a 3D model in 2026: AI generation (Meshy-6 / Tripo / Rodin Gen-2.5 / Hunyuan 3D), phone scanning, Text-to-CAD, and manual modeling. Picking wrong has concrete costs — AI-generated meshes can't be dimensionally edited, Rodin's STL exports usually need repair, and Meshy's free-tier assets are public. This guide selects by what the model is actually for, with current pricing from each vendor's own page.

AI Web Scraping Tools Landscape: A Selection Guide for 34 Open-Source Projects

From MarkItDown (175k stars, MIT) to curl_cffi (6k stars), a survey of 34 open-source tools for feeding data to AI. Categorized along five axes: whole-site crawling, AI browser agents, document conversion, smart extraction, and anti-detection infrastructure. The key to selection isn't which tool is best — it's scenario matching.

aideep-dive

Uncle Bob Doesn't Read His Agents' Code: What He Runs Instead of Code Review

Uncle Bob's 4.18M-view post of 2026/7/23 isn't a manifesto — it's a reply to an engineer who started in 1983 asking whether needing to understand code psychologically makes him old-fashioned. And he doesn't skip the code entirely: his 6/1 four-stage pipeline post says 'I spot check the code,' with thresholds of crap ≤ 6 (convention is 30) and mutation runs that kill all survivors. Plus a breakdown of his open-sourced Acceptance-Pipeline-Specification and the three metric blind spots Grady Booch names.

productdeep-dive

Value Validation for Digital Products: From Assumption Maps to the M3 Retention Baseline for AI

The unit of validation is an assumption, not an idea. Kohavi's data shows the industry median experiment success rate is ~10%, which means roughly 22% of 'winning' experiments at p<0.05 are false positives. Sean Ellis's 40% threshold has no publicly available dataset. AI product retention should be baselined at M3 rather than M0, and GRR splits from 23% below $50/mo to 70% above $250/mo.

productdeep-dive

Product Builder vs PM: What the Role Is and How to Get There

A Product Builder runs the full loop from problem discovery to design to build, alone. The core difference from a PM: PMs influence execution through authority, Product Builders influence it by directly shipping working products. LinkedIn replaced its APM program with an Associate Product Builder track, and PayFit defined the role back in 2019.

aideep-dive

Where 3D Generative Models Stand: Reading the 2026 Technical Map Through Lyra 2.0

The dominant paradigm in 3D generation in 2026 is video diffusion feeding feed-forward 3D reconstruction, and Lyra 2.0 is the flagship of that line. But three Best Papers at CVPR 2026 point at what comes next: SAM 3D brings foundation-model-scale object reconstruction, D4RT rebuilds dynamic 4D scenes in seconds from a unified transformer, and O-Voxel replaces Gaussians with structured latents. 3DGS still rules, but surface primitives are challenging it, and pixel-space diffusion is pushing back against latent space.

techdeep-dive

How Content Platforms Rank Their Feeds: From Reddit's Formula to TikTok's Interest Graph

Take apart the feed ranking of ten platforms and they're all solving the same problem: how to trade off between newest, best, and letting new content be seen. What actually decides your answer isn't how clever the algorithm is, but whether your content is oversupplied or scarce — big platforms rank to filter content out, small communities want every post to be seen.

aiguide

Which AI Courses to Take in 2026: From AI-Curious to Vibe Coding to Production

Every official course platform from OpenAI, Anthropic, and Google, plus Stanford CS146S/CS336, Elements of AI, Hugging Face, MIT 6.S191 and more — scraped page by page, then re-sorted into four tiers: AI-curious, vibe coding, shipping to production, and how models actually work. Also covers self-study repos still being updated in 2026 and browser-based platforms that need no local setup, filtered by last-commit date rather than star count. The conclusion: nearly all of it is free. What is scarce is not courses, it is the judgment to pick one. And tier four will not fix your tier three problem.

learningdeep-dive

Evergreen Books Still Trending in 2026: A Reading List Built from Threads, Dcard, and Vocus Signals

24 evergreen books across productivity, life design, brain science, psychology, and money — selected using actual discussion evidence from Threads, Dcard, PTT, and Vocus between late 2025 and mid-2026. Strongest signal: Rewire by Nicole Vignola hit top 3 on both Eslite and Books.com.tw H1 2026 bestseller charts.

techdeep-dive

Is PostgreSQL Really Enough? Don't Rush to Adopt Specialized Databases

Most teams don't need five databases. PostgreSQL's extension ecosystem covers caching, queues, full-text search, and vector search — but the real decision isn't 'can it do it' but 'where does ops cost cross performance needs.'

aideep-dive

HyperFrames Deep Dive: HTML as Video, a Paradigm Shift for the Agent Era

HeyGen's open-source HyperFrames defines video timelines with HTML data attributes, uses headless Chrome for frame-accurate seek-and-capture, then encodes via FFmpeg to MP4. 33k stars in 3 months, Apache 2.0, 21 agent skills — AI agents write HTML to produce video, no React needed.

What Is the Mini Yuanta Taiwan 50 ETF Futures (SRF): Reading a Screenshot of 0050's Futures Version and Its Leverage Design

SRF is the Mini Yuanta Taiwan 50 ETF Futures, tracking the 0050 ETF itself. A NT$7,900 initial margin controls a contract worth roughly NT$110,000 — about 14x leverage. Dividends are handled through an equity adjustment, not the backwardation mechanism used by index futures.

investingdeep-dive

Understanding a Trading Post from Scratch: Warrants, Stock Futures, and Maintenance Ratio Explained

Saw a trading post about going from NT$150k to NT$2.4M in half a year. Didn't understand a word of it — warrants, stock futures, maintenance ratio — so I looked them all up.

aideep-dive

Loop Engineering: When AI No Longer Needs You to Write Prompts

Loop Engineering is the practice of designing systems that automatically prompt AI agents, rather than prompting them manually. Boris Cherny runs hundreds of agents, Addy Osmani coined the term, and Blake Crosley identified verification cost as the real bottleneck — this article covers primary sources, the five building blocks, applicability boundaries, and criticisms.

climbingdeep-dive

A Knowledge Map of Climbing Books: 60+ Books from Training Science to Mental Philosophy

Climbing books don't exist in isolation — they represent competing schools of thought, philosophical differences, and knowledge gaps. This post maps the relationships between 60+ books to help you pick the right one to read next.

Choosing a Browser MCP: CDP, Playwright MCP, or Puppeteer MCP?

It's really a two-way choice now: @playwright/mcp (cross-browser, accessibility tree, token-cheap) versus chrome-devtools-mcp (Chrome's official server, performance and memory diagnostics). @modelcontextprotocol/server-puppeteer has been archived and is no longer a candidate. The dividing line is no longer abstraction level — it's 'drive the page' versus 'diagnose Chrome'.

Chrome DevTools MCP: The MCP Server Wired Directly to CDP

chrome-devtools-mcp, maintained by the Chrome team, packages DevTools capability as an MCP server: performance traces and insights, Lighthouse audits, heap snapshots, extension management — none of which @playwright/mcp exposes. It runs on Puppeteer, so interactions auto-wait; the costs are Chrome-only support and usage statistics reported to Google by default.

@playwright/mcp: Microsoft's Official Browser Automation MCP Server

@playwright/mcp defaults to an accessibility tree (browser_snapshot) instead of screenshots, cutting token consumption sharply. Combined with Playwright's native auto-wait it's a sensible starting point for AI agents doing web automation — but note it now runs headed by default, keeps a persistent profile by default, and gates advanced tool groups behind --caps.

@modelcontextprotocol/server-puppeteer: The Official Puppeteer MCP Server

server-puppeteer is the Puppeteer wrapper in the official MCP servers monorepo — seven lean tools built around screenshots and evaluate. It has since been archived (moved to servers-archived, no longer published), so it is not a choice for new projects; if you want Puppeteer lineage in an MCP server today, look at the Chrome team's chrome-devtools-mcp.

investingdeep-dive

I Saw This 2x ETF System on Threads — It Comes From 3 Books

A 2x leveraged ETF system traces its philosophy to three books: A Random Walk Down Wall Street answers 'what to hold' (index funds), Lifecycle Investing answers 'how to accelerate' (leverage to diversify time risk), and The Four Pillars of Investing answers 'how to survive' (rebalancing discipline). Combined: 60% 2x ETF + 40% cash, Beta=1.2, ±10% rebalancing trigger.

aideep-dive

Text / Image to Lottie: A Landscape Overview of AI Animation Generation Tools

From the CLI tool kin3o to the CVPR 2026 paper OmniLottie — a survey of open-source approaches for converting text and images into Lottie animations, with performance benchmarks and selection guidance.

techdeep-dive

AI-Powered E2E Testing: How canary, Stagehand, Magnitude, and Shortest Each Solve the Problem

AI agents running tests are non-reproducible; hand-written Playwright is hard to maintain. Four tools that emerged in 2024-2025 each tackle this dilemma with very different design philosophies.

aideep-dive

The Skill Management Revolution for LLM Agents: A Complete Landscape of Skill Lifecycle from Voyager to MUSE-Autoskill

MUSE-Autoskill (2026) introduces a five-stage skill lifecycle framework. Self-created skills achieve 60.35% (+7.16%) on SkillsBench overall, and an impressive 87.94% on tasks where skill generation succeeds — surpassing the human-authored skill ceiling. This post synthesizes six arXiv papers to map the full landscape of skill evolution research.

designdeep-dive

A Guide to Design-System Color Palettes: From Tailwind to Material 3

A comparison of seven major design systems—Tailwind, Radix, Material 3, Carbon, Ant Design, Primer, and Apple HIG—covering scale structure, neutral colors, dark-mode strategies, and why new projects should prefer OKLCH over HSL.

aideep-dive

How to Rigorously Compare Before and After Agent Changes: From Golden Sets to Statistical Testing

Even with temperature=0, LLM outputs can still fluctuate by up to 15% in practice. To rigorously compare agent changes, you need a frozen golden set, at least 3 runs per query averaged out, LLM-as-judge blind evaluation (pairwise preference flip rate reaches 35%), and paired statistical tests -- not just running each version once and going by feel.

aideep-dive

Agent Observability: From OTel Traces to Catching Hallucinations, Tool Misuse, and Infinite Loops

The industry has converged on using OpenTelemetry GenAI semantic conventions to turn every LLM call and tool call into a span. Detecting the three major failure modes then splits into three tracks: faithfulness + semantic entropy for hallucinations, framework-level symbolic guardrails for tool misuse, and max steps + action hash deduplication for infinite loops — all wired into a Final / Trajectory / Single-step three-layer evaluation framework.

aideep-dive

Resource Rationality for Agents: Optimal Decisions Across Tokens, Tool Calls, and Latency

Agent decision-making under resource constraints is bounded rationality reborn: Rational Metareasoning uses VOC rewards to save 20-37% of tokens, BATS proves that adding budget without budget awareness is futile, FrugalGPT cascades cut costs by up to 98%, and Speculative Actions reduce latency by 20%. The three constraints ultimately converge into a single Pareto curve, and the overarching trend is moving from humans tuning knobs to models making resource-rational decisions on their own.

aideep-dive

The Single Crack in Agent Security: From Prompt Injection to Trust Boundaries to Multi-Agent Worms

Three seemingly distinct agent security problems — tool output injection, trust boundaries, malicious agents — share the same root cause: LLMs flatten instructions and data into a single token stream, making them architecturally unable to distinguish between the two. Understand this through-line and you can trace every attack from EchoLeak (CVE-2025-32711, zero-click) to the Morris II AI worm, and see why 'making the model behave' doesn't work — only architectural constraints (six design patterns, CaMeL) do.

aideep-dive

How Agents Decide Whether to Retrieve, What to Retrieve, and How to Merge: Three Decision Layers of Agentic RAG

Traditional RAG is a fixed pipeline of 'retrieve then answer.' Agentic RAG splits retrieval into three decision layers: when to retrieve (FLARE uses token probabilities; Adaptive-RAG uses a complexity classifier), what to retrieve (HyDE / RAG-Fusion / decomposition / Step-back), and how to fuse (RRF k=60 then cross-encoder rerank then compression -- Anthropic measured a -67% failure rate reduction). Key counter-intuitive insight: unnecessary retrieval hurts quality -- 'deciding not to retrieve' is a first-class capability.

aideep-dive

Stop Hand-Tuning Prompts: From GEPA to Tool Descriptions, Automating Agent Behavior Optimization

Automatic prompt optimization (APO) has evolved from APE/OPRO to GEPA: replacing sparse rewards with linguistic reflection, winning over GRPO by ~6pp with 4-35x fewer rollouts. Meanwhile, tool descriptions are the overlooked prompt -- small wording changes can shift tool selection rates by 10x, and Anthropic's experiments show Claude self-rewriting tool descriptions outperforms human experts. These two lines are converging: eval-driven automatic optimization is eating hand-tuned prompts.

aideep-dive

How to Build a Deep Research Agent: Multi-Turn Search Planning, Conflict Resolution, and Verifiable Conclusions

An autonomous research agent = four controllable stages: planning (decompose into sub-questions), retrieval loop (search -> read -> reflect on gaps -> search again), evidence arbitration (>=2 independent sources, typed conflict handling), and verifiable output (sentence-level citations + independent verification pass). Two approaches: training-based uses RL to learn end-to-end when to search (Search-R1 +41%); orchestration-based uses orchestrator-worker division of labor (Anthropic internal eval +90.2%, at ~15x token cost).

aideep-dive

Machine Theory of Mind: How Agents Infer Other Agents' Intentions, Knowledge, and Goals

Inferring another's beliefs/goals/intentions from observed behavior is called Machine Theory of Mind. Three lineages: symbolic BDI, Bayesian inverse planning, and deep learning ToMnet. The biggest controversy in the LLM era is that GPT-4 still trails humans by >10 points on ToMBench — are high scores genuine reasoning or statistical shortcuts?

aideep-dive

Multi-Agent Error Propagation and Recovery: Borrowing Thirty Years of Weapons from Distributed Systems

At 99% accuracy per step over 100 steps, the error-free completion rate drops to just 36% -- error compounding is a structural problem, not something prompt tuning can fix. Distributed systems' supervisor trees, bulkheads, circuit breakers, sagas, and durable execution can be mapped almost one-to-one into agent orchestration. But LLMs introduce a failure class that traditional systems never had -- semantic errors that don't crash -- which require Inspector agents (recovering 96.4%) and redundancy voting (MAKER: one million steps with zero errors) to address.

aideep-dive

Semantic Similarity ≠ Retrieval Relevance: Scenarios, Detection, and Remedies for Systematic Embedding Retrieval Failures

Cosine similarity and relevance systematically diverge across an entire class of scenarios: negation (most IR models score at or below random on NevIR), exact identifiers, numeric thresholds, and logical combinations (SoTA models achieve recall@100 < 20 on LIMIT) -- some of these hit the theoretical ceiling of the single-vector paradigm, and switching to a larger model will not help. Recommended remedy order: hybrid BM25 -> reranker (Anthropic measured -67%) -> upstream metadata routing -> domain fine-tuning / multi-vector.

aideep-dive

How to Pick the Right Tool from Hundreds: The Collapse Curve of Tool Selection and Engineering Solutions

As tools scale up, selection accuracy doesn't degrade gracefully — it collapses: 4 to 51 tools drops from 43% to 2%, 10 to 100+ drops from 78% to 13.62%. The root fix is to stop stuffing everything in at once — Anthropic's Tool Search Tool uses defer loading plus retrieval to cut 85% of tokens, pushing Opus 4.5 accuracy from 79.5% to 88.1%. Description quality has conditional payoff: negligible in simple scenarios, but correctness jumps from 44% to 50% in multi-tool chaining.

aideep-dive

A More Expensive Embedding Won't Save Your Traditional Chinese RAG: Three Layers of Failure and the Fix Order

Traditional Chinese RAG retrieval failures are a three-layer stack: embedding granularity defects (BGE/GTE from 0.1B to 7B all mis-rank on simple queries like 'fried chicken'), Simplified Chinese / English corpus dominance causing local vocabulary drift ('premium', 'exclusion clause' alignment is unreliable), and MTEB Chinese benchmarks being Simplified Chinese making model selection signals misleading. The fix is architectural: OpenCC normalization -> hybrid + jieba segmentation -> reranker -> local fine-tuning last -- and the prerequisite for all of it is building a Traditional Chinese eval set first.

aiguide

arXiv Paper Quality Assessment Guide: From Endorsement Mechanisms to a Practical Checklist

arXiv does not perform peer review, and roughly 2% of submissions are rejected. Quality judgment relies on external signals: top venue acceptance > institution + open-source reproduction > citation quality. Includes a 20-item practical checklist and a 2026 toolbox (PWC has shut down).

techdeep-dive

Bumblebee: A Design Teardown of Perplexity's Read-Only Supply Chain Endpoint Scanner

A Go read-only scanner open-sourced by Perplexity in May 2026 (v0.1.1, zero non-stdlib dependencies). It inventories npm/PyPI/Go/RubyGems/Composer/MCP/editor and browser extensions into NDJSON, matches against a custom exposure catalog, and answers the question 'which machines in my fleet are currently affected' the moment a supply chain incident hits. It deliberately never invokes any package manager and is not an EDR.

aideep-dive

Auto-Embedding on File Upload Is a Bad Default: A Survey of Adaptive / Agentic RAG and Agentic Parsing

Making 'chunk and embed every uploaded file automatically' the default behavior means making a decision for the LLM that it could have made itself. From Self-RAG (2310.11511) and Adaptive-RAG (2403.14403) to AgenticOCR (2602.24134), the academic trajectory is pushing three layers of decision-making -- whether to retrieve, whether to parse, and how to chunk -- from the ingestion pipeline back to the agent at conversation time.

aideep-dive

Assembling LLM Agent Skills / Tools / Code Interpreter for Real: A Paper Reading Map

The hard part of LLM agents is not building function calling, skills, code interpreter, and document tools individually -- it is assembling them into a system that selects the right tool, writes code when needed, decomposes tasks, verifies results, and resists prompt injection. This post organizes the key papers into six engineering decisions: function calling reliability, tool/skill selection, code-as-action, multi-step planning, skill systems, and safety plus document generation.

techdebug

Mobile Chrome Redirects Back to Login After Sign-In: Debugging an HTTP-to-HTTPS Entry Point Issue

When mobile Chrome keeps redirecting back to the login page after sign-in, the culprit isn't always OAuth or broken frontend state. In this case, the root cause was that the HTTP entry point for app-dev.daodao.so wasn't issuing a 301 redirect to HTTPS, so /auth/me requests sent with an http origin didn't include the auth_token cookie.

aideep-dive

A2UI (Agent-to-User Interface): Google's Open Protocol for Agents to Ship UI as Data

A2UI is an agent generative UI protocol open-sourced by Google on 2025-12-15: agents send declarative JSON describing UI intent, and clients render it natively using their own component catalog whitelist, layered on top of A2A. It launched at format v0.8 and iterated to v0.9 within three months.

browse.sh: Turning What Browser Agents Learn into a Skill Catalog

browse.sh, launched by Browserbase in May 2026, is two things: a browser skill catalog and the Browse CLI. The core thesis: the bottleneck for browser agents isn't reasoning — it's amnesia. By storing learned site-specific workflows as plain-text SKILL.md files, Autobrowse cut Craigslist task costs from ~$0.22 to ~$0.12 by their own metrics. Note: this has nothing to do with the 2018 Browsh text-mode browser.

aideep-dive

CodeGraph: Local Code Knowledge Graph, and the Truth About 'Walking the Graph to Save Money'

CodeGraph uses tree-sitter to extract a codebase into a local SQLite/FTS5 knowledge graph, letting AI coding agents query the graph instead of scanning files. The official end-to-end benchmark (7 repos, median of 4 runs) averages 35% cost savings and 70% fewer tool calls -- but only if the agent actually walks the graph. Delegating exploration to a file-reading subagent that ignores CodeGraph turns it into pure overhead.

aideep-dive

How Do People Read arXiv Papers? A Complete Guide to Methods and Tools

Reading papers is two problems stacked together: methodology (Keshav's three-pass method, 5-10 min / 1 hour / 4-5 hours) determines how to read, and tools (arXiv HTML, alphaXiv, NotebookLM, Connected Papers, Zotero) shorten the time for each pass. AI lowers the barrier to understanding; judging correctness always stays with the human.

Midscene.js: Betting on Pure Vision for Cross-Platform UI Automation

An MIT-licensed open-source UI automation framework from ByteDance. UI actions rely solely on feeding screenshots to a vision-language model, with no DOM parsing. A single JS API works across Web / Android / iOS / desktop. The trade-offs: each step is slower and more token-expensive, and everything hinges on the model's grounding ability. Note that Midscene retired MCP after 1.9.8 in favour of Skills + CLI.

Antigravity CLI: How Google Folded Gemini CLI Into a Unified Terminal Agent Harness

Antigravity CLI is a terminal agent Google announced at I/O on May 19, 2026. Written in Go (versus Gemini CLI's Node.js), its binary is called agy, and it shares the same agent harness as the desktop Antigravity 2.0. It is also Gemini CLI's successor — the personal-tier Gemini CLI service ends on June 18, 2026.

aideep-dive

How Claude Reads and Writes PDF / DOCX / PPTX: Deconstructing the Three-Layer Architecture of Skills + Sandbox

Claude has no docx_tool or pdf_tool -- it relies on bash + file tools, plus SKILL.md instructions and pre-installed libraries like pdfplumber / python-pptx inside the container, assembling file handling capabilities from three layers.

aideep-dive

Open Design: The Open-Source Claude Design Alternative Forked in 11 Days

Anthropic shipped Claude Design on 2026-04-17. On 4-28, nexu-io/open-design went public -- same artifact-first loop, Apache-2.0, runs on the 16 coding-agent CLIs you already have. Two weeks from 0.1 to 0.7, 40k+ stars. A paradigm shift that flattens AI design tools from vertical SaaS into a skill bundle.

aideep-dive

system_prompts_leaks Deep Dive: What Problem Does a 40k-Star AI System Prompt Archive Solve

asgeirtj/system_prompts_leaks collects the raw system prompts of 40+ AI assistants, from GPT-5.5 and Claude Opus 4.7 to Gemini 3.1 Pro, with 40.3k stars, 461 commits, and an MIT license. The value isn't in obtaining secrets -- it's in turning vendors' implicit policies into comparable engineering material. What you should study is the design decisions, not the text itself.

aideep-dive

Dissecting Anthropic's Founder's Playbook: Four Stages, Three Moats, and One Cowork Compliance Pitfall

Anthropic's 35-page startup handbook released 2026-05-14 reorganizes Idea/MVP/Launch/Scale around agentic AI. The most valuable takeaways are 'the easier it is to build, the more important validation becomes' and treating CLAUDE.md as the first MVP artifact. The part to discount: the Launch chapter puts compliance workstreams on Cowork -- but Anthropic's own docs say Cowork doesn't write audit logs.

techdebug

LLM Agent Tool Descriptions Determine Tool Selection: Three Bug Fixes

Rewriting tool descriptions from soft suggestions to hard rules (whitelist + consequence explanation) eliminated the LLM's incorrect tool selection; adding skip_signal=True fixed vector store double-indexing.

aideep-dive

Using AI Agents to Operate Video Generation Tools: A HyperFrames, HeyGen, and Runway Integration Guide

AI agents can operate video generation tools through three approaches — Skills, MCP Connectors, and direct APIs. Choosing the right integration method matters more than choosing the right tool.

aideep-dive

Code Mode: Moving Tool Definitions from Context into Code

Stop stuffing all your tool descriptions into context at session start. Let the model write code, have the runtime execute it, and let tool definitions enter context only at the import line — Anthropic's GDrive→Salesforce example dropped from ~150K tokens to 2K, and Cloudflare's 2,500-endpoint schema shrank from 1.17M to 1K.

aideep-dive

The FDE War: Why OpenAI and Anthropic Are Both Copying Palantir's Playbook

MIT research says 95% of enterprise AI pilots yield zero return. OpenAI and Anthropic announced multi-billion-dollar joint ventures in the same week, wholesale adopting the Forward Deployed Engineer model that Palantir has used for over a decade to bring AI into the enterprise battlefield.

aideep-dive

How Others Use LLMs to Write: Trade-off Notes from Karpathy's LLM-wiki to Multi-Agent Pipelines

A survey of 11 public LLM writing pipelines, distilled into three dominant patterns: multi-agent (researcher -> writer -> critic), Karpathy LLM-wiki (raw + wiki + LLM writes, humans don't), and quality guardrails (technical verifier + never fabricate + brief gate). The Princeton GEO paper (KDD 2024) quantifies the impact: inline citations +28%, adding statistics +33%, quoting source text +41%, keyword stuffing -9%.

aideep-dive

OpenAI's Codex Secure Deployment Strategy: Sandboxing, Auto-review, and Enterprise Governance

In May 2026, OpenAI published its internal Codex deployment practices: sandboxes define technical boundaries, approval policies determine when to pause, Auto-review delegates approval decisions to a sub-agent instead of a human, and Managed configuration lets enterprise admins enforce policies top-down. The core philosophy: zero friction for low-risk actions, mandatory review for high-risk ones.

aiguide

9Router: A Local 3-Tier Fallback Router That Routes Claude Code / Cursor / Cline to 40+ Providers

Spin up a local OpenAI-compatible endpoint at localhost:20128 that automatically routes requests from Claude Code / Cursor / Cline / Codex / Copilot through a Subscription → Cheap → Free 3-tier fallback to 40+ providers. Built-in RTK compresses tool_result (saving 20–40% input tokens), Caveman mode compresses output, OAuth auto-refresh, multi-account round-robin — install with npm install -g 9router and two commands.

Claude, Codex, and Gemini Are All in the Browser Now: Comparing Three AI Agent Approaches in Chrome

Three vendors originally took three routes: Anthropic built an extension, OpenAI built its own browser, Google welded AI into Chrome. By August 2026 there are only two — OpenAI's Atlas stopped working on 9 August, with its capabilities folded back into the ChatGPT desktop app and Codex. The remaining split is 'live alongside Chrome' versus 'be Chrome'.

aideep-dive

15 Walls for Building Your Own Auto-Dev Agent: Concrete Lessons from Stripe Minions

Stripe Minions says 'The walls matter more than the model,' but the case studies from four Silicon Valley companies never explained how to actually build those walls. This post breaks down the 15 walls we implemented in the daodao auto-dev agent: what each wall prevents, where the files live, and what the tradeoffs are. Tier 1 is mandatory, Tier 2 strengthens governance, Tier 3 is serious governance.

aiguide

What Is an Auto-Dev Agent? An Intro to daodao's Automated Development System

A PM checks a task card in Notion → the system syncs it to a GitHub issue → writes a plan → writes code → opens a PR for human review. This post explains what the system does, what it doesn't do, and why it's feasible now — written for people who don't write code.

aiguide

Step-by-Step: Build a Notion → PR Auto-Dev Agent — A Reproducible Version of the daodao Pipeline

Build a Notion task → GitHub issue → spec PR → code PR auto-dev agent from scratch. Using the daodao case as a template, this guide walks through every step — what to do, what to verify, and how to handle problems. Notion DB schema → bin/ scaffold → two Claude Code routines → cloud env vars → staging tests.

aideep-dive

Claude for Financial Services: Dissecting Anthropic's Multi-Agent Reference Implementation

Anthropic open-sourced 12 financial-industry Agents and 11 MCP connectors. The real takeaway isn't the Agents themselves but the layered design of 'one prompt, two runtimes' and 'pure-file extensibility.'

aiguide

From Plan to PR: Building daodao's Auto-Dev Agent in Practice

5 rounds of consensus to write the plan, then team mode with 5 workers running 12 tasks in parallel — with plenty of pitfalls along the way. Writing it down for my future self and anyone else trying the same thing.

aideep-dive

DeepSeek-OCR: The 10x Compression Experiment That Turns Long Context into Images

DeepSeek-OCR's paper is titled Contexts Optical Compression -- OCR is just the means; what it actually validates is that 'rendering text as images and feeding them to a VLM' achieves 10x compression at 97% accuracy. This is a qualitative shift for long-context LLM and RAG token costs.

aideep-dive

2026 LLM Inference Provider Free Tiers & Pricing: 40+ Services Ranked by Tier

For side projects, toy demos, and RAG prototypes, nobody wants to swipe a credit card on day one. This is a verified roundup of 40+ LLM inference providers still operating as of 2026/05, tiered by whether free resources auto-replenish or are one-time grants. Each entry notes credit-card requirements, supported models, paid starting prices, and catches. Chinese-origin providers including Zhipu GLM (permanently free), Doubao (2M tokens/day), Kimi, DashScope, and the Ollama local option are all included.

Claude Code /loop: Turning AI into a Background Worker with Native Scheduling (v2.1.72+)

/loop is Claude Code's native cron feature — set schedules in plain English and let Claude monitor, auto-fix PRs, and run recurring tasks in the background. Session-scoped and expires after 7 days; for cross-session scheduling, use Routines or Desktop scheduled tasks.

Claude Code Routines: Complete Guide to Cloud Automation — Setup, Triggers, and Real-World Examples

Routines is Claude Code's cloud automation system (formerly Cloud Scheduled Tasks). Beyond cron scheduling, you can trigger runs via API endpoint or GitHub events — scan issues, review PRs, run checks, open PRs — all while your computer is off.

aideep-dive

Claude Skills: Package Domain Knowledge into a Folder, Teach Once and It Remembers

A Skill is a folder with a SKILL.md. Three-layer progressive disclosure lets Claude load details only when needed, eliminating the need to re-explain preferences every conversation.

Local Deep Research Walkthrough: A Privacy-First Deep Research Agent

Local Deep Research is a privacy-first deep research agent built on LangChain + LangGraph, integrating 20+ search engines and 30+ research strategies. Its flagship langgraph_agent_strategy takes the LLM-autonomous tool-calling approach, offering a fundamentally different paradigm from fixed-pipeline RAG graphs.

PageIndex: RAG Without Vectors — Turning Long Documents Into a Book With a Table of Contents

PageIndex skips chunking, embedding, and vector storage entirely. Instead it relies on LLM reasoning over a tree-structured table of contents the LLM itself wrote, reporting 98.7% on FinanceBench in its own vendor-run evaluation. It solves a different problem than vector RAG — finding the right section in a well-structured long document.

techdeep-dive

Accessing Your Home Mac From Anywhere: Cloudflare Tunnel and the Alternatives in 2026

Two answers stand out for remotely accessing your home Mac in 2026: Cloudflare Tunnel if you need browser-based access with no client install, and Tailscale if you just want something simple for personal use. This post compares both, covers ZeroTier, Pangolin, NetBird, and other alternatives, and explains why Cloudflare's remotely-managed tunnel makes setup significantly easier in 2026.

Search MCP Tools for AI Agents: What to Do When WebFetch / WebSearch Gets Blocked

When using AI agents like Claude Code or Cursor, built-in WebFetch / WebSearch often gets blocked by Cloudflare, geo-restrictions, or rate limits. Connecting a search MCP server is the most direct fix. This post compares the options actually available in 2026.

aiguide

Groq Console: The Developer Platform for Running Open-Source Models on LPU Inference

Groq Console is the developer portal for Groq's in-house LPU chip, offering an OpenAI-compatible API, Playground, and free tier credits. Its selling point is running open-source models like Llama, Qwen, and DeepSeek at the fastest tokens/second on the market.

Warp: From Modern Terminal to Agentic Development Environment

Warp evolved from a Rust-powered modern terminal into an AI Agent-integrated development environment (ADE), open-sourced under AGPL in April 2026, with over 700,000 developer users.

goose: Open-Source, Cross-Platform, LLM-Agnostic Local AI Agent

goose is an open-source AI Agent maintained by the Linux Foundation's AAIF, supporting 15+ LLM providers and 70+ MCP extensions, built with Rust as a Desktop App + CLI + API. It positions itself as a vendor-neutral, self-hostable alternative to Claude Code.

Gemma on Cloudflare Workers AI: A Pragmatic Choice for Traditional Chinese Applications

For running Traditional Chinese LLM workloads on Cloudflare Workers AI, the Gemma family follows instructions more reliably than same-tier Llama models. gemma-3-12b-it was marked deprecated on 2026-05-30; the current equivalent is gemma-4-26b-a4b-it: 256K context, Vision, Function calling, at $0.10 / $0.30 per M tokens.

aideep-dive

Knowledge Management with LLMs: From Karpathy's llm-wiki to the Open-Source Ecosystem

Karpathy proposed the llm-wiki pattern in 2026, having LLMs proactively maintain a markdown wiki instead of running RAG from scratch every time. Over 100 open-source implementations now exist, ranging from local CLI tools to serverless Telegram bots.

aideep-dive

OpenAI Workspace Agents: From Custom GPTs to a Team Automation Platform

On 2026/4/22 OpenAI launched Workspace Agents — powered by Codex, capable of long-running cloud execution, and integrating with Slack/Salesforce/Google Drive. They are the enterprise successor to Custom GPTs.

Building a Legal Contract RAG in 36 Hours: Weaviate Query Agent + ColQwen Architecture Breakdown

Using Weaviate Query Agent + ColQwen multi-vector model, a single prompt built a production-grade legal contract search system in 36 hours -- this post breaks down its architecture logic, technology choices, and what you actually need to watch out for.

marketingguide

AKIRAXCLAW's Content Model: 5 Posts a Day, a Three-Tier Funnel, and Agent-Assisted Publishing

Akira runs a Threads → Blog → Docs three-tier funnel with agent-assisted publishing, building a sustainable knowledge monetization model in the Chinese-language AI content market.

aiguide

Where AI Code Review Stands Now: Lessons from Cloudflare's Multi-Agent System

Cloudflare ran a Multi-Agent Code Review system internally for 30 days — 131K reviews, median 3 minutes. This post breaks down their architecture and compares it with solutions from Anthropic, GitHub, CodeRabbit, Greptile, and others.

aiguide

Inside the Codex Agent Loop: How OpenAI Keeps AI Agents Iterating

A detailed look at OpenAI's Codex agent loop design: how prompts are constructed, how multi-turn conversations are managed, how prompt caching prevents cost explosions, and how context window auto-compaction works.

aiguide

Codex App Server: How OpenAI Turned an Agent Harness into a Universal Protocol

OpenAI wrapped the Codex harness as a JSON-RPC over stdio App Server, enabling VS Code, JetBrains, Web, and desktop apps to share a single agent loop. Three core primitives: Item, Turn, and Thread.

OpenAI Wrote 1 Million Lines of Code with Codex: Harness Engineering in Practice

An OpenAI internal team spent 5 months with 3 people and 0 lines of hand-written code, delivering a complete product using Codex. This article distills their core lessons on AGENTS.md design, repo-local knowledge bases, architecture enforcement, and entropy management.

AEO / GEO Tool Landscape: Input, Traffic, and Output Layers — From isitagentready to aeo-radar to Profound

AEO/GEO tools aren't a single category — they span three distinct layers: the input layer (is your website ready for AI to read), the traffic layer (how much are AI bots actually crawling), and the output layer (how is your brand mentioned in AI answers). This post maps out all three layers, from open-source self-hosted options to commercial SaaS.

techproject

DeerFlow: ByteDance's Open-Source Super Agent Harness for Long-Running Research Tasks

DeerFlow is ByteDance's open-source Super Agent Harness built on Python 3.12 + LangGraph. It orchestrates long-running tasks through sandboxes, long-term memory, sub-agents, skills, and a messaging gateway. It hit #1 on GitHub Trending in February 2026, now surpassing 63,000 stars, with support for Telegram/Slack/Feishu, Claude Code integration, and multiple search backends.

travelguide

2026 Travel Inconvenience Insurance Guide: New Rules, Coverage Comparison, and Where to Buy

2026/4/1 new rules: max 2 policies per trip (different insurers), flat-rate payout cap lowered to NT$6,000. Covers six key areas including flight delays and lost luggage, with a breakdown of where to buy.

Agentic Engineering: Making AI Agents Collaborate Like a Real Engineering Team

Agentic Engineering isn't about making AI write code faster — it's about making software move through the entire delivery pipeline faster, by using multi-agent collaboration to compress cross-team coordination friction.

The Memory Problem in Agentic Engineering: Types, Implementation, and Ownership

Agent memory isn't a plugin — it's part of the harness itself. Pick the right memory type, estimate data volume, then decide on the technology. And finally, figure out whether you actually own that memory.

aiguide

Multi-Engine Code Review with Codex + Gemini + Claude: Principles, Patterns, and Implementation

AI models rationalize their own code when reviewing it. Using three different CLIs for independent review effectively catches blind spots -- this post covers the design philosophy and practical workflow patterns behind the approach.

techguide

How Does the YouTube to NotebookLM Extension Work? Reverse Engineering and Cross-Tab Architecture Dissected

NotebookLM has no official API. This extension works by combining three techniques: reverse-engineered Google batchexecute RPC calls, DOM scraping, and cross-tab message passing.

techdebug

Local AI Backend API Always Returns Empty Data: Cookie Domain Isolation

The main backend runs on a remote HTTPS server, so the auth_token cookie is scoped to that domain. The browser never sends it to the local AI backend, causing the API to treat every request as unauthenticated.

aiguide

Integrating AI Agents into Your Development Workflow: A Five-Phase SDLC Breakdown

Agentic AI is not just autocomplete — it is an AI system capable of autonomously executing multi-step tasks. This article breaks down the five phases of the SDLC, explaining where to plug in agents at each phase, how to progress from CLI tools to full-pipeline automation, and the most valuable external resources to track right now.

aiguide

A Book Written by AI Itself, Teaching You How to Build Software with AI

Encyclopedia of Agentic Coding Patterns catalogues 190 patterns to help you make the right software decisions in the age of AI-written code — and the book itself is autonomously written and maintained by an AI agent.

aiguide

GitHub Copilot Coding Agent: Assign an Issue to AI and Let It Open the PR

GitHub Copilot Coding Agent lets you assign an Issue to Copilot, which then automatically creates a branch, writes code, runs CI, and opens a PR — all inside a cloud sandbox. The key to success is setting up AGENTS.md; without it, the agent tends to go off track. Best suited for well-defined medium-sized tasks; requires Pro+ (1,500 premium requests/month) or Enterprise plan.

aiguide

knowledge-pipeline: A Six-Layer Pipeline for RAG Quality Control

A six-layer deterministic pipeline that handles everything from URL ingestion to vector embedding automatically, filtering out garbage before it enters your RAG system through an eight-dimension scoring system.

MarkItDown: Convert Any File to Markdown Before Feeding It to an LLM

A lightweight open-source tool from Microsoft that converts PDF, Office, images, audio, and more into Markdown — purpose-built for LLM pipelines.

aiguide

MCP vs CLI vs API: The Real Boundaries of Agent Tool Interfaces

MCP is not going away, but its effective scope is narrower than most people think. For local development, CLI and raw API almost always beat MCP. MCP's truly irreplaceable niche is the narrow gap of 'cross-agent shared local tool layer.'

Is Your JSON-LD Invisible to AI Search Engines? A Pipeline Breakdown and AEO/GEO Strategy

Different AI engines process web pages in vastly different ways. Some only read the body; others rely on pre-built indexes. JSON-LD and schema markup are not universally effective — body content quality and structure are the only cross-platform foundations that hold.

productproject

quidproquo Blog Improvement Roadmap: Content, Technical Debt, RAG Design, and Harness Infrastructure

Using my own 30+ RAG/Agent posts to audit the blog itself, I identified a prioritized improvement list spanning content quality, site tech, RAG design fixes, harness infrastructure, and AI agent applications — no phases, just priorities.

aiguide

Lessons from the Trenches: What AI Native Teams Must Get Right

Not everyone should use a coding agent to modify code directly. AI Native teams need interface specs, test-first development, monorepo, security guardrails, human-in-the-loop, and token budget controls. Building an agent platform layer on top of coding agents and clearly redefining developer roles is the right path forward.

aiguide

Autoreason: Teaching LLMs When to Stop Self-Refining

Autoreason replaces the traditional critique-and-revise loop with a competitive multi-version evaluation mechanism (A/B/AB + blind Borda count), solving three structural problems in LLM self-refinement: prompt bias, scope creep, and lack of restraint.

aiproject

Vercel Open Agents: Moving the Coding Agent from Your Laptop to the Cloud

An open-source coding agent reference implementation from Vercel Labs. A three-layer architecture separates the web UI, agent workflow, and sandbox VM — designed as a starting point for teams that want to self-host their own Claude Code or Cursor Background Agent.

The Full Picture of Cloudflare Workers AI Binding: It's More Than Just run()

env.AI is not just run(). It also exposes toMarkdown (document-to-Markdown conversion), autorag (managed RAG), gateway (external provider proxy), and models (metadata lookup). Understanding these four method groups is what unlocks Cloudflare as a full AI platform inside Workers.

aiguide

Claude Octopus: The Consensus Plugin That Hooks 8 Models Into Claude Code Simultaneously

Claude Octopus is a Claude Code plugin that simultaneously calls Codex, Gemini, Copilot, Qwen, Ollama, Perplexity, OpenRouter, and Claude to review the same code, using a 75% consensus threshold to catch single-model blind spots. It ships with 32 personas, 48 /octo:* slash commands, 51 skills, and a Dark Factory fully autonomous spec-to-code pipeline.

aiguide

LLM Council: Karpathy's Weekend Multi-Model Parliament — Three Stages of LLM Peer Review

LLM Council is a local Web App Andrej Karpathy built over a weekend. It sends one question to multiple LLMs simultaneously, has them anonymously peer-review each other, and then a Chairman model synthesizes a final answer. Positioned as a small tool for comparing models while studying — 99% vibe coded with no plans for long-term maintenance — but the architecture itself is a minimal ensemble LLM implementation worth studying.

techguide

Better Agent Terminal: Consolidate Multiple Project Terminals and Claude Code Agents into One Window

Better Agent Terminal (BAT) is an Electron desktop app that unifies multiple project workspaces, terminals, and Claude Code Agents into a single window — solving the everyday pain of exploding iTerm tabs and the lack of a proper GUI container for agents. MIT License, available on macOS, Windows, and Linux.

aiguide

Claude Managed Agents: Letting Anthropic Handle the Agent Shell and Sandbox

Claude Managed Agents is a beta service launched by Anthropic on 2026/04/08 that provides an agent harness plus cloud container sandbox, billed per token plus $0.08/session-hour. It suits long-running async tasks and is worth exploring if you don't want to build your own agent loop and sandbox.

aiguide

Agent Skills: A Skill Framework That Makes AI Agents Work Like Senior Engineers

Agent Skills is Addy Osmani's open-source collection of 19 production-grade engineering skills that drive AI agents to follow senior engineering discipline through /spec → /plan → /build → /test → /review → /ship commands, instead of cutting corners.

aiguide

Graphify: Turn Code and Documents into a Queryable Knowledge Graph

Graphify uses tree-sitter AST to extract code structure, then applies LLM semantic analysis to documents and images, compressing an entire project into a queryable knowledge graph. It claims to save 71.5x tokens per query compared to reading raw files.

Claw Code: An Open-Source CLI Agent That Rewrites Claude Code in Rust

Claw Code is a from-scratch Rust rewrite of the Claude Code CLI, featuring 48K lines of code, 40 tools, and MIT licensing. Most remarkably, the entire project was built by multiple AI agents collaborating over just 5 days, surpassing 170K GitHub stars within a week of launch.

aiguide

clawhip: An Event Notification Router That Keeps Multi-Agent Development Under Control

clawhip is a Rust daemon that routes AI coding agent events (commits, PRs, session status) to Discord / Slack, solving the observability problem of not knowing who is doing what when multiple agents run in parallel.

aiguide

notebooklm-py: An Unofficial Python API for Google NotebookLM

notebooklm-py reverse-engineers Google's batchexecute RPC protocol, letting you programmatically control NotebookLM via Python / CLI / AI Agent — including audio, video, slides, quiz generation and more.

aiguide

oh-my-claudecode: An Enhancement Layer That Turns Claude Code into a Multi-Agent Collaboration Platform

oh-my-claudecode (OMC) adds 8 collaboration modes, 19 specialized agents, and cross-model orchestration (Claude + Codex + Gemini) on top of Claude Code, transforming a single-user CLI tool into a multi-agent development platform. Features include Deep Interview for requirement clarification, Smart Model Routing that saves 30-50% on tokens, and automatic rate limit recovery.

aiguide

oh-my-codex: A Structured Workflow Enhancement Layer on Top of OpenAI Codex CLI

oh-my-codex (OMX) doesn't replace Codex CLI — it adds a structured workflow layer on top of it. From requirements clarification and plan generation to multi-agent parallel execution, four core Skills transform scattered prompt conversations into a trackable development process.

aiguide

oh-my-openagent: A Multi-Model Agent Team Framework That Replaces Single-LLM Coding

oh-my-openagent (OmO) transforms OpenCode from a single-LLM tool into a multi-model agent team — Opus as the workhorse, GPT-5.2 as the architect, Gemini for frontend, Sonnet for documentation lookup — all triggered to run in parallel with a single ultrawork keyword. With 48K stars, it is the earliest project in the UltraWorkers ecosystem to establish the multi-agent coding pattern.

aiproject

OpenHarness: A Fully Open-Source Agent Harness Framework

An open-source Agent Harness framework from HKUDS (HKU Data Science Lab) that implements tool calling, skill loading, memory, permissions, and multi-agent collaboration as complete infrastructure, supporting Anthropic / OpenAI / GitHub Copilot API formats.

techguide

Solving Duplicate Config Files for Codex and Claude Code with a Symlink

Claude Code only reads CLAUDE.md; Codex only reads AGENTS.md. Teams using both end up maintaining two identical files. Fix: make CLAUDE.md a symlink pointing to AGENTS.md — one source of truth.

aiguide

How to Use Claude Code Agent Teams? Design Patterns from 6,400+ Agents on GitHub

There are already 6,400+ .claude/agents/*.md files on GitHub. We dissected 4 representative projects — ChemistryTimes (content production pipeline), claude-sub-agent (document-driven development pipeline), agentic (Temporal.io DAG parallel execution), and vs-copilot-multi-agent (hook-enforced memory persistence) — plus ruflo's enterprise-grade swarm architecture, distilling 6 design patterns and 5 practical trends.

From Stripe to Meta: How Silicon Valley's Top Companies Replace Keyboards with AI Agents

Top Silicon Valley companies are independently building internal AI coding agents that automate everything from a Slack message to a merged PR. This article deep-dives into architectures from Stripe, Ramp, Coinbase, and Spotify — including their 2026 growth numbers (Stripe 7,000+ PRs/week, Ramp 75% of merged PRs) — then expands to cover Google, Meta, Amazon, Uber, Shopify, PostHog, and more.

aiguide

Three Modes of LLM Knowledge Bases: Knowledge Vault, Experience Vault, and Blog

Andrej Karpathy proposed a framework for compiling personal knowledge wikis with LLMs — collect raw data, have the LLM compile it into .md wiki pages, run Q&A against the wiki, and file outputs back. This post compares three practical approaches: Karpathy's knowledge vault model, the community's experience vault model, and quidproquo's blog model.

aiguide

AI Agent Caching Goes Beyond One Layer: From Claude Code's 18 Cache Types to Multi-Layer ReAct Agent Design

After dissecting Claude Code's 18+ caching mechanisms, I found that you can't touch provider-level prompt cache, but embedding cache, tool result cache, and entity cache are not only within your reach — they deliver even better results. Includes a complete AgentCache interface design and per-tool TTL strategy.

aiguide

AI Agent Tool Descriptions Shouldn't Be Static: Dynamic prompt() Design Learned from Claude Code

Every one of Claude Code's 45 tools uses a prompt() method that dynamically adjusts based on user type, feature flags, and system capabilities. Applying this pattern to a ReAct Agent, tool descriptions are dynamically generated along three dimensions: orchestrator model capability, locale, and available tools. Small models automatically get few-shot examples; large models save tokens.

Claude Code Complete Breakdown: The Deep Reasoning King of Terminal Agents

Claude Code runs from $20/mo Pro to $200/mo Max 20x. Quota is a rolling five-hour window with weekly limits on top, shared across Claude on web, desktop, mobile, and the terminal. When you run out you can switch to usage credits at standard API rates rather than stopping.

Cursor CLI Complete Analysis: The All-Rounder Extending IDE Agent to the Terminal

Cursor CLI brings the IDE agent to the terminal with an interactive TUI and headless mode, Plan/Ask/Agent modes, Cloud Handoff, and CI/CD integration. Billing now runs on two separate usage pools: Cursor's own models (Grok 4.6/4.5, Composer 2.5) and third-party models (Pro includes $20, Pro+ $70, Ultra $400).

Google's Terminal Agent Plans: The Free Individual Path Is Gone

The paying paths on Google's side: the individual free tier and Gemini CLI access on Google AI Pro / Ultra ended 2026/6/18, leaving individuals with Antigravity CLI or their own paid API key; enterprise licenses and Google Cloud are unaffected. The zero-cost starting option now belongs to someone else.

Kiro (AWS) Complete Analysis: The Spec-Driven Agentic IDE

Kiro has five tiers: Free 50 credits, Pro $20/1,000, Pro+ $40/2,000, Pro Max $100/5,000, and Power $200/10,000, with add-on credits at $0.04. Auto mode mixes models to cut cost (the same task costs 1.3x credits via Sonnet), and the spec-driven flow turns vibe coding into traceable, structured development.

OpenAI Codex Complete Plan Analysis: Agent Integration in the ChatGPT Ecosystem

Codex rides your ChatGPT subscription (Free / Go $8 / Plus $20 / Pro 5x $100 / Pro 20x $200), and since 2026/4/2 billing is token-based credits. The model line is GPT-5.6 Sol / Terra / Luna; GPT-5.4 and 5.4 mini retire from ChatGPT-signed-in Codex on 2026/8/31.

OpenCode Full Analysis: An Open-Source Terminal Agent Supporting 75+ Model Providers

OpenCode is a free, open-source TypeScript CLI agent (MIT, ~198K GitHub stars). It supports 75+ model providers including local Ollama, allows authentication via Copilot/ChatGPT accounts, and lets you switch models mid-session without losing context. There is also a desktop app and an official Zen gateway.

Agent CLI Subscription Plans Compared: Building a Flexible Multi-Model Routing Strategy

A comparison of six agent CLI subscriptions (Claude Code, Cursor CLI, Codex, Kiro, Antigravity/Gemini CLI, OpenCode) plus the multi-model routing pattern — cheap models for simple work, strong models for hard work. Nearly every one of these changed its billing in the first half of 2026; this version was re-verified on 8/18.

aiguide

2026 Personal AI Hardware Buying Guide: DGX Spark, Mac Studio, MSI AI Edge Compared

Comparing the NVIDIA DGX Spark, Apple Mac Studio M4 Ultra, ASUS Ascent GX10, MSI AI Edge, and more — helping you find the right local inference hardware.

Multi-Model Routing Open-Source Tools & Implementation: Getting the Right Model for the Right Job

With multi-model routing, 70% of simple tasks are directed to cheap models, and only 10-15% of complex tasks use flagship models — saving 40-85% on inference costs in practice. This article covers the architecture and implementation of five major open-source tools.

productproject

Digital Ecosystem Research: Dissecting Platform Integration Strategies from LINE and Shopify to Taiwan MarTech

A breakdown of the three-layer digital ecosystem structure: LINE's super-app, Shopify App Store flywheel, and Taiwan MarTech integration strategies. The core mechanism is using APIs and data flows to create mutual dependency among participants, collectively reinforcing the moat.

techguide

Where Should AI Agent Global Skills Live? The Division of Labor Between .claude, Codex Skills, and AGENTS.md

Skill paths are almost always runtime-specific. AGENTS.md is the reliable way to share rules across agents. Put personal reusable capabilities in each agent's supported global directory; put project workflows inside the repo.

techguide

code-review-graph: Using a Knowledge Graph to Cut AI Code Review Token Usage by 8x

code-review-graph uses Tree-sitter to parse your codebase and build a persistent knowledge graph, tracks the blast radius of changes, and feeds only truly relevant context to the AI — claiming an average 8.2x reduction in token usage.

techguide

GitBook: A Documentation Platform That Turns Docs into a Product

GitBook is a Git-based documentation platform with Markdown editing, version control, and multi-user collaboration. Ideal for technical docs, API references, and internal knowledge bases. The free plan is sufficient for individuals and small teams.

techguide

NVIDIA DGX Spark: A Desktop AI Supercomputer That Fits a Petaflop on Your Desk

The NVIDIA DGX Spark is powered by the GB10 Grace Blackwell Superchip, 128 GB of unified memory, and delivers 1 petaFLOP of FP4 compute — starting at around $3,999 USD. It lets developers run 200B-parameter models locally and fine-tune 70B models, making it the most accessible NVIDIA AI development platform available today.

techguide

Documentation Platform Guide: GitBook, Docusaurus, Mintlify, and Seven Other Options

A breakdown of nine major documentation platforms — their positioning, pros, cons, and ideal use cases. Decision logic: open-source projects → Docusaurus/VitePress, API docs → Mintlify/ReadMe, internal enterprise → Confluence, fastest to launch → GitBook.

The Complete Guide to Agent CLIs: Design Logic, Tool Comparison, and Best Practices

Agent CLIs are not smarter autocomplete tools -- they are AI agents that can read your codebase, execute multi-step tasks, and operate in real environments. Claude Code, Codex CLI, Gemini CLI, OpenCode, Aider, Pi, Kiro, Amp, Cursor CLI... the tools keep multiplying, but they all share a common set of design principles -- understanding these principles is how you actually get good at using them.

aiguide

15 Agent Frameworks Worth Watching in 2026

Sorted by GitHub Stars, a survey of 15 mainstream AI Agent frameworks in 2026 — their positioning, key features, and ideal use cases. Not a ranking — it's a map.

aiguide

One Sentence to an IG Carousel — From 3 Hours Manual Work to a Fully Automated Pipeline

Use Claude Code as an orchestrator to chain Playwright screenshots, catbox.moe image hosting, Meta Graph API publishing, and Telegram notifications — generate and publish an IG carousel from a single sentence.

aiguide

llama.cpp — From Pure C++ to an LLM Inference Engine on Consumer Hardware

llama.cpp is the most widely used local LLM inference engine, implemented in pure C/C++. It supports CPU, Metal, CUDA, Vulkan, and other backends, and uses the GGUF quantization format to run multi-billion-parameter models on consumer hardware.

aiguide

TurboQuant+ — Two-Stage Quantization to Compress KV Cache to 2-bit, Running 100B Models on a MacBook

TurboQuant+ is an open-source implementation of a Google Research ICLR 2026 paper that uses PolarQuant + QJL two-stage quantization to compress the KV cache by 3.8-6.4x, enabling consumer hardware to run larger models with longer contexts.

aiguide

Small Models That Run on Phones: Choices and Constraints in 2026

The main on-device LLMs in 2026 are Gemma 3n, Qwen 3.5 Small, Llama 3.2, Phi-4-mini, Ministral 3, and SmolLM3. Sub-3B quantized models can hit 30-50 tokens/sec on phones with 8GB RAM, but RAM, thermal throttling, and context window remain hard constraints.

aiproject

2026 Q1 Open-Source LLM Landscape: From Frontier Models to On-Device, a Complete Survey

2026 Q1 saw a full-blown open-source model explosion: on the LLM front, GLM-5, Kimi K2.5, and Qwen3.5 caught up with closed-source models; Embedding and Reranker are dominated by Qwen3 and BGE; speech has Voxtral TTS and Whisper V3; image has FLUX.2; and video has Wan 2.2 rivaling Sora. This is the complete navigation map.

Claude Code: A Complete Guide to Anthropic's Terminal AI Coding Agent

Claude Code is Anthropic's agentic coding tool that runs in the terminal, IDEs, Slack, GitHub, and on the web. Its core extension system has six layers: CLAUDE.md (persistent context), Skills (on-demand workflows), Hooks (deterministic automation), Subagents (isolated delegation), MCP (external tool connections), and Agent Teams (multi-agent collaboration).

Codex CLI: A Complete Guide to OpenAI's Open-Source Terminal Coding Agent

Codex CLI is OpenAI's open source terminal coding agent (Rust, Apache-2.0, ~106.6k stars) with MCP, subagents, image input, code review, and Skills. The model line is now GPT-5.6 Sol / Terra / Luna, and the desktop app, CLI, and IDE extension share one config.toml.

Gemini CLI: Once the Most Generous Free Terminal Agent, Now Enterprise-Only

Gemini CLI is Google's open source terminal AI agent (Apache 2.0, ~106.6k stars). It once offered 60 requests per minute and 1,000 per day for free, with a 1M context window. The individual tier stopped serving on 2026/6/18 and Antigravity CLI took over. The project isn't shut down — the repo is still maintained — but it now serves only Gemini Code Assist Standard/Enterprise licenses and paid API keys.

OpenCode: A Complete Guide to the Open-Source AI Terminal Coding Agent

OpenCode is an open-source AI coding agent written in TypeScript (MIT, ~198K GitHub stars, repo at anomalyco/opencode) with a built-in TUI, 75+ LLM providers, LSP integration, a Vim-style editor, SQLite session management, and a desktop app. Free, no subscription, local or cloud models.

Pi Coding Agent: A Minimalist Open-Source Terminal Coding Harness

Pi is a minimalist coding agent by Mario Zechner (TypeScript, MIT, ~93K stars) with just 4 core tools and a very short system prompt — everything else you add yourself via Extensions, Skills, and Prompt Templates. It deliberately omits MCP, sub-agents, plan mode, and permission popups. The repo is now earendil-works/pi and the npm scope is @earendil-works.

AI-Ready Content: The Complete Guide to Making Your Website an AI-Readable Data Source

In 2025-2026, websites need to be readable not just by humans but by AI. From llms.txt and Schema Markup to GEO and RAG ingestion pipelines, this post maps out the complete technical landscape for turning your website into an AI-consumable data source.

Advanced Harness Engineering Patterns: Tool Registry, Guard System, and Checkpoint-Resume

A Harness is more than just an LLM wrapper. Tool Registry manages dynamic tool loading and selection, Guard System establishes a four-layer defense network, and Checkpoint-Resume enables long-running tasks to survive interruptions. These three patterns form the critical infrastructure of production-grade Agent systems.

aiguide

Skill vs Subagent: Comparing Two Agent Collaboration Modes in Claude Code

A Skill is a prompt template you invoke manually. A Subagent is an independent agent that Claude routes to automatically. They look similar, but differ completely in trigger mechanism, tool isolation, and context management.

aiguide

Ticketing Is Dead — Review Is the New Planning

When AI agents can turn intent into a PR in minutes, the bottleneck in software engineering flips from 'planning what to do' to 'evaluating whether the output is correct.' Artifacts of the ticketing era — sprints, story points, backlog grooming — are collapsing to zero, replaced by review as the core practice.

Claude Code Spinner Verbs: The Complete List of 185 Status Verbs Extracted from Source Code

When processing requests, Claude Code randomly displays one of 185 built-in verbs (like Thinking, Brewing, Clauding), then picks one of 8 completion verbs with elapsed time. You can customize these via spinnerVerbs in settings.json, using either replace or append mode. All data in this post is verified directly from cli.js source code.

techguide

gstack — Garry Tan's 20 Skills That Turn Claude Code into a Virtual Engineering Team

gstack is Garry Tan's open-source Claude Code skills toolkit. Its 20 specialized skills transform a solo developer into an entire engineering team — automating everything from product planning and design review to code review, QA, and deployment.

Anthropic's Harness Design: Making AI Agents Work Like Engineers

The same model produces dramatically different results under different harness designs. Anthropic uses a dual-agent architecture, cross-session state files, and a GAN-inspired generator-evaluator loop to let Claude autonomously complete hours-long software development tasks.

aiguide

Google's Eight Multi-Agent Design Patterns

Google outlined eight multi-agent design patterns: from the simplest Sequential Pipeline to the composable Composite Pattern. More complexity isn't always better — picking the right pattern matters more than stacking agents.

From Prompt to Harness: The Three Evolutions of AI Engineering

AI engineering has gone through three phases: Prompt Engineering (write better instructions) → Context Engineering (feed the right information) → Harness Engineering (design the entire working environment). Each evolution doesn't replace the previous one — it operates at a higher level of abstraction.

The OpenClaw Agent Loop: Serialization, Writer Claims, and the Fence That Stops a Stale Turn From Committing

The agent loop is a serialized per-session run. The part worth studying is how it handles concurrency: an admitted run records an activeWriterRunId claim, every transcript write supplies expectedWriterRunId, and the commit transaction verifies the match — so a superseded run cannot commit stale data.

OpenClaw Agent Runtime: The System Prompt Is Assembled, and a Cache Boundary Cuts It in Half

OpenClaw builds its own system prompt for every run; there is no runtime default prompt. What it builds is split by an internal cache boundary — the stable workspace prefix above, the per-turn channel context below — so backends with prefix caches can reuse the same prefix across channels.

OpenClaw Access Control: SecretRef Is Not Process Isolation — Here's What It Actually Solves

SecretRefs keep credentials out of plaintext config, and the model-call chain sees process-local sentinels instead of the real value. But the docs say it plainly: this is not process isolation — the real value still exists in the same process's memory, and any plaintext file the agent can read bypasses the whole mechanism.

OpenClaw Automation, Part 1: Choosing Among Six Mechanisms, and Why 'Exactly on Time' and 'Check on It' Are Different Jobs

Cron is now called Automations (openclaw cron remains an alias), and automation spans six mechanisms. The core trade-off is one line: Automations give you exact timing and isolated execution, Heartbeat gives you full main-session context on a roughly-every-30-minutes cadence.

OpenClaw Automation, Part 2: Standing Orders Are the Authorization, Automations Are the Clock

Standing orders grant an agent permanent operating authority for a defined program, written into AGENTS.md and injected into every session. They define what it may do; automations define when — and the automation prompt should reference the standing order rather than duplicate it.

OpenClaw Enterprise Channels: Slack's Three Transports, and the 'Built-in' Column That No Longer Exists

Every enterprise channel is a plugin now, including Slack and Google Chat, which used to be built in. Slack has three transports — Socket Mode, HTTP Request URLs, and relay — and the docs say plainly that the first two have reached feature parity, so you pick by deployment shape, not by features.

OpenClaw's Main Channels: Where WhatsApp, Telegram, and Discord Each Get Stuck

Each channel has one gotcha that stops you cold: WhatsApp's login is QR-only and hard to do remotely, Telegram bots ship with Privacy Mode on so they never see group messages (and you must remove and re-add the bot after changing it), and Discord needs Message Content Intent or it receives nothing from servers.

OpenClaw's Other Channels: Signal, iMessage, LINE — and Reef, Where Two People's Agents Talk Directly

The most interesting entry here is Reef — an end-to-end-encrypted side channel between OpenClaw agents owned by different people. Messages are sealed on your machine, screened in both directions by a pinned-model guard, and the relay operator can never read the content. It ships bundled.

OpenClaw Channels Overview: 31 Channels, Nearly All Plugins — and Why 'Who Can Trigger' Is Not 'What the Model Sees'

OpenClaw supports 31 chat channels, but only WebChat lives in core — even Slack and WhatsApp are plugins you install. And group safety has two independent axes: allowlists govern who can trigger the agent, not which quotes and history the model sees. That second one is contextVisibility, and it defaults to wide open.

OpenClaw Gateway, Part 1: Strict Validation Will Refuse to Boot — and the Guards That Stop You From Yourself

OpenClaw validates config strictly — one unknown key, a wrong type, or an invalid value and the Gateway refuses to start. It keeps a last-known-good copy, but neither startup nor hot reload restores it automatically; only doctor --fix does.

OpenClaw Gateway, Part 2: Binding, Auth, and That Credential Precedence Contract

The Gateway binds to loopback by default, and binding anywhere else requires auth — that is enforced, not advised. Inside a detected container the effective default is auto, unless Tailscale serve/funnel is active, which always forces loopback.

OpenClaw Installation Guide (Part 2): Four Decisions for Cloud Deployment, and the Real Traps on K8s

Deploying OpenClaw to the cloud comes down to four decisions: where the Gateway binds, where state lives, who can reach it, and how you recover. Which platform you pick is the least important of them.

OpenClaw Installation Guide (Part 1): Choosing Among Six Local Methods, and Where It Gets Stuck

OpenClaw has six local install methods, and what separates them is not the command but whether you want reproducibility, isolation, or self-updating. The real blocker is package-manager lifecycle-script policy: both npm 12 and global pnpm installs block OpenClaw's build scripts by default.

OpenClaw Models, Advanced: Two-Stage Failover, the Real Cooldown Numbers, and Prompt Caching

OpenClaw's failover runs in two stages: rotate auth profiles within the provider, then fall back to another model. But what really governs behavior is who chose the model — a model you picked yourself with /model is strict, and its failure is reported rather than answered by some other model.

OpenClaw's Model Requirements and Provider Ecosystem: Provider, Model, and Runtime Are Three Different Things

OpenClaw's hard requirement for a model is tool use plus a large enough context — onboarding only auto-suggests a local model when it confirms tool support and at least a 16K context window. The easier thing to get wrong is that provider, model, and agent runtime are three separate layers: an `openai/*` ref does not mean Codex.

OpenClaw's 60 Providers: A Category Map, and What Actually Bites When You Attach a Local Model

The official provider directory now lists 60 entries. The most common failure when attaching a local model is writing Ollama's base URL with /v1 — that breaks tool calling, and the model starts emitting raw tool-call JSON as plain text.

OpenClaw Multi-Agent: An Agent Is a Whole Persona Boundary — and Agents Can Now Ask for New Agents

An agent is a complete persona scope — its own workspace, auth profiles, model registry, and session store. But the isolation is not absolute: when a secondary agent's OAuth credential expires, OpenClaw reads through to the main agent's profile of the same id, and a workspace is only a default working directory, not a hard sandbox.

OpenClaw Nodes in Depth: Approval Binds the Plan, Not the Command You Edited Afterward

The best part of remote node execution is how approval binds: exec prepares a canonical systemRunPlan before approval, and once granted the gateway forwards that stored plan — not any later caller-edited command, cwd, or session fields — and re-validates the working directory before running.

OpenClaw Documentation Guide: 200+ Docs — Where Do You Start?

OpenClaw has 200+ docs. This article helps you see the big picture, understand what each section covers, and decide where to start based on your role.

OpenClaw Reference: Pi Has Been Absorbed — the Built-In Runtime Is Just Called openclaw Now

"OpenClaw is a Gateway shell around Pi" is obsolete. The docs now say the built-in runtime id is openclaw, that pi is a legacy alias which normalizes to it, and that no external agent framework packages remain. The only Pi-related third-party dependency left is a terminal component toolkit.

OpenClaw Desktop Platforms: Windows Now Has a Native Hub, and Node Is Non-Negotiable

Node is the required runtime because the canonical state store uses node:sqlite — Bun is only for installing dependencies. Windows changed the most: there is now a native Windows Hub companion app that installs without administrator privileges and can provision its own app-owned WSL distro for the Gateway.

OpenClaw Mobile Platforms: Phones Are Peripherals, Not Gateways — and the Apple Watch Has Its Own Transport

The iOS and Android apps are nodes, not Gateways: they do not run the Gateway service, and Telegram or WhatsApp messages land on the Gateway rather than on the phone. The Apple Watch is the exception — because watchOS blocks generic low-level networking for ordinary apps, it uses signed HTTPS polling instead.

The OpenClaw Plugin System: Treat Installs Like Running Code, and a Cold Check Proves Nothing About Runtime

The official framing is to treat plugin installs like running code — ClawHub and the bundled catalog are trusted sources, while arbitrary npm, git, and local paths require --force in noninteractive installs. And verification means inspect --runtime, because a bare inspect is only a cold manifest check.

OpenClaw Sandboxing: Four Backends, Three Independent Switches, and Thinking You're Sandboxed When You Aren't

Sandboxing is governed by three independent settings: mode (when it applies), scope (how many containers), and backend (where it runs). The most common failure is an expectation gap — `tools.exec.host` now defaults to auto, so 'unset means sandboxed' is no longer true, and the security audit has a check specifically for it.

OpenClaw Sessions and Memory: One Rolling Conversation, Plus Four Files That Get Written to Disk

By default every DM lands in one main session, and group activity and background work report back into it. Memory is entirely Markdown on disk — the model only remembers what gets saved, with no hidden state. But if more than one person can DM your agent, DM isolation is something you have to turn on.

OpenClaw's Threat Model: It Starts by Telling You What It Does Not Protect

OpenClaw's security docs open by stating the scope: this is a personal-assistant trust model, one gateway per trusted operator. It explicitly is not a security boundary for mutually adversarial users sharing one agent — and a 'not vulnerabilities by design' list pins that down.

OpenClaw Tools, Part 1: A Dedicated Browser, Three Ways to Attach, and Search Results Typed as Untrusted

OpenClaw's browser is a separate agent-only profile, fully isolated from your personal browser. And web_search's return shape carries an externalContent.untrusted marker — search results are typed as untrusted external content at the type level.

OpenClaw Tools, Part 3: Turning Off the File Tools Does Not Make exec Read-Only

exec is a mutating shell surface: disabling write, edit, and apply_patch does nothing to make it read-only. And since sandboxing is off by default, host=auto actually resolves to the gateway — if you really want the sandbox, say so explicitly and it will at least fail closed.

OpenClaw Tools, Part 2: Six Layers of Skill Precedence, and Why Sub-Agents Get No Message Tool

Skills load from six sources with the highest precedence winning on name collisions, and a per-agent list replaces rather than merges. Sub-agents get no session or message tools by default — they return plain text to the parent, and the right to speak to a human stays with the parent agent.

OpenClaw Tools, Part 4: When the Catalog Outgrows the Prompt — Code Mode, Tool Search, and MCP

When the tool catalog no longer fits in the prompt, OpenClaw offers two answers: Code Mode shows the model only exec and wait and has it write small programs against a hidden catalog, while Tool Search keeps structured search/describe/call controls. Neither bypasses tool policy.

OpenClaw Operations: Seven Commands for the First 60 Seconds, and 'It Feels Dumber' Is Usually Not the Model

The official triage flow is seven commands and two minutes to a diagnosis. And the most common symptom — the assistant feeling limited or missing tools — is usually the tool profile: minimal allows only session_status, while coding is the default for new local configs.

OpenClaw UI: A New Rail Lets You Ask What a Session Is Doing Without Interrupting It

The Control UI gained a session rail: it uses a utility model to produce a run digest and attaches a read-only companion thread, so you can ask what a session is doing without entering or interrupting the main agent run. Its contents never enter chat.history.

aiguide

Phil Schmid: Why Agent Harness Is the Most Important Thing in 2026

The model is the CPU, the harness is the operating system, and the agent is the application. No matter how powerful a model is, without a good harness it's just a demo. Phil Schmid argues that harness is the most critical infrastructure in AI engineering for 2026.

Claude Code Troubleshooting Index: Orders 33-35 on Install, Runtime, and Config Diagnosis

The original Claude Code troubleshooting collection has been split into orders 33-35: installation & login, runtime problems, and config diagnosis. This page is their index.

Complete Guide to Bypassing Cloudflare Anti-Bot for AI Agents: From Debugging to Building an MCP Server

Standard Playwright gets blocked by Cloudflare. Both playwright-extra + stealth and nodriver can bypass it. The final step is wrapping the solution into an MCP server so AI agents can use it automatically.

Claude Code Agent Teams in Practice: Team Lead, Point-to-Point Messaging, and a Shared Task Board

Agent Teams lets multiple full Claude Code sessions work as one team: a team lead assigns work while teammates each run their own context window, coordinating through point-to-point messaging and a shared task list. This post covers the three key differences from sub-agents, the trade-off between teammateMode display modes, and why token cost scales linearly with team size.

Claude Code Workflows in Practice: Official Best Practices from Explore to Commit

Anthropic's official Claude Code best practices boil down to one constraint: manage the context window. This post reorganizes their guidance into a working loop — explore, plan mode, implement with a runnable check, verify, commit — plus prompt techniques, when to /clear vs rewind, and the five failure patterns they call out.

Claude Code Channels: External Events, Reply Tools, and Sender Gating

Channels are a special kind of MCP server that push CI failures, monitoring alerts, and Telegram messages directly into a running Claude Code session — and Claude can answer back through the same channel via a reply tool. This post breaks down the channel contract, two-way replies, security gates, and install requirements.

Claude Code Checkpointing Deep Dive: Snapshots, the Rewind Menu, and Tracking Boundaries

Checkpointing is not git commits: Claude Code snapshots your files before every user prompt, keeps the 100 most recent per session, and deletes them after 30 days. This piece breaks down the five rewind menu options, the tracking boundaries (bash, subagents, symlinks), and how checkpoints divide labor with git.

How Claude Code Sees Your Browser: Chrome Integration, Console Debugging, and Form Automation

Claude Code gains browser control through the Claude in Chrome extension: read console logs and DOM state, click, type, upload files, record GIFs, and operate sites you're already signed into. The official prerequisites list Chrome, Edge, and Chromium-based browsers such as Brave, Arc, Vivaldi, and Opera, but WSL is not supported.

Claude Code in CI/CD: @claude on GitHub Actions and the GitLab MR Flow

Put Claude Code into GitHub Actions with anthropics/claude-code-action: /install-github-app sets everything up in one command, @claude in a PR or issue comment gets bugs fixed, branches pushed, and PR creation links returned; Bedrock/Vertex/Foundry backends switch via one input with OIDC and no stored keys; the GitLab CI/CD integration (beta) mirrors it as a single .gitlab-ci.yml job where every change flows through a merge request.

How Claude Code Remembers Your Project: CLAUDE.md Layers, Imports, Rules, and Auto Memory

Every Claude Code session starts with a clean context window. Three memory mechanisms carry knowledge across sessions: CLAUDE.md files loaded every session, .claude/rules/ files that load conditionally via paths frontmatter, and auto memory Claude writes itself. All CLAUDE.md layers are concatenated into context — not inherited by override. This guide covers layer behavior, @path imports, monorepo strategies with nested CLAUDE.md, and sharing one instruction file across tools via @AGENTS.md.

How to manage Claude Code's context window: startup content, per-feature costs, and the compaction trio

Claude Code loads the system prompt, MEMORY.md, CLAUDE.md, MCP tool names, and skill descriptions before you type your first word. This post breaks down the startup context, what each of six extension features costs, and how to control auto-compaction with /compact, /autocompact, and autoCompactWindow.

How Claude Code Sandboxing Works: Sandboxed Bash, Network Allowlists, and the Threat Model of Six Isolation Approaches

Claude Code's built-in sandboxed Bash restricts every command at the OS level: writes are limited to the working directory plus session temp, while reads default to the entire machine; network traffic goes through a proxy allowlist that starts with zero domains. The switches live in the /sandbox panel and sandbox.enabled — there is no --sandbox flag. This post also compares sandbox runtime, dev containers, Docker, VMs, and Claude Code on the web to show when each heavier isolation tier earns its setup cost.

Headless and the Agent SDK: From claude -p to Programmatic Agents

claude -p turns Claude Code from an interactive terminal session into a command you can embed in scripts and CI: pipe data in, get structured JSON out with --output-format json, and skip most auto-discovery with --bare. This post focuses on CLI usage, then covers the four signals that mean it's time to switch to the Python or TypeScript Agent SDK.

How Claude Code connects to external tools: MCP scopes, transports, and auth

Claude Code connects to external tools via MCP (Model Context Protocol), with manual configuration split across three scopes: team-shared .mcp.json, personal local/user entries in ~/.claude.json, and enterprise managed config — not settings.json. This post covers claude mcp add/login flows, the transport landscape (SSE is deprecated), and tool search lazy loading.

Claude Code Plugins & Marketplaces: Packaging Skills, Hooks, and MCP into One Installable Unit

Plugins add distribution, not new capabilities: they collect skills, agents, hooks, and MCP settings scattered across .claude/ into one manifest-backed directory that marketplaces can install, update, and version-pin. This post covers plugin structure, the minimal build flow, publishing a marketplace, and dependency version constraints.

Claude Code Remote Control: Pick Up a Local Session from Any Device

Remote Control turns claude.ai/code or the Claude mobile app into a remote for your local Claude Code session: code still runs on your own machine, with MCP servers and local tools fully available, while transcript sync passes through Anthropic servers. This post covers startup, reconnection, push notifications, file delivery, and the security boundary.

Claude Code settings.json Complete Guide: Five Scopes, Merge Rules, and the Keys That Matter

Claude Code reads settings from five levels — managed settings, CLI flags, project local, shared project, and user — where plain values are overridden by higher levels while list keys like permissions.allow merge across scopes. This guide covers each file's role, the allow/deny/ask rule syntax, and how to verify your settings with /status and claude doctor.

Delegating Coding Tasks from Slack: Claude Code in Slack and Claude Tag

A single @Claude in Slack turns a bug report into a cloud-run Claude Code session. But there are now two paths: Pro/Max stays on the original Claude Code in Slack (each session runs under an individual account), while new or migrating Team/Enterprise setups should look at Claude Tag (shared org identity, admin-configured access and spend). Check your plan before setting anything up.

How Claude Code Sub-agents Work: Context Isolation, Frontmatter Definitions, Background Execution, and Permission Inheritance

Sub-agents are specialized assistants that work in their own context window: a single Markdown file defines their system prompt, tools, and model. Claude delegates automatically based on the description field, or you can @-mention to force one. This post breaks down the frontmatter schema, background execution and nested spawning, permission inheritance rules, and when not to use them.

"Recommend the next route" and "Recommend something similar" are not the same thing — Intent Disambiguation in RAG Recommendation Systems

In a climbing RAG system, 'recommend the next route' (progression) and 'recommend a similar route' (similarity) were conflated by a single hasSimilarRouteIntent() function, causing recommendation quality to collapse. The fix is a two-stage intent classification with a Regex Fast Path + LLM Fallback.

RAG Multi-Entity Queries: When the User Lists Five Routes and the System Only Sees the First

The RAG system's extractRouteReference() used a for...return pattern that grabbed only the first match — so when a user provided five completed routes, only one was used. The fix evolves through three layers: rule-based multi-entity extraction, user profile aggregation, and embedding centroid.

When Vector Search Matches by Name Instead of Grade: Attribute Conflation in RAG Systems

Query: 'I just sent Beauty in the Mirror 5.11b — recommend routes of similar difficulty.' The results came back full of routes with similar-sounding names, not similar grades. Root cause: dense embeddings compress multiple attributes into a single vector, and the rarity of the route name drowns out the grade signal. The fix: three layers of defense — metadata pre-filtering, query rewriting, and score fusion.

aiguide

LangGraph: Managing Agent Workflows with Graph Structures

LangGraph models LLM workflows as directed graphs, solving the pain points of multi-turn iteration, conditional branching, and parallel execution that are difficult to handle with linear pipelines.

techguide

Biome: Replacing ESLint + Prettier with Rust

Biome does the work of ESLint + Prettier in a single tool, running 10–20x faster with far less configuration. DaoDao uses it across an entire monorepo — lint and format in one pass.

AEO Guide: Answer Engine Optimization — Getting AI Search Engines to Cite Your Content

AEO (Answer Engine Optimization) is a content strategy aimed at AI search engines like Perplexity, ChatGPT Search, and Google AI Overview. The core idea is to make your content the easiest source for AI to cite — not just another link in the results page.

A Complete Guide to Blog SEO — From Meta Tags to Structured Data

SEO is more than keywords. Structured data (JSON-LD), Open Graph, hreflang, and robots.txt are the technical optimizations that actually help search engines understand your content. This guide walks through a complete implementation using an Astro blog as the example.

techguide

BullMQ: The Most Mature Redis-Backed Job Queue for Node.js

BullMQ is the most mature job queue in the Node.js ecosystem, backed by Redis, with support for priorities, retries, scheduling, and delayed jobs. DaoDao uses it to handle notification delivery and practice auto-completion scheduling.

techguide

Celery: The Standard Distributed Task Queue for Python

Celery is Python's go-to distributed task queue, using Redis or RabbitMQ as a broker to offload long-running work to the background. DaoDao's AI service uses it to handle async tasks like LLM feedback generation.

Claude Code Global Skills Not Found in New Sessions? Understanding Skill Discovery and How to Debug It

Global skills live in ~/.claude/skills/, but they go missing in new sessions or the Desktop App? The problem usually isn't a missing file — it's that the skill descriptions aren't being loaded into context. This post clarifies the CLI vs Desktop App differences, the role of settings.json, and the most reliable fix.

techguide

ClickHouse: When PostgreSQL Analytics Queries Start Slowing Down, You Need OLAP

ClickHouse is a column-oriented OLAP database that scans hundreds of millions of rows in seconds. DaoDao uses it to record user behavior events for the AI recommendation engine's feature engineering, letting PostgreSQL focus on transactional data.

Cloudflare D1: SQLite Relational Database at the Edge

D1 is Cloudflare's serverless SQLite database that binds directly to Workers, supports full SQL (JOINs, transactions), and handles automatic backups. It's well-suited for small-to-medium relational data needs — NobodyClimb uses it as its primary database.

Cloudflare KV: A Global Edge Key-Value Store

KV is Cloudflare's globally distributed key-value store. Reads are served from the nearest edge node with extremely low latency. It's ideal for caching, feature flags, and ephemeral data — but writes are eventually consistent.

Cloudflare R2: An S3 Alternative with Zero Egress Fees

R2 is Cloudflare's object storage service — S3-compatible API, zero egress fees, and native Workers binding. Stop worrying about bandwidth bills for media-heavy applications.

Cloudflare Workers: Not Lambda, Not Containers — It's V8 Isolates

Cloudflare Workers uses V8 Isolates instead of containers — no cold starts, global edge deployment, and direct access to D1, R2, KV, and AI via Bindings. Great for APIs, SSR, and lightweight backends; not suited for CPU-heavy work.

techguide

Docker in Practice: Containerizing from Development to Deployment

Docker lets you bundle your application together with its environment, eliminating the 'works on my machine' problem. Combined with multi-stage builds and Compose, it's an essential tool for modern backend deployment.

techguide

Expo + React Native: What It's Actually Like to Ship One Codebase for iOS and Android

Expo turns React Native development from 'environment setup hell' into a state where you can just start writing logic. Expo Router brings file-based routing that dramatically lowers the barrier for web developers making the switch. Both DaoDao and NobodyClimb use it to ship across iOS and Android.

techguide

Express.js: The Default Answer for Node.js Backends, and Why It Still Makes Sense

Express is the most mature Web framework for Node.js, with a rich middleware ecosystem and abundant learning resources. Paired with TypeScript and a clear layered architecture, it remains a justifiable choice in 2026.

techguide

FastAPI: The Go-To Framework for Python AI Services

FastAPI is a modern Python web framework built on type hints — it auto-generates OpenAPI docs, supports native async, and delivers performance close to Node.js. It's the top choice for AI/ML services and the most worthwhile framework to learn in the Python backend ecosystem.

techguide

GitHub Actions: A CI/CD Primer and Monorepo Strategy

GitHub Actions is the lowest-friction CI/CD tool available today, ideal for small-to-medium projects. The key to monorepos is using path filters so only affected apps trigger a build.

Hono: The Lightweight Web Framework Built for Edge Runtimes

Hono is a web framework designed specifically for edge runtimes like Cloudflare Workers, Deno, and Bun. It's an order of magnitude lighter than Express, natively supports Web Standard APIs, and is the go-to choice for edge environments.

techguide

Next.js 15 + App Router: What Server Components and use cache Actually Do

Next.js 15 + React 19's App Router shifts rendering responsibility from the client to the server. use cache ties caching logic directly to data functions instead of scattering it across fetch options. Both DaoDao and NobodyClimb chose this stack for very practical reasons.

@opennextjs/cloudflare: Running Next.js on Cloudflare Workers

@opennextjs/cloudflare enables Next.js App Router deployments on Cloudflare Workers — dynamic SSR runs in a Worker, static assets are served from Cloudflare Assets. Zero server management, but with clear feature limitations.

techguide

PM2: The Practical Choice for Node.js Process Management

PM2 keeps your Node.js app running on a server — auto-restarts on crash, supports cluster mode to max out CPU cores, and handles log management. Nearly every Node.js app deployed on a VM or VPS needs it.

techguide

Prisma ORM: Type-Safe Database Access for TypeScript Projects

Prisma's schema-first design gives you versioned migrations, full TypeScript types on every query, and intuitive relation handling. The tradeoff is a learning curve and the inherent limits of any ORM abstraction — but for most TypeScript projects, it's a worthwhile deal.

techguide

React Hook Form + Zod: The Best Combo for Form Handling

React Hook Form handles form performance, Zod defines the validation schema — together they eliminate nearly all form boilerplate. Share a single Zod schema across a monorepo and you get one source of truth for both frontend and backend validation.

techguide

Redis Essentials: Caching, Sessions, and Pub/Sub in One Go

Redis is an in-memory key-value store that's blazingly fast. DaoDao uses it to handle three responsibilities at once — API caching, session storage, and BullMQ job queues — all from a single Redis instance.

techguide

shadcn/ui: Not a Package — It's Copy-Pasted Component Source Code

shadcn/ui is not an npm package — it copies component source code directly into your project, giving you full ownership. DaoDao uses it to build packages/ui, a shared component library used across three Next.js apps.

techguide

TailwindCSS: Utility-First Is a CSS Management Strategy, Not Just a Style Preference

TailwindCSS's core value is solving CSS's global namespace pollution and dead code problems. Utility classes keep styles co-located with components, and unused classes are automatically purged at build time — production CSS bundles typically come in at just a few dozen KB. Both DaoDao and NobodyClimb use it for web styling.

techguide

Tamagui: A React Native UI Framework — Why NobodyClimb Chose It Over NativeWind

Tamagui is a UI framework built for React Native with a complete design token system, theme support, and compile-time optimization that moves style computation to build time. NobodyClimb chose it over NativeWind primarily because its cross-platform token system is more robust.

techguide

TanStack Query: The Standard Solution for Server State

Managing API data with useState + useEffect means reinventing the wheel — and doing it worse. TanStack Query handles caching, background updates, and loading/error states so you can focus on UI logic.

techguide

Turborepo + pnpm Workspaces: The Standard Approach to Monorepos

Turborepo solves monorepo build speed problems; pnpm workspaces solves dependency sharing. Together they are the best choice for JS/TS monorepos today.

techguide

Zod: Runtime Type Validation for TypeScript

TypeScript types only exist at compile time — they vanish at runtime. Zod lets you validate external data at runtime while inferring TypeScript types from the same schema. One definition, two jobs done.

techguide

Zustand: The Lightest Global State Management for React

No Provider, no reducer — global state in just a few lines. NobodyClimb uses it for auth and UI state, paired with TanStack Query for server state.

A One-Person Full-Stack Team: AI-Driven Development Workflow from OpenSpec to Auto-Deploy

Use OpenSpec to break requirements into engineering tasks, Claude Code to implement them, hooks to auto-format and protect, local review before committing, three AI reviewers running in parallel on PR, and auto-deploy after merge. This entire workflow lets one person maintain quality across six sub-projects.

Claude Code Hooks: A Complete Guide to Event-Driven AI Control

Hooks are Claude Code's event system. They trigger shell commands, HTTP requests, MCP tools, or LLM evaluations automatically before/after tool execution, when a prompt is submitted, or when a task ends. Use them to block dangerous operations, run automated reviews, inject context, or write audit logs.

Claude Code Skills: A Complete Guide to Turning Repetitive Workflows into Single Commands

A Skill is an SOP written for AI. Define the steps in a Markdown file and Claude follows them. No coding required, no frameworks to learn — just write down what an experienced person would do.

Turning Debug Sessions into GitHub Issues with a Claude Code Skill: Designing /file-bug-issue

Stuck mid-debug and can't fix it right now? Use /file-bug-issue to package the error analysis, reproduction steps, and attempted fixes from your conversation into a well-structured GitHub issue. Pair it with a Remote Agent to let AI automatically take over the fix.

Let AI Pick Up Issues, Write Code, and Open PRs: Hands-Off Development with Claude Code Remote Agent

Using Claude Code's Scheduled Remote Agent, automatically scan GitHub issues every 2 hours, implement features, open PRs, and address review feedback — no human intervention required. Humans only write issues and click merge. Pair it with the custom /publish-tasks skill to push OpenSpec engineering tasks directly to GitHub issues.

aiguide

Langfuse Complete Guide: LLM Application Observability from Scratch

Langfuse is currently the most mature open-source LLM Observability platform. This post covers four core capabilities — Tracing, Prompt Management, Evaluation, and Datasets — showing you how to use them in real projects.

Claude Code's Three-Layer Quality Defense: Hooks, Skills, and Instruction Files

Hooks are automated safety nets (blocking bad commits), Skills are interactive workflows (running checks + auto-fixing), and instruction files (CLAUDE.md / AGENTS.md) are behavioral guidelines. Each layer operates independently, but together they enable an AI agent to automatically run lint, typecheck, and build checks before every commit.

techguide

How to Classify Code Review Comments? From Conventional Comments to AI Review Tool Taxonomies

Three main classification systems dominate: Conventional Comments (label-based), Google's severity prefixes (Nit/Optional/FYI), and SonarQube's four quadrants (Bug/Vulnerability/Code Smell/Hotspot). AI review tools have each developed their own taxonomies, but the core dimensions consistently converge on four areas: correctness, security, performance, and maintainability.

Context Engineering: Why Your AI Agent's Problem Is Information, Not the Model

Context Engineering is the core concept that replaced Prompt Engineering in 2025: the focus shifted from 'how to ask' to 'what information to provide.' Delivering the right information at the right time into the context window is more effective than upgrading to a stronger model. This post covers the definition, four key strategies, practical techniques, and common failure modes.

techguide

From Mock to Real AI: Integrating Cloudflare Workers AI into action-maker

Upgraded action-maker from hardcoded mock data to live Cloudflare Workers AI generation. The architecture splits into Worker (AI only), Server (data storage), and Frontend (orchestration). Hit two gotchas along the way: Qwen3's thinking block and the Workers AI response format.

aiguide

MCP (Model Context Protocol): The Standardized Protocol for AI Agent Tool Invocation

Every AI tool has its own calling format, making integration costly. MCP (Model Context Protocol) is an open standard proposed by Anthropic that unifies the communication protocol between AI Agents and external tools/data sources, enabling tools to be reused across Agents.

techguide

False positives in Node.js image vulnerability scans? Separate app packages from npm built-ins first

When reviewing vulnerability scan results for a Node.js Docker image, you can't just look at package names. First distinguish between project dependencies and the packages bundled with npm inside the base image — otherwise you'll fix the wrong thing.

techguide

What Is Vulnerability Scanning? A Quick Intro to Docker and Package Scanning with Trivy

Vulnerability scanning isn't just about generating reports — it helps you discover known risks in your system before they become incidents. This post uses Trivy as a hands-on example to explain what scanners actually look for, how to read the results, and how to get started.

techguide

Turning a Scraper Script into an MCP Server for Claude to Use Directly

Wrap a local Python script into an MCP Server using FastMCP so Claude Code can call it directly — no more manually running pipelines.

techdebug

MCP Tool Returns 1M Characters: The Token Explosion in search_local_jobs

The MCP tool was returning a description field that caused 1,033 job listings to exceed the token limit. The fix: exclude description by default and add pagination.

aiguide

Agent Memory Systems: From RAG to Read-Write Memory Evolution

RAG is read-only. Agent Memory lets AI not only read but also write and persist information. Three memory types: Procedural (behavior patterns), Episodic (temporal events), and Semantic (factual knowledge) form a complete cognitive memory system.

aideep-dive

Complete Guide to AI Agent Architecture Patterns: From Three Pillars to Multi-Agent Systematic Navigation

AI Agent is not a single technology -- it is an entire architecture system. This article is a systematic navigation: starting from the Agent Three Pillars (Context/Cognition/Action), through the three-stage evolution of AI engineering (Prompt -> Context -> Harness), to eight Multi-Agent design patterns and production-grade Harness infrastructure. Each topic links to a dedicated deep-dive article.

aiguide

The Three Core Pillars of AI Agents: Context, Cognition, Action

An AI agent is not a black box — it is built from three layers: what it knows (Context), how it thinks (Cognition), and what it can do (Action). Understanding these three layers is the key to grasping why agents are sometimes brilliant and sometimes go off the rails, and how to design a truly effective agent system.

techguide

docker restart Does Not Re-apply Volumes — Debugging a Bind Mount Failure

docker restart does not recreate the container, so changes to volumes in docker-compose.yml only take effect after running docker-compose down && up.

Multi-Agent RAG: Distributed Retrieval Architecture with Specialized Agent Collaboration

A single RAG Agent handling all queries hits knowledge boundaries and performance bottlenecks. Multi-Agent RAG dispatches retrieval tasks to multiple specialized Agents, each with its own knowledge base and retrieval strategy, coordinated by a central Orchestrator that merges results.

Claude Code Permission Modes Explained: Five Modes from Default to Auto

Claude Code has five permission modes: default (confirm each step), acceptEdits (auto-accept edits), plan (read-only planning), auto (background AI classifier review), and bypassPermissions (YOLO, skip everything). Switch with Shift+Tab or configure via settings.json. Auto mode is the sweet spot — no step-by-step confirmations, but with safety guardrails.

techguide

nginx 502: Debugging Cross-Compose Container DNS Resolution

Service names aren't resolvable across Compose projects — you need to add a network alias so nginx can find the container.

techguide

Installing and Verifying Superpowers for GitHub Copilot CLI: Implementation, Diagnostics, and Validation

A hands-on log of installing Superpowers (packaged by DwainTR) for Copilot CLI on a local machine — including the diagnostic process when skills didn't appear after installation, the fix, and practical tips.

techguide

Docker DNS Resolution: container_name vs network alias

Cross-project DNS resolution requires container_name or a network alias — and only aliases support horizontal scaling.

LongRAG: Rethinking RAG Chunking Strategy with Long-Context Models

Traditional RAG splits documents into small chunks for retrieval, but this causes information fragmentation. LongRAG leverages 100K+ token long-context models to retrieve larger document segments (entire sections or even whole documents), reducing fragmentation while maintaining retrieval efficiency.

Speculative RAG: Small Models Draft in Parallel, Large Model Verifies at Once

Speculative RAG uses small specialist models to generate multiple answer drafts from different document subsets in parallel, then a large model verifies and selects the best answer in one pass. The paper reports +12.97 points accuracy and -50.83% latency on PubHealth — but that is the best cell in the table; other benchmarks gain far less.

techguide

nginx Restarted Fine, but Cloudflare Keeps Returning 502 — Even Though the Origin Is Healthy

A brief error during nginx restart caused Cloudflare to mark the origin as unhealthy and stop forwarding requests, returning 502 on its own. The key clues: localhost hits to the origin return 200, and nginx access logs are completely empty. Just wait for Cloudflare to automatically re-check the origin — it recovers on its own.

techguide

Managing Multi-Service Reverse Proxy with nginx conf.d: A Daodao Case Study

A monolithic nginx.conf becomes unwieldy as services grow. Splitting it into per-service files under conf.d/ via include is the standard solution.

techguide

nginx First Request Always 502, All Subsequent Requests Fine

When nginx uses the `set $variable` pattern for dynamic upstreams, the DNS cache expires every 30 seconds — the first request after expiry hits a 502 because no IP is available. Upgrading to nginx 1.27.3 and switching to an upstream block with the resolve parameter fixes this: DNS updates happen asynchronously in the background.

techguide

Downloading Files from a VPS Using SSH Config Aliases

Once SSH config is set up, scp works directly with aliases — no need to type out the full IP every time

aiguide

The Complete Ollama Guide: Run LLMs Locally with One Command

Ollama wraps llama.cpp in a Docker-style CLI + REST API, letting you run LLMs locally with a single command. This post covers core concepts, installation, API, hardware requirements, Modelfile customization, and what this tool is — and isn't — good for.

The Complete Guide to RAG System Patterns: A Ten-Generation Evolution from Naive to Multi-Agent with Practical Navigation

RAG has evolved far beyond simple 'search + generate' into a technology ecosystem spanning ten generations — and since 2025 into an Agentic/Reasoning era. This article is a systematic navigation guide: from Naive RAG to Multi-Agent/LongRAG across ten generations, the post-ten Agentic Era (Search-R1/RL search, MCP, GraphRAG 3.x, vision-native retrieval), retrieval strategies, chunking, embedding, reranking, evaluation frameworks, observability, and cost optimization. Each topic has a dedicated deep-dive article.

vLLM — From PagedAttention to a Production-Grade LLM Inference Engine

vLLM uses PagedAttention to eliminate KV cache memory waste, combining continuous batching and prefix caching to become the most widely adopted open-source LLM inference engine today.

techguide

Ghostty vs cmux: A Guide to Choosing Your Modern Terminal

Ghostty is a fast, native, general-purpose terminal emulator. cmux is a terminal built on top of Ghostty, specifically designed for AI coding agents. They're not competitors — they operate at different layers.

aiguide

Complete Chatbot Development Guide: State Management, Memory Strategies, and Tech Stack Selection

Building a chatbot is more than just calling an API. Conversation state management, memory mechanisms, streaming, guardrails, observability, and tech stack selection — every layer affects the user experience.

aiguide

Prompt Engineering in Practice: Iteration Methodology, Common Mistakes, and Few-shot Optimization

Good prompts aren't written in one go — they're iterated into existence. Start with the simplest prompt, test with real cases, classify error types, and make targeted fixes. This article covers the three-part System Prompt structure, reasoning framework selection, few-shot optimization, token budget management, and six common mistakes.

Cloudflare Free Plan Maintenance Page: Custom Error Pages Unavailable, Use a Worker Instead

Cloudflare Custom Error Pages require a paid plan. On the Free Plan, use a Worker with inline HTML to intercept 5xx responses instead.

techguide

Managing Personal and Work GitHub Accounts with Git Conditional Includes

Use includeIf + SSH Host aliases to let Git automatically switch accounts based on directory path — no more manual switching.

Astro + Cloudflare Workers: Native Modules Break the Build Even on Prerendered Routes

Even when a route has prerender = true, Cloudflare Workers' Rollup bundler still attempts to bundle native modules, causing the build to fail. The fix is to move any native module work into a postbuild script.

techdebug

Astro Scoped CSS Not Applied to MDX-Rendered Content

Astro scoped CSS appends a scope hash to each selector, but elements rendered by <Content /> don't receive that hash — causing all prose styles to silently break.

techguide

Conversation as Documentation: Turning Debug Sessions into Blog Posts with Claude Code

After finishing a debug session, just say 'write this up as a post' — Claude Code extracts content from the conversation, applies a template, generates frontmatter, and commits it to the repo. No extra writing required.

Agentic RAG: Letting the LLM Decide When to Search Again

For complex multi-hop questions, a single RAG search isn't enough. Agentic RAG lets the LLM evaluate whether retrieved results are sufficient — if not, it rewrites the query and searches again, forming a ReAct loop.

BGE-M3: Why This Embedding Model Works Well for Traditional Chinese RAG

Your choice of embedding model directly determines RAG search quality. BGE-M3's multilingual training, 1024-dimensional vectors, and matching Reranker make it a practical pick for Traditional Chinese RAG.

Chunking Strategies: How You Split Text Determines Whether RAG Can Find the Answer

Chunks too large and retrieval loses precision; too small and you lose context; hit a table and retrieval falls apart entirely. Chunking is the most underrated part of RAG — pick the wrong strategy and no amount of downstream optimization will save you.

ColBERT: The Third Way in Vector Search

Bi-Encoders are too coarse, Cross-Encoders are too slow — ColBERT's Late Interaction finds the sweet spot: token-level comparison between query and document, but with document vectors that can be precomputed.

Contextual Retrieval: Giving Every Chunk Its "What This Is About" Context

When you split a document into chunks, each chunk loses its place in the original document. Contextual Retrieval solves the isolated-chunk problem by generating a per-chunk context from the whole document and prepending it at index time.

CRAG: Automatically Relaxing Filters When Retrieval Comes Up Empty

Filters too strict and getting zero results? CRAG automatically relaxes them and retries — far better than letting the LLM hallucinate an answer from general knowledge.

Cross-Encoder Reranking: Surfacing the Most Relevant Documents

Vector search similarity scores don't equal relevance. Cross-Encoders use pairwise comparison to reorder results and push the truly relevant documents to the top.

GraphRAG: Structuring Knowledge as a Graph for Relationship-Based Reasoning

Vector search finds similarity; graph search traverses relationships. When a question requires reasoning across multiple entities — crag → route → sender → grade distribution — GraphRAG outperforms standard RAG.

Hybrid Search: Using BM25 + Vector Search to Cover Each Other's Blind Spots

Vector search handles semantics; BM25 handles keywords. Combining them with RRF is what lets you handle both fuzzy queries and exact terms at the same time.

HyDE: Boosting Vector Search Recall with Hypothetical Answers

Have an LLM generate an 'ideal answer' first, then embed that hypothetical document for search — it outperforms searching with the raw query.

RAG Personalization: Learning User Preferences from Conversations

After each conversation, asynchronously extract likely user preferences and skill level, then automatically personalize search parameters on the next query — no manual setup required.

MMR + Popularity Weighting: Recommendations That Are Both Relevant and Diverse

Ranking purely by relevance leaves you with five documents all describing the same route. MMR strikes a balance between relevance and diversity, and layering in popularity weighting makes results even more useful.

Modular RAG Pipeline: Designing RAG as a Composable DAG

RAG doesn't have to be a rigid three-step process. It's a set of steps that can be dynamically enabled, skipped, or reordered. Pipeline as Code lets the system adapt its behavior without redeployment.

Multi-Query Expansion: Search One Question from Multiple Angles

A single vector search on a complex query often misses relevant documents. Let the LLM rewrite the query into 3-5 sub-queries, run them in parallel, and recall improves significantly.

Multimodal RAG: Bringing Images into the Knowledge Base

Climbing routes carry a ton of visual information (topos, wall photos) that text-only RAG misses entirely. Multimodal RAG makes images searchable and understandable.

Three Generations of RAG: From Naive to Modular

Naive RAG works but has real problems. Advanced RAG patches those problems. Modular RAG rearchitects the whole system to be composable and configurable. Understanding all three generations is the key to understanding why modern RAG systems look the way they do.

Plan-and-Execute: A RAG Pattern That Plans Before It Acts

For complex queries, have the LLM map out what information is needed and in how many steps — then execute that plan. More systematic than thinking on the fly.

Query Classification: Teaching Your RAG System How to Answer Each Question

Not every question needs full RAG. Classify queries with an LLM first, then route to the right execution path — saving cost and improving accuracy.

RAG A/B Testing: A Scientific Approach to Comparing Pipeline Configurations

"Adding a Cross-Encoder feels better" is not a scientific evaluation. A/B testing tells you whether a change actually works, how much it helps, and which query types benefit.

RAG Cold Start: Building a Useful System When You Have No Data

A RAG system needs data to answer questions, but data only accumulates as the system gets used. Cold-start strategy is what bridges the gap from empty to useful.

RAG Cost Optimization: Minimizing the Cost of Every Query

RAG system costs come from LLM tokens, Embedding APIs, and vector search. Every stage has room for cost reduction, but you need to verify that optimizations don't sacrifice too much quality.

RAG Evaluation Frameworks and Tool Selection: Promptfoo, RAGAS, DeepEval, and TruLens

No industry standard mandates one RAG evaluation tool. Measure retrieval, generation, and operations separately, then choose Promptfoo, RAGAS, DeepEval, or TruLens for the actual stack.

RAG Common Failure Modes: 10 Problems and Their Solutions

When a RAG system breaks, 90% of the time it's one of these 10 failure modes. Identify which one first, then apply the matching fix — far more effective than optimizing blindly.

RAG Guardrails: Adding a Defense Layer to Inputs and Outputs

The attacks RAG systems face go beyond the technical level — Prompt Injection and Jailbreak are real threats. Both inputs and outputs need independent protection layers.

RAG Observability Tool Landscape: Choices in 2026

Rolling your own traces is good enough, but open-source tools save you a lot of work. Langfuse, Phoenix, and LangSmith each have their niche — the right choice depends on your trade-offs around self-hosting, open source, and integration complexity.

RAG Observability: Per-Node Tracing to Turn the Black Box Transparent

The hardest part of a RAG system isn't building it — it's figuring out why a particular answer went wrong. Pipeline Tracing records every step's decisions and data so debugging has a clear trail to follow.

RAG Prompt Engineering: How to Design System Prompts and Context

Search found the right documents, but the LLM's answers are still poor — often the problem lies in prompt design. System prompt structure, context formatting, and instruction placement all affect output quality.

RAG Streaming: Using SSE to Display LLM Responses as They Generate

LLM generation takes 3-5 seconds, and waiting for the full response before displaying it makes for a terrible experience. SSE pushes tokens as they're generated, reducing time-to-first-character from 5 seconds to under 1 second.

RAG Quota System: Controlling LLM Costs with Dual Limits

Limiting request count alone is not enough — a single long query can consume ten times the tokens of a normal one. Dual quotas (request count + token count) are what truly control costs.

RAG vs Fine-tuning: It's Not Either/Or

RAG and Fine-tuning solve different problems. RAG gives the model new knowledge; Fine-tuning changes the model's behavior and style. In most cases you use both, not pick one.

RRF: How to Merge Multi-Source Results in RAG Systems

BM25, vector search, HyDE, and Multi-Query each produce separate result sets -- how do you merge them sensibly? RRF uses ranks instead of scores, sidestepping the fundamental problem that scores from different systems are incomparable.

Self-Reflection + LLM-as-Judge: Having AI Evaluate Its Own Answers

Use another LLM to evaluate answer accuracy and quality — if the score is too low, regenerate, and automatically add appropriate disclaimers.

Semantic Caching: Run the RAG Pipeline Only Once for Semantically Similar Queries

Caching doesn't have to match exact query strings -- semantically similar questions can hit the cache too, skipping the entire RAG pipeline execution.

SPLADE: Smarter Sparse Vector Search Beyond BM25

BM25 only recognizes words that appear in the query. SPLADE infers related terms and adds them to the search, gaining partial semantic capability while preserving the precision of keyword search.

Text-to-SQL Router: Precise Queries That Skip RAG

Questions like 'how many routes did I complete this year' will never be answered well by RAG semantic search — querying the database directly is far more accurate. Let the LLM identify intent, extract parameters, and execute predefined SQL templates.

Vector Database Selection: How to Choose Between Pinecone, Weaviate, Qdrant, and Vectorize

Vector database selection is more constrained by deployment platform than LLM selection. Determine your platform and scale requirements first, then evaluate features — don't just look at benchmarks.

educationguide

Why Your Learning Goals Always Fizzle Out — And How DaoDao Wants to Fix It

The core reason self-directed learning fails isn't lack of motivation — it's the absence of a co-learning environment. DaoDao turns 'wanting to learn' into 'actually learning' through themed practices, inspiration feeds, group challenges, and learner connections, while turning your growth journey into tangible proof of competence.

productproject

The Next Frontier in Online Learning: Why Completion Rate Is the Real Problem

MOOC completion rates hover at just 5–15%, and the problem isn't course quality — it's the execution gap. DaoDao positions itself as a 'Learning OS,' using public commitments, community interaction, and AI recommendations to make learning visible and sustainable.

productproject

From 'Want to Learn' to 'Actually Learning': The Product Design Thinking Behind DaoDao

DaoDao is not a content platform -- it's a learning connector. Using anti-perfectionism design, community co-learning, and zero-decision recommendations, it helps learners bridge the execution gap -- from vague ideas to actionable plans.

Why Does a Climbing Community Need AI? NobodyClimb's Experiment and What We Learned

NobodyClimb uses RAG to tackle scattered climbing route information, ties quota limits to community engagement, and leverages Cloudflare Workers AI to bring inference costs close to zero.

NobodyClimb: Why the Climbing Community Needs Its Own Platform

The climbing community doesn't lack the will to share — it lacks a place to connect and preserve its culture.

productproject

Building a Low-Friction Blog from Scratch with Astro + Cloudflare Workers

To consolidate scattered notes and showcase diverse interests, I built a personal blog using Astro + Cloudflare Workers D1, paired with a Claude post skill for zero-friction writing.

The Correct Way to Bind a Custom Domain in Cloudflare Workers

In wrangler.jsonc, use custom_domain: true in routes with only the hostname as the pattern — no /* wildcard

techdeep-dive

DaoDao Tech Architecture: Monorepo, Multi-Language Backend, and AI Recommendation System

Next.js + Expo frontend, Node.js + Python dual backend, PostgreSQL + Redis core — plus a social notification system and LLM recommendation engine. Here's how DaoDao builds a learning community platform with a modern tech stack.

NobodyClimb: Building a Climbing Community Platform Entirely on Cloudflare

A climbing community platform where the web app, mobile app, and AI Q&A all run on Cloudflare — no dedicated servers.

NobodyClimb AI Architecture: Building a 20-Node RAG Pipeline on Cloudflare Workers

A dynamically composable RAG pipeline built on Cloudflare Workers AI (gemma-3-12b-it + bge-m3): 14 base steps + 6 LangGraph-specific nodes, with three strategy graphs (Baseline / Agentic / Plan-Execute) selected at runtime.

techguide

What You Need to Know Before Switching Astro Blog Templates

Switching templates means replacing the entire project foundation. Figure out what you actually need first, then choose between AstroPaper, Cactus, or AstroWind.

techguide

What Tools Power This Blog

Astro + the full Cloudflare suite — static-first, edge-computed, zero maintenance cost