An MCU has roughly 256–320kB of SRAM and 1MB of Flash, tens of thousands of times less than a phone. Even an int8 MobileNetV2 needs 5x more peak memory than that. Lecture 10 answers with MCUNet: TinyNAS picks a search space before searching for a subnet, and MCUNetV2's patch-based inference cuts MobileNetV2's peak SRAM from 1372kB to 172kB. The lecture closes with tinyML applications in vision, audio, and anomaly detection.
There are two reasons to train on the device: the model has to adapt to each user's new data, and that data should not leave the device. Lecture 21 first shows that sharing only gradients is not safe either: Deep Leakage from Gradients recovers the original images and sentences from them. Then it tackles memory. Training costs more than inference because activations must be stored, not because of the parameters. TinyTL fine-tunes only biases plus a lightweight residual and saves 6.5x memory; SparseBP updates only the important layers and channels; QAS lets real int8 training match fp32; and PockEngine does autodiff at compile time, bringing training memory on a 256KB MCU down to 141KB.
Apple is giving App Store Small Business Program developers free access to AFM 3 models on Private Cloud Compute if their apps have fewer than two million first-time downloads. The five-model family includes the sparse 20B-parameter AFM 3 Core Advanced, which activates only 1–4B parameters on-device, and AFM 3 Cloud Pro on Google Cloud NVIDIA GPUs, refined with outputs from Gemini.
Apple Foundation Models (AFM) is Apple's closed-ecosystem AI family. It evolved from a 3B dense model with LoRA adapters in 2024 into five models in 2026. AFM 3 Core Advanced runs a 20B IFP sparse architecture on phones while activating only 1–4B parameters; Cloud Pro runs on Google Cloud NVIDIA GPUs and is refined through Gemini distillation. There is no public API price or third-party benchmark, and access is limited to Apple's Foundation Models framework.
The main on-device LLMs in 2026 are Gemma 3n, Qwen 3.5 Small, Llama 3.2, Phi-4-mini, Ministral 3, and SmolLM3. Sub-3B quantized models can hit 30-50 tokens/sec on phones with 8GB RAM, but RAM, thermal throttling, and context window remain hard constraints.
2026 Q1 saw a full-blown open-source model explosion: on the LLM front, GLM-5, Kimi K2.5, and Qwen3.5 caught up with closed-source models; Embedding and Reranker are dominated by Qwen3 and BGE; speech has Voxtral TTS and Whisper V3; image has FLUX.2; and video has Wan 2.2 rivaling Sora. This is the complete navigation map.