Skip to content
All tags

#training

16 posts

WebDancer & WebThinker: Training a Deep Research Agent from Scratch

Two NeurIPS 2025 papers answer the same question: how to train a web research agent from scratch? WebThinker chooses 'bolt on web capability to existing reasoning models,' WebDancer chooses 'rebuild everything from data construction to RL training.' Two philosophies, four stages, one core insight: training beats prompting.

Taiwan's Security Certification Ecosystem — Regulations, Training, Resources, and Exam Logistics

Taiwan's Cybersecurity Act 2.0 passed and FSC's three-tier system is in place, but neither mandates specific certifications — real demand comes from the job market. Taiwan has 7+ training providers: Uuu has the widest catalog (only one offering both SecAI+ and COASP classroom courses), WUSON is the CISSP legend (monthly cohorts sold out through mid-2027), and DEVCORE exclusively brings OffSec factory instructors. Test centers in Taipei and Kaohsiung; most certs also support online proctoring from home.

Training Small LLMs From Scratch in the Chinese Community: Corpus, Tokenizers, and Three Open-Source Projects

A close look at three open-source projects training LLMs from scratch in the Chinese community — baby-llama2-chinese (218M, 63.4B tokens), ChatLM-mini-Chinese (0.2B T5, 10.23M dialogues), and Steel-LLM (1.12B, 1T tokens, 8 months) — comparing corpus strategy, tokenizer decisions, and community ecosystem. Honest evaluation included: baby-llama2 scored a bottom-ranking 21 in MiniMind's side-by-side test; ChatLM has the strongest knowledge (62) but weak coding.

Karpathy's nano lineage: nanoGPT, llm.c, and nanochat

A tour of Karpathy's three teaching repos: nanoGPT (2022, ~300 lines each to reproduce GPT-2 124M), llm.c (pure C/CUDA training), and nanochat (2025-10, one speedrun.sh from tokenizer to WebUI). $100 and 4 hours on 8×H100 buys a chatty model; GPT-2-grade capability is now down to about 2 hours and $48.

The framework for training from scratch: LitGPT's no-abstraction rewrites, pretrain flow, and the TinyLlama track record

LitGPT (Lightning AI, ~13,600 stars, Apache 2.0) rewrites 20+ mainstream LLMs — Llama 3, Qwen2.5, Phi 4 — from scratch as single-file, no-abstraction implementations, with a full pretrain / finetune / evaluate / serve CLI. TinyLlama (1.1B parameters, 3T tokens) was trained on this codebase. This post breaks down how it differs from MiniMind, how to actually use it, and where it stops.

MiniMind: Train an LLM From Scratch for $0.40

MiniMind is an open-source project for training LLMs from scratch: a 64M Dense model and a 198M-A64M MoE model that run the entire chain — Pretrain → SFT → LoRA → DPO → PPO/GRPO/CISPO → Agentic RL — in ~2 hours on a single RTX 3090 at roughly 3 RMB (~$0.40). Every core algorithm is implemented natively in PyTorch with no high-level wrappers.

Running MiniMind on RunPod: From Zero to a Chatting Model

The hands-on installment of the series: rent an RTX 3090 on RunPod (Secure Cloud $0.5/hr, Community Cloud $0.22/hr), follow the MiniMind README through pretrain (~1.21h) + SFT (~1.10h), spend roughly $0.55–1.50 USD total, and chat with your own 64M model trained from scratch in the terminal.

Training an LLM from Scratch: A Project Spectrum from $0.4 to 65B

Open-source projects have pushed the cost of training an LLM from scratch absurdly low: MiniMind runs the full PreTrain-to-RL pipeline for about $0.4 (2 hours on a single RTX 3090), while at the other end OLMo 3 and LLM360 K2 publish everything — data, code, and stage-by-stage checkpoints of 65B models. This series walks the whole project spectrum from $0.4 to 65B in 11 articles, and flags where the map is biased.

When to Train an LLM From Scratch: A Decision Tree and the Full Cost Ladder

Training from scratch only makes sense in three cases: you want to learn how training works, you have 10B+ clean tokens no open model has seen, or you need a fully transparent training process for research. Otherwise fine-tuning or RAG is almost always cheaper. This post collapses the series' main routes into one cost ladder and a decision tree.

YuLan-Mini: Squeezing a 2.4B Flagship-Scale Small Model Out of 1.08T Tokens

YuLan-Mini is a 2.4B open-source model from Renmin University's AI Box lab, trained on 48 A800 GPUs with only 1.08T tokens — scoring 37.8 on MATH-500 and 64.0 on HumanEval, beating Qwen2/Qwen2.5 peer models trained on 7T–18T tokens at math and code. What's public isn't a slogan: per-phase data mixes, pre-annealing optimizer states, and even W&B logs of the ablation studies.

How Models Improve Themselves: Gradient Descent and the Training Loop

A model uses loss to know how wrong it is and gradients to know which direction to adjust. Gradient descent repeats three things: compute loss, compute gradients, update parameters. The learning rate controls step size — too large and you overshoot, too small and training takes forever.

How a Model Knows It's Wrong: Loss Functions and Cross-Entropy

Every time a model predicts the next token, it assigns a probability to every candidate word. A loss function measures how far that probability distribution is from the correct answer — the further off, the higher the loss, the more the model knows it got it wrong. Cross-entropy is the standard formula; perplexity is its human-readable translation.

Scaling Laws: How Big Should a Model Be, and Why Bigger Isn't Always Better

Scaling laws show that loss decreases predictably with more parameters, data, and compute — following power-law relationships. The Chinchilla paper's key finding: most models were too large and undertrained. Given the same compute budget, training a smaller model on more data produces better results. This reshaped the entire industry's training strategy.

Pre-training, SFT, RLHF: Three Stages That Turn a Text Predictor into a Useful Assistant

Every LLM goes through three training stages: pre-training reads the internet to learn language, SFT uses example conversations to learn the format, and RLHF uses human preferences to learn what a good answer looks like. The gap between a base model and a chat model is what the last two stages do.

CS336 Lecture 9: Scaling Laws Are Extrapolation Tools, Not Crystal Balls

Lecture 9 begins with log-log linear relationships between data and error, then uses scaling laws to compare architectures, optimizers, batches, and model-data allocations. The Chinchilla dispute shows how fitting methods, observed ranges, and deployment objectives change the answer.

climbingdeep-dive

A Knowledge Map of Climbing Books: 60+ Books from Training Science to Mental Philosophy

Climbing books don't exist in isolation — they represent competing schools of thought, philosophical differences, and knowledge gaps. This post maps the relationships between 60+ books to help you pick the right one to read next.