Table of Contents
🌏 中文版
Units 08–12 ask how an architecture becomes a deployable, evaluated system through pre-training, fine-tuning, prompting, and post-training. Generation and evaluation are not appendices: different decoding policies turn the same logits into different quality, cost, and risk profiles.
Keep the stages and methods separate
Pre-training learns general predictive capability from broad data. Fine-tuning changes task behavior with narrower data. Prompting supplies context and output constraints without parameter updates. Post-training is a broader behavior-shaping layer, not merely “training again.” For each stage, record data provenance, objective, updated parameters, and held-out evidence.
The A2 bonus creates a runnable mini-study: pre-train your own tokenizer and Transformer, then compare a fine-tuned QA classifier with prompting. A tiny model does not establish a universal LLM law, but it is useful for controlled experiments.
Generation is a decision
Greedy decoding, sampling, and top-k variants change exploration. Stopping, length, and temperature alter the output distribution. Fix the checkpoint and prompts, change one decoding parameter at a time, and retain raw outputs rather than cherry-picked examples.
Evaluation joins models, data, and people
The schedule places inference/evaluation next to experimental design and human annotation. A benchmark score needs a defined dataset, split, and annotation process. Write an evaluation contract—task, data timestamp, metric, failure taxonomy, and human-review sample—before running comparisons.
Recordings and classroom discussion are restricted, so this article does not infer instructor opinions about particular benchmarks.
References
Loading...