Skip to content
All tags

#experimentation

6 posts

CS224U Methods and Metrics II: Datasets, Data Splits, and Comparing Models

The second half of CS224U's 'NLP methods and metrics' unit skips metric formulas. It asks whether your experiment holds up. Naturalistic or crowdsourced data, adversarial or common cases: the course answers 'both' each time. Lock the test set away. Pick baselines when you write the hypothesis. Compare two models with confidence intervals, Wilcoxon, or McNemar, and run several random initializations. The slides, three videos, and two notebooks are all public. Kawin Ethayarajh's guest session 'Real-world NLP assessments' has no public slides or video.

CS224U Final Project Workflow: Lit Review and Experiment Protocol

The first two deliverables of the CS224U final project are a literature review and an experiment protocol. The lit review covers 5, 7, or 9 papers depending on team size, under five suggested sections. The protocol has seven required sections, and its core is a hypothesis you can state. The course supplies a six-step paper-search loop, a rule that AI-assistant output must be quoted, and a worked example: a student's final project that became a Findings of EMNLP paper. The Gradescope format and rubric slides, and past exemplary papers, are behind a login.

Growth & Experimentation Interview Guide: From Growth Loops to Experiment Design

Growth interviews don't test whether you can growth hack — they test whether you have systematic growth thinking. Core skills: growth loop design (the acquisition → activation → retention → referral flywheel), experiment design (the full hypothesis → metric → experiment → analysis process), retention strategy (finding the aha moment, designing habit loops), and using data to decide what's worth continued investment.

Metrics & Analytics Interview Guide: From North Star to Experiment Design

Metrics interviews test whether you can make decisions with numbers, not how much statistics you know. Core skills: north star metric selection logic (why this one and not that one), metric tree decomposition (finding actionable levers), funnel analysis (which step's drop-off is most worth fixing), A/B testing design and pitfalls, and judgment when facing counterintuitive data.

productdeep-dive

Value Validation for Digital Products: From Assumption Maps to the M3 Retention Baseline for AI

The unit of validation is an assumption, not an idea. Kohavi's data shows the industry median experiment success rate is ~10%, which means roughly 22% of 'winning' experiments at p<0.05 are false positives. Sean Ellis's 40% threshold has no publicly available dataset. AI product retention should be baselined at M3 rather than M0, and GRR splits from 23% below $50/mo to 70% above $250/mo.

RAG A/B Testing: A Scientific Approach to Comparing Pipeline Configurations

"Adding a Cross-Encoder feels better" is not a scientific evaluation. A/B testing tells you whether a change actually works, how much it helps, and which query types benefit.