Skip to content
All tags

#metrics

13 posts

NTHU NLP Guide 10: Why the Same Model Sounds Stiff or Rambles — Decoding Strategies and NLG Evaluation

A guide to the decoding and evaluation unit in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The first half covers how to pick a word once the model outputs a probability distribution: greedy decoding can't take back a mistake, beam search keeps several candidates but favors short outputs, and top-k / top-p trade determinism for diversity (missing from the slides; the professor covers it verbally in class). The second half covers scoring generated text: BLEU's modified precision and brevity penalty, ROUGE-N and ROUGE-L, perplexity, and what GLUE, SQuAD 2.0, MTEB, and MMLU each measure.

CS224U Methods and Metrics I: A Classifier with 0.81 Accuracy and 0.43 Macro F1 — What Classifier and Generation Metrics Each Encode

The CS224U slides compute two numbers from one three-class confusion matrix: accuracy 0.81 and macro F1 0.43. One says the system is good; the other says it gets the two small classes almost entirely wrong. The unit's claim is that different metrics encode different values, and it goes through the bounds, values, and weaknesses of accuracy, the three F-score averages, perplexity, word error rate, and BLEU. Final projects are graded on whether the metrics fit, not on how high the scores are.

Product Builder Interview Daily — 2026-09-29: Metrics & Analytics

The most common way Metrics & Analytics interviews go wrong isn't failing to compute a number — it's accepting a plausible-looking metric and jumping straight into optimizing it without first asking, 'if this number goes up, is that necessarily good for the business?' Today's scenario is a metric-evaluation question: 'Netflix wants to raise average session length. Is that the right metric?' The answer framework is a metric tree: define the North Star metric, break it into driver metrics and guardrail metrics, then check whether a single metric's rise could come from either a good or a bad cause. The case study is Netflix itself — in its Q2 2026 shareholder letter, it admitted 'engagement is not just the quantity of view hours, but also refers to the quality and variety of our offering,' and shifted its biannual viewership report (running since December 2023) to an annual one, effectively demonstrating in public that a metric going up doesn't mean the story is getting better.

Product Builder Interview Daily — 2026-09-22: Metrics & Analytics

The most common trap in metrics interviews is treating 'this number went up' as proof that a decision was correct, without first separating what the north-star metric, driver metrics, and guardrail metrics are each protecting. Today's practice question: an A/B test shows a new checkout flow raised revenue but customer satisfaction (CSAT) dropped — should you ship it? The framework is to sort every number into a metric tree first, then dig into the mechanism behind the CSAT drop (did returns or complaints move too), instead of ruling on a single metric's direction. The case study is Netflix's former CPO Neil Hunt picking exactly one north-star metric — 'grab the viewer's attention within 90 seconds' — for artwork A/B tests, so a single title's engagement or overall watch hours couldn't hijack the final call. Swapping in a cover image of a kid playing golf with his caddy for 'The Short Game' alone lifted click-through by 14%.

Product Builder Interview Daily — 2026-09-15: Metrics & Analytics

The easiest way to lose points on a Metrics question isn't failing to name a cause — it's naming one cause and stopping, instead of systematically ruling possibilities out. Today we use the six-category root-cause taxonomy Google's own interviewers have publicly described — production bug, UI/UX change, rollout issue, user behavior shift, seasonality, external factor — on a real Google PM interview question: "YouTube comment engagement dropped in the last 24 hours, what's your next step," paired with AARRR to first locate where in the funnel this metric sits. The case study is Airbnb's S-1 disclosure explaining why its core operating metric is Nights and Experiences Booked, not visitor count or listing count — because that number can't be inflated from just one side of the marketplace; it only moves when a guest and a host both actually got value.

AI-Native SDLC Playbook L13: Closing the Loop with Monitoring

Stage 6 is both the endgame and the starting point of the AI-Native SDLC: a monitoring script detects an anomaly → Claude writes a diagnosis as intent.md → it flows through the entire development pipeline. Humans shift from 'starting work' to 'triaging and reviewing work.'

Product Builder Interview Daily — 2026-09-08: Metrics & Analytics

The most common way to fumble a Metrics question isn't failing to name a metric — it's naming one that sounds reasonable but you can't answer 'what does it mean when this goes up.' Today we use Lenny Rachitsky's six-category north star metric framework to narrow the candidates, then layer a metric tree on top to split the north star into output/input/guardrail, practicing the Meta PM analytical thinking prompt: 'define a north star metric for Instagram Reels, and explain the trade-off if resources tilt toward Reels at Stories' expense.' The case study is Netflix's three north star pivots — from next-day DVD delivery rate, to the share of members streaming 15+ minutes a month, to median viewing hours per month — showing how a metric should track the strategy stage.

Product Builder Interview Daily — 2026-09-01: Metrics & Analytics

Metrics questions rarely fail because you picked the wrong metric — they fail because you can't say why that metric represents user value, or you mistake correlation for causation. Exponent's latest 2026 real-interview roundup includes a Meta-style execution question: comments are up but watch time is down, what do you do. Today we break it down with a metric tree, using Facebook's famous '7 friends in 10 days' north star metric as the case study — it found Facebook's growth lever, and it also became one of Silicon Valley's most-cited correlation-causation traps.

How to Use Cloudflare Observability: Workers Logs, Traces, and Analytics Engine

Workers Observability is for debugging and request tracing; Workers Analytics Engine is for high-cardinality product events and custom metrics; GraphQL Analytics API is for querying existing Cloudflare product data. Keeping those roles separate prevents logs from becoming a database and keeps billing, monitoring, and product analytics from blending together.

Product Builder Interview Daily — 2026-08-25: Metrics & Analytics

Analytics interviews don't test whether you can write SQL — they test whether you can untangle contradictory signals like 'DAU is rising but advertisers are fleeing.' In a real Google hiring committee debrief, a candidate was rejected for treating 'DAU' as the North Star metric for News — the committee wanted a metric tied to business risk, not the prettiest number on the dashboard. Today we use a metric tree to break down exactly this kind of problem, with the legendary 'Google changed a font color and made a billion dollars' as our case study.

Metrics & Analytics Interview Guide: From North Star to Experiment Design

Metrics interviews test whether you can make decisions with numbers, not how much statistics you know. Core skills: north star metric selection logic (why this one and not that one), metric tree decomposition (finding actionable levers), funnel analysis (which step's drop-off is most worth fixing), A/B testing design and pitfalls, and judgment when facing counterintuitive data.

RAG A/B Testing: A Scientific Approach to Comparing Pipeline Configurations

"Adding a Cross-Encoder feels better" is not a scientific evaluation. A/B testing tells you whether a change actually works, how much it helps, and which query types benefit.

RAG Evaluation Frameworks and Tool Selection: Promptfoo, RAGAS, DeepEval, and TruLens

No industry standard mandates one RAG evaluation tool. Measure retrieval, generation, and operations separately, then choose Promptfoo, RAGAS, DeepEval, or TruLens for the actual stack.