Skip to content
All tags

#evals

1 posts

AI-Native SDLC Playbook L9: Continuous Evals in CI

Evals are the AI-native equivalent of stage-gate QA — collect 20–50 real tasks as test cases, run them automatically whenever CLAUDE.md, skills, or hooks change, and block the merge if the pass rate drops. Every production incident becomes a permanent eval.