Stanford CS224V Lecture 4: Task-Agent Evaluation Beyond Human-Like Answers
CS224V splits task-agent evaluation into state updates and complete interaction: isolate the semantic parser, then test task completion, grounded queries, and valid actions with real users.