Skip to content
所有標籤

#experiments

2 篇文章
ai deep-dive

Galileo 深入介紹:Experiments、Evaluators 與 Agent Observability 的完整評估迴圈

Galileo 把 dataset experiment、LLM/code/Luna evaluator、production traces 與 runtime guardrail 接成一個迴圈;適合需要企業級可觀測性與線上介入的團隊,但舊 Protect 名稱已 deprecated,評分模型也不能取代人工校準與應用安全。

ai deep-dive

Patronus AI 深入介紹:從 Evaluator、Experiment 到 Production Monitoring

Patronus AI 把 evaluator 當成可重用評分單元,再接到離線 experiment 與線上 trace;它適合要現成幻覺、安全與多模態評估模型的團隊,但 judge 分數不能取代人工標註、確定性測試或真正的資安驗證。