Models: volume, latency, token use, error rate, and evaluation scores.
Evaluations: success rate and score trends by agent or environment.
Tools: call volume, failure rate, duration, and repeat calls.
Hooks and policies: decisions, denials, and policy match rate.
Custom workflow metrics: results from your own event fields and SQL.
Start with a saved query, validate its result, then add it as a tile. Keep production and development filters explicit so test traffic does not hide a regression.
Give every dashboard an owner and a response question, such as “Is checkout-agent reliability worse than last week?”