Rolled-up evaluation health for a filtered slice.
Takes the same filters as /evaluations and returns totals, a status
breakdown, per-score-key stats, and a time-bucketed timeline — computed over
the whole matching set, not a sample of it.
bucket is a request, not a guarantee: a fine bucket over a long range is
stepped up until the timeline fits 1500 points, and timeline.bucket_unit
reports the granularity actually used. series is opt-in and capped at 24
series; when the cap bites, series_truncated comes back true rather than
the response quietly containing fewer lines than you asked for.
Authorizations
A scoped AgentEye API key. Mint one in the dashboard under Settings → API keys, or with POST /v1/keys. Each endpoint names the permission it requires; a key without it gets 403 and a required_permission field naming what was missing.
Query Parameters
Exact session id.
Exact agent id.
Comma-separated environments.
One of done, error, timeout. Anything else is a 400.
RFC 3339 lower bound on completed_at (inclusive). Also sets the default bucket width.
RFC 3339 upper bound on completed_at (inclusive). Defaults to now.
Comma-separated key:min..max ranges over the top-level numeric scores keys, at most 20. Same grammar as on /evaluations.
Same grammar and 20-entry cap, matched against the nested metrics at scores.metrics.<key>.
Aggregate only the most recent evaluation per session. Default false.
Comma-separated score keys whose per-bucket average the timeline should carry. De-duplicated and capped at 6; keys past the sixth are dropped.
Opt-in chart series, as a JSON array of [agent_id, environment, key, agg] quads — e.g. [["*","*","helpfulness","p90"]]. "*" in the agent or environment slot means all of them combined. agg is one of avg, min, max, p50, p75, p90, p95, p99, stddev, mode. Malformed entries and unknown aggregations are dropped rather than failing the request; at most 24 survive.
Where series reads from: scores (default, the 0-1 rate scores) or metrics (the nested magnitude metrics).
Timeline granularity: minute, 5m, 15m, hour, 6h, day, week, or auto (default, derived from the range). Coarsened when it would exceed 1500 buckets — read timeline.bucket_unit for what was used.
Response
Totals, status counts, per-score-key stats and the timeline. series and series_truncated are present only when series was supplied.

