Skip to main content
Jev helps at two points in an agent run: score a finished session against known answers, or review a tool call in the context of what you asked the agent to do.
Use a Jev eval when a finished session can be scored against a question with a few known answers, such as “Did the customer ask for a refund? Answer yes or no.” It helps you find patterns across sessions.

Create an eval

In the Cloud dashboard, open Analyze → eval authoring → new eval. Enter one fixed-answer question, select draft, and check that it chose a classifier score. Test it on real sessions, then deploy it.The shared eval authoring form where you describe a question, review the draft, and deploy it. This screenshot shows a code draft; use a fixed-answer question for Jev.

Read the scores

After a new session completes, open Observe → Evaluations or use the Cloud CLI:
The CLI reads scores; creating a Jev eval currently uses the dashboard. See Jev evaluations for question types and examples.