Skip to main content
The Evaluator SDK runs evaluations on your own infrastructure. Your worker registers its evaluations with Failproof AI, claims sessions as they finish, scores them, and submits the results, all over outbound HTTPS: nothing connects in to it. Use it for what hosted Python cannot do — LLM judges, model calls, packages, secrets, and network access. Its results appear beside hosted ones on the evaluations page, tagged customer. It ships in failproofai-sdk, under failproofai_sdk.evaluator; importing the tracing SDK does not load it.

Write evaluations

  • @app.eval(key, version=...) registers an evaluation. The key is what its results chart under; change the version whenever the logic changes, and each result keeps the version that produced it. One worker holds up to 100 evaluations.
  • result_kind is "score" unless you say otherwise. For a "metric" or "assertion" evaluation, name one metrics or assertions entry after the key: that entry is its result.
  • when decides whether a session applies. Return ConditionResult(False, "<reason>") to skip one, and the reason is recorded.
  • An evaluation can be a plain function or async, and timeout_seconds bounds it.
  • Payload keys — tool_name, response, and content above — are whatever your agents send, so read them off a real session.

Run the worker

Put a key with the evaluations:run permission, created under Administration → Keys, in FAILPROOFAI_EVALUATOR_TOKEN — set it from your secret store rather than typing it into a command — and start the worker:
Without the __main__ block, python -m failproofai_sdk.evaluator evaluator:app does the same.
FAILPROOFAI_EVALUATOR_ALLOW_INSECURE_HTTP sends everything in cleartext. The worker carries FAILPROOFAI_EVALUATOR_TOKEN as an Authorization: Bearer header on every request, and the transcripts it fetches are the sessions themselves — so anyone on the path reads both, and the token they read runs evaluations until you rotate it. Use it only on an isolated development network. Everywhere else the URL must be HTTPS; loopback needs no flag.

Result types

An EvalResult carries at least one score, metric, or assertion, and at most 25, each under a unique key.

The session

Each event carries id, ts, event_type, and payload.

The legacy evaluator

The earlier Evaluator SDK — an HTTP service Failproof AI called at EVALUATOR_ENDPOINT, answering /evaluate and polled through JobPending — is retired. Build new evaluators on this worker; operators of a self-hosted instance running a legacy service can keep it through the transition.