failproofai-sdk, under failproofai_sdk.evaluator; importing the tracing SDK does not load it.
Write evaluations
@app.eval(key, version=...)registers an evaluation. The key is what its results chart under; change the version whenever the logic changes, and each result keeps the version that produced it. One worker holds up to 100 evaluations.result_kindis"score"unless you say otherwise. For a"metric"or"assertion"evaluation, name onemetricsorassertionsentry after the key: that entry is its result.whendecides whether a session applies. ReturnConditionResult(False, "<reason>")to skip one, and the reason is recorded.- An evaluation can be a plain function or
async, andtimeout_secondsbounds it. - Payload keys —
tool_name,response, andcontentabove — are whatever your agents send, so read them off a real session.
Run the worker
Put a key with theevaluations:run permission, created under Administration → Keys, in FAILPROOFAI_EVALUATOR_TOKEN — set it from your secret store rather than typing it into a command — and start the worker:
__main__ block, python -m failproofai_sdk.evaluator evaluator:app does the same.
Result types
An
EvalResult carries at least one score, metric, or assertion, and at most 25, each under a unique key.
The session
Each event carries
id, ts, event_type, and payload.
The legacy evaluator
The earlier Evaluator SDK — an HTTP service Failproof AI called atEVALUATOR_ENDPOINT, answering /evaluate and polled through JobPending — is retired. Build new evaluators on this worker; operators of a self-hosted instance running a legacy service can keep it through the transition.
