Skip to main content

Customer test packs

This page is the procedure for one routing task. The complete example is the public fixture examples/contact-routing (pack id coyos.fs.contact-routing, version 0.3.1). Every case in that directory is invented. After you have a pack, the commands below are how you run it again on your own. The demonstration you can start from etalon pack init is a different pack. The commercial contact-routing pack is a private delivery and is not in this repository. Those differences are in Packs. A decision is the pack policy reading the scores for one recorded deployment. QUALIFIED is evidence for that pack version and that endpoint fingerprint. It does not authorize use of the model without human review. It does not determine that a regulation, contract, or internal control has been met.

Define the task and the expected results

Write the task in pack.yaml intended_use, and write the reply the endpoint must produce. The routing example asks a self-hosted model to read one synthetic support message and return one JSON object:
label is one of billing, access, technical, security, spam. behaviour is one of classify, abstain, escalate. The system prompt in examples/contact-routing/system_prompt.txt states the same rules: abstain when the message is ambiguous, and escalate a security ticket while keeping the security label. Each case records the referee answer the scores compare against. This is Q-0001 from examples/contact-routing/corpus/qualification.jsonl:
expected.label and expected.required_behaviour are the results the reference evaluators score. coverage names one declared value on each dimension in pack.yaml. The values have to be listed under coverage.dimensions. provenance is synthetic, expert-authored, customer-derived, or public-benchmark. etalon pack validate checks the line. A missing field names the property and lists the required case fields. A coverage value that is not in pack.yaml is reported as coverage value ... is not declared. When you change the label set, change it in the same places together: pack.yaml coverage.dimensions, the system prompt, expected.label on every case, and critical_labels in evaluators.yaml and the critical rubric. The runner does not infer the label set from the prompt.

Identify critical errors

A critical error is the failure the pack refuses to average away. On the routing example it is a security ticket sent to a public queue (billing, access, technical, or spam). The rule is in examples/contact-routing/qualification.yaml:
The rubric rubrics/critical_misroute.yaml states the same condition. The metric counts qualification cases whose referee label is security and whose output contains a label outside that set. One such case fails a maximum of zero. The Wilson interval is reported and is not used to soften the rule. An output that omits label is a schema miss, not a critical misroute. Other misses are mandatory requirements, not this critical rule: label accuracy, schema compliance, and abstention correctness. They use a 95% Wilson bound. A schema failure can still make the run NOT_QUALIFIED when its interval is on the failing side of its threshold. It is not the critical event count. The demonstration pack uses the same kind of rule on a different label. There, a complaint routed outside the complaint set is the critical error. That is etalon.demo.ticket-routing, not this routing task.

Keep development cases separate from acceptance cases

The runner scores three files, named in pack.yaml artifacts: Development cases do not have a fourth slot. Put them in a file the manifest does not name, for example corpus/development.jsonl. etalon pack validate does not read that file, and qualify does not send it. Pointing artifacts.corpus at the development file makes those cases the acceptance set. Promote a case by moving it into corpus/qualification.jsonl and deleting it from the development file. Case ids must be unique across the three artifact files. The same input text under two ids is still one situation; the runner does not detect copied text. A validation error names both files when an id is reused: move the case, do not copy it. Calibration rows are referee scores of a stored candidate. They are not a place to park endpoint cases you are still editing. A calibration row without metadata.human_passed or metadata.candidate_output fails validation with the field name.

Select acceptance limits and the number of cases

Limits live in two files that must match, plus a heading in methodology.md:
  • thresholds.yaml holds the metric, the minimum sample size, the threshold, and the written rationale.
  • qualification.yaml repeats the operator, the threshold, and rationale_ref for each mandatory requirement, and holds the critical maximum.
etalon pack validate rejects a requirement whose numbers differ from thresholds.yaml, and a rationale_ref whose anchor is missing from methodology.md. Changing a threshold or its rationale is a new pack version. The routing fixture’s numbers are demonstration bounds for this synthetic corpus. Copying them into a customer pack is not a risk appetite. Proportion requirements use the 95% Wilson interval, not the point estimate alone. For a perfect score on n cases the lower bound is n / (n + 3.841). That bound is at least 0.90 when n is 35 or more, and at least 0.80 when n is 16 or more. A point estimate of 1.0 on fewer cases straddles the threshold, and the run is INDETERMINATE. On this fixture: minimum_sample_size and the Wilson bound are both required. Meeting the sample-size minimum with a perfect score can still be INDETERMINATE when the interval crosses the threshold. Coverage is value-presence. The fixture declares intent (5 values), severity (3), input character (4), and expected behaviour (3). Realised coverage is the unweighted mean of the fraction of those values that appear at least once on a qualification case. The minimum is 0.80. A value that never appears lowers only its dimension. Coverage is not a count of cases and not a cross-product of cells. The critical rule also needs enough security cases to meet its minimum of 5, and every other declared coverage value has to appear if you want realised coverage of 1. Size the acceptance file for the strictest of those counts, then add the cases the Wilson bound needs for each proportion. Record the reason in thresholds.yaml and under the matching heading in methodology.md.

Run the tests

From a checkout, with the runner installed as in the quickstart:
Leave the fixture running. In a second terminal:
The fixture model ids for this pack are in Packs. etalon-contact-baseline returns the referee labels. etalon-contact-security-leak sends security tickets to billing. This pack declares a TypeSafe judge (transport: typesafe). A live judge needs pip install 'etalon[typesafe]' and TYPESAFE_API_KEY. Without the key the judge does not return verdicts. qualify still writes a bundle, and the decision is INDETERMINATE. That remains true when reference metrics already show a critical miss: a judge that did not run keeps the run inconclusive. Details are in TypeSafe judge. For a completed decision with no external service, use the demonstration quickstart. --model etalon-demo-baseline is QUALIFIED, --model etalon-demo-small is NOT_QUALIFIED on the complaint critical rule, and --model etalon-demo-error is INDETERMINATE. qualify exits 0 when the capture finished and a bundle was written. The decision may be QUALIFIED, NOT_QUALIFIED, or INDETERMINATE. Exit 2 means the endpoint did not complete the capture. A bundle is still written, and the decision is INDETERMINATE.

Examine failures and results that do not support a decision

Use the run directory qualify printed.
inspect prints the decision, realised coverage, the critical-event row, and each requirement with its point estimate and verdict (met, missed, straddle, or insufficient). verify checks the bundle bytes. Then read these files in the run directory: NOT_QUALIFIED means a completed capture had a hard miss: a mandatory requirement on the failing side of its interval, or critical events above the maximum. The failing cases are in failures.json. Fix the endpoint or, when the limit itself is wrong, change the pack version. Do not average the miss away in the report. INDETERMINATE means the run does not support a decision. decision.json indeterminate_conditions names which: An incomplete capture stays INDETERMINATE even when critical events were already counted. Read what_would_resolve and run again after that condition is addressed. Do not treat the run as NOT_QUALIFIED or QUALIFIED.

Repeat the tests after a deployment change

A saved bundle stays current until a declared trigger changes. valid_until in decision.json is that sentence. It is not a date. etalon check re-reads the endpoint and the inputs you pass. It does not re-score the corpus. Re-supply every recorded input. For a routing run:
Exit 0 means each captured trigger value was supplied again and matches. A difference is printed with its severity. On this pack, examples/contact-routing/triggers.yaml puts model weights, model revision, quantization, the system prompt, retrieval_corpus, pack version, threshold changes, and evaluator version on always_requalify. Engine version, serving config, sampling, chat template, and runner version are review_required. The operator name is no_material_impact. This routing fixture does not qualify a retrieval component; retrieval_corpus is still an always_requalify trigger on the pack. When a trigger fired, or when you changed the acceptance cases, run qualify again. Write a new run directory. Do not --resume into the previous run, and do not reuse its --run-id. Resuming continues one interrupted capture of one configuration.
compare prints both decisions, each metric’s interval, the critical-event counts, and realised coverage. Fingerprint fields that differ are mapped to the trigger classes above. A change to the corpus, rubrics, thresholds or their rationale, the qualification policy, evaluator configuration, methodology, or triggers is a new pack version before you treat the new bundle as the acceptance record. The run stores the pack version and the content hash.

Not in this release

This repository does not ship a personal-data test pack, and the runner does not collect product telemetry. The packs here declare an OpenAI-compatible chat-completions interface. Other protocols are not added in this tree. Adapting the routing example above is the path for a new routing task.