Customer test packs
This page is the procedure for one routing task. The complete example is the public fixtureexamples/contact-routing (pack id
coyos.fs.contact-routing, version 0.3.1). Every case in that directory is
invented. After you have a pack, the commands below are how you run it again
on your own.
The demonstration you can start from etalon pack init is a different pack.
The commercial contact-routing pack is a private delivery and is not in this
repository. Those differences are in Packs.
A decision is the pack policy reading the scores for one recorded deployment.
QUALIFIED is evidence for that pack version and that endpoint fingerprint.
It does not authorize use of the model without human review. It does not
determine that a regulation, contract, or internal control has been met.
Define the task and the expected results
Write the task inpack.yaml intended_use, and write the reply the endpoint
must produce.
The routing example asks a self-hosted model to read one synthetic support
message and return one JSON object:
label is one of billing, access, technical, security, spam.
behaviour is one of classify, abstain, escalate. The system prompt in
examples/contact-routing/system_prompt.txt states the same rules: abstain
when the message is ambiguous, and escalate a security ticket while keeping
the security label.
Each case records the referee answer the scores compare against. This is
Q-0001 from examples/contact-routing/corpus/qualification.jsonl:
expected.label and expected.required_behaviour are the results the
reference evaluators score. coverage names one declared value on each
dimension in pack.yaml. The values have to be listed under
coverage.dimensions. provenance is synthetic, expert-authored,
customer-derived, or public-benchmark.
etalon pack validate checks the line. A missing field names the property and
lists the required case fields. A coverage value that is not in pack.yaml
is reported as coverage value ... is not declared.
When you change the label set, change it in the same places together:
pack.yaml coverage.dimensions, the system prompt, expected.label on
every case, and critical_labels in evaluators.yaml and the critical
rubric. The runner does not infer the label set from the prompt.
Identify critical errors
A critical error is the failure the pack refuses to average away. On the routing example it is a security ticket sent to a public queue (billing,
access, technical, or spam).
The rule is in examples/contact-routing/qualification.yaml:
rubrics/critical_misroute.yaml states the same condition. The
metric counts qualification cases whose referee label is security and whose
output contains a label outside that set. One such case fails a maximum of
zero. The Wilson interval is reported and is not used to soften the rule.
An output that omits label is a schema miss, not a critical misroute.
Other misses are mandatory requirements, not this critical rule: label
accuracy, schema compliance, and abstention correctness. They use a 95%
Wilson bound. A schema failure can still make the run NOT_QUALIFIED when
its interval is on the failing side of its threshold. It is not the critical
event count.
The demonstration pack uses the same kind of rule on a different label.
There, a complaint routed outside the complaint set is the critical error.
That is etalon.demo.ticket-routing, not this routing task.
Keep development cases separate from acceptance cases
The runner scores three files, named inpack.yaml artifacts:
Development cases do not have a fourth slot. Put them in a file the manifest
does not name, for example
corpus/development.jsonl. etalon pack validate
does not read that file, and qualify does not send it. Pointing
artifacts.corpus at the development file makes those cases the acceptance
set.
Promote a case by moving it into corpus/qualification.jsonl and deleting it
from the development file. Case ids must be unique across the three artifact
files. The same input text under two ids is still one situation; the runner
does not detect copied text. A validation error names both files when an id
is reused: move the case, do not copy it.
Calibration rows are referee scores of a stored candidate. They are not a
place to park endpoint cases you are still editing. A calibration row without
metadata.human_passed or metadata.candidate_output fails validation with
the field name.
Select acceptance limits and the number of cases
Limits live in two files that must match, plus a heading inmethodology.md:
thresholds.yamlholds the metric, the minimum sample size, the threshold, and the written rationale.qualification.yamlrepeats the operator, the threshold, andrationale_reffor each mandatory requirement, and holds the critical maximum.
etalon pack validate rejects a requirement whose numbers differ from
thresholds.yaml, and a rationale_ref whose anchor is missing from
methodology.md. Changing a threshold or its rationale is a new pack version.
The routing fixture’s numbers are demonstration bounds for this synthetic
corpus. Copying them into a customer pack is not a risk appetite.
Proportion requirements use the 95% Wilson interval, not the point estimate
alone. For a perfect score on n cases the lower bound is n / (n + 3.841).
That bound is at least 0.90 when n is 35 or more, and at least 0.80 when
n is 16 or more. A point estimate of 1.0 on fewer cases straddles the
threshold, and the run is INDETERMINATE.
On this fixture:
minimum_sample_size and the Wilson bound are both required. Meeting the
sample-size minimum with a perfect score can still be INDETERMINATE when
the interval crosses the threshold.
Coverage is value-presence. The fixture declares intent (5 values), severity
(3), input character (4), and expected behaviour (3). Realised coverage is the
unweighted mean of the fraction of those values that appear at least once on
a qualification case. The minimum is 0.80. A value that never appears lowers
only its dimension. Coverage is not a count of cases and not a cross-product
of cells.
The critical rule also needs enough security cases to meet its minimum of 5,
and every other declared coverage value has to appear if you want realised
coverage of 1. Size the acceptance file for the strictest of those counts,
then add the cases the Wilson bound needs for each proportion. Record the
reason in thresholds.yaml and under the matching heading in methodology.md.
Run the tests
From a checkout, with the runner installed as in the quickstart:etalon-contact-baseline returns the referee labels.
etalon-contact-security-leak sends security tickets to billing.
This pack declares a TypeSafe judge (transport: typesafe). A live judge
needs pip install 'etalon[typesafe]' and TYPESAFE_API_KEY. Without the
key the judge does not return verdicts. qualify still writes a bundle, and
the decision is INDETERMINATE. That remains true when reference metrics
already show a critical miss: a judge that did not run keeps the run
inconclusive. Details are in TypeSafe judge.
For a completed decision with no external service, use the demonstration
quickstart. --model etalon-demo-baseline is QUALIFIED,
--model etalon-demo-small is NOT_QUALIFIED on the complaint critical rule,
and --model etalon-demo-error is INDETERMINATE.
qualify exits 0 when the capture finished and a bundle was written. The
decision may be QUALIFIED, NOT_QUALIFIED, or INDETERMINATE. Exit 2
means the endpoint did not complete the capture. A bundle is still written,
and the decision is INDETERMINATE.
Examine failures and results that do not support a decision
Use the run directoryqualify printed.
inspect prints the decision, realised coverage, the critical-event row, and
each requirement with its point estimate and verdict (met, missed,
straddle, or insufficient). verify checks the bundle bytes.
Then read these files in the run directory:
NOT_QUALIFIED means a completed capture had a hard miss: a mandatory
requirement on the failing side of its interval, or critical events above the
maximum. The failing cases are in failures.json. Fix the endpoint or, when
the limit itself is wrong, change the pack version. Do not average the miss
away in the report.
INDETERMINATE means the run does not support a decision. decision.json
indeterminate_conditions names which:
An incomplete capture stays
INDETERMINATE even when critical events were
already counted. Read what_would_resolve and run again after that condition
is addressed. Do not treat the run as NOT_QUALIFIED or QUALIFIED.
Repeat the tests after a deployment change
A saved bundle stays current until a declared trigger changes.valid_until
in decision.json is that sentence. It is not a date.
etalon check re-reads the endpoint and the inputs you pass. It does not
re-score the corpus. Re-supply every recorded input. For a routing run:
examples/contact-routing/triggers.yaml puts model weights, model revision,
quantization, the system prompt, retrieval_corpus, pack version, threshold
changes, and evaluator version on always_requalify. Engine version, serving
config, sampling, chat template, and runner version are review_required.
The operator name is no_material_impact. This routing fixture does not
qualify a retrieval component; retrieval_corpus is still an
always_requalify trigger on the pack.
When a trigger fired, or when you changed the acceptance cases, run qualify
again. Write a new run directory. Do not --resume into the previous run,
and do not reuse its --run-id. Resuming continues one interrupted capture
of one configuration.
compare prints both decisions, each metric’s interval, the critical-event
counts, and realised coverage. Fingerprint fields that differ are mapped to
the trigger classes above.
A change to the corpus, rubrics, thresholds or their rationale, the
qualification policy, evaluator configuration, methodology, or triggers is a
new pack version before you treat the new bundle as the acceptance record.
The run stores the pack version and the content hash.