> ## Documentation Index
> Fetch the complete documentation index at: https://docs.etalon.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Customer test packs

> Build and run a routing test pack, using the public contact-routing fixture as the complete example.

# Customer test packs

This page is the procedure for one routing task. The complete example is the
public fixture `examples/contact-routing` (pack id
`coyos.fs.contact-routing`, version `0.3.1`). Every case in that directory is
invented. After you have a pack, the commands below are how you run it again
on your own.

The demonstration you can start from `etalon pack init` is a different pack.
The commercial contact-routing pack is a private delivery and is not in this
repository. Those differences are in [Packs](/docs/packs).

A decision is the pack policy reading the scores for one recorded deployment.
`QUALIFIED` is evidence for that pack version and that endpoint fingerprint.
It does not authorize use of the model without human review. It does not
determine that a regulation, contract, or internal control has been met.

## Define the task and the expected results

Write the task in `pack.yaml` `intended_use`, and write the reply the endpoint
must produce.

The routing example asks a self-hosted model to read one synthetic support
message and return one JSON object:

```json theme={null}
{"label": "billing", "behaviour": "classify"}
```

`label` is one of `billing`, `access`, `technical`, `security`, `spam`.
`behaviour` is one of `classify`, `abstain`, `escalate`. The system prompt in
`examples/contact-routing/system_prompt.txt` states the same rules: abstain
when the message is ambiguous, and escalate a security ticket while keeping
the security label.

Each case records the referee answer the scores compare against. This is
`Q-0001` from `examples/contact-routing/corpus/qualification.jsonl`:

```json theme={null}
{
  "case_id": "Q-0001",
  "category": "short_security",
  "severity": "low",
  "input": {
    "messages": [
      {
        "role": "user",
        "content": "Expert-authored fixture, not a customer record. Synthetic security fixture: a staged sign-in from an unexpected fixture network was logged. Reference Q-0001."
      }
    ]
  },
  "expected": {
    "label": "security",
    "required_behaviour": "escalate",
    "reference_output": null,
    "required_facts": [],
    "prohibited_claims": []
  },
  "coverage": {
    "intent": "security",
    "severity": "low",
    "input_character": "short",
    "expected_behaviour": "escalate"
  },
  "provenance": "expert-authored",
  "metadata": {"language": "en", "source_type": "expert_fixture"},
  "sensitive": false,
  "notes": ""
}
```

`expected.label` and `expected.required_behaviour` are the results the
reference evaluators score. `coverage` names one declared value on each
dimension in `pack.yaml`. The values have to be listed under
`coverage.dimensions`. `provenance` is `synthetic`, `expert-authored`,
`customer-derived`, or `public-benchmark`.

`etalon pack validate` checks the line. A missing field names the property and
lists the required case fields. A coverage value that is not in `pack.yaml`
is reported as `coverage value ... is not declared`.

When you change the label set, change it in the same places together:
`pack.yaml` `coverage.dimensions`, the system prompt, `expected.label` on
every case, and `critical_labels` in `evaluators.yaml` and the critical
rubric. The runner does not infer the label set from the prompt.

## Identify critical errors

A critical error is the failure the pack refuses to average away. On the
routing example it is a security ticket sent to a public queue (`billing`,
`access`, `technical`, or `spam`).

The rule is in `examples/contact-routing/qualification.yaml`:

```yaml theme={null}
critical_failures:
  metric: critical_misroute_rate
  maximum: 0.00
```

The rubric `rubrics/critical_misroute.yaml` states the same condition. The
metric counts qualification cases whose referee label is `security` and whose
output contains a label outside that set. One such case fails a maximum of
zero. The Wilson interval is reported and is not used to soften the rule.
An output that omits `label` is a schema miss, not a critical misroute.

Other misses are mandatory requirements, not this critical rule: label
accuracy, schema compliance, and abstention correctness. They use a 95%
Wilson bound. A schema failure can still make the run `NOT_QUALIFIED` when
its interval is on the failing side of its threshold. It is not the critical
event count.

The demonstration pack uses the same kind of rule on a different label.
There, a complaint routed outside the complaint set is the critical error.
That is `etalon.demo.ticket-routing`, not this routing task.

## Keep development cases separate from acceptance cases

The runner scores three files, named in `pack.yaml` `artifacts`:

| File | Role in the routing example |
| - | - |
| `corpus/qualification.jsonl` | Acceptance set. 48 cases. Sent to the endpoint. These scores decide the run. |
| `corpus/calibration.jsonl` | Judge holdout. 12 cases. Not sent to the endpoint. Each row carries `metadata.human_passed`, `metadata.human_failure_type` when the referee failed the candidate, and `metadata.candidate_output`. |
| `corpus/challenge.jsonl` | 8 cases. Sent to the endpoint and reported separately. They do not decide the run unless `qualification.yaml` sets `challenge_critical_failures`. |

Development cases do not have a fourth slot. Put them in a file the manifest
does not name, for example `corpus/development.jsonl`. `etalon pack validate`
does not read that file, and `qualify` does not send it. Pointing
`artifacts.corpus` at the development file makes those cases the acceptance
set.

Promote a case by moving it into `corpus/qualification.jsonl` and deleting it
from the development file. Case ids must be unique across the three artifact
files. The same input text under two ids is still one situation; the runner
does not detect copied text. A validation error names both files when an id
is reused: move the case, do not copy it.

Calibration rows are referee scores of a stored candidate. They are not a
place to park endpoint cases you are still editing. A calibration row without
`metadata.human_passed` or `metadata.candidate_output` fails validation with
the field name.

## Select acceptance limits and the number of cases

Limits live in two files that must match, plus a heading in `methodology.md`:

* `thresholds.yaml` holds the metric, the minimum sample size, the threshold, and the written rationale.
* `qualification.yaml` repeats the operator, the threshold, and `rationale_ref` for each mandatory requirement, and holds the critical maximum.

`etalon pack validate` rejects a requirement whose numbers differ from
`thresholds.yaml`, and a `rationale_ref` whose anchor is missing from
`methodology.md`. Changing a threshold or its rationale is a new pack version.

The routing fixture's numbers are demonstration bounds for this synthetic
corpus. Copying them into a customer pack is not a risk appetite.

Proportion requirements use the 95% Wilson interval, not the point estimate
alone. For a perfect score on `n` cases the lower bound is `n / (n + 3.841)`.
That bound is at least 0.90 when `n` is 35 or more, and at least 0.80 when
`n` is 16 or more. A point estimate of 1.0 on fewer cases straddles the
threshold, and the run is `INDETERMINATE`.

On this fixture:

| Metric | Threshold | Minimum scored cases | Who is counted | Cases in the fixture |
| - | - | - | - | - |
| `label_accuracy` | 0.90 | 30 | Qualification cases that produced a label | 48 qualification cases. A perfect score clears 0.90 (lower bound about 0.926). 30 perfect cases do not (lower bound about 0.886). |
| `schema_compliance` | 0.90 | 30 | Qualification cases | Same 48 cases. A threshold of 1.00 cannot be cleared by a Wilson interval on a finite sample. |
| `abstention_correctness` | 0.80 | 10 | Cases whose referee behaviour is abstain or escalate, and whose output contains a behaviour | 10 abstain and 10 escalate. 16 perfect applicable cases are what clears 0.80. The minimum of 10 is a separate gate. |
| `critical_misroute_rate` | maximum 0 | 5 | Qualification cases whose referee label is security and whose output contains a label | 10 security cases. Fewer than 5 scored security cases is `INDETERMINATE`, not a pass. One public-queue label is a critical failure. |

`minimum_sample_size` and the Wilson bound are both required. Meeting the
sample-size minimum with a perfect score can still be `INDETERMINATE` when
the interval crosses the threshold.

Coverage is value-presence. The fixture declares intent (5 values), severity
(3), input character (4), and expected behaviour (3). Realised coverage is the
unweighted mean of the fraction of those values that appear at least once on
a qualification case. The minimum is 0.80. A value that never appears lowers
only its dimension. Coverage is not a count of cases and not a cross-product
of cells.

The critical rule also needs enough security cases to meet its minimum of 5,
and every other declared coverage value has to appear if you want realised
coverage of 1. Size the acceptance file for the strictest of those counts,
then add the cases the Wilson bound needs for each proportion. Record the
reason in `thresholds.yaml` and under the matching heading in `methodology.md`.

## Run the tests

From a checkout, with the runner installed as in the [quickstart](/docs/quickstart):

```bash theme={null}
etalon pack validate examples/contact-routing

python3 examples/mock_endpoint.py --port 8000 --pack examples/contact-routing
```

Leave the fixture running. In a second terminal:

```bash theme={null}
etalon qualify \
  --endpoint http://127.0.0.1:8000/v1 \
  --pack examples/contact-routing \
  --serving-config examples/contact-routing/serving.yaml \
  --system-prompt examples/contact-routing/system_prompt.txt \
  --model etalon-contact-baseline \
  --operator "Demo Operator" \
  --output ./evidence
```

The fixture model ids for this pack are in [Packs](/docs/packs).
`etalon-contact-baseline` returns the referee labels.
`etalon-contact-security-leak` sends security tickets to billing.

This pack declares a TypeSafe judge (`transport: typesafe`). A live judge
needs `pip install 'etalon[typesafe]'` and `TYPESAFE_API_KEY`. Without the
key the judge does not return verdicts. `qualify` still writes a bundle, and
the decision is `INDETERMINATE`. That remains true when reference metrics
already show a critical miss: a judge that did not run keeps the run
inconclusive. Details are in [TypeSafe judge](/docs/typesafe-judge).

For a completed decision with no external service, use the demonstration
quickstart. `--model etalon-demo-baseline` is `QUALIFIED`,
`--model etalon-demo-small` is `NOT_QUALIFIED` on the complaint critical rule,
and `--model etalon-demo-error` is `INDETERMINATE`.

`qualify` exits 0 when the capture finished and a bundle was written. The
decision may be `QUALIFIED`, `NOT_QUALIFIED`, or `INDETERMINATE`. Exit 2
means the endpoint did not complete the capture. A bundle is still written,
and the decision is `INDETERMINATE`.

## Examine failures and results that do not support a decision

Use the run directory `qualify` printed.

```bash theme={null}
etalon inspect ./evidence/<run-id>
etalon verify ./evidence/<run-id>
```

`inspect` prints the decision, realised coverage, the critical-event row, and
each requirement with its point estimate and verdict (`met`, `missed`,
`straddle`, or `insufficient`). `verify` checks the bundle bytes.

Then read these files in the run directory:

| File | What to look at |
| - | - |
| `decision.json` | `status`, `requirements`, `critical_failures`, `indeterminate_conditions`, `what_would_resolve`. |
| `failures.json` | `groups[].examples`: `case_id`, `test_id`, `expected`, `actual`, `detail`, `input_ref`, `output_ref`. |
| `metrics.json` | `n`, the Wilson interval, `unscorable`, and `minimum_sample_size` on each metric. |
| `coverage.json` | Which declared values were absent. |
| `report.html` | The same view in a browser. `etalon report ./evidence/<run-id>` prints the path. |

`NOT_QUALIFIED` means a completed capture had a hard miss: a mandatory
requirement on the failing side of its interval, or critical events above the
maximum. The failing cases are in `failures.json`. Fix the endpoint or, when
the limit itself is wrong, change the pack version. Do not average the miss
away in the report.

`INDETERMINATE` means the run does not support a decision. `decision.json`
`indeterminate_conditions` names which:

| Condition | What happened |
| - | - |
| `confidence_interval_straddles_threshold` | The Wilson interval crosses the limit. The point estimate can still look acceptable. |
| `insufficient_sample` | A metric has fewer scored cases than its minimum. |
| `realised_coverage_below_minimum` | A declared coverage value is missing, or too many are. |
| `repeat_run_variance_above_max` | The repeated subset moved more than the pack allows. |
| `required_evaluator_failed` | An evaluator did not score. A missing TypeSafe key is this condition. |
| `calibration_invalid` | Judge–referee agreement, or the calibration sample, missed the declared record. |
| `judge_uncertainty` | A verdict's uncertainty is at or above the pack maximum. |
| `mandatory_metadata_missing` | A required fingerprint field was not captured. `model.identifier` is mandatory on the routing example. |
| `endpoint_errors` | One or more requests failed. `qualify` exits 2. |

An incomplete capture stays `INDETERMINATE` even when critical events were
already counted. Read `what_would_resolve` and run again after that condition
is addressed. Do not treat the run as `NOT_QUALIFIED` or `QUALIFIED`.

## Repeat the tests after a deployment change

A saved bundle stays current until a declared trigger changes. `valid_until`
in `decision.json` is that sentence. It is not a date.

`etalon check` re-reads the endpoint and the inputs you pass. It does not
re-score the corpus. Re-supply every recorded input. For a routing run:

```bash theme={null}
etalon check ./evidence/<run-id> \
  --endpoint http://127.0.0.1:8000/v1 \
  --pack examples/contact-routing \
  --serving-config examples/contact-routing/serving.yaml \
  --system-prompt examples/contact-routing/system_prompt.txt \
  --model etalon-contact-baseline \
  --operator "Demo Operator"
```

Exit 0 means each captured trigger value was supplied again and matches.
A difference is printed with its severity. On this pack,
`examples/contact-routing/triggers.yaml` puts model weights, model revision,
quantization, the system prompt, `retrieval_corpus`, pack version, threshold
changes, and evaluator version on `always_requalify`. Engine version, serving
config, sampling, chat template, and runner version are `review_required`.
The operator name is `no_material_impact`. This routing fixture does not
qualify a retrieval component; `retrieval_corpus` is still an
`always_requalify` trigger on the pack.

When a trigger fired, or when you changed the acceptance cases, run `qualify`
again. Write a new run directory. Do not `--resume` into the previous run,
and do not reuse its `--run-id`. Resuming continues one interrupted capture
of one configuration.

```bash theme={null}
etalon qualify \
  --endpoint http://127.0.0.1:8000/v1 \
  --pack examples/contact-routing \
  --serving-config examples/contact-routing/serving.yaml \
  --system-prompt examples/contact-routing/system_prompt.txt \
  --model etalon-contact-baseline \
  --operator "Demo Operator" \
  --output ./evidence

etalon compare ./evidence/<previous-run-id> ./evidence/<new-run-id>
```

`compare` prints both decisions, each metric's interval, the critical-event
counts, and realised coverage. Fingerprint fields that differ are mapped to
the trigger classes above.

A change to the corpus, rubrics, thresholds or their rationale, the
qualification policy, evaluator configuration, methodology, or triggers is a
new pack version before you treat the new bundle as the acceptance record.
The run stores the pack version and the content hash.

## Not in this release

This repository does not ship a personal-data test pack, and the runner does
not collect product telemetry. The packs here declare an OpenAI-compatible
chat-completions interface. Other protocols are not added in this tree.
Adapting the routing example above is the path for a new routing task.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.