> ## Documentation Index
> Fetch the complete documentation index at: https://docs.etalon.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# How qualification works

> How a pack, a run, and a decision produce an evidence bundle.

# How qualification works

A pack is the specification. The runner executes that pack against one
endpoint and writes one evidence bundle. The packs checked in here are
public fixtures. They are described in [Packs](/docs/packs.md).

The runner does not determine that a regulation, contract, or internal
control has been met. The decision states what was evaluated, under which
recorded configuration, and how the pack policy read the scores.

The same decision, coverage, and check rules are stated in
[Overview](/docs/overview.md). This page follows them through a run.

## What `qualify` does

`etalon qualify` loads the pack, captures a fingerprint, and sends every
qualification case and every challenge case to the endpoint. It then repeats
the pack's variance subset of qualification cases `repeat_runs - 1` more
times, so the bundle can report variance. Calibration cases are not sent.
When the pack declares a judge, that judge scores the captured outputs on the
first pass and, separately, scores each calibration case from the candidate
output stored on the case.

The endpoint client sets `trust_env=False` and does not read proxy variables.
It talks only to `--endpoint`. When `--judge-endpoint` names a different URL
and the pack uses the `endpoint` judge transport, a second client is opened
for that URL. The `typesafe` transport calls TypeSafe with its own client.
That call is not the endpoint client. See [Judges](/docs/judges.md).

A credential for the endpoint, when one is set, comes from `ETALON_API_KEY` or
`OPENAI_API_KEY`. It is not read from the pack and it is not written into the
bundle. The TypeSafe key is `TYPESAFE_API_KEY`, also an environment variable,
also not written into the bundle.

Sampling flags (`--temperature`, `--top-p`, `--seed`, `--max-tokens`) replace
the pack's sampling values for that run. When a flag is omitted, the pack
value is sent. The fingerprint records the sampling that was sent.

The run directory is `<output>/<run-id>`. `--run-id` sets the directory name.
Without it, the runner uses a UTC timestamp and a short random suffix. An
existing run directory is refused.

## Evidence bundle

`hashes.json` is not the list of files that belong in a run. The runner's
expected file set is. `etalon verify` fails when an expected file is missing,
an expected file is absent from `hashes.json`, an extra file is in the
directory, a digest is malformed, the manifest decision does not match
`decision.json`, the fingerprint is ambiguous, or file bytes do not match.

Hashed files:

| File                | What it records                                                                                                                                 |
| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| `manifest.json`     | Run id, pack id and version, pack content hash, runner version, operator, endpoint host, whether the capture completed, and a decision summary. |
| `fingerprint.json`  | The recorded configuration. Fields that were not obtained are `unavailable` with a reason. They are not invented.                               |
| `inputs.jsonl`      | Request bodies sent to the endpoint.                                                                                                            |
| `outputs.jsonl`     | Response bodies, status codes, and errors.                                                                                                      |
| `results.jsonl`     | Per-case evaluator results.                                                                                                                     |
| `metrics.json`      | Qualification metrics, challenge metrics, repeat-run variance inputs, and the judge record when a judge ran.                                    |
| `coverage.json`     | Realised coverage under value-presence.                                                                                                         |
| `failures.json`     | Grouped failure examples from the first pass of each case.                                                                                      |
| `decision.json`     | Status, requirement rows, critical-event count, indeterminate conditions, and `valid_until`.                                                    |
| `triggers.json`     | The pack's requalification triggers. Values are read from the fingerprint.                                                                      |
| `pack_summary.json` | Pack identity, methodology text, metric definitions, thresholds, and evaluator ids.                                                             |
| `report.json`       | The report view, with case redaction applied.                                                                                                   |
| `report.html`       | The same view as a static HTML file.                                                                                                            |

`hashes.json` lists the SHA-256 of each of those files. It is not itself one
of the hashed files. `signature.json` is optional, is not hashed, and signs
the raw bytes of `hashes.json`. Unsigned bundles still verify. See
[Signing](/docs/signing.md).

Tokens and credentials that match the runner's secret patterns are removed
from the bundle and replaced with `[secret-excluded]`. That walk includes
reasoning traces (`reasoning`, `reasoning_content`, and `thinking`) because
they are strings in the stored response. Cases marked `sensitive: true` stay
in `inputs.jsonl` and `outputs.jsonl`. In `report.html` and `report.json`
those cases are quoted as `[redacted]` unless the report was rendered with
case redaction off. Case redaction does not remove reasoning traces from the
evidence files. `strip_reasoning: true` in the redaction policy, or
`--strip-reasoning`, replaces those fields with `[reasoning-stripped]` after
secrets have been counted. Token counts on the output row are kept.

## Fingerprint

Each metadata field is an envelope. `known` and `externally_supplied` are
captured values. `unavailable` records a reason. `not_applicable` records why
the field does not apply.

This version records retrieval as `not_applicable` with the reason that the
pack does not qualify a retrieval component. Weight files are not hashed as
one blob. Pass `--model-index` with per-file hashes (`files` or
`file_hashes`, each value `sha256:` and 64 hex characters) when you have
them. An empty or malformed index stays `unavailable`. GPU, driver, and CUDA
fields come from `nvidia-smi` when it returns them. On macOS, when
`nvidia-smi` has no GPU, the runner reads the display name and Metal family
from `system_profiler SPDisplaysDataType`. CUDA stays unavailable on that
host. A field that still cannot be read is `unavailable` with a reason.

When the endpoint serves Ollama's `/api/tags`, the model digest is recorded
as `model.runtime_digest` (`known`, source `ollama /api/tags`). A missing or
rejected digest is `unavailable` with a reason. The digest does not replace
a per-file model index, and it is not a requalification trigger.

`valid_until` is the sentence `material change to any declared requalification trigger`.
It is not a date. `etalon check` compares a later fingerprint with the one in
the bundle. See the [CLI reference](/docs/cli.md).

## Decision

`QUALIFIED`, `NOT_QUALIFIED`, and `INDETERMINATE` come only from the pack
policy. Proportion requirements use the 95% Wilson interval. A
critical-failure maximum of zero fails on a single observed event. The
interval is not used to soften that rule.

A completed capture with any hard miss is `NOT_QUALIFIED` even when another
metric straddles its threshold. A hard miss is a mandatory requirement on the
failing side of its interval, or critical events above the declared maximum.
An incomplete capture is `INDETERMINATE` even when critical events were
already counted. The decision summary states that critical-event count
whenever it was produced.

A missing compared field is one schema failure (or `truncated` when
`finish_reason` is `length` and the required object is absent). That case is
not scored on the label, behaviour, or critical metrics. The metric records
how many cases were omitted. The report's diagnosis section is a confusion
matrix and per-class recall for label and behaviour on the primary
qualification pass, plus capture wall-clock, effective concurrency, and token
medians when the endpoint reported them. Those figures do not gate the
decision.

`INDETERMINATE` is also the result when the interval crosses a threshold,
realised coverage is below the pack minimum, repeat-run variance is above the
pack maximum, a required evaluator did not score, a metric has fewer scored
cases than its minimum sample size, or a mandatory fingerprint field is
missing. A judge that does not return a usable verdict is
`INDETERMINATE`, not a model miss. When the pack lists `judge_uncertainty`,
uncertainty at or above the pack maximum is `INDETERMINATE` unless a hard
miss on a completed capture is already `NOT_QUALIFIED`. That uncertainty is
still written on the decision. Agreement below the pack's declared
judge–referee record is `calibration_invalid`.

Reference metrics stay the qualification decision. A judge is an overlay.
[Judges](/docs/judges.md) describes the transports.

## Coverage

v0.1 realised coverage is value-presence. A declared dimension value counts
when it appears at least once on a qualification case. The realised figure is
the unweighted mean of those fractions. This version does not expand a
cross-product of coverage cells, and it does not weight a value by how often
it appears. A pack that declares another coverage model is rejected by
`etalon pack validate`.

`qualify`, `inspect`, `compare`, and the HTML report state that limit next to
the number.
