Etalon
Operator steps, the command reference, and the public fixtures are in the documentation index. This page states how a decision is reached. How qualification works follows the same rules through the files a run writes. Etalon produces qualification evidence for one self-hosted model deployment. A pack is the specification: corpus, rubrics, evaluators, thresholds, decision policy, and requalification triggers. The runner executes that pack inside the customer’s network, records the system fingerprint with explicit metadata states, and writes an evidence bundle plus a static HTML report. The report says what was evaluated and under which recorded configuration. It does not determine that a regulation or internal control has been met. The packs checked in here are public fixtures only.examples/demo-pack is a synthetic public fixture for tests and the offline
demo. examples/contact-routing is a second public fixture, with a judge
overlay. See Packs.
Layout
Decisions
QUALIFIED, NOT_QUALIFIED, and INDETERMINATE come only from the pack
policy. Proportion requirements use the 95% Wilson interval. A critical-failure
maximum of zero fails on a single observed event. A completed capture with any
hard miss is NOT_QUALIFIED even when another metric straddles its threshold.
An incomplete capture is INDETERMINATE even when critical events were already
counted. The decision summary states that critical-event count whenever it was
produced. INDETERMINATE is also a normal outcome when the interval crosses a
threshold, coverage is short, variance is high, a required evaluator did not
score, or mandatory fingerprint fields are missing.
Qualification stays current until a declared trigger changes. valid_until is
that sentence, never a date.
Judge and calibration
A pack may declare a rubric judge. The runner records the judge model, the prompt hash, the prompt version, and the rubric version. Every judge verdict records uncertainty. A missing or unusable verdict is a judge failure and the run isINDETERMINATE. When the pack lists judge_uncertainty, uncertainty at
or above the pack maximum is also INDETERMINATE, unless a hard miss on a
completed capture is already NOT_QUALIFIED. That uncertainty is still written
on the decision.
typesafe (alias jev) is the TypeSafe Jev transport. Uncertainty there is
the Noul probability that the verdict is uncertain. Choice confidence is
evidence and does not fill a missing Noul. A live run needs TYPESAFE_API_KEY.
See docs/typesafe-judge.md. Critical dimensions stay on reference evaluators.
Calibration cases are not sent to the system under test. They carry a human
score and a candidate output. The runner reports judge–referee agreement and
the known-failure cases. Agreement below the declared record is
calibration_invalid.
Reference evaluators remain the qualification metrics. The report, inspect,
and compare show the reference route and the judge together when a judge
ran. Failure classes for this route are critical misroute, soft ambiguity,
abstain miss, and judge failure.
Coverage
v0.1 realised coverage is value-presence. A declared dimension value counts when it appears at least once on a qualification case. The realised figure is the unweighted mean of those fractions. This version does not expand a cross-product of coverage cells, and it does not weight a value by how often it appears.etalon qualify, etalon inspect, etalon compare, and the HTML report state that
limit next to the number. A pack that declares another coverage model is
rejected. A fraction of 1 means every declared value appeared at least once on
its own dimension. It does not mean every combination of values was measured.
Check and compare
etalon check prints a short legend for exit 0, 1, and 2, then three kinds of row:
- fired — the re-supplied value differs from the stored projection
- not re-observed — a captured value was not re-supplied; this changes the exit by that trigger’s severity
- unverified gaps — the stored field is
unavailable; the row is printed at that trigger’s severity and does not by itself change the exit
not_observed. Each row prints the severity, the trigger name, the stored projection, and the live projection. An ambiguous fingerprint (a trigger field is missing, or an envelope contradicts itself) is listed on its own. It is not treated as an unavailable gap, and it prevents exit 0.
etalon compare prints both decision statuses, each metric’s point estimate and
confidence interval, critical-event counts, and realised coverage, and it
states that realised coverage is value-presence. Fingerprint fields that
differ are mapped to always_requalify, review_required, or
no_material_impact. The line qualification evidence changed is about the
decision, the intervals, the critical-event count, and realised coverage.
Signatures, gates, and review exports
An operator can signhashes.json with an Ed25519 seed held offline.
signature.json sits beside the bundle and is not one of the hashed files.
Unsigned bundles from earlier versions still verify. Verification with a
separately held public key fails when the evidence bytes change, including
when hashes.json is rewritten to match the changed bytes.
examples/verify_bundle.py repeats that check without installing Etalon. A
GitHub Action runs the same script and fails the job when the bundle is
missing, altered, unsigned, or signed by a different key. etalon export
writes an HTML review file and a manifest for a person. That package is
qualification evidence. It does not determine that a regulation, contract,
or internal control has been met.
Key generation, custody, and the gate are described in docs/signing.md.