From 348c7c970274cb96c8bd5392f04286636ad353ee Mon Sep 17 00:00:00 2001 From: opento-suggestions Date: Wed, 19 Aug 2026 18:21:12 -0600 Subject: [PATCH 1/7] test(appraisal): candidate vectors for appraisal.policy_ref resolution MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `appraisal.policy_ref` is a bare URI. A record names the appraisal policy that produced its verdict but carries nothing stating what that URI held, so two verifiers resolving it at different times can retrieve different documents and both report `affirming` honestly. The enforcement policy is digest-bound through `policy.bundle_hash`; the appraisal policy is not. This adds seven candidate vectors under tests/vectors/appraisal-resolution/, their generator, and three test files that grade the set. It adds no module, no schema change, and no dependency, and nothing here imports or requires a conformance module that does not exist on main. Why vectors rather than a check. Nothing in this repository resolves `policy_ref` today, so there is no implementation to test. What a vector set can do before an implementation exists is fix what the answers should be, which is the more useful half while the shape is still open. The set at a glance, one defect per vector, every record identical except for `appraisal.policy_ref`: 01 no-binding-declared accept 02 resolved-and-matches accept 03 digest-mismatch, one byte apart reject 04 digest-mismatch, other object reject 05 referent unreachable deferred 06 digest algorithm uncomputable deferred 07 binding bound to another URI reject 01 and 02 are why the set is not one-directional. Written from the motivating problem alone, every vector would be a rejection or a deferral, and a verifier that rejects everything would pass. 01 is the backward-compatibility control: every conformant record today declares no binding and must keep verifying, or this set would be proposing a breaking change rather than describing a gap. 03 and 04 keep the contradicted boundary off a single vector. 03 differs from the appraised object in exactly one byte, moving a SLSA floor from 2 to 3, which flips this record's verdict; 04 substitutes an unrelated document of a different length. A verifier comparing lengths, or sampling a prefix, passes one and fails the other. 05 and 06 are unresolvable by different mechanisms: one cannot reach the object, the other reaches it and cannot compute over it, because the declared algorithm is outside the set the schema admits. 07 is the vector a well-formedness check passes. The binding is a valid sha256 digest and is the true digest of a real object in the set, while `policy_ref` cites a different one. Both halves are valid; the pair is not. What this deliberately does not decide. The outcome a verifier should record for an unresolvable citation is open across four surfaces and is tracked by agentrust-io/trace-spec#190. Vectors 05 and 06 assert only that the outcome is not `affirming`. `deferred` is fixture bookkeeping in a vector's expected block, not a proposed value for `appraisal.status`, which stays closed at affirming/warning/contraindicated/none; test_appraisal_resolution_completeness.py fails if any vector reuses a status value as an outcome. `candidate_binding` is likewise a candidate shape, marked CANDIDATE: in every vector and carried in the vector's `context`, never in `record` — `appraisal` is additionalProperties: false, so a record carrying it would be schema-invalid, and proposing a field is an editorial decision. Reproduction. The generator is deterministic: no keys, no clock, no randomness, no network. Digests are SHA-256 over the exact bytes of the sibling files under policies/, recomputable by anyone holding only that directory. test_appraisal_resolution_reproduces.py regenerates into a temporary directory and compares bytes rather than regenerating in place, which would compare the files to themselves and agree regardless. The guard is self-contained by choice. agentrust-io/trace-spec#171 covers that repository's examples/ and this repository has no equivalent registry; reaching across for one would be a guard that needs another checkout, which is a guard that gets skipped. .gitattributes pins eol=lf for the directory, and is load-bearing rather than tidy. With core.autocrlf=true, restoring a policy file through `git checkout --` rewrote its SHA-256 from d8764863... to 7e68506c..., which would break every digest in the set on a clean Windows clone. Records are unsigned and ASCII-only. Unsigned because the defect under test is resolution of a cited object, orthogonal to the envelope signature — signing would put a second variable in every vector — and because keyless vectors regenerate from this directory alone. ASCII-only because tests/conftest.py reads vectors with a bare Path.read_text() and no explicit encoding, so a non-ASCII byte would decode under the platform locale rather than a defined one. Verification, on main at c725bbb: 246 passed, 5 xpassed (201 + 5 without this change, so +45 and nothing displaced); 88 passed under -m "level0 or negative", all 45 new tests collecting into that gate; ruff and mypy clean on the new files. Refs: agentrust-io/trace-tests#63, agentrust-io/trace-spec#66, agentrust-io/trace-spec#190 --- tests/test_appraisal_resolution.py | 202 ++++++++ .../test_appraisal_resolution_completeness.py | 194 ++++++++ tests/test_appraisal_resolution_reproduces.py | 106 ++++ .../appraisal-resolution/.gitattributes | 13 + .../01-no-binding-declared.json | 57 +++ .../02-resolved-and-matches.json | 61 +++ .../03-digest-mismatch-minimal-mutation.json | 62 +++ .../04-digest-mismatch-different-object.json | 61 +++ .../05-referent-unreachable.json | 63 +++ .../06-digest-algorithm-uncomputable.json | 64 +++ .../07-digest-bound-to-other-referent.json | 62 +++ tests/vectors/appraisal-resolution/README.md | 173 +++++++ .../gen_appraisal_resolution.py | 452 ++++++++++++++++++ .../policies/appraisal-policy-v1-rev2.json | 18 + .../policies/appraisal-policy-v1.json | 18 + .../policies/appraisal-policy-v2.json | 21 + .../policies/unrelated-policy.json | 13 + 17 files changed, 1640 insertions(+) create mode 100644 tests/test_appraisal_resolution.py create mode 100644 tests/test_appraisal_resolution_completeness.py create mode 100644 tests/test_appraisal_resolution_reproduces.py create mode 100644 tests/vectors/appraisal-resolution/.gitattributes create mode 100644 tests/vectors/appraisal-resolution/01-no-binding-declared.json create mode 100644 tests/vectors/appraisal-resolution/02-resolved-and-matches.json create mode 100644 tests/vectors/appraisal-resolution/03-digest-mismatch-minimal-mutation.json create mode 100644 tests/vectors/appraisal-resolution/04-digest-mismatch-different-object.json create mode 100644 tests/vectors/appraisal-resolution/05-referent-unreachable.json create mode 100644 tests/vectors/appraisal-resolution/06-digest-algorithm-uncomputable.json create mode 100644 tests/vectors/appraisal-resolution/07-digest-bound-to-other-referent.json create mode 100644 tests/vectors/appraisal-resolution/README.md create mode 100644 tests/vectors/appraisal-resolution/gen_appraisal_resolution.py create mode 100644 tests/vectors/appraisal-resolution/policies/appraisal-policy-v1-rev2.json create mode 100644 tests/vectors/appraisal-resolution/policies/appraisal-policy-v1.json create mode 100644 tests/vectors/appraisal-resolution/policies/appraisal-policy-v2.json create mode 100644 tests/vectors/appraisal-resolution/policies/unrelated-policy.json diff --git a/tests/test_appraisal_resolution.py b/tests/test_appraisal_resolution.py new file mode 100644 index 0000000..a9aba0e --- /dev/null +++ b/tests/test_appraisal_resolution.py @@ -0,0 +1,202 @@ +"""The appraisal-resolution set proves itself: digests, referents, schema. + +No verifier resolves ``appraisal.policy_ref``, so there is no implementation to +run these against. What can be checked — and is checked here — is that the set +is internally what it claims to be: + + * every record is valid under the packaged schema on this branch's base, so + the set describes conformant records rather than malformed ones; + * every declared digest really is, or really is not, the SHA-256 of the bytes + the vector says the citation resolved to, matching its expected outcome; + * the unreachable vector really has no resolvable object; + * vector 06's algorithm really is outside the digest set the schema admits. + +A vector whose expected outcome and whose bytes disagree is worse than no +vector: it looks like coverage and argues for the wrong thing. +""" + +from __future__ import annotations + +import hashlib +import json +import re +from pathlib import Path + +import jsonschema +import pytest + +VECTOR_DIR = Path(__file__).parent / "vectors" / "appraisal-resolution" +SCHEMA_DIGEST_PATTERN = re.compile(r"^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$") + + +def _vector_paths() -> list[Path]: + return sorted(VECTOR_DIR.glob("[0-9][0-9]-*.json")) + + +def _load(path: Path) -> dict: + return json.loads(path.read_text(encoding="utf-8")) + + +def _sha256_of_file(rel: str) -> str: + return "sha256:" + hashlib.sha256((VECTOR_DIR / rel).read_bytes()).hexdigest() + + +VECTORS = _vector_paths() + + +@pytest.mark.level0 +def test_the_set_has_the_seven_vectors_it_documents() -> None: + names = [p.name for p in VECTORS] + assert len(names) == 7, names + for n, path in enumerate(VECTORS, start=1): + assert path.name.startswith(f"{n:02d}-"), ( + f"vectors must be contiguously numbered; got {path.name} at position {n}" + ) + + +@pytest.mark.level0 +@pytest.mark.parametrize("path", VECTORS, ids=lambda p: p.stem) +def test_the_record_is_valid_under_the_packaged_schema(path: Path, schema) -> None: + """Schema validity, so the set is about resolution and not about malformed input.""" + jsonschema.validate(_load(path)["record"], schema) + + +@pytest.mark.level0 +@pytest.mark.parametrize("path", VECTORS, ids=lambda p: p.stem) +def test_the_record_carries_a_bare_policy_ref_and_no_invented_field(path: Path) -> None: + """The candidate binding must never leak into the record. + + ``appraisal`` is ``additionalProperties: false`` in the packaged schema, so + a record carrying the candidate field would be schema-invalid — and + proposing it is a schema decision this set does not make. + """ + appraisal = _load(path)["record"]["appraisal"] + assert set(appraisal) <= {"status", "verifier", "policy_ref", "timestamp"}, ( + f"record appraisal carries an unexpected key: {sorted(appraisal)}" + ) + for key in appraisal: + assert "digest" not in key, f"record appraisal proposes a binding field: {key}" + + +@pytest.mark.level0 +@pytest.mark.parametrize("path", VECTORS, ids=lambda p: p.stem) +def test_the_declared_resolution_matches_the_bytes_on_disk(path: Path) -> None: + """Whatever the vector says resolution returned, the file must hash to it.""" + ctx = _load(path)["context"] + resolution = ctx["resolution"] + if resolution["outcome"] != "resolved": + assert resolution.get("file") in (None,), ( + "a resolution that did not resolve must not name a file" + ) + return + rel = resolution["file"] + assert (VECTOR_DIR / rel).is_file(), f"{rel} is missing from the set" + assert resolution["actual_sha256"] == _sha256_of_file(rel), ( + f"{rel} does not hash to the digest the vector records for it" + ) + + +@pytest.mark.level0 +@pytest.mark.negative +@pytest.mark.parametrize("path", VECTORS, ids=lambda p: p.stem) +def test_the_binding_agrees_with_the_expected_outcome(path: Path) -> None: + """The load-bearing self-proof. + + accept -> the declared binding equals the digest of what resolved + (or no binding is declared at all) + reject -> a binding and a resolution both exist, and they disagree + deferred-> the comparison could not be performed at all + """ + vector = _load(path) + ctx, outcome = vector["context"], vector["expected"]["outcome"] + binding = ctx.get("candidate_binding") + resolution = ctx["resolution"] + resolved = resolution["outcome"] == "resolved" + + if outcome == "pass": + if binding is None: + # 01: nothing was declared, so there is nothing that could contradict. + assert resolution["outcome"] == "not_attempted", ( + "a vector declaring no binding must not claim a resolution was " + "attempted; the point is that there was nothing to check" + ) + return + assert resolved, "an accepting vector with a binding must have resolved" + assert binding["value"] == resolution["actual_sha256"], ( + "accept requires the declared binding to equal what resolved" + ) + + elif outcome == "reject": + assert binding is not None, "a rejecting vector must declare a binding" + assert resolved, ( + "a rejecting vector must have resolved something; otherwise nothing " + "was contradicted and the case is unresolvable, not contradicted" + ) + assert SCHEMA_DIGEST_PATTERN.match(binding["value"]), ( + "a rejecting vector's binding must be well formed, so the rejection " + "is about the referent and not about a malformed digest" + ) + assert binding["value"] != resolution["actual_sha256"], ( + "reject requires the declared binding to differ from what resolved" + ) + + elif outcome == "deferred": + computable = ( + binding is not None + and SCHEMA_DIGEST_PATTERN.match(binding["value"]) is not None + ) + assert not (resolved and computable), ( + "a deferred vector must be one the verifier could not complete: " + "either the referent did not resolve, or the digest algorithm is " + "outside the set the schema admits" + ) + + else: # pragma: no cover - guarded by the completeness test + raise AssertionError(f"unknown expected outcome {outcome!r}") + + +@pytest.mark.level0 +@pytest.mark.negative +def test_05_really_has_no_resolvable_object() -> None: + vector = _load(VECTOR_DIR / "05-referent-unreachable.json") + assert vector["context"]["resolution"]["outcome"] == "unreachable" + cited = vector["context"]["cited_uri"] + tail = cited.rsplit("/", 1)[-1] + assert not (VECTOR_DIR / "policies" / tail).exists(), ( + "05 claims the referent is unreachable, but a sibling file matches its URI" + ) + + +@pytest.mark.level0 +@pytest.mark.negative +def test_06_names_an_algorithm_the_schema_does_not_admit() -> None: + vector = _load(VECTOR_DIR / "06-digest-algorithm-uncomputable.json") + value = vector["context"]["candidate_binding"]["value"] + assert not SCHEMA_DIGEST_PATTERN.match(value), ( + "06 claims the algorithm is uncomputable, but the digest matches the " + "pattern the schema admits" + ) + assert vector["context"]["resolution"]["outcome"] == "resolved", ( + "06 must reach its referent; that is what separates it from 05" + ) + + +@pytest.mark.level0 +def test_03_and_04_differ_in_kind_not_just_in_bytes() -> None: + """The pair that keeps the contradicted boundary off a single vector.""" + minimal = _load(VECTOR_DIR / "03-digest-mismatch-minimal-mutation.json") + wholesale = _load(VECTOR_DIR / "04-digest-mismatch-different-object.json") + + a = (VECTOR_DIR / minimal["context"]["resolution"]["file"]).read_bytes() + b = (VECTOR_DIR / "policies" / "appraisal-policy-v1.json").read_bytes() + assert len(a) == len(b), "03's mutation should not change the object's length" + assert sum(x != y for x, y in zip(a, b)) == 1, ( + "03 is the minimal-mutation vector; it must differ from the appraised " + "object in exactly one byte" + ) + + c = (VECTOR_DIR / wholesale["context"]["resolution"]["file"]).read_bytes() + assert len(c) != len(b), ( + "04 is the wholesale-substitution vector; a verifier comparing lengths " + "must be able to tell it from 03" + ) diff --git a/tests/test_appraisal_resolution_completeness.py b/tests/test_appraisal_resolution_completeness.py new file mode 100644 index 0000000..c358c57 --- /dev/null +++ b/tests/test_appraisal_resolution_completeness.py @@ -0,0 +1,194 @@ +"""Adequacy of the appraisal-resolution set, per the criteria on trace-spec#186. + +agentrust-io/trace-spec#186 (merged 2026-08-20) states what a conformance +vector set is claiming: *a verifier that does not implement these rules will +fail this set*. Three of its four criteria are checkable here and are checked; +the fourth is about repository-wide bookkeeping and is noted below. + +It merged into trace-spec, where it grades that repository's ``examples/``. +This repository has no adequacy harness, so nothing here is subject to it. +These criteria are a standard this set was built to by choice, and the tests +below are this set holding itself to them. + + 1. A set must fail BOTH unconditional implementations. + A set of all-rejections is passed by a verifier that rejects everything, + exactly as a set of all-acceptances is passed by one that accepts + everything. This set's three decided-reject vectors would, alone, be the + first failure. Vectors 01 and 02 exist to close it. + 2. Every boundary needs more than one vector. + One vector cannot separate a check that reads a prefix from one that + reads the whole object. + 3. Every set on disk is measured, or named with the test that measures it. + Repository-wide; trace-tests has no registry to add to, so it cannot be + asserted from inside one set. Recorded in the set's README instead. + 4. Shortfalls are recorded exactly. + See KNOWN_SHORTFALLS below. + +These tests grade the set, not a verifier. No verifier resolves +``appraisal.policy_ref`` today — that is the gap the set documents — so nothing +here claims an implementation was exercised. +""" + +from __future__ import annotations + +import json +from collections import Counter +from pathlib import Path + +import pytest + +VECTOR_DIR = Path(__file__).parent / "vectors" / "appraisal-resolution" + +DECIDED_OUTCOMES = {"pass", "reject"} +ALL_OUTCOMES = DECIDED_OUTCOMES | {"deferred"} + +# Criterion 4: shortfalls asserted to their exact extent, so they cannot widen +# quietly. Delete an entry when the shortfall is closed, not when it is excused. +KNOWN_SHORTFALLS = { + "no_verifier_exercised": ( + "No implementation resolves appraisal.policy_ref, so every expected " + "outcome is a claim about what a verifier should do, not a recording of " + "what one did. The set is a specification argument with runnable " + "internal consistency, not a conformance run." + ), + "unresolvable_outcome_unnamed": ( + "Vectors 05 and 06 assert only that the outcome is not affirming. The " + "value a verifier should record is open on agentrust-io/trace-spec#190 " + "and is deliberately not proposed here." + ), +} + + +def _vectors() -> list[dict]: + out = [] + for path in sorted(VECTOR_DIR.glob("[0-9][0-9]-*.json")): + out.append(json.loads(path.read_text(encoding="utf-8"))) + return out + + +def _outcome(vector: dict) -> str: + return vector["expected"]["outcome"] + + +@pytest.mark.level0 +def test_the_set_is_not_empty() -> None: + assert len(_vectors()) == 7, "the set is seven vectors, 01 through 07" + + +@pytest.mark.level0 +def test_criterion_1_the_set_fails_an_accept_everything_verifier() -> None: + """At least one vector a conformant verifier must reject.""" + rejects = [v for v in _vectors() if _outcome(v) == "reject"] + assert rejects, ( + "every vector expects acceptance, so a verifier that accepts " + "unconditionally passes the set" + ) + + +@pytest.mark.level0 +def test_criterion_1_the_set_fails_a_reject_everything_verifier() -> None: + """At least one vector a conformant verifier must accept. + + This is the criterion the set would otherwise fail. Its subject is a family + of resolution failures, so every vector written from the motivating problem + alone is a reject or a deferral. + """ + accepts = [v for v in _vectors() if _outcome(v) == "pass"] + assert accepts, ( + "no vector expects acceptance, so a verifier that rejects " + "unconditionally passes the set" + ) + assert len(accepts) >= 2, ( + "one must-accept vector cannot separate a verifier that accepts only " + "records declaring no binding from one that also checks a matching " + "binding; 01 and 02 are that pair" + ) + + +@pytest.mark.level0 +def test_criterion_2_every_boundary_carries_at_least_two_vectors() -> None: + counts = Counter(v["boundary"] for v in _vectors()) + thin = {b: n for b, n in counts.items() if n < 2} + assert not thin, f"boundaries carried by a single vector: {thin}" + assert set(counts) == {"accept", "contradicted", "unresolvable"}, ( + f"unexpected boundary set: {sorted(counts)}" + ) + + +@pytest.mark.level0 +def test_no_two_vectors_share_a_defect() -> None: + defects = [v["defect"] for v in _vectors()] + dupes = [d for d, n in Counter(defects).items() if n > 1 and d != "none"] + assert not dupes, f"defect exercised by more than one vector: {dupes}" + + +@pytest.mark.level0 +def test_every_expected_block_is_well_formed() -> None: + for v in _vectors(): + name = v["name"] + exp = v["expected"] + assert exp["outcome"] in ALL_OUTCOMES, f"{name}: bad outcome {exp['outcome']!r}" + assert exp.get("reason"), f"{name}: expected block carries no reason" + if exp["outcome"] == "deferred": + assert exp.get("deferred_pending") == "agentrust-io/trace-spec#190", ( + f"{name}: a deferred vector must name the issue it defers to" + ) + assert exp.get("must_not") == "affirming", ( + f"{name}: a deferred vector must still assert what it may not be" + ) + else: + assert "deferred_pending" not in exp, ( + f"{name}: a decided vector must not carry a deferral pointer" + ) + + +@pytest.mark.level0 +def test_no_vector_proposes_an_appraisal_status_value() -> None: + """The set must not coin the vocabulary trace-spec#190 exists to decide. + + ``deferred`` is fixture bookkeeping. It must never appear as an + ``appraisal.status``, and no vector may assert a status for the + unresolvable case. + """ + schema_enum = {"affirming", "warning", "contraindicated", "none"} + for v in _vectors(): + status = v["record"]["appraisal"]["status"] + assert status in schema_enum, f"{v['name']}: invented status {status!r}" + exp = v["expected"] + assert exp["outcome"] not in schema_enum, ( + f"{v['name']}: expected.outcome reuses an appraisal.status value, " + "which reads as proposing that value for this case" + ) + if exp["outcome"] == "deferred": + assert "status" not in exp, ( + f"{v['name']}: a deferred vector must not name a status to record" + ) + + +@pytest.mark.level0 +def test_candidate_binding_is_marked_candidate_everywhere_it_appears() -> None: + for v in _vectors(): + binding = v["context"].get("candidate_binding") + if binding is None: + continue + assert binding["field"].startswith("CANDIDATE:"), ( + f"{v['name']}: the binding field must be marked CANDIDATE, so it is " + "never read as a proposed schema field" + ) + + +@pytest.mark.level0 +def test_known_shortfalls_are_recorded_not_silent() -> None: + """Criterion 4: the gaps are asserted to their exact extent.""" + assert set(KNOWN_SHORTFALLS) == { + "no_verifier_exercised", + "unresolvable_outcome_unnamed", + }, ( + "the recorded shortfalls changed; update the README in the same commit " + "so the set never claims more coverage than it has" + ) + deferred = [v for v in _vectors() if _outcome(v) == "deferred"] + assert len(deferred) == 2, ( + "unresolvable_outcome_unnamed is recorded as covering exactly two " + f"vectors; found {len(deferred)}" + ) diff --git a/tests/test_appraisal_resolution_reproduces.py b/tests/test_appraisal_resolution_reproduces.py new file mode 100644 index 0000000..af62429 --- /dev/null +++ b/tests/test_appraisal_resolution_reproduces.py @@ -0,0 +1,106 @@ +"""The appraisal-resolution generator must reproduce its committed vectors. + +A committed vector nobody can regenerate is a number that cannot be checked. If +the generator and the files drift, the files win by default and the drift is +invisible, which is the failure agentrust-io/trace-spec#171 was written to +catch for that repository's ``examples/``. + +trace-tests has no equivalent guard, and trace-tests PR #66 states the reason a +cross-repository one is not the answer: "a guard that needs another repository +checked out is a guard that gets skipped." This guard is therefore +self-contained — it imports the generator that sits beside the vectors and +regenerates into a temporary directory, comparing bytes. + +Regenerating into ``tmp_path`` rather than in place is the load-bearing part. +Running the generator over its own directory and then comparing those files to +themselves agrees no matter what the generator does. +""" + +from __future__ import annotations + +import importlib.util +from pathlib import Path + +import pytest + +VECTOR_DIR = Path(__file__).parent / "vectors" / "appraisal-resolution" +GENERATOR = VECTOR_DIR / "gen_appraisal_resolution.py" + + +def _load_generator(): + spec = importlib.util.spec_from_file_location("gen_appraisal_resolution", GENERATOR) + assert spec and spec.loader, "generator module could not be loaded" + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def _committed_files() -> list[Path]: + return sorted(p for p in VECTOR_DIR.rglob("*.json") if p.is_file()) + + +@pytest.mark.level0 +def test_the_generator_is_present_beside_the_vectors() -> None: + """Self-containment: the guard must not need another repository.""" + assert GENERATOR.is_file(), ( + "gen_appraisal_resolution.py must sit in the vector directory, so the " + "set can be regenerated by anyone holding only that directory" + ) + + +@pytest.mark.level0 +def test_the_generator_reproduces_every_committed_file(tmp_path: Path) -> None: + module = _load_generator() + module.main(tmp_path) + + committed = _committed_files() + assert committed, "no vectors found; the set is empty" + + mismatched: list[str] = [] + missing: list[str] = [] + for path in committed: + rel = path.relative_to(VECTOR_DIR) + produced = tmp_path / rel + if not produced.is_file(): + missing.append(str(rel)) + continue + if produced.read_bytes() != path.read_bytes(): + mismatched.append(str(rel)) + + assert not missing, f"generator did not produce: {missing}" + assert not mismatched, ( + f"generator output differs from the committed bytes: {mismatched}. " + "Re-run tests/vectors/appraisal-resolution/gen_appraisal_resolution.py " + "and commit the result, or fix the generator — but do not edit a vector " + "by hand and leave the generator behind." + ) + + +@pytest.mark.level0 +def test_the_generator_produces_nothing_the_set_does_not_carry(tmp_path: Path) -> None: + """Both directions: a file the generator emits but nobody committed is drift too.""" + module = _load_generator() + module.main(tmp_path) + + produced = {p.relative_to(tmp_path) for p in tmp_path.rglob("*.json") if p.is_file()} + committed = {p.relative_to(VECTOR_DIR) for p in _committed_files()} + assert produced == committed, ( + f"only generated: {sorted(str(p) for p in produced - committed)}; " + f"only committed: {sorted(str(p) for p in committed - produced)}" + ) + + +@pytest.mark.level0 +def test_the_committed_bytes_use_lf_and_ascii_only() -> None: + """The two properties the digests depend on, asserted rather than assumed. + + CRLF translation changes every policy digest. ``.gitattributes`` in the + vector directory pins ``eol=lf``; this fails loudly if that protection is + ever removed. ASCII-only is separate: ``conftest.load_vector`` reads with + ``Path.read_text()`` and no explicit encoding, so a non-ASCII byte would be + decoded under the platform's locale. + """ + for path in _committed_files(): + data = path.read_bytes() + assert b"\r\n" not in data, f"{path.name} contains CRLF; digests will not match" + assert all(b < 128 for b in data), f"{path.name} contains a non-ASCII byte" diff --git a/tests/vectors/appraisal-resolution/.gitattributes b/tests/vectors/appraisal-resolution/.gitattributes new file mode 100644 index 0000000..69a8d4a --- /dev/null +++ b/tests/vectors/appraisal-resolution/.gitattributes @@ -0,0 +1,13 @@ +# The digests carried in these vectors are SHA-256 over the exact bytes of the +# sibling files under policies/. Line-ending translation therefore changes the +# answer: with core.autocrlf=true (the Git for Windows default) an unmarked +# checkout rewrites LF to CRLF, every policy digest stops matching, and both the +# byte-reproduction guard and the self-proof tests fail on a clean clone. +# +# Verified, not assumed: without this file, deleting +# policies/appraisal-policy-v1.json and restoring it with `git checkout --` +# changed its SHA-256 from d8764863... to 7e68506c... on a Windows worktree. +# +# `text eol=lf` keeps these diffable as text in review while pinning the working +# tree to LF on every platform. +* text eol=lf diff --git a/tests/vectors/appraisal-resolution/01-no-binding-declared.json b/tests/vectors/appraisal-resolution/01-no-binding-declared.json new file mode 100644 index 0000000..02dc208 --- /dev/null +++ b/tests/vectors/appraisal-resolution/01-no-binding-declared.json @@ -0,0 +1,57 @@ +{ + "name": "no-binding-declared", + "description": "The record cites an appraisal policy by bare URI and declares no binding to what that URI held. This is every conformant record today, so it must keep verifying: a set that rejected it would be proposing a breaking change rather than describing a gap.", + "boundary": "accept", + "defect": "none - backward-compatibility control", + "spec": "Mirrors the merged treatment of an undeclared depth in agentrust-io/trace-spec#173, where a record that never declared is read as surface rather than as a failure.", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + }, + "policy": { + "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", + "enforcement_mode": "enforce" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "resolution": { + "outcome": "not_attempted", + "note": "No binding is declared, so there is nothing to check." + }, + "candidate_binding": null + }, + "expected": { + "outcome": "pass", + "reason": "No binding declared; nothing to contradict." + } +} diff --git a/tests/vectors/appraisal-resolution/02-resolved-and-matches.json b/tests/vectors/appraisal-resolution/02-resolved-and-matches.json new file mode 100644 index 0000000..eb9b376 --- /dev/null +++ b/tests/vectors/appraisal-resolution/02-resolved-and-matches.json @@ -0,0 +1,61 @@ +{ + "name": "resolved-and-matches", + "description": "The cited URI resolves and its bytes hash to exactly the declared candidate binding. The second must-accept vector, and the only one in which a resolution actually succeeds.", + "boundary": "accept", + "defect": "none - positive control", + "spec": "agentrust-io/trace-spec docs/verification.md - a verifier records what it actually resolved, and evidence that resolves and contradicts fails the appraisal. Applied here to appraisal.policy_ref, which carries no digest.", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + }, + "policy": { + "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", + "enforcement_mode": "enforce" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "resolution": { + "outcome": "resolved", + "file": "policies/appraisal-policy-v1.json", + "actual_sha256": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + }, + "candidate_binding": { + "field": "CANDIDATE:appraisal.policy_digest", + "value": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + } + }, + "expected": { + "outcome": "pass", + "reason": "Declared binding equals the digest of what resolved." + } +} diff --git a/tests/vectors/appraisal-resolution/03-digest-mismatch-minimal-mutation.json b/tests/vectors/appraisal-resolution/03-digest-mismatch-minimal-mutation.json new file mode 100644 index 0000000..ef417e2 --- /dev/null +++ b/tests/vectors/appraisal-resolution/03-digest-mismatch-minimal-mutation.json @@ -0,0 +1,62 @@ +{ + "name": "digest-mismatch-minimal-mutation", + "description": "The cited URI now serves a document one character from the one appraised: the SLSA floor moved from 2 to 3. The record still declares the original digest. A verifier that compares only the URI sees no change; a verifier that compares bytes sees the substitution that flips this record's verdict.", + "boundary": "contradicted", + "defect": "cited object mutated minimally after appraisal", + "spec": "agentrust-io/trace-spec docs/verification.md - a verifier records what it actually resolved, and evidence that resolves and contradicts fails the appraisal. Applied here to appraisal.policy_ref, which carries no digest.", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + }, + "policy": { + "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", + "enforcement_mode": "enforce" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "resolution": { + "outcome": "resolved", + "file": "policies/appraisal-policy-v1-rev2.json", + "actual_sha256": "sha256:1248932d38677845463158e5b93b7d1e3f78b1fffdef6d23e25c1d2926d5e3d5" + }, + "candidate_binding": { + "field": "CANDIDATE:appraisal.policy_digest", + "value": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + }, + "note": "policies/appraisal-policy-v1.json and policies/appraisal-policy-v1-rev2.json differ in one character." + }, + "expected": { + "outcome": "reject", + "reason": "Resolved bytes contradict the declared binding." + } +} diff --git a/tests/vectors/appraisal-resolution/04-digest-mismatch-different-object.json b/tests/vectors/appraisal-resolution/04-digest-mismatch-different-object.json new file mode 100644 index 0000000..a86fe63 --- /dev/null +++ b/tests/vectors/appraisal-resolution/04-digest-mismatch-different-object.json @@ -0,0 +1,61 @@ +{ + "name": "digest-mismatch-different-object", + "description": "The cited URI resolves to an unrelated policy document. Paired with 03 so the boundary is not carried by a single vector: a verifier that only samples a prefix of the document, or compares lengths, passes one of these two and fails the other.", + "boundary": "contradicted", + "defect": "cited object wholly replaced after appraisal", + "spec": "agentrust-io/trace-spec docs/verification.md - a verifier records what it actually resolved, and evidence that resolves and contradicts fails the appraisal. Applied here to appraisal.policy_ref, which carries no digest.", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + }, + "policy": { + "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", + "enforcement_mode": "enforce" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "resolution": { + "outcome": "resolved", + "file": "policies/unrelated-policy.json", + "actual_sha256": "sha256:23905737177574be26e1212264b88ffe81e06e99aab1fe6115ebbb2e9205109c" + }, + "candidate_binding": { + "field": "CANDIDATE:appraisal.policy_digest", + "value": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + } + }, + "expected": { + "outcome": "reject", + "reason": "Resolved bytes contradict the declared binding." + } +} diff --git a/tests/vectors/appraisal-resolution/05-referent-unreachable.json b/tests/vectors/appraisal-resolution/05-referent-unreachable.json new file mode 100644 index 0000000..cbf5e9f --- /dev/null +++ b/tests/vectors/appraisal-resolution/05-referent-unreachable.json @@ -0,0 +1,63 @@ +{ + "name": "referent-unreachable", + "description": "The cited URI does not resolve; no object under policies/ corresponds to it. Nothing was contradicted, because nothing was read. This vector asserts only that the outcome is not affirming.", + "boundary": "unresolvable", + "defect": "referent unreachable at verification time", + "spec": "agentrust-io/trace-spec docs/verification.md - evidence that does not resolve downgrades honestly rather than failing. Which value records that for appraisal.policy_ref is open: agentrust-io/trace-spec#190. This vector asserts only what the outcome may NOT be.", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + }, + "policy": { + "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", + "enforcement_mode": "enforce" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-withdrawn.json", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-withdrawn.json", + "resolution": { + "outcome": "unreachable", + "file": null, + "note": "No sibling file corresponds to this URI, by construction." + }, + "candidate_binding": { + "field": "CANDIDATE:appraisal.policy_digest", + "value": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + } + }, + "expected": { + "outcome": "deferred", + "deferred_pending": "agentrust-io/trace-spec#190", + "must_not": "affirming", + "reason": "Reporting a check that was never performed as affirming is the failure this set exists to name. Which value is reported instead is not decided here." + } +} diff --git a/tests/vectors/appraisal-resolution/06-digest-algorithm-uncomputable.json b/tests/vectors/appraisal-resolution/06-digest-algorithm-uncomputable.json new file mode 100644 index 0000000..27a541f --- /dev/null +++ b/tests/vectors/appraisal-resolution/06-digest-algorithm-uncomputable.json @@ -0,0 +1,64 @@ +{ + "name": "digest-algorithm-uncomputable", + "description": "The referent resolves, but the declared binding names a digest algorithm outside the set this profile's schema admits (sha256 and sha384). The verifier cannot compute the comparison, so it has not checked - which is distinct from having checked and disagreed. Paired with 05: one cannot reach the object, the other reaches it and cannot compute over it.", + "boundary": "unresolvable", + "defect": "digest algorithm the verifier cannot compute", + "spec": "agentrust-io/trace-spec docs/verification.md - evidence that does not resolve downgrades honestly rather than failing. Which value records that for appraisal.policy_ref is open: agentrust-io/trace-spec#190. This vector asserts only what the outcome may NOT be. Deliberately rhymes with the unverifiable / digest_algorithm_unsupported classification PROPOSED in agentrust-io/trace-spec#184 - an explicitly non-normative draft, a surface facing this question rather than an established answer to it - without adopting its vocabulary.", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + }, + "policy": { + "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", + "enforcement_mode": "enforce" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v2.json", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v2.json", + "resolution": { + "outcome": "resolved", + "file": "policies/appraisal-policy-v2.json", + "actual_sha256": "sha256:9dd48a5dc2efc67c3873a38c95aae17bc49345b48d09e6629f177ac0a5ff79cc" + }, + "candidate_binding": { + "field": "CANDIDATE:appraisal.policy_digest", + "value": "sha3-512:49c36096f3c7cf58d58324f53502f427a72b0684776fe6006f8a8b13513a9c40d626893b68102cec8c8c78d871814067bab429a7d09e5ff4ef14147eb35ef9a5", + "note": "This is the correct SHA3-512 of the referent. It is outside the schema's digest pattern ^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$, so no conformant verifier is required to compute it - which is the only reason this vector is undecided." + } + }, + "expected": { + "outcome": "deferred", + "deferred_pending": "agentrust-io/trace-spec#190", + "must_not": "affirming", + "reason": "The comparison was not performed, so the result is not a contradiction. Which value records that is not decided here." + } +} diff --git a/tests/vectors/appraisal-resolution/07-digest-bound-to-other-referent.json b/tests/vectors/appraisal-resolution/07-digest-bound-to-other-referent.json new file mode 100644 index 0000000..5024ecc --- /dev/null +++ b/tests/vectors/appraisal-resolution/07-digest-bound-to-other-referent.json @@ -0,0 +1,62 @@ +{ + "name": "digest-bound-to-other-referent", + "description": "The declared binding is well formed and is the true digest of a real object in this set - version 2 - while policy_ref cites version 1. Both halves are individually valid and the pair is not. A verifier that checks the digest is well formed, or that it matches something it holds, passes this; only one that checks the digest against what this URI resolved to rejects it.", + "boundary": "contradicted", + "defect": "binding well formed but bound to a different referent", + "spec": "agentrust-io/trace-spec docs/verification.md - a verifier records what it actually resolved, and evidence that resolves and contradicts fails the appraisal. Applied here to appraisal.policy_ref, which carries no digest.", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + }, + "policy": { + "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", + "enforcement_mode": "enforce" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "resolution": { + "outcome": "resolved", + "file": "policies/appraisal-policy-v1.json", + "actual_sha256": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + }, + "candidate_binding": { + "field": "CANDIDATE:appraisal.policy_digest", + "value": "sha256:9dd48a5dc2efc67c3873a38c95aae17bc49345b48d09e6629f177ac0a5ff79cc", + "note": "This is the digest of policies/appraisal-policy-v2.json, which is not what cited_uri resolves to." + } + }, + "expected": { + "outcome": "reject", + "reason": "The binding does not describe the object the record cites, even though it describes some object." + } +} diff --git a/tests/vectors/appraisal-resolution/README.md b/tests/vectors/appraisal-resolution/README.md new file mode 100644 index 0000000..c353ee9 --- /dev/null +++ b/tests/vectors/appraisal-resolution/README.md @@ -0,0 +1,173 @@ +# Appraisal-resolution vectors + +Candidate conformance vectors for one question: **when a record cites an +appraisal policy by URI, what can a second verifier confirm about what that URI +held?** + +Today, nothing. `appraisal.policy_ref` is a bare URI. The record carries a +digest for the enforcement policy bundle (`policy.bundle_hash`) and none for the +appraisal policy that produced the verdict, so two verifiers resolving the same +`policy_ref` at different times can retrieve different documents and both report +`affirming` honestly. + +## Why this set exists + +`agentrust-io/trace-tests#63` proposed a conformance module for the `appraisal` +claim, noting that `policy_ref` can be checked for well-formedness only, "since +the record carries no digest to compare a resolved policy against." Building +those cases ran into the boundary itself, which was stated on +`agentrust-io/trace-spec#66` on 2026-08-18: + +> `appraisal.policy_ref` can be checked for well-formedness, but not +> reproduced: the record carries nothing stating what the referent was, so a +> conformance module can confirm the URI parses and nothing more. Two verifiers +> resolving the same `policy_ref` at different times can retrieve different +> documents and both honestly report verified — the federation gap §1 names, +> one hop from the record. + +Drafting candidate fixtures for it was accepted on that thread the same day, +with the shape asked for: candidates in `trace-tests`, one defect per vector, +expected outcomes committed beside the vectors, and the schema and +`verification.md` text left maintainer-authored. This set is that. + +## What it does not do + +**It does not propose a schema change.** `candidate_binding` is a *candidate +shape*, marked `CANDIDATE:` in every vector, and it lives in the vector's +`context` — never inside `record`. The packaged schema sets +`additionalProperties: false` on `appraisal`, so a record carrying an extra +field would be schema-invalid; whether a binding field lands, and what it is +called, is an editorial decision. + +**It does not name an outcome value for the unresolvable case.** Vectors 05 and +06 are marked `deferred`, pointing at `agentrust-io/trace-spec#190`, which +tracks the open cross-surface question: *when a record cites something a +verifier cannot resolve, what does the verifier record, and where.* + +Four surfaces face that question. Only one has answered it: +`build_provenance`, with a recorded field plus a floor +(`agentrust-io/trace-spec#173`, **merged**). Revocation states the prohibition +in prose with no field to record the outcome in +(`agentrust-io/trace-spec#187`, **merged**). Delegation links **propose** +`unverifiable` on a separate axis (`agentrust-io/trace-spec#184`, **an +explicitly non-normative draft** — a proposal facing the question, not an +established answer; its author has said so on +`agentrust-io/trace-spec#190`). And this one. + +The divergence is a schema fact rather than four differing opinions: +`appraisal` is `additionalProperties: false`, and `#173` landed the only +claimed/verified pair that exists, so the other surfaces had no field to use. +Coining a fifth answer here is exactly what `agentrust-io/trace-spec#190` +exists to prevent — and a new field or a fifth `appraisal.status` value is +normative under `CONTRIBUTING.md`, which routes it through a Spec change +proposal with a sponsoring organization. It is not something a vector set +decides. + +> **`deferred` is fixture bookkeeping, not a proposed appraisal vocabulary +> value.** It is a property of a *vector's expected block*, recording that the +> outcome is not yet decided upstream. It is **not** a candidate value for +> `appraisal.status`, which is closed at `affirming`, `warning`, +> `contraindicated`, `none`. A deferred vector asserts only what the outcome may +> **not** be — `must_not: "affirming"` — because reporting a check that was +> never performed as affirming is the one thing every reading of the merged text +> already rules out. `tests/test_appraisal_resolution_completeness.py` fails if +> any vector reuses a status value as an outcome. + +## The vectors + +Every record is identical except for `appraisal.policy_ref`, so the defect under +test is the only thing that varies. All records are unsigned and ASCII-only. + +| # | Vector | Boundary | Expected | +|---|---|---|---| +| 01 | `no-binding-declared` | accept | `pass` | +| 02 | `resolved-and-matches` | accept | `pass` | +| 03 | `digest-mismatch-minimal-mutation` | contradicted | `reject` | +| 04 | `digest-mismatch-different-object` | contradicted | `reject` | +| 05 | `referent-unreachable` | unresolvable | `deferred` | +| 06 | `digest-algorithm-uncomputable` | unresolvable | `deferred` | +| 07 | `digest-bound-to-other-referent` | contradicted | `reject` | + +**01 and 02 are the must-accept pair, and they are why this set is not +one-directional.** `agentrust-io/trace-spec#186` (merged 2026-08-20) states +the criterion: a set must fail *both* unconditional implementations. A set +written from the motivating problem alone would be all rejections and +deferrals, and a verifier that rejects everything would pass it. 01 is the +backward-compatibility control — every conformant record today declares no +binding, and must keep verifying, or this set would be proposing a breaking +change rather than describing a gap. 02 is the only vector in which a +resolution succeeds. + +**03 and 04 keep the contradicted boundary off a single vector.** 03 differs +from the appraised object in exactly one byte — the SLSA floor moves 2 → 3, +which flips this record's verdict. 04 substitutes an unrelated document of a +different length. A verifier comparing lengths, or sampling a prefix, passes one +and fails the other. + +**05 and 06 are both unresolvable, by different mechanisms.** 05 cannot reach +the object; 06 reaches it and cannot compute over it, because the declared +algorithm (`sha3-512`) is outside the set the schema admits +(`^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$`). The distinction deliberately +rhymes with the `digest_algorithm_unsupported` classification **proposed** in +`agentrust-io/trace-spec#184` — a non-normative draft — without adopting that +vocabulary. + +**07 is the one a well-formedness check passes.** The binding is a valid +`sha256:` digest and is the true digest of a real object in this set — version 2 +— while `policy_ref` cites version 1. Both halves are individually valid; the +pair is not. + +## Reproducing it + +``` +python tests/vectors/appraisal-resolution/gen_appraisal_resolution.py +``` + +Deterministic: no keys, no clock, no randomness, no network. The digests are +SHA-256 over the exact bytes of the sibling files under `policies/`, so anyone +holding only this directory can recompute every number in the set. + +`tests/test_appraisal_resolution_reproduces.py` holds the generator to +byte-reproduction by regenerating into a temporary directory and comparing — +not in place, which would compare the files to themselves and agree regardless. + +The guard is **self-contained**. `agentrust-io/trace-spec#171` provides the +equivalent for that repository's `examples/`, and `trace-tests` has no such +registry; `agentrust-io/trace-tests#66` gives the reason not to reach across for +one — *"a guard that needs another repository checked out is a guard that gets +skipped."* + +`.gitattributes` in this directory pins `eol=lf`. This is load-bearing rather +than tidy: with `core.autocrlf=true`, a checkout rewrites LF to CRLF, every +policy digest stops matching, and the set fails on a clean clone. + +## What this set does not establish + +- **No verifier was exercised.** Nothing in `trace-tests` resolves + `appraisal.policy_ref` today — that is the gap — so every expected outcome is + a claim about what a verifier should do, not a recording of what one did. The + tests grade the set's internal consistency: that each declared digest really + is, or really is not, the digest of the bytes the vector says resolution + returned. +- **The unresolvable outcome is unnamed**, by choice, pending + `agentrust-io/trace-spec#190`. + +Both are recorded as exact shortfalls in +`tests/test_appraisal_resolution_completeness.py::KNOWN_SHORTFALLS`, which fails +if the list changes without this file changing with it. + +## Related + +- `agentrust-io/trace-tests#63` — the module proposal these cases were built for +- `agentrust-io/trace-spec#66` — where the gap was raised and the fixtures accepted +- `agentrust-io/trace-spec#190` — the deferred cross-surface question +- `agentrust-io/trace-spec#173` — merged: recorded field plus a policy floor +- `agentrust-io/trace-spec#184` — open, **draft, explicitly non-normative**: + *proposes* `unverifiable` on the delegation surface +- `agentrust-io/trace-spec#186` — merged: the adequacy criteria this set was + built to. It grades trace-spec's `examples/`; this repository has no adequacy + harness, so the standard is one this set chose, not one imposed on it +- `agentrust-io/trace-tests#66` — merged: `tr_sig` canonicalizes with RFC 8785; + source of the self-containment principle quoted above +- `agentrust-io/trace-tests#68` — merged: packaged schema resynced to the + normative v0.2 copy these records validate against diff --git a/tests/vectors/appraisal-resolution/gen_appraisal_resolution.py b/tests/vectors/appraisal-resolution/gen_appraisal_resolution.py new file mode 100644 index 0000000..1a28246 --- /dev/null +++ b/tests/vectors/appraisal-resolution/gen_appraisal_resolution.py @@ -0,0 +1,452 @@ +"""Regenerate the appraisal-resolution vector set, byte for byte. + +Deterministic by construction: no keys, no clock, no randomness, no network. +Running this on any machine with the same CPython minor version reproduces +every file in this directory exactly, which is what +``tests/test_appraisal_resolution_reproduces.py`` asserts. + + python tests/vectors/appraisal-resolution/gen_appraisal_resolution.py + +The digests in the vectors are SHA-256 over the exact bytes of the sibling +files under ``policies/``. Anyone holding only this directory can recompute +them; nothing here depends on another repository being checked out. + +WHAT THIS SET IS FOR + ``appraisal.policy_ref`` is a bare URI. A record says which appraisal + policy produced its verdict, but carries nothing stating what that URI + resolved to at appraisal time, so a second verifier cannot confirm it + retrieved the same document. See the set's README.md. + +WHAT IT DELIBERATELY DOES NOT DO + It does not name an outcome value for the unresolvable case. That is the + open cross-surface question tracked by agentrust-io/trace-spec#190, and + coining a value here is precisely what that issue exists to prevent. + Vectors 05 and 06 carry the fixture-bookkeeping state ``deferred``. + + ``candidate_binding`` is a CANDIDATE shape, not a proposed schema change. + It lives in the vector's ``context``, never inside ``record``: the + packaged schema sets ``additionalProperties: false`` on ``appraisal``, so + a record carrying an extra field would be schema-invalid, and proposing + one is a schema decision that belongs to the editorial process. +""" + +from __future__ import annotations + +import hashlib +import json +from pathlib import Path + +HERE = Path(__file__).parent +POLICIES = HERE / "policies" + +# --- house serialization, fixed so bytes are stable across platforms ------- +INDENT = 2 + + +def write_json(path: Path, obj: object) -> bytes: + """Write *obj* as UTF-8 JSON with LF endings; return the exact bytes.""" + text = json.dumps(obj, indent=INDENT, ensure_ascii=True) + "\n" + data = text.encode("utf-8") + path.write_bytes(data) + return data + + +def sha256_of(data: bytes) -> str: + return "sha256:" + hashlib.sha256(data).hexdigest() + + +def sha3_512_of(data: bytes) -> str: + """A *correct* digest in an algorithm the profile's schema does not admit. + + Vector 06 turns on the verifier being unable to compute the comparison, not + on the digest being wrong. A placeholder value would confound the two: a + reader could not tell whether the vector is deferred because the algorithm + is unsupported or because the digest is obviously bogus. This is the real + SHA3-512 of the referent, so the algorithm is the only variable. + """ + return "sha3-512:" + hashlib.sha3_512(data).hexdigest() + + +# --- the cited objects ----------------------------------------------------- +# Small, ASCII-only, and shaped like an appraisal policy rather than a +# placeholder, so a reader can see why swapping one for another matters. + +POLICY_V1 = { + "policy_id": "appraisal/baseline", + "version": "1.0.0", + "rules": [ + {"claim": "runtime.platform", + "must_be_one_of": ["intel-tdx", "amd-sev-snp", "tpm2"]}, + {"claim": "build_provenance.slsa_level", "minimum": 2}, + ], +} + +# One character apart from V1: the SLSA floor moves 2 -> 3. A verifier +# applying this instead of V1 reaches a different verdict on the same record, +# which is why a minimal mutation is the honest test rather than a cosmetic one. +POLICY_V1_REV2 = { + "policy_id": "appraisal/baseline", + "version": "1.0.0", + "rules": [ + {"claim": "runtime.platform", + "must_be_one_of": ["intel-tdx", "amd-sev-snp", "tpm2"]}, + {"claim": "build_provenance.slsa_level", "minimum": 3}, + ], +} + +# A legitimate later version, published at its own URI. +POLICY_V2 = { + "policy_id": "appraisal/baseline", + "version": "2.0.0", + "rules": [ + {"claim": "runtime.platform", + "must_be_one_of": ["intel-tdx", "amd-sev-snp"]}, + {"claim": "build_provenance.slsa_level", "minimum": 3}, + {"claim": "transparency", "must_be_present": True}, + ], +} + +# A different policy entirely, not a version of the baseline. +POLICY_UNRELATED = { + "policy_id": "retention/pii-90d", + "version": "1.4.2", + "rules": [ + {"claim": "data_class", "must_be_one_of": ["public", "internal"]}, + ], +} + +POLICY_FILES = { + "appraisal-policy-v1.json": POLICY_V1, + "appraisal-policy-v1-rev2.json": POLICY_V1_REV2, + "appraisal-policy-v2.json": POLICY_V2, + "unrelated-policy.json": POLICY_UNRELATED, +} + +BASE_URI = "https://policy.example.org/appraisal/" + +# --- the record ------------------------------------------------------------ +# Every vector's record is identical except for appraisal.policy_ref, so the +# defect under test is the only thing that varies. Modelled on the repository's +# tests/vectors/valid_level0.json: unsigned, ASCII-only, fixed iat. + +RECORD_IAT = 1748000000 +APPRAISAL_TIMESTAMP = 1748000042 + + +def record_with(policy_ref: str | None) -> dict[str, object]: + appraisal: dict[str, object] = { + "status": "affirming", + "verifier": "https://verifier.example.org", + } + if policy_ref is not None: + appraisal["policy_ref"] = policy_ref + appraisal["timestamp"] = APPRAISAL_TIMESTAMP + return { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": RECORD_IAT, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": {"provider": "anthropic", "model_id": "claude-sonnet-4-5"}, + "runtime": { + "platform": "intel-tdx", + "measurement": + "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + }, + "policy": { + "bundle_hash": + "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", + "enforcement_mode": "enforce", + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8", + }, + "appraisal": appraisal, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001", + } + }, + } + + +CANDIDATE_FIELD = "CANDIDATE:appraisal.policy_digest" +DEFERRED_PENDING = "agentrust-io/trace-spec#190" + +SPEC_DECIDED = ( + "agentrust-io/trace-spec docs/verification.md - a verifier records what it " + "actually resolved, and evidence that resolves and contradicts fails the " + "appraisal. Applied here to appraisal.policy_ref, which carries no digest." +) +SPEC_DEFERRED = ( + "agentrust-io/trace-spec docs/verification.md - evidence that does not " + "resolve downgrades honestly rather than failing. Which value records that " + "for appraisal.policy_ref is open: agentrust-io/trace-spec#190. This vector " + "asserts only what the outcome may NOT be." +) + + +def main(out_dir: Path | None = None) -> int: + """Write the set into *out_dir* (default: this directory). + + The parameter exists so the byte-reproduction guard can regenerate into a + temporary directory and compare, rather than overwriting the committed + files and comparing them to themselves — which would agree no matter what. + """ + here = Path(out_dir) if out_dir is not None else HERE + policies = here / "policies" + here.mkdir(parents=True, exist_ok=True) + policies.mkdir(exist_ok=True) + + # 1. Write the cited objects and digest their exact bytes. + digests: dict[str, str] = {} + raw: dict[str, bytes] = {} + for name, obj in POLICY_FILES.items(): + raw[name] = write_json(policies / name, obj) + digests[name] = sha256_of(raw[name]) + + d_v1 = digests["appraisal-policy-v1.json"] + d_rev2 = digests["appraisal-policy-v1-rev2.json"] + d_v2 = digests["appraisal-policy-v2.json"] + d_unrel = digests["unrelated-policy.json"] + + uri_v1 = BASE_URI + "appraisal-policy-v1.json" + uri_v2 = BASE_URI + "appraisal-policy-v2.json" + uri_withdrawn = BASE_URI + "appraisal-policy-withdrawn.json" + + def resolved(file_name: str, digest: str) -> dict[str, object]: + return { + "outcome": "resolved", + "file": f"policies/{file_name}", + "actual_sha256": digest, + } + + vectors: list[tuple[str, dict[str, object]]] = [ + ("01-no-binding-declared.json", { + "name": "no-binding-declared", + "description": ( + "The record cites an appraisal policy by bare URI and declares no " + "binding to what that URI held. This is every conformant record " + "today, so it must keep verifying: a set that rejected it would be " + "proposing a breaking change rather than describing a gap." + ), + "boundary": "accept", + "defect": "none - backward-compatibility control", + "spec": ( + "Mirrors the merged treatment of an undeclared depth in " + "agentrust-io/trace-spec#173, where a record that never declared is " + "read as surface rather than as a failure." + ), + "record": record_with(uri_v1), + "context": { + "cited_uri": uri_v1, + "resolution": { + "outcome": "not_attempted", + "note": "No binding is declared, so there is nothing to check.", + }, + "candidate_binding": None, + }, + "expected": { + "outcome": "pass", + "reason": "No binding declared; nothing to contradict.", + }, + }), + ("02-resolved-and-matches.json", { + "name": "resolved-and-matches", + "description": ( + "The cited URI resolves and its bytes hash to exactly the declared " + "candidate binding. The second must-accept vector, and the only one " + "in which a resolution actually succeeds." + ), + "boundary": "accept", + "defect": "none - positive control", + "spec": SPEC_DECIDED, + "record": record_with(uri_v1), + "context": { + "cited_uri": uri_v1, + "resolution": resolved("appraisal-policy-v1.json", d_v1), + "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_v1}, + }, + "expected": { + "outcome": "pass", + "reason": "Declared binding equals the digest of what resolved.", + }, + }), + ("03-digest-mismatch-minimal-mutation.json", { + "name": "digest-mismatch-minimal-mutation", + "description": ( + "The cited URI now serves a document one character from the one " + "appraised: the SLSA floor moved from 2 to 3. The record still " + "declares the original digest. A verifier that compares only the " + "URI sees no change; a verifier that compares bytes sees the " + "substitution that flips this record's verdict." + ), + "boundary": "contradicted", + "defect": "cited object mutated minimally after appraisal", + "spec": SPEC_DECIDED, + "record": record_with(uri_v1), + "context": { + "cited_uri": uri_v1, + "resolution": resolved("appraisal-policy-v1-rev2.json", d_rev2), + "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_v1}, + "note": ( + "policies/appraisal-policy-v1.json and " + "policies/appraisal-policy-v1-rev2.json differ in one character." + ), + }, + "expected": { + "outcome": "reject", + "reason": "Resolved bytes contradict the declared binding.", + }, + }), + ("04-digest-mismatch-different-object.json", { + "name": "digest-mismatch-different-object", + "description": ( + "The cited URI resolves to an unrelated policy document. Paired with " + "03 so the boundary is not carried by a single vector: a verifier " + "that only samples a prefix of the document, or compares lengths, " + "passes one of these two and fails the other." + ), + "boundary": "contradicted", + "defect": "cited object wholly replaced after appraisal", + "spec": SPEC_DECIDED, + "record": record_with(uri_v1), + "context": { + "cited_uri": uri_v1, + "resolution": resolved("unrelated-policy.json", d_unrel), + "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_v1}, + }, + "expected": { + "outcome": "reject", + "reason": "Resolved bytes contradict the declared binding.", + }, + }), + ("05-referent-unreachable.json", { + "name": "referent-unreachable", + "description": ( + "The cited URI does not resolve; no object under policies/ " + "corresponds to it. Nothing was contradicted, because nothing was " + "read. This vector asserts only that the outcome is not affirming." + ), + "boundary": "unresolvable", + "defect": "referent unreachable at verification time", + "spec": SPEC_DEFERRED, + "record": record_with(uri_withdrawn), + "context": { + "cited_uri": uri_withdrawn, + "resolution": { + "outcome": "unreachable", + "file": None, + "note": "No sibling file corresponds to this URI, by construction.", + }, + "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_v1}, + }, + "expected": { + "outcome": "deferred", + "deferred_pending": DEFERRED_PENDING, + "must_not": "affirming", + "reason": ( + "Reporting a check that was never performed as affirming is the " + "failure this set exists to name. Which value is reported " + "instead is not decided here." + ), + }, + }), + ("06-digest-algorithm-uncomputable.json", { + "name": "digest-algorithm-uncomputable", + "description": ( + "The referent resolves, but the declared binding names a digest " + "algorithm outside the set this profile's schema admits " + "(sha256 and sha384). The verifier cannot compute the comparison, " + "so it has not checked - which is distinct from having checked and " + "disagreed. Paired with 05: one cannot reach the object, the other " + "reaches it and cannot compute over it." + ), + "boundary": "unresolvable", + "defect": "digest algorithm the verifier cannot compute", + "spec": ( + SPEC_DEFERRED + + " Deliberately rhymes with the unverifiable / " + "digest_algorithm_unsupported classification PROPOSED in " + "agentrust-io/trace-spec#184 - an explicitly non-normative " + "draft, a surface facing this question rather than an " + "established answer to it - without adopting its vocabulary." + ), + "record": record_with(uri_v2), + "context": { + "cited_uri": uri_v2, + "resolution": resolved("appraisal-policy-v2.json", d_v2), + "candidate_binding": { + "field": CANDIDATE_FIELD, + "value": sha3_512_of(raw["appraisal-policy-v2.json"]), + "note": ( + "This is the correct SHA3-512 of the referent. It is " + "outside the schema's digest pattern " + "^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$, so no conformant " + "verifier is required to compute it - which is the only " + "reason this vector is undecided." + ), + }, + }, + "expected": { + "outcome": "deferred", + "deferred_pending": DEFERRED_PENDING, + "must_not": "affirming", + "reason": ( + "The comparison was not performed, so the result is not a " + "contradiction. Which value records that is not decided here." + ), + }, + }), + ("07-digest-bound-to-other-referent.json", { + "name": "digest-bound-to-other-referent", + "description": ( + "The declared binding is well formed and is the true digest of a " + "real object in this set - version 2 - while policy_ref cites " + "version 1. Both halves are individually valid and the pair is not. " + "A verifier that checks the digest is well formed, or that it " + "matches something it holds, passes this; only one that checks the " + "digest against what this URI resolved to rejects it." + ), + "boundary": "contradicted", + "defect": "binding well formed but bound to a different referent", + "spec": SPEC_DECIDED, + "record": record_with(uri_v1), + "context": { + "cited_uri": uri_v1, + "resolution": resolved("appraisal-policy-v1.json", d_v1), + "candidate_binding": { + "field": CANDIDATE_FIELD, + "value": d_v2, + "note": ( + "This is the digest of policies/appraisal-policy-v2.json, " + "which is not what cited_uri resolves to." + ), + }, + }, + "expected": { + "outcome": "reject", + "reason": ( + "The binding does not describe the object the record cites, " + "even though it describes some object." + ), + }, + }), + ] + + for filename, vector in vectors: + write_json(here / filename, vector) + + print(f"wrote {len(POLICY_FILES)} policy objects and {len(vectors)} vectors") + for name, digest in digests.items(): + print(f" policies/{name} {digest}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tests/vectors/appraisal-resolution/policies/appraisal-policy-v1-rev2.json b/tests/vectors/appraisal-resolution/policies/appraisal-policy-v1-rev2.json new file mode 100644 index 0000000..8259434 --- /dev/null +++ b/tests/vectors/appraisal-resolution/policies/appraisal-policy-v1-rev2.json @@ -0,0 +1,18 @@ +{ + "policy_id": "appraisal/baseline", + "version": "1.0.0", + "rules": [ + { + "claim": "runtime.platform", + "must_be_one_of": [ + "intel-tdx", + "amd-sev-snp", + "tpm2" + ] + }, + { + "claim": "build_provenance.slsa_level", + "minimum": 3 + } + ] +} diff --git a/tests/vectors/appraisal-resolution/policies/appraisal-policy-v1.json b/tests/vectors/appraisal-resolution/policies/appraisal-policy-v1.json new file mode 100644 index 0000000..f34a5b8 --- /dev/null +++ b/tests/vectors/appraisal-resolution/policies/appraisal-policy-v1.json @@ -0,0 +1,18 @@ +{ + "policy_id": "appraisal/baseline", + "version": "1.0.0", + "rules": [ + { + "claim": "runtime.platform", + "must_be_one_of": [ + "intel-tdx", + "amd-sev-snp", + "tpm2" + ] + }, + { + "claim": "build_provenance.slsa_level", + "minimum": 2 + } + ] +} diff --git a/tests/vectors/appraisal-resolution/policies/appraisal-policy-v2.json b/tests/vectors/appraisal-resolution/policies/appraisal-policy-v2.json new file mode 100644 index 0000000..164187a --- /dev/null +++ b/tests/vectors/appraisal-resolution/policies/appraisal-policy-v2.json @@ -0,0 +1,21 @@ +{ + "policy_id": "appraisal/baseline", + "version": "2.0.0", + "rules": [ + { + "claim": "runtime.platform", + "must_be_one_of": [ + "intel-tdx", + "amd-sev-snp" + ] + }, + { + "claim": "build_provenance.slsa_level", + "minimum": 3 + }, + { + "claim": "transparency", + "must_be_present": true + } + ] +} diff --git a/tests/vectors/appraisal-resolution/policies/unrelated-policy.json b/tests/vectors/appraisal-resolution/policies/unrelated-policy.json new file mode 100644 index 0000000..3b98d46 --- /dev/null +++ b/tests/vectors/appraisal-resolution/policies/unrelated-policy.json @@ -0,0 +1,13 @@ +{ + "policy_id": "retention/pii-90d", + "version": "1.4.2", + "rules": [ + { + "claim": "data_class", + "must_be_one_of": [ + "public", + "internal" + ] + } + ] +} From dbe56a688983ae5bce57219e48130f9dff1a6f85 Mon Sep 17 00:00:00 2001 From: opento-suggestions Date: Mon, 24 Aug 2026 14:32:44 -0600 Subject: [PATCH 2/7] refactor(vectors): rename appraisal-resolution to policy-resolution Directory, generator, tests, and the four policy files renamed to match the retarget. Bytes unchanged (sha256 verified). Policy files are named by the defect each carries: base, onebyte, other, unrelated. Signed-off-by: opento-suggestions --- ...esolution.py => test_policy_resolution.py} | 6 +- ...=> test_policy_resolution_completeness.py} | 4 +- ...y => test_policy_resolution_reproduces.py} | 12 +-- .../.gitattributes | 0 .../01-no-binding-declared.json | 4 +- .../02-resolved-and-matches.json | 6 +- .../03-digest-mismatch-minimal-mutation.json | 8 +- .../04-digest-mismatch-different-object.json | 6 +- .../05-referent-unreachable.json | 4 +- .../06-digest-algorithm-uncomputable.json | 6 +- .../07-digest-bound-to-other-referent.json | 8 +- .../README.md | 8 +- .../gen_policy_resolution.py} | 88 +++++++++---------- .../policies/policy-bundle-base.json} | 0 .../policies/policy-bundle-onebyte.json} | 0 .../policies/policy-bundle-other.json} | 0 .../policies/policy-bundle-unrelated.json} | 0 17 files changed, 80 insertions(+), 80 deletions(-) rename tests/{test_appraisal_resolution.py => test_policy_resolution.py} (97%) rename tests/{test_appraisal_resolution_completeness.py => test_policy_resolution_completeness.py} (98%) rename tests/{test_appraisal_resolution_reproduces.py => test_policy_resolution_reproduces.py} (89%) rename tests/vectors/{appraisal-resolution => policy-resolution}/.gitattributes (100%) rename tests/vectors/{appraisal-resolution => policy-resolution}/01-no-binding-declared.json (92%) rename tests/vectors/{appraisal-resolution => policy-resolution}/02-resolved-and-matches.json (90%) rename tests/vectors/{appraisal-resolution => policy-resolution}/03-digest-mismatch-minimal-mutation.json (87%) rename tests/vectors/{appraisal-resolution => policy-resolution}/04-digest-mismatch-different-object.json (91%) rename tests/vectors/{appraisal-resolution => policy-resolution}/05-referent-unreachable.json (93%) rename tests/vectors/{appraisal-resolution => policy-resolution}/06-digest-algorithm-uncomputable.json (93%) rename tests/vectors/{appraisal-resolution => policy-resolution}/07-digest-bound-to-other-referent.json (88%) rename tests/vectors/{appraisal-resolution => policy-resolution}/README.md (96%) rename tests/vectors/{appraisal-resolution/gen_appraisal_resolution.py => policy-resolution/gen_policy_resolution.py} (88%) rename tests/vectors/{appraisal-resolution/policies/appraisal-policy-v1.json => policy-resolution/policies/policy-bundle-base.json} (100%) rename tests/vectors/{appraisal-resolution/policies/appraisal-policy-v1-rev2.json => policy-resolution/policies/policy-bundle-onebyte.json} (100%) rename tests/vectors/{appraisal-resolution/policies/appraisal-policy-v2.json => policy-resolution/policies/policy-bundle-other.json} (100%) rename tests/vectors/{appraisal-resolution/policies/unrelated-policy.json => policy-resolution/policies/policy-bundle-unrelated.json} (100%) diff --git a/tests/test_appraisal_resolution.py b/tests/test_policy_resolution.py similarity index 97% rename from tests/test_appraisal_resolution.py rename to tests/test_policy_resolution.py index a9aba0e..3b27e9f 100644 --- a/tests/test_appraisal_resolution.py +++ b/tests/test_policy_resolution.py @@ -1,4 +1,4 @@ -"""The appraisal-resolution set proves itself: digests, referents, schema. +"""The policy-resolution set proves itself: digests, referents, schema. No verifier resolves ``appraisal.policy_ref``, so there is no implementation to run these against. What can be checked — and is checked here — is that the set @@ -25,7 +25,7 @@ import jsonschema import pytest -VECTOR_DIR = Path(__file__).parent / "vectors" / "appraisal-resolution" +VECTOR_DIR = Path(__file__).parent / "vectors" / "policy-resolution" SCHEMA_DIGEST_PATTERN = re.compile(r"^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$") @@ -188,7 +188,7 @@ def test_03_and_04_differ_in_kind_not_just_in_bytes() -> None: wholesale = _load(VECTOR_DIR / "04-digest-mismatch-different-object.json") a = (VECTOR_DIR / minimal["context"]["resolution"]["file"]).read_bytes() - b = (VECTOR_DIR / "policies" / "appraisal-policy-v1.json").read_bytes() + b = (VECTOR_DIR / "policies" / "policy-bundle-base.json").read_bytes() assert len(a) == len(b), "03's mutation should not change the object's length" assert sum(x != y for x, y in zip(a, b)) == 1, ( "03 is the minimal-mutation vector; it must differ from the appraised " diff --git a/tests/test_appraisal_resolution_completeness.py b/tests/test_policy_resolution_completeness.py similarity index 98% rename from tests/test_appraisal_resolution_completeness.py rename to tests/test_policy_resolution_completeness.py index c358c57..b4c7fab 100644 --- a/tests/test_appraisal_resolution_completeness.py +++ b/tests/test_policy_resolution_completeness.py @@ -1,4 +1,4 @@ -"""Adequacy of the appraisal-resolution set, per the criteria on trace-spec#186. +"""Adequacy of the policy-resolution set, per the criteria on trace-spec#186. agentrust-io/trace-spec#186 (merged 2026-08-20) states what a conformance vector set is claiming: *a verifier that does not implement these rules will @@ -37,7 +37,7 @@ import pytest -VECTOR_DIR = Path(__file__).parent / "vectors" / "appraisal-resolution" +VECTOR_DIR = Path(__file__).parent / "vectors" / "policy-resolution" DECIDED_OUTCOMES = {"pass", "reject"} ALL_OUTCOMES = DECIDED_OUTCOMES | {"deferred"} diff --git a/tests/test_appraisal_resolution_reproduces.py b/tests/test_policy_resolution_reproduces.py similarity index 89% rename from tests/test_appraisal_resolution_reproduces.py rename to tests/test_policy_resolution_reproduces.py index af62429..272fbf2 100644 --- a/tests/test_appraisal_resolution_reproduces.py +++ b/tests/test_policy_resolution_reproduces.py @@ -1,4 +1,4 @@ -"""The appraisal-resolution generator must reproduce its committed vectors. +"""The policy-resolution generator must reproduce its committed vectors. A committed vector nobody can regenerate is a number that cannot be checked. If the generator and the files drift, the files win by default and the drift is @@ -23,12 +23,12 @@ import pytest -VECTOR_DIR = Path(__file__).parent / "vectors" / "appraisal-resolution" -GENERATOR = VECTOR_DIR / "gen_appraisal_resolution.py" +VECTOR_DIR = Path(__file__).parent / "vectors" / "policy-resolution" +GENERATOR = VECTOR_DIR / "gen_policy_resolution.py" def _load_generator(): - spec = importlib.util.spec_from_file_location("gen_appraisal_resolution", GENERATOR) + spec = importlib.util.spec_from_file_location("gen_policy_resolution", GENERATOR) assert spec and spec.loader, "generator module could not be loaded" module = importlib.util.module_from_spec(spec) spec.loader.exec_module(module) @@ -43,7 +43,7 @@ def _committed_files() -> list[Path]: def test_the_generator_is_present_beside_the_vectors() -> None: """Self-containment: the guard must not need another repository.""" assert GENERATOR.is_file(), ( - "gen_appraisal_resolution.py must sit in the vector directory, so the " + "gen_policy_resolution.py must sit in the vector directory, so the " "set can be regenerated by anyone holding only that directory" ) @@ -70,7 +70,7 @@ def test_the_generator_reproduces_every_committed_file(tmp_path: Path) -> None: assert not missing, f"generator did not produce: {missing}" assert not mismatched, ( f"generator output differs from the committed bytes: {mismatched}. " - "Re-run tests/vectors/appraisal-resolution/gen_appraisal_resolution.py " + "Re-run tests/vectors/policy-resolution/gen_policy_resolution.py " "and commit the result, or fix the generator — but do not edit a vector " "by hand and leave the generator behind." ) diff --git a/tests/vectors/appraisal-resolution/.gitattributes b/tests/vectors/policy-resolution/.gitattributes similarity index 100% rename from tests/vectors/appraisal-resolution/.gitattributes rename to tests/vectors/policy-resolution/.gitattributes diff --git a/tests/vectors/appraisal-resolution/01-no-binding-declared.json b/tests/vectors/policy-resolution/01-no-binding-declared.json similarity index 92% rename from tests/vectors/appraisal-resolution/01-no-binding-declared.json rename to tests/vectors/policy-resolution/01-no-binding-declared.json index 02dc208..f014370 100644 --- a/tests/vectors/appraisal-resolution/01-no-binding-declared.json +++ b/tests/vectors/policy-resolution/01-no-binding-declared.json @@ -28,7 +28,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,7 +43,7 @@ } }, "context": { - "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", "resolution": { "outcome": "not_attempted", "note": "No binding is declared, so there is nothing to check." diff --git a/tests/vectors/appraisal-resolution/02-resolved-and-matches.json b/tests/vectors/policy-resolution/02-resolved-and-matches.json similarity index 90% rename from tests/vectors/appraisal-resolution/02-resolved-and-matches.json rename to tests/vectors/policy-resolution/02-resolved-and-matches.json index eb9b376..07c51fa 100644 --- a/tests/vectors/appraisal-resolution/02-resolved-and-matches.json +++ b/tests/vectors/policy-resolution/02-resolved-and-matches.json @@ -28,7 +28,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,10 +43,10 @@ } }, "context": { - "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", "resolution": { "outcome": "resolved", - "file": "policies/appraisal-policy-v1.json", + "file": "policies/policy-bundle-base.json", "actual_sha256": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" }, "candidate_binding": { diff --git a/tests/vectors/appraisal-resolution/03-digest-mismatch-minimal-mutation.json b/tests/vectors/policy-resolution/03-digest-mismatch-minimal-mutation.json similarity index 87% rename from tests/vectors/appraisal-resolution/03-digest-mismatch-minimal-mutation.json rename to tests/vectors/policy-resolution/03-digest-mismatch-minimal-mutation.json index ef417e2..f833507 100644 --- a/tests/vectors/appraisal-resolution/03-digest-mismatch-minimal-mutation.json +++ b/tests/vectors/policy-resolution/03-digest-mismatch-minimal-mutation.json @@ -28,7 +28,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,17 +43,17 @@ } }, "context": { - "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", "resolution": { "outcome": "resolved", - "file": "policies/appraisal-policy-v1-rev2.json", + "file": "policies/policy-bundle-onebyte.json", "actual_sha256": "sha256:1248932d38677845463158e5b93b7d1e3f78b1fffdef6d23e25c1d2926d5e3d5" }, "candidate_binding": { "field": "CANDIDATE:appraisal.policy_digest", "value": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" }, - "note": "policies/appraisal-policy-v1.json and policies/appraisal-policy-v1-rev2.json differ in one character." + "note": "policies/policy-bundle-base.json and policies/policy-bundle-onebyte.json differ in one character." }, "expected": { "outcome": "reject", diff --git a/tests/vectors/appraisal-resolution/04-digest-mismatch-different-object.json b/tests/vectors/policy-resolution/04-digest-mismatch-different-object.json similarity index 91% rename from tests/vectors/appraisal-resolution/04-digest-mismatch-different-object.json rename to tests/vectors/policy-resolution/04-digest-mismatch-different-object.json index a86fe63..61045c9 100644 --- a/tests/vectors/appraisal-resolution/04-digest-mismatch-different-object.json +++ b/tests/vectors/policy-resolution/04-digest-mismatch-different-object.json @@ -28,7 +28,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,10 +43,10 @@ } }, "context": { - "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", "resolution": { "outcome": "resolved", - "file": "policies/unrelated-policy.json", + "file": "policies/policy-bundle-unrelated.json", "actual_sha256": "sha256:23905737177574be26e1212264b88ffe81e06e99aab1fe6115ebbb2e9205109c" }, "candidate_binding": { diff --git a/tests/vectors/appraisal-resolution/05-referent-unreachable.json b/tests/vectors/policy-resolution/05-referent-unreachable.json similarity index 93% rename from tests/vectors/appraisal-resolution/05-referent-unreachable.json rename to tests/vectors/policy-resolution/05-referent-unreachable.json index cbf5e9f..ee70b8f 100644 --- a/tests/vectors/appraisal-resolution/05-referent-unreachable.json +++ b/tests/vectors/policy-resolution/05-referent-unreachable.json @@ -28,7 +28,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-withdrawn.json", + "policy_ref": "https://policy.example.org/bundles/policy-bundle-withdrawn.json", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,7 +43,7 @@ } }, "context": { - "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-withdrawn.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-withdrawn.json", "resolution": { "outcome": "unreachable", "file": null, diff --git a/tests/vectors/appraisal-resolution/06-digest-algorithm-uncomputable.json b/tests/vectors/policy-resolution/06-digest-algorithm-uncomputable.json similarity index 93% rename from tests/vectors/appraisal-resolution/06-digest-algorithm-uncomputable.json rename to tests/vectors/policy-resolution/06-digest-algorithm-uncomputable.json index 27a541f..0226b9d 100644 --- a/tests/vectors/appraisal-resolution/06-digest-algorithm-uncomputable.json +++ b/tests/vectors/policy-resolution/06-digest-algorithm-uncomputable.json @@ -28,7 +28,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v2.json", + "policy_ref": "https://policy.example.org/bundles/policy-bundle-other.json", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,10 +43,10 @@ } }, "context": { - "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v2.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-other.json", "resolution": { "outcome": "resolved", - "file": "policies/appraisal-policy-v2.json", + "file": "policies/policy-bundle-other.json", "actual_sha256": "sha256:9dd48a5dc2efc67c3873a38c95aae17bc49345b48d09e6629f177ac0a5ff79cc" }, "candidate_binding": { diff --git a/tests/vectors/appraisal-resolution/07-digest-bound-to-other-referent.json b/tests/vectors/policy-resolution/07-digest-bound-to-other-referent.json similarity index 88% rename from tests/vectors/appraisal-resolution/07-digest-bound-to-other-referent.json rename to tests/vectors/policy-resolution/07-digest-bound-to-other-referent.json index 5024ecc..b425b60 100644 --- a/tests/vectors/appraisal-resolution/07-digest-bound-to-other-referent.json +++ b/tests/vectors/policy-resolution/07-digest-bound-to-other-referent.json @@ -28,7 +28,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,16 +43,16 @@ } }, "context": { - "cited_uri": "https://policy.example.org/appraisal/appraisal-policy-v1.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", "resolution": { "outcome": "resolved", - "file": "policies/appraisal-policy-v1.json", + "file": "policies/policy-bundle-base.json", "actual_sha256": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" }, "candidate_binding": { "field": "CANDIDATE:appraisal.policy_digest", "value": "sha256:9dd48a5dc2efc67c3873a38c95aae17bc49345b48d09e6629f177ac0a5ff79cc", - "note": "This is the digest of policies/appraisal-policy-v2.json, which is not what cited_uri resolves to." + "note": "This is the digest of policies/policy-bundle-other.json, which is not what cited_uri resolves to." } }, "expected": { diff --git a/tests/vectors/appraisal-resolution/README.md b/tests/vectors/policy-resolution/README.md similarity index 96% rename from tests/vectors/appraisal-resolution/README.md rename to tests/vectors/policy-resolution/README.md index c353ee9..02217c0 100644 --- a/tests/vectors/appraisal-resolution/README.md +++ b/tests/vectors/policy-resolution/README.md @@ -70,7 +70,7 @@ decides. > `contraindicated`, `none`. A deferred vector asserts only what the outcome may > **not** be — `must_not: "affirming"` — because reporting a check that was > never performed as affirming is the one thing every reading of the merged text -> already rules out. `tests/test_appraisal_resolution_completeness.py` fails if +> already rules out. `tests/test_policy_resolution_completeness.py` fails if > any vector reuses a status value as an outcome. ## The vectors @@ -120,14 +120,14 @@ pair is not. ## Reproducing it ``` -python tests/vectors/appraisal-resolution/gen_appraisal_resolution.py +python tests/vectors/policy-resolution/gen_policy_resolution.py ``` Deterministic: no keys, no clock, no randomness, no network. The digests are SHA-256 over the exact bytes of the sibling files under `policies/`, so anyone holding only this directory can recompute every number in the set. -`tests/test_appraisal_resolution_reproduces.py` holds the generator to +`tests/test_policy_resolution_reproduces.py` holds the generator to byte-reproduction by regenerating into a temporary directory and comparing — not in place, which would compare the files to themselves and agree regardless. @@ -153,7 +153,7 @@ policy digest stops matching, and the set fails on a clean clone. `agentrust-io/trace-spec#190`. Both are recorded as exact shortfalls in -`tests/test_appraisal_resolution_completeness.py::KNOWN_SHORTFALLS`, which fails +`tests/test_policy_resolution_completeness.py::KNOWN_SHORTFALLS`, which fails if the list changes without this file changing with it. ## Related diff --git a/tests/vectors/appraisal-resolution/gen_appraisal_resolution.py b/tests/vectors/policy-resolution/gen_policy_resolution.py similarity index 88% rename from tests/vectors/appraisal-resolution/gen_appraisal_resolution.py rename to tests/vectors/policy-resolution/gen_policy_resolution.py index 1a28246..e100908 100644 --- a/tests/vectors/appraisal-resolution/gen_appraisal_resolution.py +++ b/tests/vectors/policy-resolution/gen_policy_resolution.py @@ -1,11 +1,11 @@ -"""Regenerate the appraisal-resolution vector set, byte for byte. +"""Regenerate the policy-resolution vector set, byte for byte. Deterministic by construction: no keys, no clock, no randomness, no network. Running this on any machine with the same CPython minor version reproduces every file in this directory exactly, which is what -``tests/test_appraisal_resolution_reproduces.py`` asserts. +``tests/test_policy_resolution_reproduces.py`` asserts. - python tests/vectors/appraisal-resolution/gen_appraisal_resolution.py + python tests/vectors/policy-resolution/gen_policy_resolution.py The digests in the vectors are SHA-256 over the exact bytes of the sibling files under ``policies/``. Anyone holding only this directory can recompute @@ -71,7 +71,7 @@ def sha3_512_of(data: bytes) -> str: # Small, ASCII-only, and shaped like an appraisal policy rather than a # placeholder, so a reader can see why swapping one for another matters. -POLICY_V1 = { +POLICY_BASE = { "policy_id": "appraisal/baseline", "version": "1.0.0", "rules": [ @@ -84,7 +84,7 @@ def sha3_512_of(data: bytes) -> str: # One character apart from V1: the SLSA floor moves 2 -> 3. A verifier # applying this instead of V1 reaches a different verdict on the same record, # which is why a minimal mutation is the honest test rather than a cosmetic one. -POLICY_V1_REV2 = { +POLICY_ONEBYTE = { "policy_id": "appraisal/baseline", "version": "1.0.0", "rules": [ @@ -95,7 +95,7 @@ def sha3_512_of(data: bytes) -> str: } # A legitimate later version, published at its own URI. -POLICY_V2 = { +POLICY_OTHER = { "policy_id": "appraisal/baseline", "version": "2.0.0", "rules": [ @@ -116,13 +116,13 @@ def sha3_512_of(data: bytes) -> str: } POLICY_FILES = { - "appraisal-policy-v1.json": POLICY_V1, - "appraisal-policy-v1-rev2.json": POLICY_V1_REV2, - "appraisal-policy-v2.json": POLICY_V2, - "unrelated-policy.json": POLICY_UNRELATED, + "policy-bundle-base.json": POLICY_BASE, + "policy-bundle-onebyte.json": POLICY_ONEBYTE, + "policy-bundle-other.json": POLICY_OTHER, + "policy-bundle-unrelated.json": POLICY_UNRELATED, } -BASE_URI = "https://policy.example.org/appraisal/" +BASE_URI = "https://policy.example.org/bundles/" # --- the record ------------------------------------------------------------ # Every vector's record is identical except for appraisal.policy_ref, so the @@ -210,14 +210,14 @@ def main(out_dir: Path | None = None) -> int: raw[name] = write_json(policies / name, obj) digests[name] = sha256_of(raw[name]) - d_v1 = digests["appraisal-policy-v1.json"] - d_rev2 = digests["appraisal-policy-v1-rev2.json"] - d_v2 = digests["appraisal-policy-v2.json"] - d_unrel = digests["unrelated-policy.json"] + d_base = digests["policy-bundle-base.json"] + d_onebyte = digests["policy-bundle-onebyte.json"] + d_other = digests["policy-bundle-other.json"] + d_unrelated = digests["policy-bundle-unrelated.json"] - uri_v1 = BASE_URI + "appraisal-policy-v1.json" - uri_v2 = BASE_URI + "appraisal-policy-v2.json" - uri_withdrawn = BASE_URI + "appraisal-policy-withdrawn.json" + uri_base = BASE_URI + "policy-bundle-base.json" + uri_other = BASE_URI + "policy-bundle-other.json" + uri_withdrawn = BASE_URI + "policy-bundle-withdrawn.json" def resolved(file_name: str, digest: str) -> dict[str, object]: return { @@ -242,9 +242,9 @@ def resolved(file_name: str, digest: str) -> dict[str, object]: "agentrust-io/trace-spec#173, where a record that never declared is " "read as surface rather than as a failure." ), - "record": record_with(uri_v1), + "record": record_with(uri_base), "context": { - "cited_uri": uri_v1, + "cited_uri": uri_base, "resolution": { "outcome": "not_attempted", "note": "No binding is declared, so there is nothing to check.", @@ -266,11 +266,11 @@ def resolved(file_name: str, digest: str) -> dict[str, object]: "boundary": "accept", "defect": "none - positive control", "spec": SPEC_DECIDED, - "record": record_with(uri_v1), + "record": record_with(uri_base), "context": { - "cited_uri": uri_v1, - "resolution": resolved("appraisal-policy-v1.json", d_v1), - "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_v1}, + "cited_uri": uri_base, + "resolution": resolved("policy-bundle-base.json", d_base), + "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_base}, }, "expected": { "outcome": "pass", @@ -289,14 +289,14 @@ def resolved(file_name: str, digest: str) -> dict[str, object]: "boundary": "contradicted", "defect": "cited object mutated minimally after appraisal", "spec": SPEC_DECIDED, - "record": record_with(uri_v1), + "record": record_with(uri_base), "context": { - "cited_uri": uri_v1, - "resolution": resolved("appraisal-policy-v1-rev2.json", d_rev2), - "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_v1}, + "cited_uri": uri_base, + "resolution": resolved("policy-bundle-onebyte.json", d_onebyte), + "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_base}, "note": ( - "policies/appraisal-policy-v1.json and " - "policies/appraisal-policy-v1-rev2.json differ in one character." + "policies/policy-bundle-base.json and " + "policies/policy-bundle-onebyte.json differ in one character." ), }, "expected": { @@ -315,11 +315,11 @@ def resolved(file_name: str, digest: str) -> dict[str, object]: "boundary": "contradicted", "defect": "cited object wholly replaced after appraisal", "spec": SPEC_DECIDED, - "record": record_with(uri_v1), + "record": record_with(uri_base), "context": { - "cited_uri": uri_v1, - "resolution": resolved("unrelated-policy.json", d_unrel), - "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_v1}, + "cited_uri": uri_base, + "resolution": resolved("policy-bundle-unrelated.json", d_unrelated), + "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_base}, }, "expected": { "outcome": "reject", @@ -344,7 +344,7 @@ def resolved(file_name: str, digest: str) -> dict[str, object]: "file": None, "note": "No sibling file corresponds to this URI, by construction.", }, - "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_v1}, + "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_base}, }, "expected": { "outcome": "deferred", @@ -377,13 +377,13 @@ def resolved(file_name: str, digest: str) -> dict[str, object]: "draft, a surface facing this question rather than an " "established answer to it - without adopting its vocabulary." ), - "record": record_with(uri_v2), + "record": record_with(uri_other), "context": { - "cited_uri": uri_v2, - "resolution": resolved("appraisal-policy-v2.json", d_v2), + "cited_uri": uri_other, + "resolution": resolved("policy-bundle-other.json", d_other), "candidate_binding": { "field": CANDIDATE_FIELD, - "value": sha3_512_of(raw["appraisal-policy-v2.json"]), + "value": sha3_512_of(raw["policy-bundle-other.json"]), "note": ( "This is the correct SHA3-512 of the referent. It is " "outside the schema's digest pattern " @@ -416,15 +416,15 @@ def resolved(file_name: str, digest: str) -> dict[str, object]: "boundary": "contradicted", "defect": "binding well formed but bound to a different referent", "spec": SPEC_DECIDED, - "record": record_with(uri_v1), + "record": record_with(uri_base), "context": { - "cited_uri": uri_v1, - "resolution": resolved("appraisal-policy-v1.json", d_v1), + "cited_uri": uri_base, + "resolution": resolved("policy-bundle-base.json", d_base), "candidate_binding": { "field": CANDIDATE_FIELD, - "value": d_v2, + "value": d_other, "note": ( - "This is the digest of policies/appraisal-policy-v2.json, " + "This is the digest of policies/policy-bundle-other.json, " "which is not what cited_uri resolves to." ), }, diff --git a/tests/vectors/appraisal-resolution/policies/appraisal-policy-v1.json b/tests/vectors/policy-resolution/policies/policy-bundle-base.json similarity index 100% rename from tests/vectors/appraisal-resolution/policies/appraisal-policy-v1.json rename to tests/vectors/policy-resolution/policies/policy-bundle-base.json diff --git a/tests/vectors/appraisal-resolution/policies/appraisal-policy-v1-rev2.json b/tests/vectors/policy-resolution/policies/policy-bundle-onebyte.json similarity index 100% rename from tests/vectors/appraisal-resolution/policies/appraisal-policy-v1-rev2.json rename to tests/vectors/policy-resolution/policies/policy-bundle-onebyte.json diff --git a/tests/vectors/appraisal-resolution/policies/appraisal-policy-v2.json b/tests/vectors/policy-resolution/policies/policy-bundle-other.json similarity index 100% rename from tests/vectors/appraisal-resolution/policies/appraisal-policy-v2.json rename to tests/vectors/policy-resolution/policies/policy-bundle-other.json diff --git a/tests/vectors/appraisal-resolution/policies/unrelated-policy.json b/tests/vectors/policy-resolution/policies/policy-bundle-unrelated.json similarity index 100% rename from tests/vectors/appraisal-resolution/policies/unrelated-policy.json rename to tests/vectors/policy-resolution/policies/policy-bundle-unrelated.json From f34dd5120279a32f66be024a7bc526dbeb93150e Mon Sep 17 00:00:00 2001 From: opento-suggestions Date: Mon, 24 Aug 2026 14:36:02 -0600 Subject: [PATCH 3/7] feat(modules): per-code table for when UNVERIFIED fails the run Adds modules/unverified.py: UNVERIFIED_FAILS_FROM_LEVEL and unverified_fails(). cli.py and report.py call it in place of the blanket level >= 1. Unlisted codes default to level >= 1 (fail-closed, previous behaviour). result.py restates the UNVERIFIED contract in general terms: the check could not be executed against the evidence the record cites. The table lives under modules/ so test_docs_match_the_modules can see it, and the error-codes.md row for TR-POL-003 lands here because the table names the code and the guard reads string literals. Test totals unchanged. Signed-off-by: opento-suggestions --- docs/error-codes.md | 1 + src/trace_tests/cli.py | 15 ++++++---- src/trace_tests/modules/unverified.py | 40 +++++++++++++++++++++++++++ src/trace_tests/report.py | 15 +++++----- src/trace_tests/result.py | 7 +++-- 5 files changed, 63 insertions(+), 15 deletions(-) create mode 100644 src/trace_tests/modules/unverified.py diff --git a/docs/error-codes.md b/docs/error-codes.md index 4fb6759..db1a054 100644 --- a/docs/error-codes.md +++ b/docs/error-codes.md @@ -35,6 +35,7 @@ All TRACE test failures emit a structured error code of the form `TR--= 1: - failures += unverified + # Defense in depth: an unverified finding must fail the run from the level + # its code is registered at, even if a module forgot to emit a hard FAIL. + # The level is per-code rather than blanket; an unregistered code fails + # from level 1, which is what the blanket rule did for all of them. + failures += unverified_failing - total = passes + failures + skips + (unverified if level == 0 else 0) + total = passes + failures + skips + (unverified - unverified_failing) click.echo("") if failures == 0: if unverified: diff --git a/src/trace_tests/modules/unverified.py b/src/trace_tests/modules/unverified.py new file mode 100644 index 0000000..a5ed3a2 --- /dev/null +++ b/src/trace_tests/modules/unverified.py @@ -0,0 +1,40 @@ +"""Which unverified findings fail a run, and from which conformance level. + +An unverified finding means the check could not be executed against the +evidence the record cites. That is a statement about reachability, not about +the record being wrong: the evidence may be perfectly good and simply out of +reach. It is held apart from a skip so that a consumer can never read it as a +benign omission. + +Whether such a finding fails the run is **per code**, read from the table +below, rather than one rule applied to every unverified finding at once. +Different checks lose their evidence for different reasons and at different +levels. A record with no signature cannot be called conformant anywhere that +requires one. A policy bundle that did not resolve is a weaker statement, and +the level at which it stops being tolerable is a property of that check rather +than of the status. + +A code absent from the table fails from level 1, which is what the blanket +rule did for every code before this table existed. That default is deliberate: +a new code that nobody remembered to register still fails closed, and the +registration guard turns red rather than the run turning quietly permissive. +""" + +from __future__ import annotations + +#: Lowest conformance level at which an unverified finding under this code +#: counts as a failure. Registered here and in ``docs/levels.md``; the two are +#: held equal by ``tests/test_docs_match_the_modules.py``. +UNVERIFIED_FAILS_FROM_LEVEL: dict[str, int] = { + "TR-SIG-005": 1, + "TR-POL-003": 2, +} + +#: Applied to any code the table does not name. Fail-closed on purpose; see +#: the module docstring. +DEFAULT_FAILS_FROM_LEVEL = 1 + + +def unverified_fails(code: str, level: int) -> bool: + """Return True when an unverified finding under *code* must fail at *level*.""" + return level >= UNVERIFIED_FAILS_FROM_LEVEL.get(code, DEFAULT_FAILS_FROM_LEVEL) diff --git a/src/trace_tests/report.py b/src/trace_tests/report.py index 686235c..081ea5a 100644 --- a/src/trace_tests/report.py +++ b/src/trace_tests/report.py @@ -25,6 +25,7 @@ from dataclasses import dataclass from typing import Any +from trace_tests.modules.unverified import unverified_fails from trace_tests.result import Finding, Status __all__ = [ @@ -107,13 +108,13 @@ def verdict(self) -> str: def _tally(results: dict[str, list[Finding]], level: int) -> tuple[int, int]: failures = sum(1 for fs in results.values() for f in fs if f.failed()) - unverified = sum(1 for fs in results.values() for f in fs if f.unverified()) - # Mirrors the CLI: an unverified finding is a failure wherever signatures are - # required. A report that called an unverified record "PASS" at Level 1 would - # be worse than no report. - if level >= 1: - failures += unverified - return failures, unverified + unverified_findings = [f for fs in results.values() for f in fs if f.unverified()] + # Mirrors the CLI: an unverified finding is a failure from the level its code + # is registered at. A report that called such a record "PASS" at a level that + # required the check would be worse than no report. Per-code rather than + # blanket; an unregistered code fails from level 1, as the blanket rule did. + failures += sum(1 for f in unverified_findings if unverified_fails(f.code, level)) + return failures, len(unverified_findings) def build( diff --git a/src/trace_tests/result.py b/src/trace_tests/result.py index 82b9b1a..cf7e79f 100644 --- a/src/trace_tests/result.py +++ b/src/trace_tests/result.py @@ -10,9 +10,10 @@ class Status(StrEnum): PASS = "pass" FAIL = "fail" SKIP = "skip" - # No cryptographic verification was possible. Distinct from SKIP so callers - # can never mistake an unverified record for a benign omission. Treated as - # a failure at any conformance level that requires signatures (level >= 1). + # The check could not be executed against the evidence the record cites. + # Distinct from SKIP so callers can never mistake an unverified check for a + # benign omission. Whether it fails the run is per-code, from the table in + # modules/unverified.py, rather than one rule over every unverified finding. UNVERIFIED = "unverified" From 7dbbde436837111fda8102eb673aa102bdb2a595 Mon Sep 17 00:00:00 2001 From: opento-suggestions Date: Mon, 24 Aug 2026 14:47:24 -0600 Subject: [PATCH 4/7] feat(tr-pol): TR-POL-003 resolves policy_uri and matches bundle_hash policy_resolver: Callable[[str], bytes] | None = None on tr_pol.check and runner.run, threaded explicitly - the fifth caller-supplied input of this shape after agentrust-io/trace-tests#79's receipt. None -> SKIP, so offline runs are unchanged. policy_uri that is not an absolute URI -> FAIL, checked before the resolver is consulted. Resolver raises -> UNVERIFIED, message carries the exception text. Digest recomputed with the algorithm bundle_hash names (sha256 or sha384); match -> PASS, mismatch -> FAIL. The CLI summary line for unverified findings is reworded: TR-SIG-005 is no longer the only code that can be unverified. The TR-SIG-005 finding text is unchanged. Signed-off-by: opento-suggestions --- src/trace_tests/cli.py | 3 +- src/trace_tests/modules/tr_pol.py | 130 ++++++++++++++++++- src/trace_tests/runner.py | 4 +- tests/unit/test_cli.py | 6 +- tests/unit/test_tr_pol_resolution.py | 184 +++++++++++++++++++++++++++ 5 files changed, 322 insertions(+), 5 deletions(-) create mode 100644 tests/unit/test_tr_pol_resolution.py diff --git a/src/trace_tests/cli.py b/src/trace_tests/cli.py index e17c4f8..8da3b5c 100644 --- a/src/trace_tests/cli.py +++ b/src/trace_tests/cli.py @@ -77,7 +77,8 @@ def _print_report(path: str, fmt: str, level: int, results: dict[str, list[Any]] if unverified: click.echo( f"Result: PASS ({total} checks, {skips} skipped, {unverified} UNVERIFIED " - f"-- record is NOT cryptographically verified)" + f"-- {unverified} check(s) could not be executed against the evidence " + f"this record cites)" ) else: click.echo(f"Result: PASS ({total} checks, {skips} skipped)") diff --git a/src/trace_tests/modules/tr_pol.py b/src/trace_tests/modules/tr_pol.py index 7848505..beb932c 100644 --- a/src/trace_tests/modules/tr_pol.py +++ b/src/trace_tests/modules/tr_pol.py @@ -2,19 +2,143 @@ from __future__ import annotations +import hashlib import re +from collections.abc import Callable from typing import Any +from urllib.parse import urlsplit from trace_tests.result import Finding, Status _DIGEST_RE = re.compile(r"^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$") +#: Digest algorithms this module can compute, keyed by the prefix a record uses. +#: Kept in step with `_DIGEST_RE`: a prefix the pattern admits and this map does +#: not would be accepted by TR-POL-001 and uncomputable by TR-POL-003. +_DIGEST_ALGOS: dict[str, Any] = {"sha256:": hashlib.sha256, "sha384:": hashlib.sha384} +#: RFC 3986 scheme: ALPHA *( ALPHA / DIGIT / "+" / "-" / "." ) +_SCHEME_RE = re.compile(r"^[A-Za-z][A-Za-z0-9+.\-]*$") #: Mirrors `policy.enforcement_mode` in the packaged schema; `test_enum_parity` fails if #: it drifts from that copy. _VALID_ENFORCEMENT = frozenset({"enforce", "advisory", "silent", "declared"}) -def check(trace: dict[str, Any]) -> list[Finding]: - """Return TR-POL findings for the policy bundle claim.""" +def _not_absolute_uri(policy_uri: str) -> str | None: + """Say why *policy_uri* is not an absolute URI, or None when it is one. + + Two rules, kept apart because they catch different mistakes. A reference + with no scheme is a relative reference: the packaged schema asks for + ``format: "uri"``, which is the absolute form, and a reader who writes the + ``uri-reference`` form gets something no verifier can dereference on its + own. A reference carrying whitespace or a control character is a + transcription accident, and one that survives a diff unnoticed. + """ + for ch in policy_uri: + if ch.isspace() or ord(ch) < 0x20 or ord(ch) == 0x7F: + return f"contains whitespace or a control character ({ch!r})" + try: + scheme = urlsplit(policy_uri).scheme + except ValueError as exc: # pragma: no cover - urlsplit is total for str + return f"cannot be parsed as a URI ({exc})" + if not scheme: + return "has no scheme, so it is a relative reference rather than an absolute URI" + if not _SCHEME_RE.match(scheme): + return f"has a scheme that RFC 3986 does not admit ({scheme!r})" + return None + + +def _resolution_finding( + policy: dict[str, Any], policy_resolver: Callable[[str], bytes] | None +) -> Finding: + """Return the TR-POL-003 finding for *policy*. + + The order below is the whole of the check's meaning, so it is worth stating. + A reference the record got wrong is a defect in the record, visible with no + network and no resolver, exactly like the digest shape TR-POL-001 tests. A + referent that could not be fetched is weather. The malformed check therefore + runs before the resolver is consulted: running an offline verification must + not mean being blind to a defect the record carries on its face. + """ + policy_uri = policy.get("policy_uri") + if policy_uri is None: + return Finding( + "TR-POL-003", Status.SKIP, + "policy.policy_uri not present (optional); no bundle to resolve", + ) + if not isinstance(policy_uri, str): + return Finding( + "TR-POL-003", Status.FAIL, + f"TR-POL-003: policy.policy_uri must be a string, got {type(policy_uri).__name__}", + ) + + malformed = _not_absolute_uri(policy_uri) + if malformed is not None: + return Finding( + "TR-POL-003", Status.FAIL, + f"TR-POL-003: policy.policy_uri {malformed}: {policy_uri!r}", + ) + + bundle_hash = str(policy.get("bundle_hash", "")) + if not _DIGEST_RE.match(bundle_hash): + return Finding( + "TR-POL-003", Status.SKIP, + "policy.bundle_hash is not a well-formed digest, so there is nothing to " + "compare the resolved bundle against; reported by TR-POL-001", + ) + + if policy_resolver is None: + return Finding( + "TR-POL-003", Status.SKIP, + "policy.policy_uri not resolved; no resolver supplied", + ) + + try: + resolved = policy_resolver(policy_uri) + except Exception as exc: + # The exception text is carried deliberately. The status says only that + # the bundle could not be read, which is the same word for a withdrawn + # referent and a mistyped path; without the reason the second is + # indistinguishable from weather. + return Finding( + "TR-POL-003", Status.UNVERIFIED, + f"TR-POL-003: policy.policy_uri could not be resolved, so " + f"policy.bundle_hash was not checked against it: {policy_uri!r} " + f"({type(exc).__name__}: {exc})", + ) + + if not isinstance(resolved, bytes): + return Finding( + "TR-POL-003", Status.UNVERIFIED, + "TR-POL-003: the policy resolver violated its contract by returning " + f"{type(resolved).__name__} rather than bytes, so policy.bundle_hash " + f"was not checked against policy.policy_uri {policy_uri!r}", + ) + + prefix, _, _ = bundle_hash.partition(":") + algo = _DIGEST_ALGOS[f"{prefix}:"] + actual = f"{prefix}:{algo(resolved).hexdigest().lower()}" + if actual == bundle_hash.lower(): + return Finding( + "TR-POL-003", Status.PASS, + f"policy.policy_uri resolves to the bundle policy.bundle_hash declares " + f"({len(resolved)} bytes, {prefix})", + ) + return Finding( + "TR-POL-003", Status.FAIL, + f"TR-POL-003: policy.bundle_hash does not describe what policy.policy_uri " + f"resolves to; declared {bundle_hash}, resolved {actual}", + ) + + +def check( + trace: dict[str, Any], *, policy_resolver: Callable[[str], bytes] | None = None +) -> list[Finding]: + """Return TR-POL findings for the policy bundle claim. + + *policy_resolver* is supplied by the caller and never derived from the + record: a record that could name its own resolver could name one that + agrees with it. When it is None the resolution check skips, so offline + verification stays a first-class use rather than a degraded one. + """ findings: list[Finding] = [] policy = trace.get("policy") @@ -39,4 +163,6 @@ def check(trace: dict[str, Any]) -> list[Finding]: f"TR-POL-002: policy.enforcement_mode must be one of {sorted(_VALID_ENFORCEMENT)}, got {enforcement!r}", )) + findings.append(_resolution_finding(policy, policy_resolver)) + return findings diff --git a/src/trace_tests/runner.py b/src/trace_tests/runner.py index f8f318e..01431a9 100644 --- a/src/trace_tests/runner.py +++ b/src/trace_tests/runner.py @@ -2,6 +2,7 @@ from __future__ import annotations +from collections.abc import Callable from typing import Any from trace_tests.loader import extract_trace @@ -23,6 +24,7 @@ def run( max_age_seconds: int = tr_env.DEFAULT_MAX_AGE_SECONDS, expected_nonce: str | None = None, receipt: dict[str, Any] | None = None, + policy_resolver: Callable[[str], bytes] | None = None, ) -> dict[str, list[Finding]]: """Run all modules required for *level* and return findings keyed by module ID.""" if level not in _LEVEL_MODULES: @@ -40,7 +42,7 @@ def run( results["TR-SIG"] = tr_sig.check(trace, record, fmt, level) if "TR-POL" in active: - results["TR-POL"] = tr_pol.check(trace) + results["TR-POL"] = tr_pol.check(trace, policy_resolver=policy_resolver) if "TR-RTE" in active: results["TR-RTE"] = tr_rte.check(trace, level, expected_nonce=expected_nonce) diff --git a/tests/unit/test_cli.py b/tests/unit/test_cli.py index 3c20483..be93fd2 100644 --- a/tests/unit/test_cli.py +++ b/tests/unit/test_cli.py @@ -43,7 +43,11 @@ def test_unsigned_record_level_0_reports_unverified(fresh_level0_path): result = CliRunner().invoke(main, ["verify", "--record", fresh_level0_path, "--level", "0"]) assert result.exit_code == 0, result.output assert "UNVERIFIED" in result.output - assert "NOT cryptographically verified" in result.output + # The summary no longer says "cryptographically": under per-code + # registration a policy bundle that did not resolve is unverified too, and + # a line naming only signatures would be a published claim the code stopped + # making the moment TR-POL-003 existed. + assert "could not be executed against the evidence this record cites" in result.output def test_stale_record_fails(tmp_path): diff --git a/tests/unit/test_tr_pol_resolution.py b/tests/unit/test_tr_pol_resolution.py new file mode 100644 index 0000000..ef6fbb2 --- /dev/null +++ b/tests/unit/test_tr_pol_resolution.py @@ -0,0 +1,184 @@ +"""TR-POL-003: every branch of the resolution check, one test each. + +The resolver here is a dict, not a filesystem or a network. The module's +contract is ``Callable[[str], bytes]`` and nothing more, so the tests exercise +that contract directly: a lookup that answers, one that raises, and one that +returns the wrong type. What the caller does to obtain the bytes is the +caller's business, which is the point of the seam. +""" + +from __future__ import annotations + +import hashlib + +import pytest + +from trace_tests.modules.tr_pol import check +from trace_tests.result import Status + +BUNDLE = b'{"policy_id":"appraisal/agent-v1","version":"1.0.0"}' +OTHER = b'{"policy_id":"retention/pii-90d","version":"1.4.2"}' +URI = "https://policy.example.org/bundles/policy-bundle-base.json" + +SHA256 = "sha256:" + hashlib.sha256(BUNDLE).hexdigest() +SHA384 = "sha384:" + hashlib.sha384(BUNDLE).hexdigest() + + +def resolver(mapping: dict[str, bytes]): + """A resolver over *mapping*; a URI it does not hold raises, as a fetch would.""" + def _resolve(uri: str) -> bytes: + return mapping[uri] + return _resolve + + +def trace_with(**policy): + base = {"bundle_hash": SHA256, "enforcement_mode": "enforce"} + base.update(policy) + return {"policy": base} + + +def pol003(findings): + return next(f for f in findings if f.code == "TR-POL-003") + + +def test_no_policy_uri_skips(): + f = pol003(check(trace_with(), policy_resolver=resolver({URI: BUNDLE}))) + assert f.status is Status.SKIP + assert "not present" in f.message + + +def test_policy_uri_of_the_wrong_type_fails(): + f = pol003(check(trace_with(policy_uri=42), policy_resolver=resolver({}))) + assert f.status is Status.FAIL + assert "must be a string" in f.message + + +@pytest.mark.parametrize( + "bad,why", + [ + ("bundles/policy-bundle-base.json", "no scheme"), + ("//policy.example.org/bundles/base.json", "no scheme"), + ("https://policy.example.org/bundles/policy bundle base.json", "whitespace"), + ("https://policy.example.org/bundles/base.json\n", "whitespace"), + ("1https://policy.example.org/x.json", "no scheme"), + ], +) +def test_a_reference_that_is_not_an_absolute_uri_fails(bad, why): + """A defect in the record, so it is a failure and not weather.""" + f = pol003(check(trace_with(policy_uri=bad), policy_resolver=resolver({}))) + assert f.status is Status.FAIL, why + assert "policy.policy_uri" in f.message + + +def test_a_malformed_reference_fails_with_no_resolver_at_all(): + """The check that matters most: offline must not mean blind. + + A reference the record got wrong needs no network to detect. If this + skipped when no resolver was supplied, an offline run would report nothing + about a record that is wrong on its face. + """ + f = pol003(check(trace_with(policy_uri="bundles/base.json"), policy_resolver=None)) + assert f.status is Status.FAIL + + +def test_malformed_reference_outranks_a_malformed_digest(): + f = pol003(check(trace_with(policy_uri="bundles/base.json", bundle_hash="not-a-digest"))) + assert f.status is Status.FAIL + + +def test_unusable_bundle_hash_skips_toward_tr_pol_001(): + f = pol003(check(trace_with(policy_uri=URI, bundle_hash="sha256:short"), + policy_resolver=resolver({URI: BUNDLE}))) + assert f.status is Status.SKIP + assert "TR-POL-001" in f.message + + +def test_no_resolver_skips_rather_than_reporting_unverified(): + """Offline verification is a first-class use, not a degraded one.""" + f = pol003(check(trace_with(policy_uri=URI), policy_resolver=None)) + assert f.status is Status.SKIP + assert "no resolver supplied" in f.message + + +def test_a_resolver_that_raises_is_unverified_and_says_why(): + f = pol003(check(trace_with(policy_uri=URI), policy_resolver=resolver({}))) + assert f.status is Status.UNVERIFIED + assert "KeyError" in f.message, "the reason must survive into the message" + + +def test_a_resolver_that_raises_oserror_is_unverified_and_says_why(): + """A mistyped manifest path and a withdrawn referent share a status. + + Only the message distinguishes them, which is why it carries the exception. + """ + def raising(uri: str) -> bytes: + raise FileNotFoundError(2, "No such file or directory", "policies/missing.json") + + f = pol003(check(trace_with(policy_uri=URI), policy_resolver=raising)) + assert f.status is Status.UNVERIFIED + assert "FileNotFoundError" in f.message + assert "missing.json" in f.message + + +def test_a_resolver_returning_non_bytes_names_the_contract_it_broke(): + f = pol003(check(trace_with(policy_uri=URI), + policy_resolver=lambda uri: BUNDLE.decode())) + assert f.status is Status.UNVERIFIED + assert "violated its contract" in f.message + assert "str" in f.message + + +def test_resolved_bundle_matching_sha256_passes(): + f = pol003(check(trace_with(policy_uri=URI), policy_resolver=resolver({URI: BUNDLE}))) + assert f.status is Status.PASS + + +def test_resolved_bundle_matching_sha384_passes(): + """A set that never exercises sha384 passes a verifier that hardcodes sha256.""" + f = pol003(check(trace_with(policy_uri=URI, bundle_hash=SHA384), + policy_resolver=resolver({URI: BUNDLE}))) + assert f.status is Status.PASS + + +def test_resolved_bundle_contradicting_the_declared_digest_fails(): + f = pol003(check(trace_with(policy_uri=URI), policy_resolver=resolver({URI: OTHER}))) + assert f.status is Status.FAIL + assert "does not describe what" in f.message + + +def test_sha384_mismatch_fails_rather_than_passing_unchecked(): + """The fail-open shape: a verifier that cannot compute sha384 must not pass.""" + f = pol003(check(trace_with(policy_uri=URI, bundle_hash=SHA384), + policy_resolver=resolver({URI: OTHER}))) + assert f.status is Status.FAIL + + +def test_an_uppercase_declared_digest_still_compares_equal(): + upper = "sha256:" + hashlib.sha256(BUNDLE).hexdigest().upper() + f = pol003(check(trace_with(policy_uri=URI, bundle_hash=upper), + policy_resolver=resolver({URI: BUNDLE}))) + # TR-POL-001's pattern is lowercase-only, so an uppercase digest never + # reaches the comparison in practice; this pins the behaviour if it changes. + assert f.status in (Status.PASS, Status.SKIP) + + +def test_the_resolver_is_never_called_when_no_policy_uri_is_present(): + calls: list[str] = [] + + def counting(uri: str) -> bytes: + calls.append(uri) + return BUNDLE + + check(trace_with(), policy_resolver=counting) + assert calls == [] + + +def test_the_resolver_is_never_called_for_a_malformed_reference(): + calls: list[str] = [] + + def counting(uri: str) -> bytes: + calls.append(uri) + return BUNDLE + + check(trace_with(policy_uri="bundles/base.json"), policy_resolver=counting) + assert calls == [], "a record defect must not cost a fetch" From 8d22b72416099ccb802b0d1d057d9596b6816036 Mon Sep 17 00:00:00 2001 From: opento-suggestions Date: Mon, 24 Aug 2026 14:56:04 -0600 Subject: [PATCH 5/7] feat(cli): --policy-dir supplies a manifest-backed policy resolver Same shape as --receipt: DIR/resolutions.json is form-checked at load and malformed input exits 2. Existence is a resolve-time fact, so a missing key or a missing file surfaces as UNVERIFIED, not as a CLI error. Signed-off-by: opento-suggestions --- src/trace_tests/cli.py | 101 +++++++++++++++++++++- tests/unit/test_cli_policy_dir.py | 136 ++++++++++++++++++++++++++++++ 2 files changed, 236 insertions(+), 1 deletion(-) create mode 100644 tests/unit/test_cli_policy_dir.py diff --git a/src/trace_tests/cli.py b/src/trace_tests/cli.py index 8da3b5c..5029bdb 100644 --- a/src/trace_tests/cli.py +++ b/src/trace_tests/cli.py @@ -6,7 +6,9 @@ import json import importlib.metadata import pathlib +import re import sys +from collections.abc import Callable from typing import Any import click @@ -110,6 +112,71 @@ def _load_receipt(path: str | None) -> dict | None: sys.exit(2) return data + +def _load_policy_resolver(policy_dir: str | None) -> Callable[[str], bytes] | None: + """Build a policy-bundle resolver from DIR/resolutions.json, or None when not supplied. + + The manifest is checked for *form* only: an object, string to string, + relative paths, no parent traversal. Whether a mapped file is actually + there is deliberately not checked here. Existence is a resolve-time fact, + and a manifest that refused to load because one bundle had gone missing + would be the manifest-level version of treating a lost referent as a wrong + reference, which is the confusion TR-POL-003 exists to keep apart. A + missing file surfaces as an unverified finding for the record that cites + it, and leaves every other record in the run readable. + """ + if policy_dir is None: + return None + root = pathlib.Path(policy_dir) + manifest_path = root / "resolutions.json" + try: + with open(manifest_path, encoding="utf-8") as fh: + data = json.load(fh) + except OSError as exc: + click.echo(f"Error: cannot read policy manifest {manifest_path}: {exc}", err=True) + sys.exit(2) + except json.JSONDecodeError as exc: + click.echo(f"Error: policy manifest {manifest_path} is not valid JSON: {exc}", err=True) + sys.exit(2) + if not isinstance(data, dict): + click.echo( + f"Error: policy manifest {manifest_path} must be a JSON object, " + f"got {type(data).__name__}", + err=True, + ) + sys.exit(2) + for uri, rel in data.items(): + if not isinstance(rel, str): + click.echo( + f"Error: policy manifest {manifest_path} maps {uri!r} to " + f"{type(rel).__name__}, expected a relative path string", + err=True, + ) + sys.exit(2) + if _is_unsafe_relative(rel): + click.echo( + f"Error: policy manifest {manifest_path} maps {uri!r} to {rel!r}; " + "entries must be relative paths inside the directory, with no " + "parent traversal", + err=True, + ) + sys.exit(2) + + def _resolve(uri: str) -> bytes: + # A URI the manifest does not hold raises, exactly as a fetch would; + # so does a mapped file that is not there. Both reach TR-POL-003 as + # unverified, and the message carries which happened. + return (root / data[uri]).read_bytes() + + return _resolve + + +def _is_unsafe_relative(rel: str) -> bool: + """True when *rel* escapes the manifest's own directory, or tries to.""" + if not rel or rel.startswith(("/", "\\")) or re.match(r"^[A-Za-z]:", rel): + return True + return ".." in pathlib.PurePosixPath(rel.replace("\\", "/")).parts + @click.group() @click.version_option(__version__) def main() -> None: @@ -148,7 +215,25 @@ def main() -> None: "says where the anchor lives, the receipt is what proves the record is in it." ), ) -def verify(record: str, level: int, max_age: int, expected_nonce: str | None, receipt: str | None) -> None: +@click.option( + "--policy-dir", + "policy_dir", + default=None, + type=click.Path(), + help=( + "Directory holding resolutions.json, a map from policy_uri to a relative " + "path inside it. Supplying it lets TR-POL-003 resolve policy.policy_uri and " + "compare the bundle against policy.bundle_hash; without it that check skips." + ), +) +def verify( + record: str, + level: int, + max_age: int, + expected_nonce: str | None, + receipt: str | None, + policy_dir: str | None, +) -> None: """Verify a TRACE trust record against the conformance suite.""" try: data, fmt = load_record(record) @@ -157,6 +242,7 @@ def verify(record: str, level: int, max_age: int, expected_nonce: str | None, re sys.exit(2) receipt_data = _load_receipt(receipt) + policy_resolver = _load_policy_resolver(policy_dir) results = run( data, @@ -165,6 +251,7 @@ def verify(record: str, level: int, max_age: int, expected_nonce: str | None, re max_age_seconds=max_age, expected_nonce=expected_nonce, receipt=receipt_data, + policy_resolver=policy_resolver, ) exit_code = _print_report(record, fmt, level, results) sys.exit(exit_code) @@ -210,6 +297,15 @@ def verify(record: str, level: int, max_age: int, expected_nonce: str | None, re type=click.Path(), help="Path to the anchor receipt (JSON). Required for TR-ANC-002 at Level 2.", ) +@click.option( + "--policy-dir", + "policy_dir", + default=None, + type=click.Path(), + help="Directory holding resolutions.json, a map from policy_uri to a relative " + "path inside it. Required for TR-POL-003 to resolve the bundle; without it " + "that check skips.", +) def report( record: str, max_level: int, @@ -220,6 +316,7 @@ def report( fail_under: int | None, expected_nonce: str | None, receipt: str | None, + policy_dir: str | None, ) -> None: """Produce a conformance report you can hand to someone else. @@ -234,6 +331,7 @@ def report( sys.exit(2) receipt_data = _load_receipt(receipt) + policy_resolver = _load_policy_resolver(policy_dir) results_by_level = { level: run( @@ -243,6 +341,7 @@ def report( max_age_seconds=max_age, expected_nonce=expected_nonce, receipt=receipt_data, + policy_resolver=policy_resolver, ) for level in range(max_level + 1) } diff --git a/tests/unit/test_cli_policy_dir.py b/tests/unit/test_cli_policy_dir.py new file mode 100644 index 0000000..ac1af69 --- /dev/null +++ b/tests/unit/test_cli_policy_dir.py @@ -0,0 +1,136 @@ +"""--policy-dir: the manifest's shape-check, and what it deliberately does not check.""" + +from __future__ import annotations + +import hashlib +import json +import pathlib +import time + +import pytest +from click.testing import CliRunner + +from trace_tests.cli import _is_unsafe_relative, main + +VECTORS = pathlib.Path(__file__).resolve().parents[1] / "vectors" +BUNDLE = b'{"policy_id":"appraisal/agent-v1","version":"1.0.0"}\n' +URI = "https://policy.example.org/bundles/policy-bundle-base.json" + + +def _record(tmp_path, **policy): + rec = json.loads((VECTORS / "valid_level0.json").read_text(encoding="utf-8")) + rec["iat"] = int(time.time()) + rec["policy"].update(policy) + p = tmp_path / "record.json" + p.write_text(json.dumps(rec), encoding="utf-8") + return str(p) + + +def _policy_dir(tmp_path, manifest, *, write_bundle=True): + d = tmp_path / "policies" + d.mkdir(exist_ok=True) + if write_bundle: + (d / "policy-bundle-base.json").write_bytes(BUNDLE) + (d / "resolutions.json").write_text(json.dumps(manifest), encoding="utf-8") + return str(d) + + +def _run(record, policy_dir=None, level="0"): + args = ["verify", "--record", record, "--level", level, "--max-age", "3153600000"] + if policy_dir is not None: + args += ["--policy-dir", policy_dir] + return CliRunner().invoke(main, args) + + +def test_a_resolvable_bundle_passes_tr_pol_003(tmp_path): + digest = "sha256:" + hashlib.sha256(BUNDLE).hexdigest() + rec = _record(tmp_path, policy_uri=URI, bundle_hash=digest) + d = _policy_dir(tmp_path, {URI: "policy-bundle-base.json"}) + result = _run(rec, d) + assert "TR-POL-003" not in result.output or "PASS" in result.output + assert result.exit_code == 0, result.output + + +def test_a_uri_the_manifest_does_not_hold_is_unverified(tmp_path): + digest = "sha256:" + hashlib.sha256(BUNDLE).hexdigest() + rec = _record(tmp_path, policy_uri=URI, bundle_hash=digest) + d = _policy_dir( + tmp_path, + {"https://policy.example.org/bundles/other.json": "policy-bundle-base.json"}, + ) + result = _run(rec, d) + assert "UNVERIFIED" in result.output + assert "KeyError" in result.output + + +def test_a_mapped_file_that_is_not_there_is_unverified_not_a_load_error(tmp_path): + """The manifest loads; the missing bundle surfaces per record, as weather. + + Refusing the whole manifest because one bundle had gone missing would be + the manifest-level version of treating a lost referent as a wrong + reference, and it would make every other record in the run unreadable too. + """ + digest = "sha256:" + hashlib.sha256(BUNDLE).hexdigest() + rec = _record(tmp_path, policy_uri=URI, bundle_hash=digest) + d = _policy_dir(tmp_path, {URI: "policy-bundle-gone.json"}) + result = _run(rec, d) + assert result.exit_code != 2, "a missing bundle must not be a manifest load error" + assert "UNVERIFIED" in result.output + assert "FileNotFoundError" in result.output or "Errno 2" in result.output + + +def test_a_missing_manifest_exits_2(tmp_path): + rec = _record(tmp_path) + result = _run(rec, str(tmp_path / "nowhere")) + assert result.exit_code == 2 + assert "cannot read policy manifest" in result.output + + +def test_a_manifest_that_is_not_json_exits_2(tmp_path): + rec = _record(tmp_path) + d = tmp_path / "policies" + d.mkdir() + (d / "resolutions.json").write_text("{not json", encoding="utf-8") + result = _run(rec, str(d)) + assert result.exit_code == 2 + assert "is not valid JSON" in result.output + + +def test_a_manifest_that_is_not_an_object_exits_2(tmp_path): + rec = _record(tmp_path) + d = _policy_dir(tmp_path, ["a", "list"]) + result = _run(rec, d) + assert result.exit_code == 2 + assert "must be a JSON object" in result.output + + +def test_a_manifest_value_that_is_not_a_string_exits_2(tmp_path): + rec = _record(tmp_path) + d = _policy_dir(tmp_path, {URI: 7}) + result = _run(rec, d) + assert result.exit_code == 2 + assert "expected a relative path string" in result.output + + +@pytest.mark.parametrize( + "bad", ["/etc/passwd", "../outside.json", "a/../../outside.json", "C:/Windows/x.json", ""] +) +def test_a_manifest_entry_escaping_the_directory_exits_2(tmp_path, bad): + rec = _record(tmp_path) + d = _policy_dir(tmp_path, {URI: bad}) + result = _run(rec, d) + assert result.exit_code == 2 + assert "must be relative paths inside the directory" in result.output + + +@pytest.mark.parametrize("ok", ["a.json", "sub/a.json", "./a.json"]) +def test_ordinary_relative_paths_are_accepted(ok): + assert _is_unsafe_relative(ok) is False + + +def test_without_policy_dir_the_check_skips_rather_than_failing(tmp_path): + digest = "sha256:" + hashlib.sha256(BUNDLE).hexdigest() + rec = _record(tmp_path, policy_uri=URI, bundle_hash=digest) + result = _run(rec) + assert result.exit_code == 0, result.output + assert "no resolver supplied" in result.output From ef928bdc5bc3cabe9915101b796ea3981231bea2 Mon Sep 17 00:00:00 2001 From: opento-suggestions Date: Mon, 24 Aug 2026 15:08:14 -0600 Subject: [PATCH 6/7] test(vectors): retarget onto policy.bundle_hash and policy_uri Eleven vectors, every record identical except policy.*; runtime.nonce added so level 1 carries exactly one constant failure. Boundaries: accept 3, contradicted 4, unresolvable 2, malformed 2. Each vector's context carries its anchor and tier; candidate_binding and the deferred markers are removed, both fields being merged. Adds the CLI differential (failure deltas [0, 0, 1] across levels, exit 0 at level 0 both ways) and the report-path test. No hand-built Findings, and no expectation derived from the table: a test that read its expected delta from the registration table stayed green under every row mutation and was removed. Byte-reproduction re-established over vectors, bundles, and manifest. Signed-off-by: opento-suggestions --- tests/test_policy_resolution.py | 316 +++++----- tests/test_policy_resolution_cli.py | 169 ++++++ tests/test_policy_resolution_completeness.py | 196 ++++--- ...ng-declared.json => 01-no-policy-uri.json} | 28 +- .../02-resolved-and-matches.json | 31 +- .../03-digest-mismatch-minimal-mutation.json | 36 +- .../04-digest-mismatch-different-object.json | 33 +- .../05-referent-unreachable-no-route.json | 62 ++ .../05-referent-unreachable.json | 63 -- .../06-digest-algorithm-uncomputable.json | 64 --- .../06-resolved-and-matches-sha384.json | 62 ++ .../07-digest-bound-to-other-referent.json | 32 +- ...08-policy-uri-is-a-relative-reference.json | 62 ++ .../09-sha384-bound-to-other-referent.json | 62 ++ .../10-policy-uri-carries-a-space.json | 62 ++ .../11-referent-unreachable-route-fails.json | 62 ++ tests/vectors/policy-resolution/README.md | 277 ++++----- .../gen_policy_resolution.py | 542 ++++++++++-------- .../policy-resolution/resolutions.json | 7 + 19 files changed, 1396 insertions(+), 770 deletions(-) create mode 100644 tests/test_policy_resolution_cli.py rename tests/vectors/policy-resolution/{01-no-binding-declared.json => 01-no-policy-uri.json} (54%) create mode 100644 tests/vectors/policy-resolution/05-referent-unreachable-no-route.json delete mode 100644 tests/vectors/policy-resolution/05-referent-unreachable.json delete mode 100644 tests/vectors/policy-resolution/06-digest-algorithm-uncomputable.json create mode 100644 tests/vectors/policy-resolution/06-resolved-and-matches-sha384.json create mode 100644 tests/vectors/policy-resolution/08-policy-uri-is-a-relative-reference.json create mode 100644 tests/vectors/policy-resolution/09-sha384-bound-to-other-referent.json create mode 100644 tests/vectors/policy-resolution/10-policy-uri-carries-a-space.json create mode 100644 tests/vectors/policy-resolution/11-referent-unreachable-route-fails.json create mode 100644 tests/vectors/policy-resolution/resolutions.json diff --git a/tests/test_policy_resolution.py b/tests/test_policy_resolution.py index 3b27e9f..7caf4d4 100644 --- a/tests/test_policy_resolution.py +++ b/tests/test_policy_resolution.py @@ -1,202 +1,246 @@ -"""The policy-resolution set proves itself: digests, referents, schema. +"""The policy-resolution set proves itself: digests, referents, and TR-POL-003. -No verifier resolves ``appraisal.policy_ref``, so there is no implementation to -run these against. What can be checked — and is checked here — is that the set -is internally what it claims to be: - - * every record is valid under the packaged schema on this branch's base, so - the set describes conformant records rather than malformed ones; - * every declared digest really is, or really is not, the SHA-256 of the bytes - the vector says the citation resolved to, matching its expected outcome; - * the unreachable vector really has no resolvable object; - * vector 06's algorithm really is outside the digest set the schema admits. - -A vector whose expected outcome and whose bytes disagree is worse than no -vector: it looks like coverage and argues for the wrong thing. +Every vector states one expected value — the status of the TR-POL-003 finding +— and every test here obtains that status by running the module, never by +constructing a ``Finding`` by hand. A test that builds its own finding and +hands it to the reporting layer tests whichever code it happened to name; it +stays green while the thing under test is wrong. """ from __future__ import annotations import hashlib import json -import re from pathlib import Path import jsonschema import pytest -VECTOR_DIR = Path(__file__).parent / "vectors" / "policy-resolution" -SCHEMA_DIGEST_PATTERN = re.compile(r"^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$") +from trace_tests.runner import run +VECTOR_DIR = Path(__file__).parent / "vectors" / "policy-resolution" +SCHEMA_PATH = Path(__file__).parent.parent / "schemas" / "trace-claim.json" +MANIFEST = VECTOR_DIR / "resolutions.json" -def _vector_paths() -> list[Path]: - return sorted(VECTOR_DIR.glob("[0-9][0-9]-*.json")) +VECTOR_PATHS = sorted(VECTOR_DIR.glob("[0-9][0-9]-*.json")) +IDS = [p.name[:2] for p in VECTOR_PATHS] def _load(path: Path) -> dict: return json.loads(path.read_text(encoding="utf-8")) -def _sha256_of_file(rel: str) -> str: - return "sha256:" + hashlib.sha256((VECTOR_DIR / rel).read_bytes()).hexdigest() +@pytest.fixture(scope="module") +def schema() -> dict: + return json.loads(SCHEMA_PATH.read_text(encoding="utf-8")) -VECTORS = _vector_paths() +@pytest.fixture(scope="module") +def manifest() -> dict: + return _load(MANIFEST) -@pytest.mark.level0 -def test_the_set_has_the_seven_vectors_it_documents() -> None: - names = [p.name for p in VECTORS] - assert len(names) == 7, names - for n, path in enumerate(VECTORS, start=1): - assert path.name.startswith(f"{n:02d}-"), ( - f"vectors must be contiguously numbered; got {path.name} at position {n}" - ) +@pytest.fixture(scope="module") +def resolver(manifest): + """The set's own resolver: the manifest, then the bytes on disk. + + A URI the manifest does not hold raises, and so does a mapped file that is + not there. Both are how a fetch fails, and TR-POL-003 must reach the same + status by either road while saying which one happened. + """ + def _resolve(uri: str) -> bytes: + return (VECTOR_DIR / manifest[uri]).read_bytes() + return _resolve + + +def _pol003(record: dict, resolver, level: int = 0): + results = run(record, "trace", level, max_age_seconds=10**9, policy_resolver=resolver) + return next(f for f in results["TR-POL"] if f.code == "TR-POL-003") @pytest.mark.level0 -@pytest.mark.parametrize("path", VECTORS, ids=lambda p: p.stem) -def test_the_record_is_valid_under_the_packaged_schema(path: Path, schema) -> None: - """Schema validity, so the set is about resolution and not about malformed input.""" - jsonschema.validate(_load(path)["record"], schema) +def test_the_set_has_the_vectors_its_readme_documents() -> None: + readme = (VECTOR_DIR / "README.md").read_text(encoding="utf-8") + assert len(VECTOR_PATHS) == 11, "the set is eleven vectors, 01 through 11" + for path in VECTOR_PATHS: + assert path.name in readme, f"{path.name} is not documented in the set's README" @pytest.mark.level0 -@pytest.mark.parametrize("path", VECTORS, ids=lambda p: p.stem) -def test_the_record_carries_a_bare_policy_ref_and_no_invented_field(path: Path) -> None: - """The candidate binding must never leak into the record. +@pytest.mark.parametrize("path", VECTOR_PATHS, ids=IDS) +def test_the_record_is_valid_under_the_packaged_schema(path: Path, schema) -> None: + """Including the malformed-URI vectors. - ``appraisal`` is ``additionalProperties: false`` in the packaged schema, so - a record carrying the candidate field would be schema-invalid — and - proposing it is a schema decision this set does not make. + ``format: "uri"`` is inert here: no format checker is passed at any + validate call site in this repository, and ``uri`` is unregistered without + an optional dependency that is not installed. So 08 and 10 are + schema-valid and are caught by TR-POL-003 instead, which is the whole + reason the check has something to do. """ - appraisal = _load(path)["record"]["appraisal"] - assert set(appraisal) <= {"status", "verifier", "policy_ref", "timestamp"}, ( - f"record appraisal carries an unexpected key: {sorted(appraisal)}" - ) - for key in appraisal: - assert "digest" not in key, f"record appraisal proposes a binding field: {key}" + jsonschema.validate(_load(path)["record"], schema) @pytest.mark.level0 -@pytest.mark.parametrize("path", VECTORS, ids=lambda p: p.stem) -def test_the_declared_resolution_matches_the_bytes_on_disk(path: Path) -> None: - """Whatever the vector says resolution returned, the file must hash to it.""" - ctx = _load(path)["context"] - resolution = ctx["resolution"] - if resolution["outcome"] != "resolved": - assert resolution.get("file") in (None,), ( - "a resolution that did not resolve must not name a file" - ) - return - rel = resolution["file"] - assert (VECTOR_DIR / rel).is_file(), f"{rel} is missing from the set" - assert resolution["actual_sha256"] == _sha256_of_file(rel), ( - f"{rel} does not hash to the digest the vector records for it" +@pytest.mark.parametrize("path", VECTOR_PATHS, ids=IDS) +def test_the_record_carries_no_invented_field(path: Path) -> None: + """The set argues from the fields that exist, not from ones it wishes existed.""" + policy = _load(path)["record"]["policy"] + assert set(policy) <= {"bundle_hash", "enforcement_mode", "version", "policy_uri"}, ( + f"{path.name} carries a policy field the packaged schema does not define" ) @pytest.mark.level0 -@pytest.mark.negative -@pytest.mark.parametrize("path", VECTORS, ids=lambda p: p.stem) -def test_the_binding_agrees_with_the_expected_outcome(path: Path) -> None: - """The load-bearing self-proof. - - accept -> the declared binding equals the digest of what resolved - (or no binding is declared at all) - reject -> a binding and a resolution both exist, and they disagree - deferred-> the comparison could not be performed at all +@pytest.mark.parametrize("path", VECTOR_PATHS, ids=IDS) +def test_the_declared_resolution_matches_the_manifest(path: Path, manifest) -> None: + """The vector's story about what it cites must match the one mapping on disk. + + There is a single manifest, so a vector cannot hold a private idea of what + its URI resolves to. The three unreachable-by-design cases are named here + rather than skipped, so that "not in the manifest" stays a deliberate state + instead of an oversight. """ vector = _load(path) - ctx, outcome = vector["context"], vector["expected"]["outcome"] - binding = ctx.get("candidate_binding") - resolution = ctx["resolution"] - resolved = resolution["outcome"] == "resolved" - - if outcome == "pass": - if binding is None: - # 01: nothing was declared, so there is nothing that could contradict. - assert resolution["outcome"] == "not_attempted", ( - "a vector declaring no binding must not claim a resolution was " - "attempted; the point is that there was nothing to check" - ) - return - assert resolved, "an accepting vector with a binding must have resolved" - assert binding["value"] == resolution["actual_sha256"], ( - "accept requires the declared binding to equal what resolved" - ) + cited = vector["context"]["cited_uri"] + outcome = vector["context"]["resolution"]["outcome"] - elif outcome == "reject": - assert binding is not None, "a rejecting vector must declare a binding" - assert resolved, ( - "a rejecting vector must have resolved something; otherwise nothing " - "was contradicted and the case is unresolvable, not contradicted" + if outcome == "not_attempted": + assert cited is None or cited not in manifest, ( + f"{path.name} says resolution was not attempted, but the manifest " + "offers a route; the vector and the manifest disagree" ) - assert SCHEMA_DIGEST_PATTERN.match(binding["value"]), ( - "a rejecting vector's binding must be well formed, so the rejection " - "is about the referent and not about a malformed digest" + return + if outcome == "unreachable_no_route": + assert cited not in manifest, ( + f"{path.name} claims no route, but the manifest maps {cited}" ) - assert binding["value"] != resolution["actual_sha256"], ( - "reject requires the declared binding to differ from what resolved" + return + if outcome == "unreachable_route_fails": + assert cited in manifest, f"{path.name} claims a route the manifest does not hold" + assert not (VECTOR_DIR / manifest[cited]).exists(), ( + f"{path.name} claims the route fails, but {manifest[cited]} is present" ) + return - elif outcome == "deferred": - computable = ( - binding is not None - and SCHEMA_DIGEST_PATTERN.match(binding["value"]) is not None - ) - assert not (resolved and computable), ( - "a deferred vector must be one the verifier could not complete: " - "either the referent did not resolve, or the digest algorithm is " - "outside the set the schema admits" - ) + assert outcome == "resolved", f"{path.name}: unknown resolution outcome {outcome!r}" + assert cited in manifest, f"{path.name} says it resolved, but no route is mapped" + assert (VECTOR_DIR / manifest[cited]).exists() - else: # pragma: no cover - guarded by the completeness test - raise AssertionError(f"unknown expected outcome {outcome!r}") + +@pytest.mark.level0 +@pytest.mark.parametrize("path", VECTOR_PATHS, ids=IDS) +def test_the_module_produces_the_status_the_vector_expects(path: Path, resolver) -> None: + """The load-bearing test: the module is run, and its finding is read.""" + vector = _load(path) + finding = _pol003(vector["record"], resolver) + assert finding.status.value == vector["expected"]["tr_pol_003"], ( + f"{path.name}: expected {vector['expected']['tr_pol_003']}, " + f"got {finding.status.value} — {finding.message}" + ) @pytest.mark.level0 -@pytest.mark.negative -def test_05_really_has_no_resolvable_object() -> None: - vector = _load(VECTOR_DIR / "05-referent-unreachable.json") - assert vector["context"]["resolution"]["outcome"] == "unreachable" - cited = vector["context"]["cited_uri"] - tail = cited.rsplit("/", 1)[-1] - assert not (VECTOR_DIR / "policies" / tail).exists(), ( - "05 claims the referent is unreachable, but a sibling file matches its URI" +@pytest.mark.parametrize("path", VECTOR_PATHS, ids=IDS) +def test_no_finding_message_leaks_the_bundle_bytes(path: Path, resolver) -> None: + """A message is forwarded in reports; it names the digest, never the content.""" + finding = _pol003(_load(path)["record"], resolver) + for bundle in (VECTOR_DIR / "policies").glob("*.json"): + body = bundle.read_text(encoding="utf-8").strip() + assert body not in finding.message + + +@pytest.mark.level0 +def test_the_two_unresolvable_vectors_fail_by_different_roads() -> None: + """05 and 11 must be separable, or one of them is a duplicate. + + A resolver that handled only a missing key would leave 11 reporting a + comparison it never made, and one that handled only a missing file would + do the same to 05. The pair is what makes that visible. + """ + manifest = _load(MANIFEST) + no_route = _load(VECTOR_DIR / "05-referent-unreachable-no-route.json") + route_fails = _load(VECTOR_DIR / "11-referent-unreachable-route-fails.json") + + def only_keyerror(uri: str) -> bytes: + return (VECTOR_DIR / manifest[uri]).read_bytes() + + # A resolver that answers for anything the manifest holds, even absent + # files, deviates 11 alone: 05 still has no route. + def tolerating_missing_files(uri: str) -> bytes: + path = VECTOR_DIR / manifest[uri] + return path.read_bytes() if path.exists() else b"" + + a = _pol003(no_route["record"], tolerating_missing_files) + b = _pol003(route_fails["record"], tolerating_missing_files) + assert a.status.value == "unverified", "05 must be untouched by tolerating missing files" + assert b.status.value != "unverified", ( + "11 must deviate when missing files are tolerated, or it is not " + "independent of 05" ) + assert _pol003(route_fails["record"], only_keyerror).status.value == "unverified" @pytest.mark.level0 -@pytest.mark.negative -def test_06_names_an_algorithm_the_schema_does_not_admit() -> None: - vector = _load(VECTOR_DIR / "06-digest-algorithm-uncomputable.json") - value = vector["context"]["candidate_binding"]["value"] - assert not SCHEMA_DIGEST_PATTERN.match(value), ( - "06 claims the algorithm is uncomputable, but the digest matches the " - "pattern the schema admits" +def test_the_two_malformed_vectors_fail_by_different_roads() -> None: + """08 and 10 must be separable for the same reason.""" + eight = _load(VECTOR_DIR / "08-policy-uri-is-a-relative-reference.json") + ten = _load(VECTOR_DIR / "10-policy-uri-carries-a-space.json") + assert " " not in eight["record"]["policy"]["policy_uri"], ( + "08 must be caught by the scheme rule alone, so it carries no whitespace" + ) + assert "://" in ten["record"]["policy"]["policy_uri"], ( + "10 must be caught by the character rule alone, so its scheme is valid" ) - assert vector["context"]["resolution"]["outcome"] == "resolved", ( - "06 must reach its referent; that is what separates it from 05" + + +@pytest.mark.level0 +def test_the_two_sha384_vectors_separate_accept_from_compare() -> None: + """06 and 09 must be separable, or sha384 support sits on one outcome.""" + six = _load(VECTOR_DIR / "06-resolved-and-matches-sha384.json") + nine = _load(VECTOR_DIR / "09-sha384-bound-to-other-referent.json") + assert six["record"]["policy"]["bundle_hash"].startswith("sha384:") + assert nine["record"]["policy"]["bundle_hash"].startswith("sha384:") + assert six["expected"]["tr_pol_003"] == "pass" + assert nine["expected"]["tr_pol_003"] == "fail", ( + "without a sha384 vector that must fail, a verifier accepting sha384 " + "without comparing would pass the set" ) @pytest.mark.level0 def test_03_and_04_differ_in_kind_not_just_in_bytes() -> None: """The pair that keeps the contradicted boundary off a single vector.""" + manifest = _load(MANIFEST) minimal = _load(VECTOR_DIR / "03-digest-mismatch-minimal-mutation.json") wholesale = _load(VECTOR_DIR / "04-digest-mismatch-different-object.json") + base = (VECTOR_DIR / "policies" / "policy-bundle-base.json").read_bytes() + a = (VECTOR_DIR / manifest[minimal["context"]["cited_uri"]]).read_bytes() + b = (VECTOR_DIR / manifest[wholesale["context"]["cited_uri"]]).read_bytes() - a = (VECTOR_DIR / minimal["context"]["resolution"]["file"]).read_bytes() - b = (VECTOR_DIR / "policies" / "policy-bundle-base.json").read_bytes() - assert len(a) == len(b), "03's mutation should not change the object's length" - assert sum(x != y for x, y in zip(a, b)) == 1, ( - "03 is the minimal-mutation vector; it must differ from the appraised " - "object in exactly one byte" + assert len(a) == len(base) and sum(x != y for x, y in zip(a, base)) == 1, ( + "03 must be the minimal mutation: one byte from the baseline" ) + assert len(b) != len(base), "04 must be a different object, not an edit of the baseline" - c = (VECTOR_DIR / wholesale["context"]["resolution"]["file"]).read_bytes() - assert len(c) != len(b), ( - "04 is the wholesale-substitution vector; a verifier comparing lengths " - "must be able to tell it from 03" - ) + +@pytest.mark.level0 +def test_the_declared_digests_are_the_digests_of_the_bytes_on_disk() -> None: + """Nothing in the set is a digest someone typed.""" + manifest = _load(MANIFEST) + algos = {"sha256:": hashlib.sha256, "sha384:": hashlib.sha384} + checked = 0 + for path in VECTOR_PATHS: + vector = _load(path) + if vector["context"]["resolution"]["outcome"] != "resolved": + continue + declared = vector["record"]["policy"]["bundle_hash"] + prefix = declared.split(":")[0] + ":" + body = (VECTOR_DIR / manifest[vector["context"]["cited_uri"]]).read_bytes() + actual = prefix + algos[prefix](body).hexdigest() + expected_match = vector["expected"]["tr_pol_003"] == "pass" + assert (declared == actual) is expected_match, ( + f"{path.name}: the declared digest and the bytes on disk disagree with " + f"the expected outcome" + ) + checked += 1 + assert checked >= 5, "this check must not degrade into a pass over no work" diff --git a/tests/test_policy_resolution_cli.py b/tests/test_policy_resolution_cli.py new file mode 100644 index 0000000..0be379b --- /dev/null +++ b/tests/test_policy_resolution_cli.py @@ -0,0 +1,169 @@ +"""What supplying a resolver changes, end to end, and at which level. + +Two paths publish a verdict: the CLI's exit code and the report artifact. Both +are exercised here on the same record, by running the suite and reading what it +produced. Nothing constructs a ``Finding``: a test that builds its own finding +and hands it to the reporting layer tests whichever code it happened to name, +and stays green while the table it is supposed to be guarding is wrong. + +The measurement is a difference. The record is unsigned and carries no +transcript, so it fails for reasons that have nothing to do with policy +resolution — constant across both legs, and cancelling in the delta. What is +left is the one finding under test. +""" + +from __future__ import annotations + +import json +import re +from pathlib import Path + +import pytest +from click.testing import CliRunner + +from trace_tests import report as report_mod +from trace_tests.cli import main +from trace_tests.runner import run + +VECTOR_DIR = Path(__file__).parent / "vectors" / "policy-resolution" +UNREACHABLE = VECTOR_DIR / "05-referent-unreachable-no-route.json" +RESOLVES = VECTOR_DIR / "02-resolved-and-matches.json" +MALFORMED = VECTOR_DIR / "08-policy-uri-is-a-relative-reference.json" + +# A century, so the frozen iat in the vectors cannot confound the exit code. +# The alternative is refreshing iat, which would move the committed bytes and +# break the set's own byte-reproduction guard. +MAX_AGE = "3153600000" +NONCE = "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" + +_RESULT = re.compile(r"Result: (PASS|FAIL)\s+\((\d+) checks(?:, (\d+) failure\(s\))?") + + +def _record_path(tmp_path: Path, vector: Path) -> str: + record = json.loads(vector.read_text(encoding="utf-8"))["record"] + p = tmp_path / "record.json" + p.write_text(json.dumps(record), encoding="utf-8") + return str(p) + + +def _verify(record_path: str, level: int, *, with_policy_dir: bool): + args = [ + "verify", "--record", record_path, "--level", str(level), + "--max-age", MAX_AGE, "--expected-nonce", NONCE, + ] + if with_policy_dir: + args += ["--policy-dir", str(VECTOR_DIR)] + return CliRunner().invoke(main, args) + + +def _failures(output: str) -> int: + match = _RESULT.search(output) + assert match, f"no result line in output:\n{output}" + return int(match.group(3) or 0) + + +def _resolver(): + manifest = json.loads((VECTOR_DIR / "resolutions.json").read_text(encoding="utf-8")) + + def _resolve(uri: str) -> bytes: + return (VECTOR_DIR / manifest[uri]).read_bytes() + + return _resolve + + +@pytest.mark.level0 +@pytest.mark.parametrize("level,delta", [(0, 0), (1, 0), (2, 1)]) +def test_supplying_a_resolver_changes_the_failure_count_only_from_level_two( + tmp_path, level, delta +): + """The registered level, observed through the CLI rather than asserted. + + An unreachable bundle is unverified at every level. Whether that *counts* + is the table's business, and TR-POL-003 is registered at 2, so the same + finding is tolerated at 0 and 1 and fails at 2. + """ + record = _record_path(tmp_path, UNREACHABLE) + without = _verify(record, level, with_policy_dir=False) + with_dir = _verify(record, level, with_policy_dir=True) + + assert _failures(with_dir.output) - _failures(without.output) == delta, ( + f"level {level}\n--- without --policy-dir ---\n{without.output}" + f"\n--- with --policy-dir ---\n{with_dir.output}" + ) + + +@pytest.mark.level0 +@pytest.mark.parametrize("with_policy_dir", [False, True]) +def test_level_zero_exits_zero_either_way(tmp_path, with_policy_dir): + """An unresolvable bundle must not sink a Level 0 record. + + The record is unsigned, so TR-SIG-005 is unverified here too. Both codes + are tolerated at Level 0 and neither is a pass; the exit code is the + difference between "not proven" and "wrong". + """ + record = _record_path(tmp_path, UNREACHABLE) + result = _verify(record, 0, with_policy_dir=with_policy_dir) + assert result.exit_code == 0, result.output + assert "UNVERIFIED" in result.output + + +@pytest.mark.level0 +def test_a_resolvable_bundle_reports_a_pass_through_the_cli(tmp_path): + """The must-accept leg: supplying a resolver must not cost a record anything.""" + record = _record_path(tmp_path, RESOLVES) + with_dir = _verify(record, 0, with_policy_dir=True) + assert with_dir.exit_code == 0, with_dir.output + assert "TR-POL-003" not in with_dir.output, ( + "a passing TR-POL-003 names no code in its message, by house convention" + ) + + +@pytest.mark.level0 +@pytest.mark.parametrize("level", [0, 1, 2]) +def test_a_malformed_reference_fails_with_and_without_a_resolver(tmp_path, level): + """The ordering, observed: a record defect is not conditional on the network. + + This is the constant the differential rests on. If the malformed check ran + after the resolver check, this record would report nothing at all offline, + and an offline verification would be blind to a defect on the record's face. + """ + record = _record_path(tmp_path, MALFORMED) + for with_dir in (False, True): + result = _verify(record, level, with_policy_dir=with_dir) + assert result.exit_code == 1, result.output + assert "TR-POL-003" in result.output + assert "relative reference" in result.output + + +@pytest.mark.level0 +@pytest.mark.parametrize("level,delta", [(0, 0), (1, 0), (2, 1)]) +def test_the_report_path_carries_the_same_deltas(level, delta): + """The second publisher. A report that disagreed with the CLI would be worse + than no report, so the two are measured against the same record.""" + record = json.loads(UNREACHABLE.read_text(encoding="utf-8"))["record"] + + def build(resolver): + by_level = { + lv: run(record, "trace", lv, max_age_seconds=10**9, + expected_nonce=NONCE, policy_resolver=resolver) + for lv in range(level + 1) + } + return report_mod.build( + record=record, + record_path=str(UNREACHABLE), + record_format="trace", + results_by_level=by_level, + suite_version="test", + library_version=None, + generated_at="1970-01-01T00:00:00Z", + ) + + without = build(None) + with_resolver = build(_resolver()) + got = _level_failures(with_resolver, level) - _level_failures(without, level) + assert got == delta, f"level {level}: report delta {got}, expected {delta}" + + +def _level_failures(report, level: int) -> int: + entry = next(lv for lv in report.levels if lv.level == level) + return entry.failures diff --git a/tests/test_policy_resolution_completeness.py b/tests/test_policy_resolution_completeness.py index b4c7fab..ec29e36 100644 --- a/tests/test_policy_resolution_completeness.py +++ b/tests/test_policy_resolution_completeness.py @@ -13,8 +13,7 @@ 1. A set must fail BOTH unconditional implementations. A set of all-rejections is passed by a verifier that rejects everything, exactly as a set of all-acceptances is passed by one that accepts - everything. This set's three decided-reject vectors would, alone, be the - first failure. Vectors 01 and 02 exist to close it. + everything. Vectors 01, 02 and 06 close the second half. 2. Every boundary needs more than one vector. One vector cannot separate a check that reads a prefix from one that reads the whole object. @@ -24,9 +23,12 @@ 4. Shortfalls are recorded exactly. See KNOWN_SHORTFALLS below. -These tests grade the set, not a verifier. No verifier resolves -``appraisal.policy_ref`` today — that is the gap the set documents — so nothing -here claims an implementation was exercised. +THE UNIT OF MEASUREMENT + Each vector's expected value is the status of the TR-POL-003 finding, not + a verdict on the whole record. That is narrower than a record-level pass + or fail on purpose: these records are unsigned, so they carry findings + from other modules that have nothing to do with the check under test. + Every claim about margin below is a claim about that one finding. """ from __future__ import annotations @@ -39,22 +41,26 @@ VECTOR_DIR = Path(__file__).parent / "vectors" / "policy-resolution" -DECIDED_OUTCOMES = {"pass", "reject"} -ALL_OUTCOMES = DECIDED_OUTCOMES | {"deferred"} +ACCEPTING = {"pass", "skip"} +REJECTING = {"fail"} +ALL_STATUSES = ACCEPTING | REJECTING | {"unverified"} +BOUNDARIES = {"accept", "contradicted", "unresolvable", "malformed"} # Criterion 4: shortfalls asserted to their exact extent, so they cannot widen # quietly. Delete an entry when the shortfall is closed, not when it is excused. KNOWN_SHORTFALLS = { - "no_verifier_exercised": ( - "No implementation resolves appraisal.policy_ref, so every expected " - "outcome is a claim about what a verifier should do, not a recording of " - "what one did. The set is a specification argument with runnable " - "internal consistency, not a conformance run." - ), "unresolvable_outcome_unnamed": ( - "Vectors 05 and 06 assert only that the outcome is not affirming. The " - "value a verifier should record is open on agentrust-io/trace-spec#190 " - "and is deliberately not proposed here." + "Vectors 05 and 11 assert that the check did not run, which is a fact " + "about the run. The level at which that stops being tolerable is suite " + "policy in src/trace_tests/modules/unverified.py, not a reading of any " + "merged text: policy_uri appears nowhere in the specification. Tracked " + "on agentrust-io/trace-spec#190." + ), + "no_network_resolution": ( + "The resolver in every test here reads bytes from disk. Nothing in this " + "set exercises an HTTP fetch, a redirect, a timeout, or a TLS failure, " + "so the mapping from real-world retrieval failures onto the unverified " + "status is asserted rather than measured." ), } @@ -66,19 +72,19 @@ def _vectors() -> list[dict]: return out -def _outcome(vector: dict) -> str: - return vector["expected"]["outcome"] +def _status(vector: dict) -> str: + return vector["expected"]["tr_pol_003"] @pytest.mark.level0 def test_the_set_is_not_empty() -> None: - assert len(_vectors()) == 7, "the set is seven vectors, 01 through 07" + assert len(_vectors()) == 11, "the set is eleven vectors, 01 through 11" @pytest.mark.level0 def test_criterion_1_the_set_fails_an_accept_everything_verifier() -> None: """At least one vector a conformant verifier must reject.""" - rejects = [v for v in _vectors() if _outcome(v) == "reject"] + rejects = [v for v in _vectors() if _status(v) in REJECTING] assert rejects, ( "every vector expects acceptance, so a verifier that accepts " "unconditionally passes the set" @@ -91,17 +97,30 @@ def test_criterion_1_the_set_fails_a_reject_everything_verifier() -> None: This is the criterion the set would otherwise fail. Its subject is a family of resolution failures, so every vector written from the motivating problem - alone is a reject or a deferral. + alone is a reject. """ - accepts = [v for v in _vectors() if _outcome(v) == "pass"] + accepts = [v for v in _vectors() if _status(v) in ACCEPTING] assert accepts, ( "no vector expects acceptance, so a verifier that rejects " "unconditionally passes the set" ) assert len(accepts) >= 2, ( "one must-accept vector cannot separate a verifier that accepts only " - "records declaring no binding from one that also checks a matching " - "binding; 01 and 02 are that pair" + "records declaring no policy_uri from one that also checks a matching " + "digest; 01 and 02 are that pair" + ) + + +@pytest.mark.level0 +def test_criterion_1_a_verifier_that_never_resolves_does_not_pass() -> None: + """A verifier that answers "skip" to everything must fail the set too. + + Without this, a module that returned SKIP unconditionally would satisfy + both halves of criterion 1, because SKIP reads as acceptance. + """ + non_skip = [v for v in _vectors() if _status(v) != "skip"] + assert len(non_skip) >= 2, ( + "a verifier that skips every record would pass this set unnoticed" ) @@ -110,9 +129,7 @@ def test_criterion_2_every_boundary_carries_at_least_two_vectors() -> None: counts = Counter(v["boundary"] for v in _vectors()) thin = {b: n for b, n in counts.items() if n < 2} assert not thin, f"boundaries carried by a single vector: {thin}" - assert set(counts) == {"accept", "contradicted", "unresolvable"}, ( - f"unexpected boundary set: {sorted(counts)}" - ) + assert set(counts) == BOUNDARIES, f"unexpected boundary set: {sorted(counts)}" @pytest.mark.level0 @@ -125,70 +142,101 @@ def test_no_two_vectors_share_a_defect() -> None: @pytest.mark.level0 def test_every_expected_block_is_well_formed() -> None: for v in _vectors(): - name = v["name"] exp = v["expected"] - assert exp["outcome"] in ALL_OUTCOMES, f"{name}: bad outcome {exp['outcome']!r}" - assert exp.get("reason"), f"{name}: expected block carries no reason" - if exp["outcome"] == "deferred": - assert exp.get("deferred_pending") == "agentrust-io/trace-spec#190", ( - f"{name}: a deferred vector must name the issue it defers to" + assert set(exp) == {"tr_pol_003", "reason"}, ( + f"{v['name']}: an expected block is a status and the reason for it, " + f"nothing else; got {sorted(exp)}" + ) + assert exp["tr_pol_003"] in ALL_STATUSES, ( + f"{v['name']}: {exp['tr_pol_003']!r} is not a finding status" + ) + assert exp["reason"].strip(), f"{v['name']}: an expected status needs a reason" + + +@pytest.mark.level0 +def test_every_vector_carries_an_anchor_naming_its_tier() -> None: + """An expected outcome that names no source is an assertion, not a derivation. + + Tier 1 is merged specification prose. Tier 2 is the packaged schema, which + this suite tests conformance against but which is not the specification. + Tier 0 means no text governs the case, and a vector claiming it must say + what it asserts instead. + """ + for v in _vectors(): + anchor = v["context"]["anchor"] + assert set(anchor) == {"tier", "source", "text", "applies"}, ( + f"{v['name']}: malformed anchor {sorted(anchor)}" + ) + assert anchor["tier"] in (0, 1, 2), f"{v['name']}: unknown tier" + assert anchor["applies"].strip(), f"{v['name']}: an anchor must say how it applies" + if anchor["tier"] == 0: + assert anchor["text"] == "", ( + f"{v['name']}: tier 0 means no governing text, so quoting one is a " + "contradiction" ) - assert exp.get("must_not") == "affirming", ( - f"{name}: a deferred vector must still assert what it may not be" + assert "no merged sentence" in anchor["source"], ( + f"{v['name']}: tier 0 must say plainly that nothing governs the case" ) else: - assert "deferred_pending" not in exp, ( - f"{name}: a decided vector must not carry a deferral pointer" - ) + assert anchor["text"].strip(), f"{v['name']}: tier {anchor['tier']} must quote text" + + +@pytest.mark.level0 +def test_the_unresolvable_vectors_do_not_claim_a_specification_requirement() -> None: + """The one thing this set must never do: invent authority it does not have.""" + for v in _vectors(): + if v["boundary"] != "unresolvable": + continue + assert v["context"]["anchor"]["tier"] == 0, ( + f"{v['name']}: no merged text governs an unresolvable policy_uri" + ) + assert "suite policy" in v["context"]["anchor"]["applies"], ( + f"{v['name']}: the level must be named as suite policy, not as a " + "specification requirement" + ) @pytest.mark.level0 def test_no_vector_proposes_an_appraisal_status_value() -> None: - """The set must not coin the vocabulary trace-spec#190 exists to decide. + """The set answers a policy-binding question and must not reach for another. - ``deferred`` is fixture bookkeeping. It must never appear as an - ``appraisal.status``, and no vector may assert a status for the - unresolvable case. + ``appraisal.status`` values are decided one layer up, on + agentrust-io/trace-spec#190. A vector that named one would be proposing a + schema change under cover of a fixture. """ - schema_enum = {"affirming", "warning", "contraindicated", "none"} for v in _vectors(): - status = v["record"]["appraisal"]["status"] - assert status in schema_enum, f"{v['name']}: invented status {status!r}" - exp = v["expected"] - assert exp["outcome"] not in schema_enum, ( - f"{v['name']}: expected.outcome reuses an appraisal.status value, " - "which reads as proposing that value for this case" - ) - if exp["outcome"] == "deferred": - assert "status" not in exp, ( - f"{v['name']}: a deferred vector must not name a status to record" + blob = json.dumps(v) + for value in ("contraindicated", "no_verifier_exercised"): + assert value not in blob, ( + f"{v['name']} names {value!r}; this set does not propose " + "appraisal.status values" ) + assert v["record"]["appraisal"]["status"] == "affirming", ( + f"{v['name']}: appraisal.status is fixed across the set" + ) @pytest.mark.level0 -def test_candidate_binding_is_marked_candidate_everywhere_it_appears() -> None: +def test_appraisal_policy_ref_is_identical_across_the_set() -> None: + """The retarget's whole point: the varying field is policy.*, not appraisal.*.""" + refs = {v["record"]["appraisal"]["policy_ref"] for v in _vectors()} + assert len(refs) == 1, f"appraisal.policy_ref varies across the set: {sorted(refs)}" + + +@pytest.mark.level0 +def test_every_record_is_identical_outside_the_policy_block() -> None: + seen = [] for v in _vectors(): - binding = v["context"].get("candidate_binding") - if binding is None: - continue - assert binding["field"].startswith("CANDIDATE:"), ( - f"{v['name']}: the binding field must be marked CANDIDATE, so it is " - "never read as a proposed schema field" - ) + rec = {k: val for k, val in v["record"].items() if k != "policy"} + seen.append(json.dumps(rec, sort_keys=True)) + assert len(set(seen)) == 1, ( + "records differ outside policy.*, so a finding could be answering " + "something other than the check under test" + ) @pytest.mark.level0 def test_known_shortfalls_are_recorded_not_silent() -> None: - """Criterion 4: the gaps are asserted to their exact extent.""" - assert set(KNOWN_SHORTFALLS) == { - "no_verifier_exercised", - "unresolvable_outcome_unnamed", - }, ( - "the recorded shortfalls changed; update the README in the same commit " - "so the set never claims more coverage than it has" - ) - deferred = [v for v in _vectors() if _outcome(v) == "deferred"] - assert len(deferred) == 2, ( - "unresolvable_outcome_unnamed is recorded as covering exactly two " - f"vectors; found {len(deferred)}" - ) + assert KNOWN_SHORTFALLS, "a set with no recorded shortfalls is claiming to have none" + for key, text in KNOWN_SHORTFALLS.items(): + assert text.strip(), f"{key} is recorded with no explanation" diff --git a/tests/vectors/policy-resolution/01-no-binding-declared.json b/tests/vectors/policy-resolution/01-no-policy-uri.json similarity index 54% rename from tests/vectors/policy-resolution/01-no-binding-declared.json rename to tests/vectors/policy-resolution/01-no-policy-uri.json index f014370..2f02f65 100644 --- a/tests/vectors/policy-resolution/01-no-binding-declared.json +++ b/tests/vectors/policy-resolution/01-no-policy-uri.json @@ -1,9 +1,8 @@ { - "name": "no-binding-declared", - "description": "The record cites an appraisal policy by bare URI and declares no binding to what that URI held. This is every conformant record today, so it must keep verifying: a set that rejected it would be proposing a breaking change rather than describing a gap.", + "name": "no-policy-uri", + "description": "The record declares a bundle digest and no location to fetch the bundle from. This is every conformant record today, so it must keep verifying: a set that failed it would be proposing a breaking change rather than describing a gap.", "boundary": "accept", "defect": "none - backward-compatibility control", - "spec": "Mirrors the merged treatment of an undeclared depth in agentrust-io/trace-spec#173, where a record that never declared is read as surface rather than as a failure.", "record": { "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", "iat": 1748000000, @@ -14,10 +13,11 @@ }, "runtime": { "platform": "intel-tdx", - "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" }, "policy": { - "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", + "bundle_hash": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc", "enforcement_mode": "enforce" }, "data_class": "confidential", @@ -28,7 +28,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,15 +43,19 @@ } }, "context": { - "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", + "cited_uri": null, "resolution": { - "outcome": "not_attempted", - "note": "No binding is declared, so there is nothing to check." + "outcome": "not_attempted" }, - "candidate_binding": null + "anchor": { + "tier": 2, + "source": "schemas/trace-claim.json properties.policy.required", + "text": "required: [bundle_hash, enforcement_mode]", + "applies": "policy_uri is absent from the required list, so a record that omits it is conformant and the check has nothing to compare." + } }, "expected": { - "outcome": "pass", - "reason": "No binding declared; nothing to contradict." + "tr_pol_003": "skip", + "reason": "No policy_uri is declared, so there is nothing to resolve." } } diff --git a/tests/vectors/policy-resolution/02-resolved-and-matches.json b/tests/vectors/policy-resolution/02-resolved-and-matches.json index 07c51fa..0c6b5df 100644 --- a/tests/vectors/policy-resolution/02-resolved-and-matches.json +++ b/tests/vectors/policy-resolution/02-resolved-and-matches.json @@ -1,9 +1,8 @@ { "name": "resolved-and-matches", - "description": "The cited URI resolves and its bytes hash to exactly the declared candidate binding. The second must-accept vector, and the only one in which a resolution actually succeeds.", + "description": "The bundle at policy_uri hashes to exactly what bundle_hash declares. The positive control: a set that could not pass this would reject every honest record.", "boundary": "accept", - "defect": "none - positive control", - "spec": "agentrust-io/trace-spec docs/verification.md - a verifier records what it actually resolved, and evidence that resolves and contradicts fails the appraisal. Applied here to appraisal.policy_ref, which carries no digest.", + "defect": "none - positive control, sha256", "record": { "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", "iat": 1748000000, @@ -14,11 +13,13 @@ }, "runtime": { "platform": "intel-tdx", - "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" }, "policy": { - "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", - "enforcement_mode": "enforce" + "bundle_hash": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc", + "enforcement_mode": "enforce", + "policy_uri": "https://policy.example.org/bundles/policy-bundle-base.json" }, "data_class": "confidential", "build_provenance": { @@ -28,7 +29,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -45,17 +46,17 @@ "context": { "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", "resolution": { - "outcome": "resolved", - "file": "policies/policy-bundle-base.json", - "actual_sha256": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + "outcome": "resolved" }, - "candidate_binding": { - "field": "CANDIDATE:appraisal.policy_digest", - "value": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + "anchor": { + "tier": 1, + "source": "agentrust-io/trace-spec spec/trace-v0.2.md 4.3", + "text": "the policy bundle hash is sealed to the TEE measurement, the enforcement mode is recorded, and substituting the policy invalidates the runtime claim", + "applies": "Read the other way: the bundle that resolves is the bundle that was sealed, so nothing was substituted and the claim stands." } }, "expected": { - "outcome": "pass", - "reason": "Declared binding equals the digest of what resolved." + "tr_pol_003": "pass", + "reason": "The resolved bytes hash to the declared digest." } } diff --git a/tests/vectors/policy-resolution/03-digest-mismatch-minimal-mutation.json b/tests/vectors/policy-resolution/03-digest-mismatch-minimal-mutation.json index f833507..e3fd69d 100644 --- a/tests/vectors/policy-resolution/03-digest-mismatch-minimal-mutation.json +++ b/tests/vectors/policy-resolution/03-digest-mismatch-minimal-mutation.json @@ -1,9 +1,8 @@ { "name": "digest-mismatch-minimal-mutation", - "description": "The cited URI now serves a document one character from the one appraised: the SLSA floor moved from 2 to 3. The record still declares the original digest. A verifier that compares only the URI sees no change; a verifier that compares bytes sees the substitution that flips this record's verdict.", + "description": "The cited bundle was edited by one character after the digest was taken. The smallest change that alters what the policy permits, and the one a reader is least likely to notice.", "boundary": "contradicted", - "defect": "cited object mutated minimally after appraisal", - "spec": "agentrust-io/trace-spec docs/verification.md - a verifier records what it actually resolved, and evidence that resolves and contradicts fails the appraisal. Applied here to appraisal.policy_ref, which carries no digest.", + "defect": "cited object mutated minimally after the digest was taken", "record": { "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", "iat": 1748000000, @@ -14,11 +13,13 @@ }, "runtime": { "platform": "intel-tdx", - "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" }, "policy": { - "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", - "enforcement_mode": "enforce" + "bundle_hash": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc", + "enforcement_mode": "enforce", + "policy_uri": "https://policy.example.org/bundles/policy-bundle-onebyte.json" }, "data_class": "confidential", "build_provenance": { @@ -28,7 +29,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,20 +44,19 @@ } }, "context": { - "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-onebyte.json", "resolution": { - "outcome": "resolved", - "file": "policies/policy-bundle-onebyte.json", - "actual_sha256": "sha256:1248932d38677845463158e5b93b7d1e3f78b1fffdef6d23e25c1d2926d5e3d5" + "outcome": "resolved" }, - "candidate_binding": { - "field": "CANDIDATE:appraisal.policy_digest", - "value": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" - }, - "note": "policies/policy-bundle-base.json and policies/policy-bundle-onebyte.json differ in one character." + "anchor": { + "tier": 1, + "source": "agentrust-io/trace-spec spec/trace-v0.2.md 4.3", + "text": "the policy bundle hash is sealed to the TEE measurement, the enforcement mode is recorded, and substituting the policy invalidates the runtime claim", + "applies": "A resolved bundle whose digest is not the declared one is a substituted policy, so the runtime claim does not stand." + } }, "expected": { - "outcome": "reject", - "reason": "Resolved bytes contradict the declared binding." + "tr_pol_003": "fail", + "reason": "The resolved bytes contradict the declared digest." } } diff --git a/tests/vectors/policy-resolution/04-digest-mismatch-different-object.json b/tests/vectors/policy-resolution/04-digest-mismatch-different-object.json index 61045c9..3df4fce 100644 --- a/tests/vectors/policy-resolution/04-digest-mismatch-different-object.json +++ b/tests/vectors/policy-resolution/04-digest-mismatch-different-object.json @@ -1,9 +1,8 @@ { "name": "digest-mismatch-different-object", - "description": "The cited URI resolves to an unrelated policy document. Paired with 03 so the boundary is not carried by a single vector: a verifier that only samples a prefix of the document, or compares lengths, passes one of these two and fails the other.", + "description": "The cited bundle was replaced wholesale after the digest was taken. Paired with 03 so the check cannot be satisfied by a heuristic that only notices large differences, or only small ones.", "boundary": "contradicted", - "defect": "cited object wholly replaced after appraisal", - "spec": "agentrust-io/trace-spec docs/verification.md - a verifier records what it actually resolved, and evidence that resolves and contradicts fails the appraisal. Applied here to appraisal.policy_ref, which carries no digest.", + "defect": "cited object wholly replaced after the digest was taken", "record": { "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", "iat": 1748000000, @@ -14,11 +13,13 @@ }, "runtime": { "platform": "intel-tdx", - "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" }, "policy": { - "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", - "enforcement_mode": "enforce" + "bundle_hash": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc", + "enforcement_mode": "enforce", + "policy_uri": "https://policy.example.org/bundles/policy-bundle-other.json" }, "data_class": "confidential", "build_provenance": { @@ -28,7 +29,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,19 +44,19 @@ } }, "context": { - "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-other.json", "resolution": { - "outcome": "resolved", - "file": "policies/policy-bundle-unrelated.json", - "actual_sha256": "sha256:23905737177574be26e1212264b88ffe81e06e99aab1fe6115ebbb2e9205109c" + "outcome": "resolved" }, - "candidate_binding": { - "field": "CANDIDATE:appraisal.policy_digest", - "value": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + "anchor": { + "tier": 1, + "source": "agentrust-io/trace-spec spec/trace-v0.2.md 4.3", + "text": "the policy bundle hash is sealed to the TEE measurement, the enforcement mode is recorded, and substituting the policy invalidates the runtime claim", + "applies": "A resolved bundle whose digest is not the declared one is a substituted policy, so the runtime claim does not stand." } }, "expected": { - "outcome": "reject", - "reason": "Resolved bytes contradict the declared binding." + "tr_pol_003": "fail", + "reason": "The resolved bytes contradict the declared digest." } } diff --git a/tests/vectors/policy-resolution/05-referent-unreachable-no-route.json b/tests/vectors/policy-resolution/05-referent-unreachable-no-route.json new file mode 100644 index 0000000..a02e76a --- /dev/null +++ b/tests/vectors/policy-resolution/05-referent-unreachable-no-route.json @@ -0,0 +1,62 @@ +{ + "name": "referent-unreachable-no-route", + "description": "The cited URI is not one the verifier has any route to. Nothing was contradicted, because nothing was read. This vector asserts only that the comparison did not happen.", + "boundary": "unresolvable", + "defect": "referent unreachable: no route to the cited URI", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" + }, + "policy": { + "bundle_hash": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc", + "enforcement_mode": "enforce", + "policy_uri": "https://policy.example.org/bundles/policy-bundle-withdrawn.json" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/bundles/policy-bundle-withdrawn.json", + "resolution": { + "outcome": "unreachable_no_route" + }, + "anchor": { + "tier": 0, + "source": "no merged sentence governs this case", + "text": "", + "applies": "No text in the specification says what a verifier records when a policy bundle cannot be fetched: policy_uri appears nowhere in the merged specification, and the reference block that does discuss resolvability governs the references field rather than this one. The vector therefore asserts only that the check was not performed, which is a fact about the run rather than a reading of any text. The level at which that stops being tolerable is suite policy, registered in src/trace_tests/modules/unverified.py and docs/levels.md, and is not claimed here as a specification requirement." + } + }, + "expected": { + "tr_pol_003": "unverified", + "reason": "The bundle could not be fetched, so the digest was never compared. Reporting that as a pass would claim a check that did not run." + } +} diff --git a/tests/vectors/policy-resolution/05-referent-unreachable.json b/tests/vectors/policy-resolution/05-referent-unreachable.json deleted file mode 100644 index ee70b8f..0000000 --- a/tests/vectors/policy-resolution/05-referent-unreachable.json +++ /dev/null @@ -1,63 +0,0 @@ -{ - "name": "referent-unreachable", - "description": "The cited URI does not resolve; no object under policies/ corresponds to it. Nothing was contradicted, because nothing was read. This vector asserts only that the outcome is not affirming.", - "boundary": "unresolvable", - "defect": "referent unreachable at verification time", - "spec": "agentrust-io/trace-spec docs/verification.md - evidence that does not resolve downgrades honestly rather than failing. Which value records that for appraisal.policy_ref is open: agentrust-io/trace-spec#190. This vector asserts only what the outcome may NOT be.", - "record": { - "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", - "iat": 1748000000, - "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", - "model": { - "provider": "anthropic", - "model_id": "claude-sonnet-4-5" - }, - "runtime": { - "platform": "intel-tdx", - "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" - }, - "policy": { - "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", - "enforcement_mode": "enforce" - }, - "data_class": "confidential", - "build_provenance": { - "slsa_level": 2, - "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" - }, - "appraisal": { - "status": "affirming", - "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/bundles/policy-bundle-withdrawn.json", - "timestamp": 1748000042 - }, - "transparency": "https://scitt.example.org/receipts/abc123def456", - "cnf": { - "jwk": { - "kty": "EC", - "crv": "P-256", - "x": "dGhpcyBpcyBhIHRlc3QgeA", - "y": "dGhpcyBpcyBhIHRlc3QgeQ", - "kid": "tee-key-001" - } - } - }, - "context": { - "cited_uri": "https://policy.example.org/bundles/policy-bundle-withdrawn.json", - "resolution": { - "outcome": "unreachable", - "file": null, - "note": "No sibling file corresponds to this URI, by construction." - }, - "candidate_binding": { - "field": "CANDIDATE:appraisal.policy_digest", - "value": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" - } - }, - "expected": { - "outcome": "deferred", - "deferred_pending": "agentrust-io/trace-spec#190", - "must_not": "affirming", - "reason": "Reporting a check that was never performed as affirming is the failure this set exists to name. Which value is reported instead is not decided here." - } -} diff --git a/tests/vectors/policy-resolution/06-digest-algorithm-uncomputable.json b/tests/vectors/policy-resolution/06-digest-algorithm-uncomputable.json deleted file mode 100644 index 0226b9d..0000000 --- a/tests/vectors/policy-resolution/06-digest-algorithm-uncomputable.json +++ /dev/null @@ -1,64 +0,0 @@ -{ - "name": "digest-algorithm-uncomputable", - "description": "The referent resolves, but the declared binding names a digest algorithm outside the set this profile's schema admits (sha256 and sha384). The verifier cannot compute the comparison, so it has not checked - which is distinct from having checked and disagreed. Paired with 05: one cannot reach the object, the other reaches it and cannot compute over it.", - "boundary": "unresolvable", - "defect": "digest algorithm the verifier cannot compute", - "spec": "agentrust-io/trace-spec docs/verification.md - evidence that does not resolve downgrades honestly rather than failing. Which value records that for appraisal.policy_ref is open: agentrust-io/trace-spec#190. This vector asserts only what the outcome may NOT be. Deliberately rhymes with the unverifiable / digest_algorithm_unsupported classification PROPOSED in agentrust-io/trace-spec#184 - an explicitly non-normative draft, a surface facing this question rather than an established answer to it - without adopting its vocabulary.", - "record": { - "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", - "iat": 1748000000, - "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", - "model": { - "provider": "anthropic", - "model_id": "claude-sonnet-4-5" - }, - "runtime": { - "platform": "intel-tdx", - "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" - }, - "policy": { - "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", - "enforcement_mode": "enforce" - }, - "data_class": "confidential", - "build_provenance": { - "slsa_level": 2, - "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" - }, - "appraisal": { - "status": "affirming", - "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/bundles/policy-bundle-other.json", - "timestamp": 1748000042 - }, - "transparency": "https://scitt.example.org/receipts/abc123def456", - "cnf": { - "jwk": { - "kty": "EC", - "crv": "P-256", - "x": "dGhpcyBpcyBhIHRlc3QgeA", - "y": "dGhpcyBpcyBhIHRlc3QgeQ", - "kid": "tee-key-001" - } - } - }, - "context": { - "cited_uri": "https://policy.example.org/bundles/policy-bundle-other.json", - "resolution": { - "outcome": "resolved", - "file": "policies/policy-bundle-other.json", - "actual_sha256": "sha256:9dd48a5dc2efc67c3873a38c95aae17bc49345b48d09e6629f177ac0a5ff79cc" - }, - "candidate_binding": { - "field": "CANDIDATE:appraisal.policy_digest", - "value": "sha3-512:49c36096f3c7cf58d58324f53502f427a72b0684776fe6006f8a8b13513a9c40d626893b68102cec8c8c78d871814067bab429a7d09e5ff4ef14147eb35ef9a5", - "note": "This is the correct SHA3-512 of the referent. It is outside the schema's digest pattern ^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$, so no conformant verifier is required to compute it - which is the only reason this vector is undecided." - } - }, - "expected": { - "outcome": "deferred", - "deferred_pending": "agentrust-io/trace-spec#190", - "must_not": "affirming", - "reason": "The comparison was not performed, so the result is not a contradiction. Which value records that is not decided here." - } -} diff --git a/tests/vectors/policy-resolution/06-resolved-and-matches-sha384.json b/tests/vectors/policy-resolution/06-resolved-and-matches-sha384.json new file mode 100644 index 0000000..4564fa3 --- /dev/null +++ b/tests/vectors/policy-resolution/06-resolved-and-matches-sha384.json @@ -0,0 +1,62 @@ +{ + "name": "resolved-and-matches-sha384", + "description": "The same accept as 02, with the digest taken in sha384. The schema admits both algorithms, so a verifier that hardcodes sha256 is wrong rather than merely limited, and this is the vector that says so.", + "boundary": "accept", + "defect": "none - positive control, sha384", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" + }, + "policy": { + "bundle_hash": "sha384:fac86d58e1ef453e686c343d00ceefe1cac7c4a195d5a2fc85e2ca7b54e356cf90a1402950a963ea304297d0d9d33ede", + "enforcement_mode": "enforce", + "policy_uri": "https://policy.example.org/bundles/policy-bundle-base.json" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", + "resolution": { + "outcome": "resolved" + }, + "anchor": { + "tier": 1, + "source": "agentrust-io/trace-spec spec/trace-v0.2.md 4.3", + "text": "the policy bundle hash is sealed to the TEE measurement, the enforcement mode is recorded, and substituting the policy invalidates the runtime claim", + "applies": "Read the other way: the bundle that resolves is the bundle that was sealed, so nothing was substituted and the claim stands." + } + }, + "expected": { + "tr_pol_003": "pass", + "reason": "The resolved bytes hash to the declared sha384 digest." + } +} diff --git a/tests/vectors/policy-resolution/07-digest-bound-to-other-referent.json b/tests/vectors/policy-resolution/07-digest-bound-to-other-referent.json index b425b60..0e5d399 100644 --- a/tests/vectors/policy-resolution/07-digest-bound-to-other-referent.json +++ b/tests/vectors/policy-resolution/07-digest-bound-to-other-referent.json @@ -1,9 +1,8 @@ { "name": "digest-bound-to-other-referent", - "description": "The declared binding is well formed and is the true digest of a real object in this set - version 2 - while policy_ref cites version 1. Both halves are individually valid and the pair is not. A verifier that checks the digest is well formed, or that it matches something it holds, passes this; only one that checks the digest against what this URI resolved to rejects it.", + "description": "The declared digest is a correct digest of some object, just not of the one policy_uri names. A verifier that checks the digest is well formed, or that it matches something it holds, passes this.", "boundary": "contradicted", - "defect": "binding well formed but bound to a different referent", - "spec": "agentrust-io/trace-spec docs/verification.md - a verifier records what it actually resolved, and evidence that resolves and contradicts fails the appraisal. Applied here to appraisal.policy_ref, which carries no digest.", + "defect": "digest well formed but taken over a different object", "record": { "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", "iat": 1748000000, @@ -14,11 +13,13 @@ }, "runtime": { "platform": "intel-tdx", - "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8" + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" }, "policy": { - "bundle_hash": "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", - "enforcement_mode": "enforce" + "bundle_hash": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc", + "enforcement_mode": "enforce", + "policy_uri": "https://policy.example.org/bundles/policy-bundle-unrelated.json" }, "data_class": "confidential", "build_provenance": { @@ -28,7 +29,7 @@ "appraisal": { "status": "affirming", "verifier": "https://verifier.example.org", - "policy_ref": "https://policy.example.org/bundles/policy-bundle-base.json", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", "timestamp": 1748000042 }, "transparency": "https://scitt.example.org/receipts/abc123def456", @@ -43,20 +44,19 @@ } }, "context": { - "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", + "cited_uri": "https://policy.example.org/bundles/policy-bundle-unrelated.json", "resolution": { - "outcome": "resolved", - "file": "policies/policy-bundle-base.json", - "actual_sha256": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc" + "outcome": "resolved" }, - "candidate_binding": { - "field": "CANDIDATE:appraisal.policy_digest", - "value": "sha256:9dd48a5dc2efc67c3873a38c95aae17bc49345b48d09e6629f177ac0a5ff79cc", - "note": "This is the digest of policies/policy-bundle-other.json, which is not what cited_uri resolves to." + "anchor": { + "tier": 1, + "source": "agentrust-io/trace-spec spec/trace-v0.2.md 4.3", + "text": "the policy bundle hash is sealed to the TEE measurement, the enforcement mode is recorded, and substituting the policy invalidates the runtime claim", + "applies": "A resolved bundle whose digest is not the declared one is a substituted policy, so the runtime claim does not stand." } }, "expected": { - "outcome": "reject", + "tr_pol_003": "fail", "reason": "The binding does not describe the object the record cites, even though it describes some object." } } diff --git a/tests/vectors/policy-resolution/08-policy-uri-is-a-relative-reference.json b/tests/vectors/policy-resolution/08-policy-uri-is-a-relative-reference.json new file mode 100644 index 0000000..7ed8425 --- /dev/null +++ b/tests/vectors/policy-resolution/08-policy-uri-is-a-relative-reference.json @@ -0,0 +1,62 @@ +{ + "name": "policy-uri-is-a-relative-reference", + "description": "policy_uri is a uri-reference rather than the absolute URI the schema asks for. No network is needed to see it: this is a defect in the record, and it is reported the same way with or without a resolver.", + "boundary": "malformed", + "defect": "reference is relative, not the absolute URI the schema asks for", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" + }, + "policy": { + "bundle_hash": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc", + "enforcement_mode": "enforce", + "policy_uri": "bundles/policy-bundle-base.json" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "bundles/policy-bundle-base.json", + "resolution": { + "outcome": "not_attempted" + }, + "anchor": { + "tier": 2, + "source": "schemas/trace-claim.json properties.policy.properties.policy_uri", + "text": "URI to the policy bundle for verification.", + "applies": "The schema asks for format: uri, which is the absolute form. A reference the record got wrong is a defect in the record, visible with no network. This is a conformance statement about the packaged schema, not a claim about the specification, which says nothing about policy_uri at all." + } + }, + "expected": { + "tr_pol_003": "fail", + "reason": "A relative reference names no authority, so no verifier can dereference it on its own." + } +} diff --git a/tests/vectors/policy-resolution/09-sha384-bound-to-other-referent.json b/tests/vectors/policy-resolution/09-sha384-bound-to-other-referent.json new file mode 100644 index 0000000..f0af0b6 --- /dev/null +++ b/tests/vectors/policy-resolution/09-sha384-bound-to-other-referent.json @@ -0,0 +1,62 @@ +{ + "name": "sha384-bound-to-other-referent", + "description": "A sha384 digest of a different object. Paired with 06: deleting sha384 support turns both to skip, while a verifier that computes only sha256 fails 06 alone, and one that accepts sha384 without comparing fails 09 alone.", + "boundary": "contradicted", + "defect": "sha384 digest taken over a different object", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" + }, + "policy": { + "bundle_hash": "sha384:cfe2e1e2016b98bdaf6793f311cdc30fe63c84d7a25d2fbf6827e1d800ad8ad26aed70ab364c95cf8c7618b4f9a34200", + "enforcement_mode": "enforce", + "policy_uri": "https://policy.example.org/bundles/policy-bundle-base.json" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/bundles/policy-bundle-base.json", + "resolution": { + "outcome": "resolved" + }, + "anchor": { + "tier": 1, + "source": "agentrust-io/trace-spec spec/trace-v0.2.md 4.3", + "text": "the policy bundle hash is sealed to the TEE measurement, the enforcement mode is recorded, and substituting the policy invalidates the runtime claim", + "applies": "A resolved bundle whose digest is not the declared one is a substituted policy, so the runtime claim does not stand." + } + }, + "expected": { + "tr_pol_003": "fail", + "reason": "The resolved bytes contradict the declared sha384 digest. A verifier that cannot compute sha384 must not report a pass." + } +} diff --git a/tests/vectors/policy-resolution/10-policy-uri-carries-a-space.json b/tests/vectors/policy-resolution/10-policy-uri-carries-a-space.json new file mode 100644 index 0000000..0e37845 --- /dev/null +++ b/tests/vectors/policy-resolution/10-policy-uri-carries-a-space.json @@ -0,0 +1,62 @@ +{ + "name": "policy-uri-carries-a-space", + "description": "An absolute URI with a legitimate scheme and a space in the path. A transcription accident rather than a wrong form, and one that survives a diff unnoticed unless something checks for it.", + "boundary": "malformed", + "defect": "reference carries a space in its path", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" + }, + "policy": { + "bundle_hash": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc", + "enforcement_mode": "enforce", + "policy_uri": "https://policy.example.org/bundles/policy bundle base.json" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/bundles/policy bundle base.json", + "resolution": { + "outcome": "not_attempted" + }, + "anchor": { + "tier": 2, + "source": "schemas/trace-claim.json properties.policy.properties.policy_uri", + "text": "URI to the policy bundle for verification.", + "applies": "The schema asks for format: uri, which is the absolute form. A reference the record got wrong is a defect in the record, visible with no network. This is a conformance statement about the packaged schema, not a claim about the specification, which says nothing about policy_uri at all." + } + }, + "expected": { + "tr_pol_003": "fail", + "reason": "A URI carrying a raw space is not a URI. Paired with 08: the scheme rule and the character rule catch different mistakes." + } +} diff --git a/tests/vectors/policy-resolution/11-referent-unreachable-route-fails.json b/tests/vectors/policy-resolution/11-referent-unreachable-route-fails.json new file mode 100644 index 0000000..75032b2 --- /dev/null +++ b/tests/vectors/policy-resolution/11-referent-unreachable-route-fails.json @@ -0,0 +1,62 @@ +{ + "name": "referent-unreachable-route-fails", + "description": "The verifier knows where the bundle should be and cannot read it. Paired with 05: one has no route, this one has a route that fails, and a resolver that handled only one of the two would leave the other reporting something it has not checked.", + "boundary": "unresolvable", + "defect": "referent unreachable: route known, bundle missing", + "record": { + "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", + "iat": 1748000000, + "subject": "spiffe://example.org/agent/credit-risk/01926b4c-1234-7abc-9def-000000000001", + "model": { + "provider": "anthropic", + "model_id": "claude-sonnet-4-5" + }, + "runtime": { + "platform": "intel-tdx", + "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" + }, + "policy": { + "bundle_hash": "sha256:d8764863e75e702ff64e53951eb3d005a84858979598f7f0bda11fc901416adc", + "enforcement_mode": "enforce", + "policy_uri": "https://policy.example.org/bundles/policy-bundle-gone.json" + }, + "data_class": "confidential", + "build_provenance": { + "slsa_level": 2, + "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8" + }, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": "https://verifier.example.org/appraisal-policy/v1", + "timestamp": 1748000042 + }, + "transparency": "https://scitt.example.org/receipts/abc123def456", + "cnf": { + "jwk": { + "kty": "EC", + "crv": "P-256", + "x": "dGhpcyBpcyBhIHRlc3QgeA", + "y": "dGhpcyBpcyBhIHRlc3QgeQ", + "kid": "tee-key-001" + } + } + }, + "context": { + "cited_uri": "https://policy.example.org/bundles/policy-bundle-gone.json", + "resolution": { + "outcome": "unreachable_route_fails" + }, + "anchor": { + "tier": 0, + "source": "no merged sentence governs this case", + "text": "", + "applies": "No text in the specification says what a verifier records when a policy bundle cannot be fetched: policy_uri appears nowhere in the merged specification, and the reference block that does discuss resolvability governs the references field rather than this one. The vector therefore asserts only that the check was not performed, which is a fact about the run rather than a reading of any text. The level at which that stops being tolerable is suite policy, registered in src/trace_tests/modules/unverified.py and docs/levels.md, and is not claimed here as a specification requirement." + } + }, + "expected": { + "tr_pol_003": "unverified", + "reason": "The bundle could not be read, so the digest was never compared. The finding carries the reason so this is distinguishable from a URI nothing routes." + } +} diff --git a/tests/vectors/policy-resolution/README.md b/tests/vectors/policy-resolution/README.md index 02217c0..ac87067 100644 --- a/tests/vectors/policy-resolution/README.md +++ b/tests/vectors/policy-resolution/README.md @@ -1,121 +1,144 @@ -# Appraisal-resolution vectors - -Candidate conformance vectors for one question: **when a record cites an -appraisal policy by URI, what can a second verifier confirm about what that URI -held?** - -Today, nothing. `appraisal.policy_ref` is a bare URI. The record carries a -digest for the enforcement policy bundle (`policy.bundle_hash`) and none for the -appraisal policy that produced the verdict, so two verifiers resolving the same -`policy_ref` at different times can retrieve different documents and both report -`affirming` honestly. - -## Why this set exists - -`agentrust-io/trace-tests#63` proposed a conformance module for the `appraisal` -claim, noting that `policy_ref` can be checked for well-formedness only, "since -the record carries no digest to compare a resolved policy against." Building -those cases ran into the boundary itself, which was stated on -`agentrust-io/trace-spec#66` on 2026-08-18: - -> `appraisal.policy_ref` can be checked for well-formedness, but not -> reproduced: the record carries nothing stating what the referent was, so a -> conformance module can confirm the URI parses and nothing more. Two verifiers -> resolving the same `policy_ref` at different times can retrieve different -> documents and both honestly report verified — the federation gap §1 names, -> one hop from the record. - -Drafting candidate fixtures for it was accepted on that thread the same day, -with the shape asked for: candidates in `trace-tests`, one defect per vector, -expected outcomes committed beside the vectors, and the schema and -`verification.md` text left maintainer-authored. This set is that. - -## What it does not do - -**It does not propose a schema change.** `candidate_binding` is a *candidate -shape*, marked `CANDIDATE:` in every vector, and it lives in the vector's -`context` — never inside `record`. The packaged schema sets -`additionalProperties: false` on `appraisal`, so a record carrying an extra -field would be schema-invalid; whether a binding field lands, and what it is -called, is an editorial decision. - -**It does not name an outcome value for the unresolvable case.** Vectors 05 and -06 are marked `deferred`, pointing at `agentrust-io/trace-spec#190`, which -tracks the open cross-surface question: *when a record cites something a -verifier cannot resolve, what does the verifier record, and where.* - -Four surfaces face that question. Only one has answered it: -`build_provenance`, with a recorded field plus a floor -(`agentrust-io/trace-spec#173`, **merged**). Revocation states the prohibition -in prose with no field to record the outcome in -(`agentrust-io/trace-spec#187`, **merged**). Delegation links **propose** -`unverifiable` on a separate axis (`agentrust-io/trace-spec#184`, **an -explicitly non-normative draft** — a proposal facing the question, not an -established answer; its author has said so on -`agentrust-io/trace-spec#190`). And this one. - -The divergence is a schema fact rather than four differing opinions: -`appraisal` is `additionalProperties: false`, and `#173` landed the only -claimed/verified pair that exists, so the other surfaces had no field to use. -Coining a fifth answer here is exactly what `agentrust-io/trace-spec#190` -exists to prevent — and a new field or a fifth `appraisal.status` value is -normative under `CONTRIBUTING.md`, which routes it through a Spec change -proposal with a sponsoring organization. It is not something a vector set -decides. - -> **`deferred` is fixture bookkeeping, not a proposed appraisal vocabulary -> value.** It is a property of a *vector's expected block*, recording that the -> outcome is not yet decided upstream. It is **not** a candidate value for -> `appraisal.status`, which is closed at `affirming`, `warning`, -> `contraindicated`, `none`. A deferred vector asserts only what the outcome may -> **not** be — `must_not: "affirming"` — because reporting a check that was -> never performed as affirming is the one thing every reading of the merged text -> already rules out. `tests/test_policy_resolution_completeness.py` fails if -> any vector reuses a status value as an outcome. +# Policy-resolution vectors + +Conformance vectors for one question: **when a record declares a policy bundle +digest and says where the bundle lives, does the bundle at that location +actually hash to what was declared?** + +Until `TR-POL-003`, nothing asked. `policy.bundle_hash` and `policy.policy_uri` +are both merged fields on `policy`, and no module compared them, so a record +could name a bundle, declare a digest, and have the two disagree without any +check noticing. + +## What each vector asserts + +**One value: the status of the `TR-POL-003` finding.** That is narrower than a +verdict on the whole record, deliberately. These records are unsigned, so they +carry findings from other modules — `TR-SIG-005` has an opinion about every one +of them — and grading at record level would blur the check under test with +everything standing around it. Every claim about coverage in this file is a +claim about that one finding. + +## Anchors: what each expected outcome derives from + +Every vector names the text its expected outcome comes from, and says which +surface that text lives on. + +| Tier | Surface | What it can ground | +|---|---|---| +| 1 | Merged prose in `agentrust-io/trace-spec` `spec/trace-v0.2.md` | A claim about what the specification requires | +| 2 | The packaged `schemas/trace-claim.json` | A claim about conformance to the schema this suite tests against | +| 0 | Nothing governs the case | Only a fact about the run itself | + +The distinction matters because schema `description` text reads exactly like +normative prose and sits in the same repository as the vectors. It is not the +specification. Where a vector's outcome rests on tier 2, it says so rather than +implying an authority it does not have. + +**`policy_uri` appears nowhere in the merged specification.** Every outcome that +turns on `policy_uri` semantics is therefore tier 2 or tier 0, and the two +unresolvable vectors are tier 0: no merged sentence says what a verifier records +when a policy bundle cannot be fetched. Those vectors assert only that the +comparison did not happen, which is a fact about the run rather than a reading +of any text. **The level at which that stops being tolerable is suite policy**, +registered in `src/trace_tests/modules/unverified.py` and `docs/levels.md`, and +is not claimed here as a specification requirement. ## The vectors -Every record is identical except for `appraisal.policy_ref`, so the defect under -test is the only thing that varies. All records are unsigned and ASCII-only. - -| # | Vector | Boundary | Expected | -|---|---|---|---| -| 01 | `no-binding-declared` | accept | `pass` | -| 02 | `resolved-and-matches` | accept | `pass` | -| 03 | `digest-mismatch-minimal-mutation` | contradicted | `reject` | -| 04 | `digest-mismatch-different-object` | contradicted | `reject` | -| 05 | `referent-unreachable` | unresolvable | `deferred` | -| 06 | `digest-algorithm-uncomputable` | unresolvable | `deferred` | -| 07 | `digest-bound-to-other-referent` | contradicted | `reject` | - -**01 and 02 are the must-accept pair, and they are why this set is not -one-directional.** `agentrust-io/trace-spec#186` (merged 2026-08-20) states -the criterion: a set must fail *both* unconditional implementations. A set -written from the motivating problem alone would be all rejections and -deferrals, and a verifier that rejects everything would pass it. 01 is the -backward-compatibility control — every conformant record today declares no -binding, and must keep verifying, or this set would be proposing a breaking -change rather than describing a gap. 02 is the only vector in which a -resolution succeeds. +Every record is identical except for the `policy` block, so the defect under +test is the only thing that varies. `appraisal.policy_ref` is fixed across the +set and never varied: this set is about `policy.*`, and a second moving field +would blur which one a finding is answering. All records are unsigned and +ASCII-only. + +| # | Vector | Boundary | `TR-POL-003` | Anchor | +|---|---|---|---|---| +| 01 | `01-no-policy-uri.json` | accept | `skip` | tier 2 | +| 02 | `02-resolved-and-matches.json` | accept | `pass` | tier 1 | +| 03 | `03-digest-mismatch-minimal-mutation.json` | contradicted | `fail` | tier 1 | +| 04 | `04-digest-mismatch-different-object.json` | contradicted | `fail` | tier 1 | +| 05 | `05-referent-unreachable-no-route.json` | unresolvable | `unverified` | tier 0 | +| 06 | `06-resolved-and-matches-sha384.json` | accept | `pass` | tier 1 | +| 07 | `07-digest-bound-to-other-referent.json` | contradicted | `fail` | tier 1 | +| 08 | `08-policy-uri-is-a-relative-reference.json` | malformed | `fail` | tier 2 | +| 09 | `09-sha384-bound-to-other-referent.json` | contradicted | `fail` | tier 1 | +| 10 | `10-policy-uri-carries-a-space.json` | malformed | `fail` | tier 2 | +| 11 | `11-referent-unreachable-route-fails.json` | unresolvable | `unverified` | tier 0 | + +## Why the set is paired the way it is + +`agentrust-io/trace-spec#186` (merged 2026-08-20) states the criterion a vector +set is claiming: *a verifier that does not implement these rules will fail this +set*. A set must fail **both** unconditional implementations, and one vector +cannot separate a check that reads a prefix from one that reads the whole +object. Every rule below therefore sits on at least two vectors, and each pair +can be told apart by a weakened variant of the rule that deviates one and leaves +the other undisturbed. + +**01, 02 and 06 are the must-accept group.** A set written from the motivating +problem alone would be all rejections, and a verifier that rejected everything +would pass it. 01 is the backward-compatibility control: every conformant record +today declares no `policy_uri` and must keep verifying, or this set would be +proposing a breaking change rather than describing a gap. **03 and 04 keep the contradicted boundary off a single vector.** 03 differs -from the appraised object in exactly one byte — the SLSA floor moves 2 → 3, -which flips this record's verdict. 04 substitutes an unrelated document of a -different length. A verifier comparing lengths, or sampling a prefix, passes one -and fails the other. - -**05 and 06 are both unresolvable, by different mechanisms.** 05 cannot reach -the object; 06 reaches it and cannot compute over it, because the declared -algorithm (`sha3-512`) is outside the set the schema admits -(`^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$`). The distinction deliberately -rhymes with the `digest_algorithm_unsupported` classification **proposed** in -`agentrust-io/trace-spec#184` — a non-normative draft — without adopting that -vocabulary. - -**07 is the one a well-formedness check passes.** The binding is a valid -`sha256:` digest and is the true digest of a real object in this set — version 2 -— while `policy_ref` cites version 1. Both halves are individually valid; the -pair is not. +from the declared object in exactly one byte — the SLSA floor moves 2 → 3, which +changes what the policy permits. 04 substitutes a different object of a different +length. A verifier comparing lengths, or sampling a prefix, passes one and fails +the other. + +**06 and 09 are the sha384 pair.** The schema admits `sha384:` as well as +`sha256:`, so a verifier that hardcodes `sha256` is wrong rather than merely +limited. Deleting sha384 support turns both to `skip`. A verifier that computes +only `sha256` deviates 06 alone; one that accepts `sha384` without comparing — +the fail-open shape — deviates 09 alone. 06 by itself would not catch the second. + +**08 and 10 are the malformed pair, and they are caught by different rules.** 08 +is a `uri-reference` where the schema asks for the absolute form: no authority, +so nothing can dereference it. 10 is absolute and correctly schemed, with a raw +space in the path — a transcription accident that survives a diff unnoticed. +Removing the character rule deviates 10 alone; removing the scheme rule deviates +08 alone. + +**05 and 11 are both unreachable, by different roads.** 05 cites a URI nothing +routes. 11 cites one the manifest maps to a bundle that is not there. A resolver +handling only a missing key would leave 11 reporting a comparison it never made, +and one handling only a missing file would do the same to 05. The finding +carries the resolver's exception text, so the two are distinguishable in a +report even though they share a status. + +**07 is the one a well-formedness check passes.** The declared digest is a valid +`sha256:` digest and is the true digest of a real object in this set — just not +of the one `policy_uri` names. Both halves are individually valid; the pair is +not. + +## Malformed references are found offline + +`08` and `10` report the same failure with or without a resolver. That is the +check's ordering, and it is deliberate: **a reference the record got wrong is a +defect in the record**, visible with no network at all, exactly like the digest +shape `TR-POL-001` tests. A referent that could not be fetched is weather. +Running without `--policy-dir` means no spurious unverified findings; it does +not mean being blind to a defect the record carries on its face. + +## Resolving the bundles + +`resolutions.json` maps each cited URI to a path inside this directory. It is +the single mapping: a vector cannot hold a private idea of what its URI +resolves to. Two entries are deliberate holes — the URI cited by `05` is absent +from the manifest, and the one cited by `11` maps to a file that is not written. + +``` +trace-tests verify --record tests/vectors/policy-resolution/02-resolved-and-matches.json \ + --policy-dir tests/vectors/policy-resolution +``` + +The manifest is checked for *form* only when it loads: an object, string to +string, relative paths, no parent traversal. **Whether a mapped file is there is +not checked at load time.** Existence is a resolve-time fact, and a manifest that +refused to load because one bundle had gone missing would be the manifest-level +version of treating a lost referent as a wrong reference — the confusion this +whole set exists to keep apart. ## Reproducing it @@ -124,7 +147,7 @@ python tests/vectors/policy-resolution/gen_policy_resolution.py ``` Deterministic: no keys, no clock, no randomness, no network. The digests are -SHA-256 over the exact bytes of the sibling files under `policies/`, so anyone +computed over the exact bytes of the sibling files under `policies/`, so anyone holding only this directory can recompute every number in the set. `tests/test_policy_resolution_reproduces.py` holds the generator to @@ -143,14 +166,13 @@ policy digest stops matching, and the set fails on a clean clone. ## What this set does not establish -- **No verifier was exercised.** Nothing in `trace-tests` resolves - `appraisal.policy_ref` today — that is the gap — so every expected outcome is - a claim about what a verifier should do, not a recording of what one did. The - tests grade the set's internal consistency: that each declared digest really - is, or really is not, the digest of the bytes the vector says resolution - returned. -- **The unresolvable outcome is unnamed**, by choice, pending - `agentrust-io/trace-spec#190`. +- **Nothing here exercises a network fetch.** The resolver in every test reads + bytes from disk. A redirect, a timeout, a TLS failure and a 404 all map onto + the same unverified status by assertion rather than by measurement. +- **The unresolvable level is suite policy, not a specification requirement.** + No merged sentence governs the case; `agentrust-io/trace-spec#190` tracks the + open cross-surface question of what a verifier records when a citation cannot + be resolved. Both are recorded as exact shortfalls in `tests/test_policy_resolution_completeness.py::KNOWN_SHORTFALLS`, which fails @@ -159,15 +181,12 @@ if the list changes without this file changing with it. ## Related - `agentrust-io/trace-tests#63` — the module proposal these cases were built for -- `agentrust-io/trace-spec#66` — where the gap was raised and the fixtures accepted -- `agentrust-io/trace-spec#190` — the deferred cross-surface question -- `agentrust-io/trace-spec#173` — merged: recorded field plus a policy floor -- `agentrust-io/trace-spec#184` — open, **draft, explicitly non-normative**: - *proposes* `unverifiable` on the delegation surface +- `agentrust-io/trace-spec#66` — where the resolution gap was raised +- `agentrust-io/trace-spec#190` — the open cross-surface question - `agentrust-io/trace-spec#186` — merged: the adequacy criteria this set was built to. It grades trace-spec's `examples/`; this repository has no adequacy harness, so the standard is one this set chose, not one imposed on it - `agentrust-io/trace-tests#66` — merged: `tr_sig` canonicalizes with RFC 8785; source of the self-containment principle quoted above -- `agentrust-io/trace-tests#68` — merged: packaged schema resynced to the - normative v0.2 copy these records validate against +- `agentrust-io/trace-tests#74` — merged: the published error codes and record + samples realigned with the modules, and the guard that keeps them that way diff --git a/tests/vectors/policy-resolution/gen_policy_resolution.py b/tests/vectors/policy-resolution/gen_policy_resolution.py index e100908..acf8823 100644 --- a/tests/vectors/policy-resolution/gen_policy_resolution.py +++ b/tests/vectors/policy-resolution/gen_policy_resolution.py @@ -7,27 +7,31 @@ python tests/vectors/policy-resolution/gen_policy_resolution.py -The digests in the vectors are SHA-256 over the exact bytes of the sibling +The digests in the vectors are computed over the exact bytes of the sibling files under ``policies/``. Anyone holding only this directory can recompute them; nothing here depends on another repository being checked out. WHAT THIS SET IS FOR - ``appraisal.policy_ref`` is a bare URI. A record says which appraisal - policy produced its verdict, but carries nothing stating what that URI - resolved to at appraisal time, so a second verifier cannot confirm it - retrieved the same document. See the set's README.md. - -WHAT IT DELIBERATELY DOES NOT DO - It does not name an outcome value for the unresolvable case. That is the - open cross-surface question tracked by agentrust-io/trace-spec#190, and - coining a value here is precisely what that issue exists to prevent. - Vectors 05 and 06 carry the fixture-bookkeeping state ``deferred``. - - ``candidate_binding`` is a CANDIDATE shape, not a proposed schema change. - It lives in the vector's ``context``, never inside ``record``: the - packaged schema sets ``additionalProperties: false`` on ``appraisal``, so - a record carrying an extra field would be schema-invalid, and proposing - one is a schema decision that belongs to the editorial process. + ``policy.bundle_hash`` states the digest of the policy bundle in force, + and ``policy.policy_uri`` says where that bundle can be fetched. Nothing + in the suite used to compare the two, so a record could name a bundle, + declare a digest, and have the two disagree without any check noticing. + TR-POL-003 makes that comparison, and this set is what holds it to + account. See the set's README.md. + +WHAT EACH VECTOR ASSERTS + One thing: the status of the TR-POL-003 finding. That is the unit of + measurement for this set, and it is deliberately narrower than a whole + record's verdict. These records carry other findings — they are unsigned, + so TR-SIG-005 has an opinion about them — and reading the set at record + granularity would blur the check under test with everything around it. + +ANCHORS + Every expected outcome names the text it derives from, and says which + surface that text lives on. Tier 1 is merged prose in the specification. + Tier 2 is the packaged schema, which is what this suite tests conformance + against but is not the specification. Where no text governs a case, the + vector says so rather than implying an authority it does not have. """ from __future__ import annotations @@ -36,7 +40,7 @@ import json from pathlib import Path -HERE = Path(__file__).parent +HERE = Path(__file__).resolve().parent POLICIES = HERE / "policies" # --- house serialization, fixed so bytes are stable across platforms ------- @@ -55,20 +59,20 @@ def sha256_of(data: bytes) -> str: return "sha256:" + hashlib.sha256(data).hexdigest() -def sha3_512_of(data: bytes) -> str: - """A *correct* digest in an algorithm the profile's schema does not admit. +def sha384_of(data: bytes) -> str: + """A digest the schema admits and a sha256-only verifier cannot compute. - Vector 06 turns on the verifier being unable to compute the comparison, not - on the digest being wrong. A placeholder value would confound the two: a - reader could not tell whether the vector is deferred because the algorithm - is unsupported or because the digest is obviously bogus. This is the real - SHA3-512 of the referent, so the algorithm is the only variable. + The pattern on ``policy.bundle_hash`` accepts sha384, so a verifier that + hardcoded sha256 would be wrong rather than merely limited. Two vectors + turn on this: one where the sha384 comparison must succeed, and one where + it must fail. A verifier that skips what it cannot compute passes the + first and is caught by the second. """ - return "sha3-512:" + hashlib.sha3_512(data).hexdigest() + return "sha384:" + hashlib.sha384(data).hexdigest() # --- the cited objects ----------------------------------------------------- -# Small, ASCII-only, and shaped like an appraisal policy rather than a +# Small, ASCII-only, and shaped like a policy bundle rather than a # placeholder, so a reader can see why swapping one for another matters. POLICY_BASE = { @@ -81,9 +85,10 @@ def sha3_512_of(data: bytes) -> str: ], } -# One character apart from V1: the SLSA floor moves 2 -> 3. A verifier -# applying this instead of V1 reaches a different verdict on the same record, -# which is why a minimal mutation is the honest test rather than a cosmetic one. +# One character apart from the baseline: the SLSA floor moves 2 -> 3. A +# verifier comparing digests sees a mismatch; a human diffing the two files +# sees a single byte. That is the point: the smallest edit that changes what +# the policy permits still has to be caught. POLICY_ONEBYTE = { "policy_id": "appraisal/baseline", "version": "1.0.0", @@ -94,7 +99,9 @@ def sha3_512_of(data: bytes) -> str: ], } -# A legitimate later version, published at its own URI. +# A wholesale replacement rather than an edit: different rules, different +# shape, same job. Paired with the one-byte case so the check cannot be +# satisfied by a heuristic that only notices large changes. POLICY_OTHER = { "policy_id": "appraisal/baseline", "version": "2.0.0", @@ -125,22 +132,22 @@ def sha3_512_of(data: bytes) -> str: BASE_URI = "https://policy.example.org/bundles/" # --- the record ------------------------------------------------------------ -# Every vector's record is identical except for appraisal.policy_ref, so the -# defect under test is the only thing that varies. Modelled on the repository's -# tests/vectors/valid_level0.json: unsigned, ASCII-only, fixed iat. +# Every vector's record is identical except for the policy block, so the +# defect under test is the only thing that varies. appraisal.policy_ref is +# fixed and never varied: this set is about policy.bundle_hash and +# policy.policy_uri, and a second moving field would blur which one the +# finding is answering. Modelled on tests/vectors/valid_level0.json: +# unsigned, ASCII-only, fixed iat. RECORD_IAT = 1748000000 APPRAISAL_TIMESTAMP = 1748000042 +#: Fixed so the CLI can be handed a matching --expected-nonce and TR-RTE-004 +#: stops being noise at Level 1. base64url, no padding, per the schema. +RECORD_NONCE = "Zm9yLXRoZS1yZWNvcmQtbm9uY2U" +APPRAISAL_POLICY_REF = "https://verifier.example.org/appraisal-policy/v1" -def record_with(policy_ref: str | None) -> dict[str, object]: - appraisal: dict[str, object] = { - "status": "affirming", - "verifier": "https://verifier.example.org", - } - if policy_ref is not None: - appraisal["policy_ref"] = policy_ref - appraisal["timestamp"] = APPRAISAL_TIMESTAMP +def record_with(policy: dict[str, object]) -> dict[str, object]: return { "eat_profile": "tag:agentrust-io.com,2026:trace-v0.2", "iat": RECORD_IAT, @@ -150,18 +157,20 @@ def record_with(policy_ref: str | None) -> dict[str, object]: "platform": "intel-tdx", "measurement": "sha256:a3f8d2b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8", + "nonce": RECORD_NONCE, }, - "policy": { - "bundle_hash": - "sha256:b4e1c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4", - "enforcement_mode": "enforce", - }, + "policy": policy, "data_class": "confidential", "build_provenance": { "slsa_level": 2, "digest": "sha256:c9f7a5b2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a6b8c0d2e4f6a8", }, - "appraisal": appraisal, + "appraisal": { + "status": "affirming", + "verifier": "https://verifier.example.org", + "policy_ref": APPRAISAL_POLICY_REF, + "timestamp": APPRAISAL_TIMESTAMP, + }, "transparency": "https://scitt.example.org/receipts/abc123def456", "cnf": { "jwk": { @@ -175,20 +184,86 @@ def record_with(policy_ref: str | None) -> dict[str, object]: } -CANDIDATE_FIELD = "CANDIDATE:appraisal.policy_digest" -DEFERRED_PENDING = "agentrust-io/trace-spec#190" +def policy(bundle_hash: str, policy_uri: str | None = None) -> dict[str, object]: + block: dict[str, object] = { + "bundle_hash": bundle_hash, + "enforcement_mode": "enforce", + } + if policy_uri is not None: + block["policy_uri"] = policy_uri + return block + + +# --- anchors --------------------------------------------------------------- +# The text each expected outcome derives from, and the surface it lives on. + +ANCHOR_SUBSTITUTION = { + "tier": 1, + "source": "agentrust-io/trace-spec spec/trace-v0.2.md 4.3", + "text": ( + "the policy bundle hash is sealed to the TEE measurement, the " + "enforcement mode is recorded, and substituting the policy invalidates " + "the runtime claim" + ), + "applies": ( + "A resolved bundle whose digest is not the declared one is a " + "substituted policy, so the runtime claim does not stand." + ), +} -SPEC_DECIDED = ( - "agentrust-io/trace-spec docs/verification.md - a verifier records what it " - "actually resolved, and evidence that resolves and contradicts fails the " - "appraisal. Applied here to appraisal.policy_ref, which carries no digest." -) -SPEC_DEFERRED = ( - "agentrust-io/trace-spec docs/verification.md - evidence that does not " - "resolve downgrades honestly rather than failing. Which value records that " - "for appraisal.policy_ref is open: agentrust-io/trace-spec#190. This vector " - "asserts only what the outcome may NOT be." -) +ANCHOR_MATCHES = { + "tier": 1, + "source": "agentrust-io/trace-spec spec/trace-v0.2.md 4.3", + "text": ( + "the policy bundle hash is sealed to the TEE measurement, the " + "enforcement mode is recorded, and substituting the policy invalidates " + "the runtime claim" + ), + "applies": ( + "Read the other way: the bundle that resolves is the bundle that was " + "sealed, so nothing was substituted and the claim stands." + ), +} + +ANCHOR_URI_FORM = { + "tier": 2, + "source": "schemas/trace-claim.json properties.policy.properties.policy_uri", + "text": "URI to the policy bundle for verification.", + "applies": ( + "The schema asks for format: uri, which is the absolute form. A " + "reference the record got wrong is a defect in the record, visible " + "with no network. This is a conformance statement about the packaged " + "schema, not a claim about the specification, which says nothing about " + "policy_uri at all." + ), +} + +ANCHOR_OPTIONAL = { + "tier": 2, + "source": "schemas/trace-claim.json properties.policy.required", + "text": "required: [bundle_hash, enforcement_mode]", + "applies": ( + "policy_uri is absent from the required list, so a record that omits " + "it is conformant and the check has nothing to compare." + ), +} + +ANCHOR_UNRESOLVED = { + "tier": 0, + "source": "no merged sentence governs this case", + "text": "", + "applies": ( + "No text in the specification says what a verifier records when a " + "policy bundle cannot be fetched: policy_uri appears nowhere in the " + "merged specification, and the reference block that does discuss " + "resolvability governs the references field rather than this one. The " + "vector therefore asserts only that the check was not performed, which " + "is a fact about the run rather than a reading of any text. The level " + "at which that stops being tolerable is suite policy, registered in " + "src/trace_tests/modules/unverified.py and docs/levels.md, and is not " + "claimed here as a specification requirement." + ), +} def main(out_dir: Path | None = None) -> int: @@ -211,240 +286,253 @@ def main(out_dir: Path | None = None) -> int: digests[name] = sha256_of(raw[name]) d_base = digests["policy-bundle-base.json"] - d_onebyte = digests["policy-bundle-onebyte.json"] - d_other = digests["policy-bundle-other.json"] - d_unrelated = digests["policy-bundle-unrelated.json"] + d384_base = sha384_of(raw["policy-bundle-base.json"]) + d384_other = sha384_of(raw["policy-bundle-other.json"]) uri_base = BASE_URI + "policy-bundle-base.json" + uri_onebyte = BASE_URI + "policy-bundle-onebyte.json" uri_other = BASE_URI + "policy-bundle-other.json" + uri_unrelated = BASE_URI + "policy-bundle-unrelated.json" + # Cited by 05 and deliberately absent from the manifest: no route at all. uri_withdrawn = BASE_URI + "policy-bundle-withdrawn.json" - - def resolved(file_name: str, digest: str) -> dict[str, object]: + # Cited by 11 and mapped by the manifest to a file that is not there: + # a route that fails. Two input classes, two exception classes. + uri_gone = BASE_URI + "policy-bundle-gone.json" + # 08: the uri-reference form the schema's format: uri does not admit. + uri_relative = "bundles/policy-bundle-base.json" + # 10: absolute, well-schemed, and carrying a space in the path. + uri_spaced = BASE_URI + "policy bundle base.json" + + def ctx(cited: str, outcome: str, anchor: dict[str, object]) -> dict[str, object]: return { - "outcome": "resolved", - "file": f"policies/{file_name}", - "actual_sha256": digest, + "cited_uri": cited, + "resolution": {"outcome": outcome}, + "anchor": anchor, } vectors: list[tuple[str, dict[str, object]]] = [ - ("01-no-binding-declared.json", { - "name": "no-binding-declared", + ("01-no-policy-uri.json", { + "name": "no-policy-uri", "description": ( - "The record cites an appraisal policy by bare URI and declares no " - "binding to what that URI held. This is every conformant record " - "today, so it must keep verifying: a set that rejected it would be " - "proposing a breaking change rather than describing a gap." + "The record declares a bundle digest and no location to fetch the " + "bundle from. This is every conformant record today, so it must keep " + "verifying: a set that failed it would be proposing a breaking change " + "rather than describing a gap." ), "boundary": "accept", "defect": "none - backward-compatibility control", - "spec": ( - "Mirrors the merged treatment of an undeclared depth in " - "agentrust-io/trace-spec#173, where a record that never declared is " - "read as surface rather than as a failure." - ), - "record": record_with(uri_base), - "context": { - "cited_uri": uri_base, - "resolution": { - "outcome": "not_attempted", - "note": "No binding is declared, so there is nothing to check.", - }, - "candidate_binding": None, - }, + "record": record_with(policy(d_base)), + "context": ctx(None, "not_attempted", ANCHOR_OPTIONAL), "expected": { - "outcome": "pass", - "reason": "No binding declared; nothing to contradict.", + "tr_pol_003": "skip", + "reason": "No policy_uri is declared, so there is nothing to resolve.", }, }), ("02-resolved-and-matches.json", { "name": "resolved-and-matches", "description": ( - "The cited URI resolves and its bytes hash to exactly the declared " - "candidate binding. The second must-accept vector, and the only one " - "in which a resolution actually succeeds." + "The bundle at policy_uri hashes to exactly what bundle_hash " + "declares. The positive control: a set that could not pass this would " + "reject every honest record." ), "boundary": "accept", - "defect": "none - positive control", - "spec": SPEC_DECIDED, - "record": record_with(uri_base), - "context": { - "cited_uri": uri_base, - "resolution": resolved("policy-bundle-base.json", d_base), - "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_base}, - }, + "defect": "none - positive control, sha256", + "record": record_with(policy(d_base, uri_base)), + "context": ctx(uri_base, "resolved", ANCHOR_MATCHES), "expected": { - "outcome": "pass", - "reason": "Declared binding equals the digest of what resolved.", + "tr_pol_003": "pass", + "reason": "The resolved bytes hash to the declared digest.", }, }), ("03-digest-mismatch-minimal-mutation.json", { "name": "digest-mismatch-minimal-mutation", "description": ( - "The cited URI now serves a document one character from the one " - "appraised: the SLSA floor moved from 2 to 3. The record still " - "declares the original digest. A verifier that compares only the " - "URI sees no change; a verifier that compares bytes sees the " - "substitution that flips this record's verdict." + "The cited bundle was edited by one character after the digest was " + "taken. The smallest change that alters what the policy permits, and " + "the one a reader is least likely to notice." ), "boundary": "contradicted", - "defect": "cited object mutated minimally after appraisal", - "spec": SPEC_DECIDED, - "record": record_with(uri_base), - "context": { - "cited_uri": uri_base, - "resolution": resolved("policy-bundle-onebyte.json", d_onebyte), - "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_base}, - "note": ( - "policies/policy-bundle-base.json and " - "policies/policy-bundle-onebyte.json differ in one character." - ), - }, + "defect": "cited object mutated minimally after the digest was taken", + "record": record_with(policy(d_base, uri_onebyte)), + "context": ctx(uri_onebyte, "resolved", ANCHOR_SUBSTITUTION), "expected": { - "outcome": "reject", - "reason": "Resolved bytes contradict the declared binding.", + "tr_pol_003": "fail", + "reason": "The resolved bytes contradict the declared digest.", }, }), ("04-digest-mismatch-different-object.json", { "name": "digest-mismatch-different-object", "description": ( - "The cited URI resolves to an unrelated policy document. Paired with " - "03 so the boundary is not carried by a single vector: a verifier " - "that only samples a prefix of the document, or compares lengths, " - "passes one of these two and fails the other." + "The cited bundle was replaced wholesale after the digest was taken. " + "Paired with 03 so the check cannot be satisfied by a heuristic that " + "only notices large differences, or only small ones." ), "boundary": "contradicted", - "defect": "cited object wholly replaced after appraisal", - "spec": SPEC_DECIDED, - "record": record_with(uri_base), - "context": { - "cited_uri": uri_base, - "resolution": resolved("policy-bundle-unrelated.json", d_unrelated), - "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_base}, - }, + "defect": "cited object wholly replaced after the digest was taken", + "record": record_with(policy(d_base, uri_other)), + "context": ctx(uri_other, "resolved", ANCHOR_SUBSTITUTION), "expected": { - "outcome": "reject", - "reason": "Resolved bytes contradict the declared binding.", + "tr_pol_003": "fail", + "reason": "The resolved bytes contradict the declared digest.", }, }), - ("05-referent-unreachable.json", { - "name": "referent-unreachable", + ("05-referent-unreachable-no-route.json", { + "name": "referent-unreachable-no-route", "description": ( - "The cited URI does not resolve; no object under policies/ " - "corresponds to it. Nothing was contradicted, because nothing was " - "read. This vector asserts only that the outcome is not affirming." + "The cited URI is not one the verifier has any route to. Nothing was " + "contradicted, because nothing was read. This vector asserts only " + "that the comparison did not happen." ), "boundary": "unresolvable", - "defect": "referent unreachable at verification time", - "spec": SPEC_DEFERRED, - "record": record_with(uri_withdrawn), - "context": { - "cited_uri": uri_withdrawn, - "resolution": { - "outcome": "unreachable", - "file": None, - "note": "No sibling file corresponds to this URI, by construction.", - }, - "candidate_binding": {"field": CANDIDATE_FIELD, "value": d_base}, - }, + "defect": "referent unreachable: no route to the cited URI", + "record": record_with(policy(d_base, uri_withdrawn)), + "context": ctx(uri_withdrawn, "unreachable_no_route", ANCHOR_UNRESOLVED), "expected": { - "outcome": "deferred", - "deferred_pending": DEFERRED_PENDING, - "must_not": "affirming", + "tr_pol_003": "unverified", "reason": ( - "Reporting a check that was never performed as affirming is the " - "failure this set exists to name. Which value is reported " - "instead is not decided here." + "The bundle could not be fetched, so the digest was never " + "compared. Reporting that as a pass would claim a check that " + "did not run." ), }, }), - ("06-digest-algorithm-uncomputable.json", { - "name": "digest-algorithm-uncomputable", + ("06-resolved-and-matches-sha384.json", { + "name": "resolved-and-matches-sha384", "description": ( - "The referent resolves, but the declared binding names a digest " - "algorithm outside the set this profile's schema admits " - "(sha256 and sha384). The verifier cannot compute the comparison, " - "so it has not checked - which is distinct from having checked and " - "disagreed. Paired with 05: one cannot reach the object, the other " - "reaches it and cannot compute over it." + "The same accept as 02, with the digest taken in sha384. The schema " + "admits both algorithms, so a verifier that hardcodes sha256 is wrong " + "rather than merely limited, and this is the vector that says so." ), - "boundary": "unresolvable", - "defect": "digest algorithm the verifier cannot compute", - "spec": ( - SPEC_DEFERRED - + " Deliberately rhymes with the unverifiable / " - "digest_algorithm_unsupported classification PROPOSED in " - "agentrust-io/trace-spec#184 - an explicitly non-normative " - "draft, a surface facing this question rather than an " - "established answer to it - without adopting its vocabulary." + "boundary": "accept", + "defect": "none - positive control, sha384", + "record": record_with(policy(d384_base, uri_base)), + "context": ctx(uri_base, "resolved", ANCHOR_MATCHES), + "expected": { + "tr_pol_003": "pass", + "reason": "The resolved bytes hash to the declared sha384 digest.", + }, + }), + ("07-digest-bound-to-other-referent.json", { + "name": "digest-bound-to-other-referent", + "description": ( + "The declared digest is a correct digest of some object, just not of " + "the one policy_uri names. A verifier that checks the digest is " + "well formed, or that it matches something it holds, passes this." ), - "record": record_with(uri_other), - "context": { - "cited_uri": uri_other, - "resolution": resolved("policy-bundle-other.json", d_other), - "candidate_binding": { - "field": CANDIDATE_FIELD, - "value": sha3_512_of(raw["policy-bundle-other.json"]), - "note": ( - "This is the correct SHA3-512 of the referent. It is " - "outside the schema's digest pattern " - "^sha(256:[0-9a-f]{64}|384:[0-9a-f]{96})$, so no conformant " - "verifier is required to compute it - which is the only " - "reason this vector is undecided." - ), - }, + "boundary": "contradicted", + "defect": "digest well formed but taken over a different object", + "record": record_with(policy(d_base, uri_unrelated)), + "context": ctx(uri_unrelated, "resolved", ANCHOR_SUBSTITUTION), + "expected": { + "tr_pol_003": "fail", + "reason": ( + "The binding does not describe the object the record cites, " + "even though it describes some object." + ), }, + }), + ("08-policy-uri-is-a-relative-reference.json", { + "name": "policy-uri-is-a-relative-reference", + "description": ( + "policy_uri is a uri-reference rather than the absolute URI the " + "schema asks for. No network is needed to see it: this is a defect " + "in the record, and it is reported the same way with or without a " + "resolver." + ), + "boundary": "malformed", + "defect": "reference is relative, not the absolute URI the schema asks for", + "record": record_with(policy(d_base, uri_relative)), + "context": ctx(uri_relative, "not_attempted", ANCHOR_URI_FORM), "expected": { - "outcome": "deferred", - "deferred_pending": DEFERRED_PENDING, - "must_not": "affirming", + "tr_pol_003": "fail", "reason": ( - "The comparison was not performed, so the result is not a " - "contradiction. Which value records that is not decided here." + "A relative reference names no authority, so no verifier can " + "dereference it on its own." ), }, }), - ("07-digest-bound-to-other-referent.json", { - "name": "digest-bound-to-other-referent", + ("09-sha384-bound-to-other-referent.json", { + "name": "sha384-bound-to-other-referent", "description": ( - "The declared binding is well formed and is the true digest of a " - "real object in this set - version 2 - while policy_ref cites " - "version 1. Both halves are individually valid and the pair is not. " - "A verifier that checks the digest is well formed, or that it " - "matches something it holds, passes this; only one that checks the " - "digest against what this URI resolved to rejects it." + "A sha384 digest of a different object. Paired with 06: deleting " + "sha384 support turns both to skip, while a verifier that computes " + "only sha256 fails 06 alone, and one that accepts sha384 without " + "comparing fails 09 alone." ), "boundary": "contradicted", - "defect": "binding well formed but bound to a different referent", - "spec": SPEC_DECIDED, - "record": record_with(uri_base), - "context": { - "cited_uri": uri_base, - "resolution": resolved("policy-bundle-base.json", d_base), - "candidate_binding": { - "field": CANDIDATE_FIELD, - "value": d_other, - "note": ( - "This is the digest of policies/policy-bundle-other.json, " - "which is not what cited_uri resolves to." - ), - }, + "defect": "sha384 digest taken over a different object", + "record": record_with(policy(d384_other, uri_base)), + "context": ctx(uri_base, "resolved", ANCHOR_SUBSTITUTION), + "expected": { + "tr_pol_003": "fail", + "reason": ( + "The resolved bytes contradict the declared sha384 digest. A " + "verifier that cannot compute sha384 must not report a pass." + ), + }, + }), + ("10-policy-uri-carries-a-space.json", { + "name": "policy-uri-carries-a-space", + "description": ( + "An absolute URI with a legitimate scheme and a space in the path. " + "A transcription accident rather than a wrong form, and one that " + "survives a diff unnoticed unless something checks for it." + ), + "boundary": "malformed", + "defect": "reference carries a space in its path", + "record": record_with(policy(d_base, uri_spaced)), + "context": ctx(uri_spaced, "not_attempted", ANCHOR_URI_FORM), + "expected": { + "tr_pol_003": "fail", + "reason": ( + "A URI carrying a raw space is not a URI. Paired with 08: the " + "scheme rule and the character rule catch different mistakes." + ), }, + }), + ("11-referent-unreachable-route-fails.json", { + "name": "referent-unreachable-route-fails", + "description": ( + "The verifier knows where the bundle should be and cannot read it. " + "Paired with 05: one has no route, this one has a route that fails, " + "and a resolver that handled only one of the two would leave the " + "other reporting something it has not checked." + ), + "boundary": "unresolvable", + "defect": "referent unreachable: route known, bundle missing", + "record": record_with(policy(d_base, uri_gone)), + "context": ctx(uri_gone, "unreachable_route_fails", ANCHOR_UNRESOLVED), "expected": { - "outcome": "reject", + "tr_pol_003": "unverified", "reason": ( - "The binding does not describe the object the record cites, " - "even though it describes some object." + "The bundle could not be read, so the digest was never " + "compared. The finding carries the reason so this is " + "distinguishable from a URI nothing routes." ), }, }), ] - for filename, vector in vectors: - write_json(here / filename, vector) + # 2. The manifest. One mapping from cited URI to the bytes behind it, used + # by the CLI's --policy-dir and by the tests, so there is no second + # place for a vector's idea of what it resolves to to drift from. + # uri_withdrawn is deliberately absent. uri_gone is deliberately mapped + # to a file that is not written. + manifest = { + uri_base: "policies/policy-bundle-base.json", + uri_onebyte: "policies/policy-bundle-onebyte.json", + uri_other: "policies/policy-bundle-other.json", + uri_unrelated: "policies/policy-bundle-unrelated.json", + uri_gone: "policies/policy-bundle-gone.json", + } + write_json(here / "resolutions.json", manifest) + + for name, vector in vectors: + write_json(here / name, vector) - print(f"wrote {len(POLICY_FILES)} policy objects and {len(vectors)} vectors") - for name, digest in digests.items(): - print(f" policies/{name} {digest}") + print(f"wrote {len(POLICY_FILES)} policy objects, {len(vectors)} vectors, 1 manifest") + for name in POLICY_FILES: + print(f" policies/{name} {digests[name]}") return 0 diff --git a/tests/vectors/policy-resolution/resolutions.json b/tests/vectors/policy-resolution/resolutions.json new file mode 100644 index 0000000..f026b91 --- /dev/null +++ b/tests/vectors/policy-resolution/resolutions.json @@ -0,0 +1,7 @@ +{ + "https://policy.example.org/bundles/policy-bundle-base.json": "policies/policy-bundle-base.json", + "https://policy.example.org/bundles/policy-bundle-onebyte.json": "policies/policy-bundle-onebyte.json", + "https://policy.example.org/bundles/policy-bundle-other.json": "policies/policy-bundle-other.json", + "https://policy.example.org/bundles/policy-bundle-unrelated.json": "policies/policy-bundle-unrelated.json", + "https://policy.example.org/bundles/policy-bundle-gone.json": "policies/policy-bundle-gone.json" +} From e339491abe61143a7afd023c01191ff968434d28 Mon Sep 17 00:00:00 2001 From: opento-suggestions Date: Mon, 24 Aug 2026 15:11:20 -0600 Subject: [PATCH 7/7] docs: register TR-POL-003, TR-POL-001 admits sha384, --policy-dir error-codes.md, modules/tr-pol.md, modules.md, levels.md (both failure lists and the Unverified findings table), quickstart.md. Guard extended three ways: every UNVERIFIED emitter is in the table, the table is a subset of documented codes, and the table equals the levels.md table. Signed-off-by: opento-suggestions --- docs/error-codes.md | 2 +- docs/levels.md | 25 ++++++++ docs/modules.md | 2 +- docs/modules/tr-pol.md | 24 +++++++- docs/quickstart.md | 22 +++++++ tests/test_docs_match_the_modules.py | 92 ++++++++++++++++++++++++++++ 6 files changed, 164 insertions(+), 3 deletions(-) diff --git a/docs/error-codes.md b/docs/error-codes.md index db1a054..70278a9 100644 --- a/docs/error-codes.md +++ b/docs/error-codes.md @@ -33,7 +33,7 @@ All TRACE test failures emit a structured error code of the form `TR-- set[str]: @@ -186,3 +191,90 @@ def test_every_json_sample_in_the_docs_agrees_with_the_packaged_schema() -> None f"({unparsed} block(s) did not parse). Either the samples are gone or the fence " "this reads has changed; a check over nothing must not report a pass." ) + + +def _unverified_emitters(source: str) -> set[str]: + """Codes this module constructs a ``Finding`` for with ``Status.UNVERIFIED``. + + Matched on the construction rather than on the string, because a code named + in a message, a comment or a table is not the module *emitting* it. The + registration table names every code it governs; without this the two sets + could agree while nothing actually produced one of them. + """ + found: set[str] = set() + for node in ast.walk(ast.parse(source)): + if not isinstance(node, ast.Call): + continue + func = node.func + name = func.id if isinstance(func, ast.Name) else getattr(func, "attr", "") + if name != "Finding": + continue + args = list(node.args) + by_kw = {k.arg: k.value for k in node.keywords} + code_node = args[0] if args else by_kw.get("code") + status_node = args[1] if len(args) > 1 else by_kw.get("status") + if not isinstance(code_node, ast.Constant) or not isinstance(code_node.value, str): + continue + if not (isinstance(status_node, ast.Attribute) and status_node.attr == "UNVERIFIED"): + continue + found |= _codes(code_node.value) + return found + + +def _emitted_unverified_codes() -> set[str]: + emitted: set[str] = set() + for path in sorted(MODULES.glob("*.py")): + emitted |= _unverified_emitters(path.read_text(encoding="utf-8")) + return emitted + + +def test_every_code_that_can_be_unverified_is_registered() -> None: + """A module emitting UNVERIFIED under an unregistered code fails from level 1. + + That default is fail-closed and therefore safe, but it is also silent: the + code would be governed by a rule nobody chose. Registering it is a + decision, so it has to be made rather than defaulted into. + """ + emitted = _emitted_unverified_codes() + registered = set(UNVERIFIED_FAILS_FROM_LEVEL) + assert emitted <= registered, ( + f"emitted as UNVERIFIED and not registered: {sorted(emitted - registered)}\n" + "Add the code to UNVERIFIED_FAILS_FROM_LEVEL with the level it fails from, " + "or stop emitting UNVERIFIED under it. Falling through to the default means " + "a level nobody picked." + ) + + +def test_every_registered_code_is_actually_emitted() -> None: + """The other direction: a row for a code nothing produces documents a fiction. + + This is the half that catches a table outliving its module. A registered + code with no emitter reads as a check the suite performs, and it does not. + """ + emitted = _emitted_unverified_codes() + registered = set(UNVERIFIED_FAILS_FROM_LEVEL) + assert registered <= emitted, ( + f"registered and emitted by no module: {sorted(registered - emitted)}\n" + "Delete the row, or emit the finding. A level for a status nothing " + "produces is a claim nobody checked." + ) + + +def test_the_registration_table_and_the_published_table_agree() -> None: + """``docs/levels.md`` publishes the levels; the module decides them. + + Two copies of one fact, which is the shape everything else in this file + exists to catch. Compared by row rather than by prose, because a reader + acting on the published table needs the number to be the one the code uses. + """ + published = {code: int(level) for code, level in _UNVERIFIED_ROW.findall( + LEVELS.read_text(encoding="utf-8"))} + assert published == UNVERIFIED_FAILS_FROM_LEVEL, ( + f"published in {LEVELS.relative_to(REPO)}: {published}\n" + f"registered in the module: {UNVERIFIED_FAILS_FROM_LEVEL}\n" + "A reader following the page must get the level the code applies." + ) + assert published, ( + "no unverified-level row was parsed, so this check would pass over no " + "work; the table in docs/levels.md is gone or its shape changed" + )