| Internet-Draft | AI Evaluation Claim Preservation | September 2026 |
| Abak | Expires 27 March 2027 | [Page] |
AI evaluation records can pass through evaluation frameworks, exporters, evidence services, independent reviewers, and systems that make decisions. A conversion can retain a score while losing whether the measurement ran, which result was selected, whether a criterion existed, or which evidence was unavailable. Authenticating the converted record does not recover these distinctions.¶
This document specifies format-neutral requirements for mapping profiles and consumers that exchange AI evaluation evidence. It separates source assertions, explicit derivations, and later policy judgments; requires claim-relevant loss and uncertainty to remain visible; and describes consumer behavior when preservation cannot be established. It includes synthetic counterexamples and guidance for composition with existing evidence mechanisms. It defines neither a wire format nor an evaluation benchmark, safety certification, authorization protocol, or new cryptographic envelope. A concrete mapping profile is needed for interoperable implementation.¶
This note is to be removed before publishing as an RFC.¶
This is an individual contribution with an Informational target. No working-group adoption or endorsement is implied. Comments may be sent to the author. The choice between a standalone application profile and guidance incorporated into existing work remains open.¶
This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79.¶
Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/.¶
Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress."¶
This Internet-Draft will expire on 27 March 2027.¶
Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved.¶
This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Code Components extracted from this document must include Revised BSD License text as described in Section 4.e of the Trust Legal Provisions and are provided without warranty as described in the Revised BSD License.¶
Consider an evaluation operator that exports results to an evidence service. A second organization retrieves a summary from that service and uses selected results in its own decision system. These boundaries can involve different native formats, parser versions, aggregation conventions, access permissions, and trust policies. A correctly authenticated summary can nevertheless misrepresent the source if a conversion changes the meaning of a result.¶
Native formats already make relevant distinctions. Inspect documents evaluation-log status separately from scoring results [INSPECT]. LightEval documents task-indexed numerical results and configuration metadata [LIGHTEVAL]. This document does not allege a defect in either implementation. Its question is what a converter and a downstream consumer need to preserve when exchanging such records.¶
For example, completion of a run does not mean that a metric met a criterion. A consumer can legitimately apply its own threshold to a source score, but that is a new assessment by that consumer, not a verdict reported by the original evaluator. Similarly, checking a signature on a model-name assertion does not establish which model actually executed.¶
General evidence-to-decision architectures and preservation principles are not new contributions of this document. RATS separates evidence appraisal from relying-party policy [RFC9334]. Related work on evidence qualification, statement graphs, and agent control intermediaries is discussed in Section 8. The contribution proposed here is a bounded set of requirements for AI-evaluation-specific mappings: result selection, run and measurement state, scores and criteria, population scope, and the qualifications that survive conversion.¶
The requirements apply to exchange across components or administrative domains, including offline exchange. They apply to evaluations of AI systems, including frontier-system use cases, without defining a frontier capability threshold. They do not require native evaluation frameworks to adopt a new format or to disclose confidential samples.¶
This document does not define scientific validity, acceptable risk, evaluator competence or independence, legal authority, or whether deployment should occur. It does not create an evidence-to-decision graph format, transparency service, trust-anchor system, or control-delivery mechanism. Accurate preservation of an unsupported or false source assertion does not make that assertion true.¶
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [BCP14] when, and only when, they appear in all capitals, as shown here.¶
These requirements apply to a mapping profile, converter, or consumer that explicitly adopts them. They do not retroactively change the specifications of native formats or evidence carriers. A claim of compliance with these requirements MUST identify the concrete profile, the implementation role, and the tested version; a bare claim of "claim-preserving" is insufficient. A profile supplies the encodings, selection rules, and processing details identified in Section 5. This document alone is not a complete wire-level interoperability specification.¶
An application can accept a limited result while withholding a stronger conclusion. Preservation failures affect the conclusions that require the missing distinction; they do not automatically make every other field unusable. No requirement here prescribes a particular deployment or business decision.¶
The source output of an evaluation process, such as a result file, sample record, log, or configuration. "Native" identifies provenance, not truth or completeness. Evidence is used here in a general sense, not as a replacement definition for RATS Evidence.¶
The particular source observation or assessment being exchanged: an identified source object and selection, together with the run, attempt, task, scorer, metric, aggregation, and population context needed to distinguish it. A run identifier alone need not identify a result unit.¶
An identified, versioned definition of source and destination interpretations, selection rules, permitted transformations, preserved qualifications, and consumer behavior. It is not a model-safety profile.¶
A particular application of a mapping profile to identified source objects. It records the converter, outputs, transformations, and material limitations of that application.¶
An assertion attributable to an identified source, including its scope and qualifications. A copied assertion does not become independently observed merely because a converter signs it.¶
A reported calculation or transformation with identified inputs and rules. A derivation may add useful information, but it is distinct from a source assertion. A policy judgment based on it is also distinct.¶
The procedure, inputs, accepted identities or trust material, and policy under which a particular verification result is produced. Different verification types establish different properties.¶
Context or uncertainty needed to interpret a stated claim, such as a denominator, missing samples, a criterion, units, access limitations, or the distinction between declared and verified identity. Materiality is relative to the claim and mapping profile, not to an arbitrary converter preference.¶
Preservation of the attribution, meaning, scope, and material qualifications of the source assertions selected for exchange. It does not prohibit explicitly identified derivations or new judgments. It prohibits representing those as stronger original assertions or silently broadening what the source supports.¶
A record asserting that an actor used identified evidence as an input to a particular decision under an identified policy. It does not, by itself, establish causation, complete deliberation, valid authority, or enforcement.¶
Claim preservation is evaluated for a specified mapping and bounded claim set. This document provides no universal decision procedure for semantic equivalence, natural-language entailment, or arbitrary converter correctness. A profile that claims preservation of a particular distinction needs an explicit rule and testable examples for that distinction.¶
The roles are source operator, converter or publisher, evidence service, consumer or verifier, and optional decision maker. One entity can perform several roles. Different names, keys, or services do not establish organizational independence.¶
Native result + context
|
identified mapping + loss record
|
evidence view -----> another mapping, if present
|
consumer
|
optional new assessment / real reliance decision
The protected unit of exchange can use existing statements, attachments, or referenced objects. Co-location in a file, matching display names, and nearby timestamps do not establish a required relationship. The profile needs an explicit binding between the source selection, mapping information, and output; otherwise a loss statement from one conversion could be substituted for that of another.¶
Failures include status coercion, task or metric substitution, omitted failed attempts, changed denominators, rounding across a threshold, criterion laundering, identity substitution, stale or revised source objects, and loss of access restrictions. Attackers can exploit the same conditions deliberately. The requirements also address honest conversion mistakes; they do not assume an honest source or converter merely because its output is signed.¶
A source may already be incomplete or contradictory. A converter can preserve that condition or report that it cannot map it. It cannot reconstruct missing observations from an absence of evidence. A consumer may need external evidence or a narrower claim, rather than a more permissive default.¶
A mapping instance MUST identify its source objects, source format and interpretation basis, destination basis, converter implementation/version, and mapping-profile revision. Where a source revision cannot be established, that limitation MUST be explicit and MUST NOT be replaced with an assumed current revision. A content digest identifies bytes; it does not by itself identify the rules used to interpret those bytes.¶
The instance MUST bind the output and material mapping information to the source objects and selections on which they depend. For each content digest, the algorithm and byte selection or canonicalization rule MUST be identified. Hashing an original JSON file and hashing a canonical representation are different operations; a profile MUST NOT silently substitute one for the other. Native source objects SHOULD remain available under the applicable access policy; their unavailability is handled under Section 4.8.¶
A mapping MUST identify the result unit being carried or assessed. It MUST preserve the context needed to distinguish task, scorer, metric, run attempt, aggregation, dataset/split, and population whenever these affect the claim. A profile MUST specify which native identifiers or selectors supply that context and how their uniqueness is scoped. Unsupported, unresolved, or ambiguous selections MUST NOT be resolved by taking an arbitrary first, last, largest, or most favorable result.¶
A selector MUST be interpreted against the identified source snapshot and under an identified selector syntax. When a summary combines results, the inputs and selection/aggregation rule MUST be identifiable. When a referenced object changes, the old selector MUST NOT silently become a reference to new content. A selector identifies data, not a measured population's completeness. JSON Pointer [RFC6901] is one possible syntax, not a required wire mechanism.¶
A mapping MUST distinguish whether a measurement executed, what value or finding was reported, and whether an identified criterion was assessed. It MUST NOT translate successful process completion into a successful metric or safety assessment. Non-execution, invalidation, error, inconclusive outcome, and absence of evidence MUST NOT silently become a measured pass. They also MUST NOT be relabeled as a measured failure unless the source semantics actually establish such a failure.¶
A score without an established criterion MUST remain a score without an attributed source verdict. "The source states that no criterion was applied" and "the available source does not establish a criterion" MUST remain distinguishable. An omitted field alone establishes neither condition unless the identified native format defines that omission unambiguously. A non-applicability assertion MUST retain its stated scope and basis, not be invented from a missing value.¶
For a selected result, a mapping MUST preserve the meaning of its value, including applicable units, scale, metric definition, comparator direction, precision, and uncertainty information. A profile MUST specify acceptable numerical conversions and any tolerance. It MUST distinguish a source uncertainty estimate from information lost through conversion. Missing uncertainty information MUST NOT be mapped to zero uncertainty or to a source claim of statistical confidence. An exact integer identifier MUST NOT be changed by a lossy numeric representation.¶
Rounding, normalization, aggregation, or recomputation MUST be identified as a derivation when it can affect interpretation. A consumer MUST NOT use a rounded presentation value to claim that the source met a threshold when the unrounded source did not. If the native representation cannot support the requested numerical comparison, that comparison MUST remain unestablished. A profile MAY permit a bounded conversion for a narrower claim; the bound and the affected claim MUST be explicit.¶
A transported criterion assessment MUST identify the criterion and revision actually attributed to its source. A consumer MAY apply a new criterion or policy to a source observation, provided it records a separate assessment with its own actor, inputs, rule/version, scope, and result. It MUST NOT describe this new judgment as the original evaluator's verdict or as a criterion fixed before the source run without supporting evidence.¶
A claim that criteria preceded execution MUST have a supported ordering relationship between the identified criteria and that run. A bare timestamp string, matching identifiers, or a hash-only commitment is not sufficient by itself. Existing preregistration mechanisms can be referenced; this document defines none. Unknown ordering MUST remain unknown. Legitimate corrections and later reassessments are permitted but MUST preserve the original assessment and explicit revision relationship.¶
A claim about an evaluation campaign or complete population MUST identify the population or reproducible inclusion rule, its source, and the accounting method. Recorded samples, unexecuted samples, errors, exclusions, invalidations, and unknown portions MUST NOT disappear solely because an aggregate counts only successful records. When totals or categories are unknown, a mapping MUST record that limitation rather than invent counts.¶
The profile MUST define how retries, repeated measurements, duplicate records, changed sample sets, and overlapping outcomes affect the denominator. It MUST NOT assume that native status categories are disjoint when they are not. A partial run MAY preserve a local criterion result; that result MUST retain its narrower scope. Preserving all supplied records does not establish that all actual attempts were supplied. Selection or deduplication that changes the population MUST be visible as a transformation.¶
A mapping MUST distinguish display names and declarations from identities established by a specified verification procedure. When a conclusion requires an exact model, checkpoint, configuration, dataset, harness, or environment, the profile MUST identify the evidence required for that conclusion. Missing required identity evidence MUST prevent that conclusion, but need not prevent carrying a more limited declared result.¶
Context material to the selected claim, such as tool access, network restrictions, safeguard state, scoring method, or system configuration, MUST retain its source and verification status. A declared container digest or model hash MUST NOT be presented as proof that those bytes executed. Attestation references MUST retain the scope and accepting verification basis of the attestation. Evaluator access and independence declarations MUST NOT be promoted to established access or independence merely through conversion.¶
A mapping MUST distinguish evidence availability from digest knowledge. It MUST distinguish a digest recomputed from available bytes by an identified actor, a digest reported by another source but not recomputed by that actor, and a digest that is unavailable. It MUST NOT fabricate a digest or use a dummy value to make an unavailable object appear content-bound.¶
Availability statements MUST be scoped to an actor and the relevant observation or exchange context. Public location, permission to access, successful retrieval, and retention are different properties. A profile MUST define how known restrictions and unknown availability are represented. It MUST NOT invent a custodian, withholding reason, or claim that an object exists. When the destination cannot express a required unavailable state, the mapper MUST provide a bound, interpretable qualification through the profile or report that the affected mapping is unsupported.¶
The mapping report MUST distinguish copied assertions, transformations, out-of-band declarations, and material losses. A profile MUST identify which source qualifications are required for each supported claim class. A converter MUST NOT omit a known material qualification merely by declaring it irrelevant. Unknown extensions that may qualify the selected claim MUST remain uninterpreted or prevent that claim from being established; a profile MAY define an explicit criticality mechanism or extension points whose contents cannot qualify its stated claim class under that profile.¶
Every transformation hop used to support an end-to-end preservation claim MUST be accounted for, either through the chain of mapping records or by a direct check against an adequately identified earlier source. A later converter MUST NOT erase an earlier material loss or present reintroduced information as though it had survived the lost hop. Corrections, contradictory records, alternative interpretations, and amendments MUST retain provenance. A signed loss report is itself an assertion; it does not prove that its loss inventory is complete. A contradictory source verdict may be retained as an attributed assertion, but MUST NOT be endorsed merely because it was copied faithfully.¶
A verification result MUST identify the actor, procedure/version, checked inputs, verification type, result, and material limitations. Structural validation, digest recomputation, signature verification, attestation appraisal, mapping verification, and substantive evaluation judgments MUST remain distinct. Successful verification of a carrier MUST NOT imply successful verification of an unsupported profile, unavailable source, or source statement not covered by that carrier.¶
A consumer MUST evaluate preservation relative to the selected claim and accepted verification basis. If an essential binding, interpretation, qualification, or verification input is unavailable, it MUST NOT report that claim as established. It MAY retain or forward an opaque record without accepting its semantics. A local policy may deny, defer, request evidence, or otherwise act on uncertainty, but the resulting action MUST NOT rewrite uncertainty as a measured source failure or success.¶
An exporter completing a conversion MUST NOT, on that basis alone, assert that a governance decision occurred. When a reliance record is supplied, it MUST identify the asserting actor, actual reported decision, applicable policy basis, affected subject, and selected evaluation inputs. A new judgment under Section 4.5 MUST remain separate from the transported source assertions.¶
The consumer MUST NOT infer valid decision authority, complete deliberation, causal influence, control delivery, enforcement, or observed effect solely from a reliance record. Authentication of its author is a separate check. Changes to evidence or policy MUST produce a distinguishable reassessment or amendment, not retroactively rewrite the evidentiary basis of an earlier decision. Existing decision and control mechanisms should carry any actual downstream records; this document defines no such protocol.¶
A concrete profile MUST identify the source and destination versions and semantic rules it supports, including its extension and unknown-value policy. An implementation MUST NOT silently treat a new or incompatible basis as the old one. A profile MUST specify observable expected behavior for negative and ambiguous cases, including cases where a limited result can be preserved but a stronger claim cannot.¶
Implementation reports MUST identify the profile and implementation versions, source cases, procedure, and observed results on which their claims rely. A schema check or a documentation crosswalk MUST NOT be described as a complete semantic mapping test. Same-author tests MUST NOT be described as independent interoperability. An interoperability claim MUST name the tested producer and consumer, exact bases, inputs, procedure, outcomes, and known limitations.¶
A mapping profile can use native fields, an existing evidence predicate, or a sidecar manifest. This document does not allocate an identifier or prescribe a new envelope. A sidecar used to qualify an output MUST be bound to that output and its source selection under the accepted integrity and attribution mechanism; an unrelated explanatory file is not sufficient.¶
The source and destination interpretation bases; claim classes supported; and how a result unit is selected without ambiguity.¶
The preservation rules for execution, observation, criteria, values, population scope, identity, and uncertainty, including explicit unsupported cases.¶
The representation and protection of mapping provenance, material loss, conflicts, access limitations, and digest knowledge.¶
The trust, authentication, freshness, retrieval, and resource-limit rules required for its claims, with unknown and failed states distinguished.¶
The consumer-visible outcomes and reproducible examples against which mapping and consumption are tested.¶
The profile MUST describe how each applicable requirement in Section 4 is satisfied and why a conditional requirement is inapplicable when it is not used. It need not require every possible identity or metadata field for every claim. However, reducing metadata cannot be used to preserve the name of a stronger claim while discarding the conditions needed to support it.¶
A loss statement need not enumerate every field outside the chosen claim set. It needs to identify which material information was not preserved and which claims consequently remain unsupported. Retention or disclosure of confidential source content is not required when a bounded commitment, restricted reference, or explicit unavailable state suffices for the narrower claim.¶
The following procedure summarizes the requirements. It is not a new verification algorithm or mandatory wire-state vocabulary. An implementation MAY combine steps if the same distinctions remain observable.¶
Select the requested claim and supported mapping profile. Bound parsing, retrieval, and decompression before processing untrusted content.¶
Identify the source snapshot and result unit. Check that selectors, references, and content/projection bindings are unambiguous and apply to those exact inputs.¶
Perform the required integrity, attribution, freshness, and other checks under the consumer's accepted basis. Do not use a statement's own declaration as a substitute for that basis.¶
Check that source status, values, criteria, coverage, identity qualifications, and material limitations were preserved or explicitly transformed under the profile. Consider previous hops or recheck against the source.¶
Report separately: which preservation checks succeeded, which found violations, and which could not be established. Preserve reasons when more than one condition applies; a summary label must not hide a failure or unknown prerequisite.¶
Only then use the supported result as an input to a separately identified assessment or decision, if any. Retain the original source assertions and the separate basis of the new judgment.¶
These outcomes concern a bounded preservation check. A demonstrated mapping violation is not proof that the model failed a benchmark. Failure to establish preservation is not a proof that a source assertion is false. Conversely, success in every mapping check does not validate the benchmark, source honesty, or deployment safety.¶
All examples in this section are synthetic. Field names outside the explicitly described native-style fragments are explanatory notation, not a wire schema, registered vocabulary, full native log, or released AIREP profile. No model was run to produce these values. The examples illustrate required distinctions; they neither demonstrate upstream implementation defects nor establish interoperability.¶
Inspect documents run status and scoring results separately [INSPECT]. The following synthetic excerpt uses those concepts. It is not a complete EvalLog. The run finished successfully; the accuracy observation is 0.734. The excerpt supplies no criterion, so it does not establish whether the full evaluation had one.¶
{
"status": "success",
"results": {
"scores": [
{
"name": "example_scorer",
"metrics": {
"accuracy": {
"value": 0.734
}
}
}
]
}
}
An output that attributes PASS to the evaluator solely from status = success violates Section 4.3. A limited view can preserve the reported completion status and selected score while leaving the criterion state unknown. Explicit source evidence that no criterion was applied would instead support that narrower absence assertion.¶
A consumer can apply its own threshold. The following explanatory record reports a new local assessment; it does not invent a historical source verdict. The source object reference is an illustrative identifier scoped to this example, not a substitute for the binding required by a deployed profile.¶
{
"actor": "consumer.example",
"source_selection": {
"object": "example-inspect-fragment",
"syntax": "RFC6901",
"pointer": "/results/scores/0/metrics/accuracy/value"
},
"source_criterion_state": "unknown",
"new_criterion": {
"id": "local-accuracy-policy-v1",
"operator": ">=",
"threshold": 0.7
},
"new_assessment": "PASS",
"criterion_preceded_source_run": "not-established"
}
This new assessment is allowed by Section 4.5. Its authority and suitability for a real decision remain separate questions. A later evidence view must not remove the attribution and turn it into "the evaluation passed its preregistered test".¶
LightEval describes a task-keyed results object with numerical metrics [LIGHTEVAL]. The following is a synthetic excerpt, using a fictional task name and an added explanatory run identifier. The values are not a real LightEval run.¶
{
"example_run_id": "run-17",
"results": {
"example/task|0": {
"em": 0.62,
"maj@8": 0.8
}
},
"versions": {
"example/task|0": 1
}
}
Under JSON Pointer [RFC6901], /results/example~1task|0/em selects 0.62; /results/example~1task|0/maj@8 selects 0.8. These are different result units. A record saying only "run-17 passed 0.75" fails to specify the metric or the assessment rule. The slash in the task key is escaped as ~1 in the pointer. This example uses a JSON Pointer string, not an assumed URI-fragment convention for application/json.¶
A deployed mapping must additionally bind the source snapshot, native task/scoring interpretation, and population needed for its claim. A JSON pointer into a mutable URL does not satisfy those requirements. A pointer to a numerical value also does not carry the surrounding context on its own.¶
{
"population": {
"source": "declared-plan-v1",
"planned_samples": 100
},
"completed_samples": 80,
"sample_criterion_met": 76,
"sample_criterion_not_met": 4,
"not_run_samples": 20,
"derived_fraction_among_completed": 0.95
}
Here the sample categories are defined by the example to be disjoint. The observed fraction is 76/80, not evidence that 95 of 100 planned samples met the criterion. A new policy could accept incomplete coverage, but it must identify that choice separately. If the plan were unavailable, the record could not infer a total of 100. If attempts rather than unique samples were counted, the deduplication and retry rules would have to be stated. Even a verified plan does not establish that an operator disclosed all real attempts.¶
{
"source_decimal": "0.94996",
"display_decimal": "0.950",
"criterion": {
"operator": ">=",
"threshold_decimal": "0.95"
},
"comparison_on_source": "FAIL",
"comparison_on_display": "PASS"
}
The strings in this example denote exact base-10 values. Rounding to three fractional places changes the result of the stated comparison. Displaying 0.950 can be useful, but treating it as the original evidence for a 0.95 threshold violates Section 4.4. A new assessment deliberately using rounded values would require its own explicit rule and attribution; it cannot replace the earlier comparison silently.¶
The publisher can preserve the reference and its limitation. It cannot supply a made-up all-zero hash, claim a recomputed digest, infer the log's contents, or conclude that the source actually retained a complete log. Whether another actor can retrieve the object is a separate observation. A destination requiring a known digest needs a supported qualification mechanism or must report the affected object as unsupported, not counterfeit compliance.¶
Suppose an initial mapper receives the partial-population record in Section 7.3 and exports only 0.95. A second mapper signs that score and reports "complete evaluation passed". The signature may authenticate the second mapper, but it neither recovers the missing population nor authenticates the first operator's actions. A consumer without adequate upstream evidence cannot establish that stronger claim.¶
A direct recheck against the identified original can restore the missing context for a new derived view. The new view must identify that recheck, rather than assert that the context survived the earlier mapping. If a corrected source result later appears, it must be distinguished from the source snapshot used in the original decision.¶
A conforming mapping can preserve a false, biased, manipulated, or selectively published source assertion. A compromised source, converter, or verifier can sign false claims. Claim preservation is not a substitute for source trust, independent observation, evaluator competence assessment, containment, or scientific review. A loss report cannot independently prove that no omitted information existed.¶
A deployed profile MUST specify the integrity and attribution bindings required for its claims, including which actor authenticated which object. An authenticated channel can protect a transfer while leaving later re-export attribution unresolved. When offline third-party verification is claimed, the profile MUST provide evidence sufficient for that verification rather than relying only on an earlier channel session. Signature validity under an unaccepted or self-declared key is not accepted source attribution.¶
Source and mapping substitution, replay, and equivocation can make a genuine result appear to concern another model, policy, population, or time. Consumers need binding and freshness checks appropriate to the requested claim. A recent signature over an old evaluation does not make the evaluation recent. Append-only history does not by itself prove all real evaluations were logged. Revocation or a changed trust basis may alter a verification conclusion without authorizing silent rewriting of historical records.¶
Implementations MUST apply bounded resource limits to parsing, archive expansion, reference traversal, and evidence graphs. A profile using JSON needs deterministic handling of duplicate names, excessive nesting, large numbers, and invalid encodings; ambiguity in claim-relevant data MUST NOT be resolved by accidental parser behavior. The interoperability cautions in [RFC8259] are relevant. A structural validator does not necessarily perform numerical, chronological, or semantic consistency checks.¶
An evidence reference is not permission to fetch or execute its target. Consumers MUST apply an explicit retrieval policy, including access control and restrictions on destinations, redirects, schemes, and content processing. Evaluation logs, prompts, and tool outputs are untrusted data, even when embedded in a valid signed record. Instructions in them MUST NOT be executed as verifier instructions merely because they appear in evidence.¶
A display layer can undo preservation performed by the converter. Interfaces that collapse "not checked", "unavailable", "invalid", and "passed" into the same indicator can mislead a decision maker. A profile MUST preserve the distinction at the consumer interface for the claims it supports, not only in an inaccessible raw attachment.¶
Evaluation evidence may include personal data, credentials, proprietary configurations, restricted test items, or information useful for misuse. A mapping SHOULD minimize transferred content to what is needed for the selected claim. Restrictions and redactions MUST remain visible when material; they need not reveal the restricted content itself.¶
Content hashes are not anonymization. Low-entropy values can be guessed; stable digests and identifiers can correlate runs, operators, and sensitive activity. Profiles should select existing confidentiality, access-control, or commitment mechanisms appropriate to their threat model. This document specifies no confidentiality-preserving proof system and does not require publication of confidential benchmark contents.¶
References can expire, access can change, and retention obligations may require deletion. A preserved reference should not be described as perpetual availability. A later consumer that lacks material evidence MUST record its actual verification limit, not rely on a publisher's earlier access statement as proof that the bytes were checked by that consumer.¶
This document has no IANA actions. It creates no registry, media type, URI scheme, namespace, or wire-level status code. Identifiers and JSON member names in examples are illustrative only.¶
This revision specifies requirements and illustrative processing behavior, not a complete mapping profile. It claims no deployed implementation, independent implementation, interoperability result, real-model evaluation, or proof of arbitrary semantic preservation. The examples and test scenarios are synthetic explanatory material, not measurements of frontier-model safety.¶
Existing AIREP profile tooling is related work, not automatically an implementation of this document. Parsing an example, checking arithmetic, or validating an existing schema does not establish compliance with all of these requirements. An implementation report needs the versioned mappings and consumer behavior required by Section 4.12.¶
The next engineering question is whether concrete mappings using existing native formats and evidence predicates can satisfy these requirements without a new carrier. Useful work includes version-pinned Inspect and LightEval mappings, explicit treatment of score-only and partial-run cases, and independent consumer checks. A mapping profile must resolve claim-dependent materiality and source ambiguity rather than hide them in a generic success flag.¶
Whether these requirements are best maintained as an independent application profile or incorporated into related work remains open. This document makes no claim about working-group adoption. Its scope is evidence interpretation at exchange boundaries, not selection of benchmarks or policy outcomes.¶
These are proposed test obligations for concrete profiles, not a claimed executed conformance suite. A profile needs actual source and destination fixtures, accepted interpretation bases, and observable consumer outputs. An expected limitation can be the correct result; a test does not need to produce a pass verdict about a model.¶
Each test report should state exactly which checks were performed. Native-format schema acceptance, arithmetic correctness, carrier verification, mapping correctness for selected cases, and independent interoperability are different claims.¶
CP-T01 (Section 4.3): Preserve completion and score; do not attribute a criterion verdict to the source.¶
CP-T02 (Section 4.3): Keep the two meanings distinct; do not infer explicit absence from omission alone.¶
CP-T03 (Section 4.5): Permit a separately attributed new assessment with its inputs and rule; do not rewrite the source verdict or criterion chronology.¶
CP-T04 (Section 4.2): Require unambiguous result selection; reject arbitrary metric choice as support for the requested claim.¶
CP-T05 (Section 4.3): Retain the native distinction; do not promote it to a measured pass or silently coerce it to a measured failure.¶
CP-T06 (Section 4.6): Retain local outcomes and population limits; do not infer complete coverage.¶
CP-T07 (Section 4.6): Apply the declared attempt/aggregation rules and expose exclusions; do not count only favorable attempts silently.¶
CP-T08 (Section 4.4): Preserve the source interpretation and explicit derivation; do not attribute the rounded comparison to the source.¶
CP-T09 (Section 4.7): Preserve the declaration; leave actual executed identity or enforced isolation unestablished.¶
CP-T10 (Section 4.8): Report actual access and digest knowledge; no invented bytes, custodian, existence proof, or dummy digest.¶
CP-T11 (Section 4.8): Preserve reported digest and its source; do not call it recomputed by the consumer.¶
CP-T12 (Section 4.9): Report the limit or recheck against adequate source evidence; a later signature does not repair the loss.¶
CP-T13 (Section 4.10): Retain the carrier result; do not claim semantic preservation or successful model assessment.¶
CP-T14 (Section 4.11): Do not manufacture a release, pause, authorization, or control event.¶
CP-T15 (Section 4.5): Preserve the commitment; do not infer that criteria preceded the run.¶
CP-T16 (Section 4.9): Preserve snapshots, attribution, and amendment relationship; do not rewrite the earlier decision basis.¶
CP-T17 (Section 4.12): Expose unsupported semantics; do not silently reuse the old interpretation.¶
CP-T18 (Section 4.1): Require binding to the actual source selection and output; do not rely on an unrelated qualification.¶
CP-T19 (Section 4.2): Handle under deterministic profile rules; parser accident must not establish the claim.¶
CP-T20 (Section 4.10): Preserve the common source assertion and both distinct judgments; disagreement is not repaired by overwriting a result.¶