Internet-Draft Claim Boundaries September 2026
Sergeev Expires 27 March 2027 [Page]
Workgroup:
Network Working Group
Internet-Draft:
draft-sergeev-claim-boundaries-01
Published:
Intended Status:
Informational
Expires:
Author:
M. Sergeev
Independent Researcher

Claim Boundaries for Execution Evidence

Abstract

Systems that act in the world produce logs, receipts, approvals, traces, attestations, provenance statements, and transparency records. These artifacts are routinely offered as evidence that an action was authorized, performed, or completed. This document states a discipline for bounding such claims: the strength of an execution-related claim is limited by what the available evidence actually observed, constitutes, or proves, and by the control and observation topology at the boundary that produced it. Message formats, signature validity, receipt validity, and registration do not create observation or independence that did not exist. No message format can supply the independent enforcement or observation dependencies that a prevention or adversary-resistant detection-coverage guarantee requires. An appendix works through a scenario in which one party creates another and may be able to act in its name. The document defines no protocol, no record format, and no registry. It collects non-inference rules, a control-topology test for prevention and detection claims, a worked example, and reporting distinctions for evidence that does not support the claim asserted over it.

Status of This Memo

This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79.

Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/.

Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress."

This Internet-Draft will expire on 27 March 2027.

Table of Contents

1. Introduction

An operator states that "the payment was executed". The artifact behind the statement is a signed log entry produced by the operator's own service. The signature verifies. The schema validates. The entry was registered with a transparency service. None of that establishes that the payment reached the payment network, only that the operator recorded and registered a well-formed statement saying so.

Gaps of this kind are routine wherever execution-related evidence is consumed: audit pipelines, compliance reporting, supply-chain attestation, and incident forensics. They appeared recently and sharply in the IETF's agent-protocol discussions of 2026 (Appendix A), in protocols for AI agents, where one party's software performs actions whose records are produced largely by the acting side itself.

The failure mode is semantic and assurance inflation: the strength of a claim silently grows as an artifact moves between parties, formats, and summaries. A record of an invocation is cited as proof of execution. A valid signature is read as truth of the signed statement. The presence of an audit trail is read as completeness of coverage. An authorization is reported as an action performed. Each step feels small, but together they produce a claim that no producing boundary ever observed.

This document states the bounding discipline in one place, in protocol-neutral terms, so that specifications, deployments, and reviews can name the exact point at which a claim outruns its evidence.

1.1. Scope and non-goals

This document is descriptive. It defines no wire format, no transport, no carrier, no record or receipt format, no authority or authorization mechanism, no registry, and no generic evidence protocol. It does not compete with, extend, or profile any of the systems discussed in Section 13. It states analytic limits that hold whichever of those systems carries the evidence. Because it specifies no protocol behavior, it does not use BCP 14 requirement keywords: statements of the form "X does not establish Y" are claims about what an artifact can support, not conformance requirements.

This document states a claim-appraisal discipline, not an evidence-design discipline. It identifies when available evidence does or does not establish a claim, from the position of a relying party deciding what to conclude from what it has. It does not prescribe the complete set of observations, bindings, retention mechanisms, or architecture that would have to be put in place in advance so that a given claim remains determinable later. Enforcement or observation at the time of an action and determination at a later time are related, but they are not the same requirement. What further material should have been captured or bound at action time, where, and by whom, is related work and outside the scope of this document.

The discipline is not new in its parts. Bounded claim statements are already in use in specific domains: a time-stamp token is evidence that a datum existed before a particular time, and the authority time-stamps only a hash of the datum and does not examine it [RFC3161]; DomainKeys Identified Mail (DKIM) distinguishes the signing domain from the purported author and limits its integrity assertion to the content covered by the signature [RFC6376]; certificate transparency describes its logs as making misissuance detectable rather than preventing it [RFC9162]; supply-chain transparency describes itself as holding issuers accountable rather than preventing dishonest ones [RFC9943]. This document collects such statements as one discipline for execution-related evidence and claims no priority for any individual rule.

2. Terminology

This document uses the following terms descriptively. It uses "claim", "evidence", and "relying party" in their general senses. In the Remote ATtestation procedureS (RATS) architecture [RFC9334], Evidence and Attestation Results are distinct artifacts, and either can serve as evidence here. A RATS Claim is a piece of asserted information inside such an artifact, while a claim here is a statement offered for reliance. Where a RATS term is meant, it is capitalized.

Claim:

A statement offered for reliance, concerning an action: that it was intended, authorized, invoked, performed, completed, or had an effect.

Evidence:

Artifacts offered in support of a claim: records, logs, receipts, signatures, attestations, provenance statements, transparency entries, and similar.

Producing boundary:

The vantage point, component, or interface at which a piece of evidence was produced, together with what that vantage point could observe and who controls it. This document uses "producing boundary" throughout as its genus term for that vantage point.

Interested party:

A party whose conduct the claim describes or whose incentives are served by the claim being accepted; typically the operator of the acting system.

Relying party:

A party deciding whether to accept the claim.

Control topology:

The arrangement of which parties control which enforcement points, observation points, keys, and delivery paths relevant to a claim.

Independence:

A property of deployment and control. A statement of independence names a component and its role: enforcement, observation, delivery, or evaluation. It can instead name a record or the grounds of a premise. It names a party or a coalition (parties acting together), a period, and a threat model. The named thing is independent of that party with respect to a capability if the party lacks that capability over it. Section 8.1 names three capabilities: forgery, concealment, and circumvention. Which of them matter depends on the claim. Independence describes control. Whether the evidence supports the claim is a separate question (Section 3). A label, a field value, a distinct key, an organizational name, or a count of components does not by itself establish independence.

3. The bounding principle

The strength of an execution-related claim is bounded by:

  1. what the available evidence actually observed, constitutes, or proves, at the boundary that produced it; and

  2. the control and observation topology at that boundary: who operated it, who could bypass it, and through whose hands its output travels.

Message format, signature validity, receipt validity, and registration do not raise either bound. A format can carry a statement about control or observation, but it cannot create control or observation that did not exist. A signature authenticates bytes under a key. By itself it establishes neither the truth of what the bytes assert nor the signer's control of the boundary they describe. A receipt's evidential force depends on the property being checked, the observations or proofs that support it, and the applicable trust and deployment assumptions. Successful verification of a registration receipt can support that a statement was registered under the service's policy. Registration alone does not establish that the event described in the statement occurred.

These limits concern what follows from the stated evidence and premises. They do not preclude an inference from complementary observations whose relationships and load-bearing assumptions are established. Cryptographic verification can support those assumptions. It does not supply a missing observation by authenticating a statement about it.

Support does not have to come from direct observation of the claimed action. Section 9 describes the other routes and their premises.

In this document, evidence supports a claim when the claim follows from the evidence under the stated premises. One practical test checks this. Try to describe two cases in which the relying party holds the same evidence under the same premises, but the claim is true in one case and false in the other. If both cases are possible under those premises, the evidence does not support the claim as stated. Failing to find such a pair does not prove that none exists. Evidence that only makes a claim more likely does not support the claim in this sense, although it can still support a weaker claim (Section 11).

4. Evidence bases: what did the evidence observe?

Execution-related artifacts answer different questions, and the differences are load-bearing. At minimum, the following bases are worth distinguishing:

Observation:

Some event concerning the action was recorded or reported. Absence of a record alone is not an observation that an event did not occur.

Intent:

A specific intention, plan, approval, or decision concerning the exact action is evidenced. This establishes neither the validity of the approval nor anything about performance.

Invocation:

The exact request crossed an observed invocation boundary. This does not establish that the invoked operation executed or had an effect.

Execution:

Performance of the exact action was evidenced at a boundary in a position to observe execution. This does not establish completion of a larger process, a durable external effect, goal satisfaction, correctness, or policy compliance. Section 9 addresses claims about an effect, its persistence, and its finality.

The observation base asks only whether some event concerning the action was recorded or reported. Elsewhere in this document, an observation is an event seen at a boundary in a position to see it (Sections 3 and 9). Evidence at the observation base need not be an observation in that sense.

These bases are distinct questions about one action. They are not levels on one scale, and they need not be mutually exclusive. A conclusion from one base to another holds only under stated premises. Evidence relevant to one base does not, merely by being assigned to that base, satisfy the evidential criteria relied on for another. Distinctions of this kind, between stages of an action's lifecycle whose evidence does not carry across, are long established in distributed-systems practice (for example the separation of a remote procedure call's invocation from its execution and its result). One executable realization that binds exact predicates to these four bases appears in [I-D.sergeev-wexp-core]. This document does not depend on it.

Other distinctions cut across the bases. Verification reach (who can check a record) is orthogonal to claim strength: a third-party-verifiable invocation record is still invocation evidence. And asserted content differs from observed content: a mediator's record may faithfully report what the mediator asserts about a downstream outcome it never observed. A useful record says at which boundary it was produced and whether each statement in it was observed there or merely asserted there.

The action referred to by a claim should be identified at the relevant boundary. A business intention, an authorization, a request attempt, an execution, and a durable effect need not have a one-to-one relation. Retries, proxies, and idempotent processing can change those relations. A composite claim depends on evidence establishing the particular bindings it uses. This document prescribes no common identifier format.

5. The claim ceiling

In this document, a kind of claim is a family of claims to which the same type of evidence is relevant, such as the bases of Section 4. For a stated kind of claim, and under stated premises, a producing boundary limits the claims its evidence can support. Call that limit its claim ceiling. The term does not presume that all claims of one kind are comparable, or that one uniquely strongest claim exists.

A boundary in a position to observe invocations, but not execution, cannot by its observations warrant execution claims. That holds no matter how many well-formed execution assertions its records carry.

The underlying idea is old and this document claims no priority for it: a conclusion cannot be stronger than the vantage from which its supporting evidence was produced. Section 13 describes prior formulations.

A claim ceiling is an exclusion rule, not an evidence source. It can prevent a stronger claim from being supported, but it never creates support, so a supported claim still requires positive evidence at its own base.

A ceiling is also only as good as the grounding of the boundary description itself. A boundary descriptor asserted by the producer is only the producer's own account. Section 8.1 describes what grounds a statement about who controls a boundary.

The accepted boundary description must apply to the time and deployment context of the claimed action. A later assessment or a current configuration does not, by itself, establish the earlier boundary's properties.

The same applies to any premise that would make the producer's records informative without observation. Where a claim rests on the producer's incentives or objectives being known, that knowledge is itself a premise requiring its own grounding. An alignment asserted by the producer is only the producer's own account.

This document defines independence in terms of control. Payment, contract, or ownership can give one party the means to direct another party's output or decision. They count here to the extent that they give that capability.

A record that requires no later cooperation from the party whose conduct is at issue is not, for that reason, independent of that party. A record it made at the time and left in place needs nothing from it afterwards and remains its own account of itself. The distinction between independence from cooperation and independence of the evidence itself follows an exchange with Douglas Wadkins; see [I-D.wadkins-agentproto-action-determinability].

A claim ceiling is relative. It is stated for one kind of claim, at one observation boundary, with the support relation proper to that kind: which evidence supports which claim of that kind. Examples are the bases of Section 4 and, separately, the prevention claims and the detection claims of Section 8.

Claims of different kinds are not ranked against one another by a ceiling. An execution claim and a detection claim are different kinds of claim, not points on one scale, and this document defines no scalar scale, level, or score on which all claims are ordered. "Stronger" and "weaker" in this document are always read within the kind of claim at issue.

6. Non-inference rules

The following rules restate Section 3 as individual limits. Each names an inference that is invalid without additional, separately established premises. None is original here; scoped versions of several appear in the documents cited in Section 13.

Each rule is justified analytically. An example or a standards reference given with a rule is an illustration, and it is not presented as a documented incident. Claims about practical occurrence or consequences require their own support.

The list is not closed. A reader who finds a promotion these rules do not name has found a gap in this section, not a license for the promotion.

  1. Successful signature verification does not by itself establish the truth of an event claim in the signed content. Establishing that correspondence requires grounds beyond signature validity. Separately, signature validity alone does not establish the signer's control of the boundary described by the claim. The same distinction is drawn for verifiable credentials, where verifiability of a credential does not imply the truth of the claims encoded in it [VC-DATA-MODEL], and for signed Decentralized Identifier (DID) documents, where proofs in the document do not by themselves necessarily prove control over the DID [DID-CORE].

  2. Attribution of an actor is not proof of actorship. That a record is attributed to a party -- by a key, an identifier, or a credential associated with that party -- does not establish that the party performed the action, where another party can obtain, invoke, or emulate the attributed capability. Where one party creates or hosts another and can reach that other's signing or authenticating capability, an action attributed to the hosted party may have been performed by the host. Where a record reaches the relying party through intermediaries, attribution established at one step of that path does not carry to the next without that step's own basis. A transport can establish which component delivered a record. A source that the deliverer names inside the record is the deliverer's assertion, unless the record carries its own verifiable binding to that source.

  3. A receipt, record, or registration is not external execution or effect. Registration establishes that a statement entered a log or service under that service's policy. Matching records held by different parties establish that their contents agree (correspondence), not that the parties agree. Correspondence does not establish which event preceded which (precedence), and it does not establish occurrence.

  4. Presence of records is not completeness of coverage. A set of records, each individually authentic, does not establish that the set is all the records there are. Detecting omission requires more than the authenticity of each record: an independently grounded expectation of the complete population, which the set itself may carry, or a recording boundary that cannot be bypassed.

  5. Absence of a record is not non-occurrence, unless the observation regime supports that inference. The inference from "no record" to "no event" is valid only where a declared, enforced recording boundary meets three conditions. It covers the event class. It cannot be bypassed by the parties in question. And gaps in the record stream are themselves detectable, for example against a pre-declared cadence or sequence carried in the records and watched by a party outside the producer. Otherwise absence is only absence.

    Such an inference also requires an identified interval, evidence that the required recording and delivery coverage held for that interval, and resolution of relevant delays and gaps. Detectability of a gap does not license a non-occurrence claim while the gap remains unresolved. The conclusion is limited to the specified event class, population, and interval under the stated assumptions.

    Bypass of the recording boundary is not the only way a record can fail to appear. A record may have been made and then withheld, whether by a retention policy that removed it or by a disclosure decision that did not produce it [AP-SCHROCK-RETENTION]. Where either is possible, the result is a statement about the evidence set examined and not about the recording boundary, and the conditions above are not met by the boundary alone.

  6. Authority granted is not action performed. Authorization evidence, however exact, single-use, and attenuated, bounds what a grant permits at the point where it is enforced. It does not establish that the authorized action occurred. What is missing is evidence on the action's own basis: an observation of invocation or of execution, at a boundary in a position to make it, or a record that constitutes the action rather than describing it (Section 4 and Section 9).

  7. Internal observation is not independent observation of an external effect. An observation made inside the acting party's boundary can establish, at most, what that boundary observed, and a record's statements about effects at other systems are assertions. In particular, a record minted by the deciding or authorizing side cannot by itself establish the order of its own decision against an effect that side does not observe. Other evidence carried in the same record is appraised under the routes and premises of Section 9, regardless of its source.

  8. Time and sequence evidence is bounded by the observation and control properties of its source, like any other evidence. A timestamp or sequence number is only as strong as the party and mechanism that produced it: one minted by the interested party over its own record orders that record, not the world. This is not a claim that time evidence is worthless. Trusted time-stamping has a distinct and stronger evidentiary role: a time-stamp token from an appropriately trusted authority is evidence that a datum existed before a particular time [RFC3161]. And a sequencing mechanism whose observational domain covers both of two events can order them. The rule is that the strength of the time or sequence claim follows the source, and must not be read past what that source observed and controlled.

  9. A missing or unverifiable load-bearing premise cannot be silently promoted. Where support for a claim depends on a premise that was not evaluated, or that only the interested party can vouch for, the claim inherits that limitation. The report names the claims that are supported, and keeps the status of each premise that was not evaluated. A conditional claim states the premise it depends on (Section 11).

  10. Current authentication, when its verified scope is limited to a present binding, does not by itself establish identity continuity or succession relative to an earlier subject. Identity continuity, succession, and the present applicability of prior authorization, standing, or reputation are distinct claims. Each requires grounds covering the specific relation or applicability asserted under the relevant identity and authorization rules. A record's evidentiary force depends on those rules and its verified properties, not merely on its statement of the claim. DID Core states the corresponding limit for persistent identifiers: absent published operational policies, requesting parties are not expected to assume that an identifier is persistent for the same subject [DID-CORE].

  11. Behavior observed under one set of conditions does not establish the same behavior under other conditions without grounds for carrying the conclusion across. A test, an audit, or a monitored run observes a system under the conditions it applies. The result describes those conditions. Carrying it to other conditions needs grounds that the two do not differ in anything the behavior depends on. Where the system can detect that it is being tested or observed, results obtained under those conditions do not by themselves supply such grounds, since the system can make its behavior depend on them. Appendix B reads the emissions case this way.

  12. Multiplicity is not independence. A number of records, keys, services, or organizational names does not by itself establish independent support. Where one party has a capability that matters to the claim over several of them, their number adds nothing against that party (Section 8.1). Independence is judged claim by claim, for the components on whose correctness the support for that claim relies.

  13. Coherence is not corroboration. The detail, internal consistency, plausibility, or fluency of an account produced by the interested party does not by itself support the claim that what it describes occurred. A fabricated account can have all of these properties, so they do not distinguish a true claim from a false one (the test of Section 3). This holds with particular force for accounts generated by software, including language models, whose output can be fluent whether or not it matches the events it describes. Support needs grounds beyond these qualities, and evidence carried with the account can provide them by one of the routes of Section 9.

Every rule here is stated against premises. The premises a conclusion rests on, their grounds, and their status should be available directly or by unambiguous reference. A premise accepted as a matter of policy is not thereby independently established. Where the premises change, whether the conclusion remains warranted has to be checked rather than assumed.

Aggregation, transformation, summarization, and re-signing remove none of these limitations. A pipeline that normalizes, merges, or re-encodes evidence inherits each relevant limitation of its inputs unless a specific limitation is specifically resolved by additional evidence. An unsupported widening of a claim at the output of such a pipeline is a defect to locate, not a result to report.

7. A claim-boundary review lens

The rules above can be applied as a short checklist when reviewing a specification, a deployment, or an evidence design. For a given claim, ask:

  1. Exact claim. What exact claim, about what exact action, is being made?

  2. Producing boundary. Which boundary produced the supporting evidence, and what could that boundary observe?

  3. Control. Who controls that boundary? Can the interested party or coalition forge its output, conceal it, or circumvent it (Section 8.1)? Which of these capabilities matter to the claim being made? Did the limits on them hold for the period the claim depends on? What grounds those limits?

  4. Ceiling. For the kind of claim at issue, what claim or claims can that boundary warrant? Does the asserted claim exceed them?

  5. Positive evidence. Is there positive evidence at the claim's own base? A ceiling supplies no evidence, and a signature or a registration supports only the predicate that its verification establishes.

  6. Prevention, detection coverage, or a particular finding. If the claim is that misbehavior is prevented, which independently controlled enforcement point does it rest on (Section 8)? If the claim is that misbehavior of a specified class cannot be concealed, which independently controlled observation and delivery dependencies does that coverage rest on? For a claim that a particular event was detected, which evaluation act (evaluator, predicate, evidence basis, window or context) is identified?

  7. Basis for a claimed effect. Which route of Section 9 supports the claim? The routes described there include observation at a boundary outside the interested party's control, a record that constitutes the action, and evidence whose verification establishes a predicate. Where the support rests on another basis, is that basis stated? What limits its scope?

  8. Completeness and absence. Does the claim depend on the record set being complete, or on an absence meaning non-occurrence? Is that supported (rules 4 and 5)? What interval and delivery horizon were evaluated, and were relevant gaps resolved?

  9. Pipeline widening. Does any step between production and reliance widen the claim beyond what the producing boundary supported?

  10. Reporting. Does the report separate whether each evaluation was made from what it found (Section 11)? Does it name the supported claims, and the reason for any evaluation that was not made? Is anything silently rounded up?

The considerations sections of [RFC3552] (security) and [RFC6973] (privacy) are the models for this genre: a document states, in its own terms, what its mechanisms do and do not establish, and what residual exposure remains. This lens applies the same documentation discipline to the semantic strength of execution-related claims. It is a review aid, not a conformance procedure.

8. The control-topology test

Claims that an architecture prevents or detects misbehavior are execution-related claims about the architecture itself, and they are bounded the same way. Farrell's (a,b,c) scenario, described in Appendix A, motivates the control-topology test used here. Stated mechanism-neutrally:

Prevention:

A claim that party B is prevented from performing or forging an action requires an enforcement point that B does not control and cannot bypass, positioned so that the action cannot complete without it. A component outside B's control that is not actually required for the action prevents nothing.

Detection coverage:

A claim that misbehavior in a specified class cannot be concealed by B requires the following.

Evidence:

adequate to distinguish that misbehavior under the stated assumptions.

Controls:

delivery or omission controls that prevent B from hiding it without an observable indication.

The claim should state its coverage, timing, and failure limits. These coverage conditions are distinct from establishing that a particular event was detected from evidence already obtained.

Independence:

The prevention and coverage properties above depend on independence in the sense of Section 8.1. The components concerned are the enforcement point, the observation point, and the delivery path. The party concerned is the stated party or coalition.

An architecture claiming prevention or adversary-resistant detection coverage should state:

Without those dependencies, that guarantee is not established against adversarial B, whatever its message formats carry. This does not invalidate a narrower finding supported by evidence that has actually been received and evaluated.

A protected observation path is not, by itself, sufficient for detection. The evidence must support a check that distinguishes the specified misconduct under the stated assumptions. An observable gap can establish a loss of coverage without establishing which underlying event occurred.

A claim that a particular event was detected requires an identifiable evaluation act. The record should name which evaluator evaluated, which predicate was checked and with what result, and on what evidence basis. Where a window or context applies, it should name that too.

Such a finding does not, by itself, establish that all events of that class were detectable or detected. A path that could have revealed an event does not show that anyone looked. This document defines no format for stating those items.

Existing mechanisms slot into this test rather than exempting themselves from it. Remote attestation [RFC9334] contributes where all of the following hold:

RFC 9334 permits its roles to be aggregated into one entity, so independence from a given party is a property of the deployment, not of the architecture.

Threshold signing [RFC9591] contributes a prevention dependency exactly where the adversarial party or coalition cannot reach the threshold without an independently controlled signer that checks the exact operation and can withhold its share. These mechanisms contribute only to the properties and dependencies actually established by their deployment.

OAuth sender-constraining mechanisms such as Demonstrating Proof of Possession (DPoP) [RFC9449] can limit the use of stolen tokens. They do not by themselves prevent B from invoking a key through a signing interface B controls. That holds whether the key was provisioned by another party or stored as non-exportable (Section 11.4 of [RFC9449]).

An isolated execution environment, including a trusted execution environment (TEE), can contribute an enforcement dependency where B cannot modify or bypass the relevant checks. Claims about effects outside that environment still require an appropriate evidence basis.

The Workload Identity in a Multi System Environment (WIMSE) architecture [I-D.ietf-wimse-arch] separates workload authentication from authorization and permits different placements of policy enforcement. The test therefore asks who controls the deployed enforcement point and which operations it covers, rather than treating workload identity or credential provisioning as evidence that the claimed operation occurred.

The independently controlled point need not be cryptographic. A preserving repository may be audited and certified against a published standard [ISO16363], by a body the repository does not control and which is itself accredited under a separate regime [CCSDS652-1]. That contributes a detection dependency if the audit observes the practice at issue and if any suppression of an adverse finding by the audited party would become visible.

That contribution is limited by the audit's scope, period, sampling, access, and reporting arrangements. Certification alone does not establish complete observation of all relevant conduct, nor the visibility of every suppressed adverse finding. The recursion terminates in an accreditation body that some party must be prepared to treat as terminal.

The test is indifferent to whether the point outside the party in question is a key, a verifier, or an accreditation regime. It asks only whether the deployment placed one there, and what that point actually covers.

8.1. Independence and the claim at issue

A statement of independence says which capabilities of the party matter to the claim at issue. It says what grounds the limit on each of them. This document does not say which grounds suffice to establish that such a limit holds.

The capabilities in question are these.

Forgery:

The party can make a relying party accept false evidence. One way is to fabricate the evidence, substitute or replay it, or alter it without detection. Another is to direct the component that produces the evidence, so that its output is false.

Concealment:

The party can act outside the recording boundary, so that the activity is not recorded there. It can also suppress, withhold, or delay evidence that exists.

Circumvention:

The party can complete the action without the enforcement point, or direct the decision of that point.

Where this document says "bypass", bypassing an enforcement point is circumvention and bypassing a recording boundary is concealment.

A claim needs limits only on the capabilities that matter to it. A capability matters to a claim if the party could use it to leave the relying party holding the same evidence while the claim is false. That is the test of Section 3, applied to what the party can do. Take a particular finding made from evidence already received. Forgery matters to it, and so does the binding of the evidence to the exact action (Section 9). A payment network's record of transfer T can support that T occurred. This holds although the operator could have withheld the record or made other transfers elsewhere. Concealment matters to completeness and to detection coverage (rules 4 and 5). The same record does not support that these are all of the operator's transfers. Concealment can also matter to a particular finding, where omitted evidence would change the finding. Circumvention matters to prevention. The limits have to hold for the period the claim depends on. A present assessment does not by itself establish that the limits held earlier.

Independence does not make an observation relevant to the claim. An observer outside the operator's control that saw a command sent has not thereby seen it executed. A separate question remains: can the observation distinguish a true claim from a false one (Section 3)?

A dependency can enter where evidence is formed, where it is delivered, and where it is evaluated. An independent evaluator can verify the interested party's own account, but the account does not become an independent observation. Control over delivery gives the capability to conceal, and it can also permit substitution or replay, but it does not give control over the source. An independent component does not make a record presented as its output authentic: the operator can present a forged confirmation from an independent payment network. The relying party therefore checks origin, binding to the exact action, and any required freshness separately (Section 9).

Two sources that take a fact from the same origin both depend on that origin for that fact. Two logs that each copy the operator's message are not two observations of execution (rule 3). Where one party has the relevant capability over several components, those components do not count separately against that party. Three services that act, record, and check, run by one operator who can change their code or their records, divide the work but not the control. The useful question is which error or fabrication the second source rules out. Independence from one party does not establish independence from a coalition that includes it.

A statement of independence is itself a claim, with a producing boundary and a ceiling. Its grounds can stop at a premise that a party accepts as a matter of policy. A claim that relies on the statement then stays conditional on that premise. The premise travels with that claim when it is passed on (Sections 6 and 11). The audit example in Section 8 shows one such premise. Section 9 describes evidence whose source need not be independent.

9. When stronger claims become supportable

The discipline does not only rule claims out. Stronger claims about a particular effect become supportable when the effect is observed at a boundary capable of observing it. That capability is judged under the stated threat and control model. The evidence of that observation must also be protected against undetected fabrication or alteration by the interested party. Coverage claims additionally depend on the observation and delivery conditions described in Section 8.

An external footprint is one important route to such observation. Examples are movement on a payment rail, and resource or configuration records in an external provider's control plane. In each, a system the interested party does not control observed something at its own boundary, and a relying party can check that system's records against the claim.

It is not the only route. The observing boundary may be any of several things:

A distinct third party is not always necessary. The relying party's own observation may provide the required evidence, subject to the same claim-specific trust and deployment assumptions.

A further route relies on a record constituting the precise action claimed, rather than describing a separate action:

Such a record supports that precise claim, because issuing or accepting it constitutes the action under the applicable rules. The evidence must establish the conditions that give the record that constitutive effect, including any applicable authority, acceptance, and commit or finality conditions. A record on a branch that was never accepted does not constitute the action merely by resembling one that would.

What it supports is bounded by the action it constitutes and extends no further. A ledger entry that is the transfer establishes the transfer, not that goods moved. A signed order that is the order establishes the order, not that it was carried out. Where a record describes an action it does not constitute, this route does not apply.

Another route relies on evidence whose verification itself establishes a predicate. Examples are a proof of computation, a proof of possession, and a response to a fresh challenge. The source of such evidence does not need to be independent of the interested party. Successful verification establishes the proved predicate, under the premises of the verification. Those premises include the soundness of the proof system, the integrity of the verifier, and that the statement verified is the one the claim names. Where soundness bounds the probability of a false acceptance rather than excluding it, treating it as exact is an idealization, and a report that relies on it should say so. A claim that does not follow from that predicate needs its own evidentiary grounds and the bindings it requires. For a proof of computation, such claims can concern whether the committed inputs match the facts they stand for. They can concern who performed the computation, when a policy was applied, or whether an external effect followed. Signature verification is the familiar case: the proved predicate is that the bytes were signed under the key, and rule 1 applies to anything beyond it.

The qualifications below travel with these routes:

10. Worked example: a payment

Consider an agent that reports having paid a supplier. Three evidence artifacts and a report combining them are on the table. The point of the example is that they support different claims, and that a report must keep separate what each boundary observed from what it merely asserts.

  1. A self-produced execution record. The operator's own service emits a signed record: "paid supplier S, amount X, at time T". Observed at that boundary: that the operator's service produced and signed this statement. Asserted, not observed there: that the payment reached S. What it supports on its own: that the operator stated a payment; not that the payment occurred (rules 1, 3, 7).

    This example assumes that the record contains only the operator's own statement. If it carries independently verifiable evidence from another source, that evidence is appraised at its own producing boundary. The party assembling the container does not determine every item's origin.

  2. A transparency registration. The signed record is registered with a transparency service, which returns a receipt carrying an inclusion proof. Observed: that this statement was registered under the service's policy. Any claim about registration time depends on the time evidence provided and its verified scope. What it does not establish: the truth of the statement or the occurrence of the payment (rule 3). Registration raises auditability, not the claim ceiling.

  3. External payment-network evidence. A record from the payment network -- a settlement entry or network-issued confirmation -- that the network observed a transfer matching the action. Observed at a boundary outside the operator's control: that the network saw a transfer with these attributes. This is the footprint of Section 9, and it must be bound to the exact action: a bare amount-and-time match is only correlation. Its strength holds against the operator, not against the payment network's own operator.

  4. A composite report. A report that draws on all three should read, in substance, as three statements. The operator asserts a payment (1). The assertion is registered and independently checkable as an assertion (2). And where (3) is present and bound to the exact action, a system outside the operator observed a matching transfer at its boundary. That last supports a payment-reached-the-network claim, to the extent the binding holds and that evidence is protected against fabrication or undetected alteration by the operator. With only artifacts (1) and (2), the supportable claim is that the operator stated and registered a payment, not that a payment occurred.

A report over these artifacts could read as follows. Evaluator V checked the origin of each artifact and the binding of the network record to the payment. V accepts one named premise: the network's record is reliable for the event it reports. Assume that the record reports that a transfer was seen, and says nothing about finality. V did not assess whether the set covers all of the operator's payments.

Suppose instead that only the amount and the time match, and the binding is not established. The third entry is then unsupported for this payment. That the network saw some transfer remains supported.

The example makes no universal assertion about what any particular service observes. What a given payment service, transparency service, or agent runtime actually observed is a fact about that deployment, to be stated from its evidence, not assumed from its role.

11. Reporting what the evidence supports

Where evidence does not support the claim asserted over it, the useful output is not a bare failure. For each claim, a report says whether the evaluation was made and, if it was, what it found. These are separate questions. An outcome-binding profile drew a related line earlier: its indeterminate state, used when required evidence is missing, unauthenticated, unpinned, or not exactly bound, is not a comparison outcome [I-D.schrock-ep-outcome-binding].

Whether the evaluation was made:

Evaluated:

The relevant assessment ran over identified evidence, in an identified appraisal context.

Not evaluated:

The assessment did not run, the identified evidence could not be obtained, or the premise was outside the evaluation's scope. This status supplies no finding for or against the claim. It does not mean the claim is false, and it does not determine what the available evidence would support if evaluated. Nor is it permission to discard evidence that positively supports a claim at its own base. A claim well supported at its own base stays supported, even while another claim is not evaluated. It is also not a reason to repeat an action whose outcome is unknown, since the first attempt may have completed. That profile states the same limit for its indeterminate state [I-D.schrock-ep-outcome-binding].

Unverifiable:

A reason for "not evaluated" that belongs to the evaluator. The statement is one that the evaluator in question cannot check from the evidence and inputs available to it at the time of evaluation. The standing example is a producer's description of its own deployment, which an evaluator may be unable to check without deployment access or other adequate evidence about that deployment. The category is relative to an evaluator and to an evidence or input set. Where they matter, it is relative to an evaluation context or time as well. A statement can pass from verifiable to unverifiable while the record itself is unchanged. The evaluator that could check it may cease to exist, or lose its access or the means to interpret what it holds. Such statements can still be worth carrying. The report names the reason and the premise that was not checked. Where a statement rests only on the producer's account, the report says so. A reader can then see which parts of a composite claim rest on trust. A report that fixes the category without fixing the evaluator and the moment says less than it appears to.

What the evaluation found:

Supported:

The evidence evaluated supports the claim as stated, under the stated premises.

Unsupported:

The assessment ran over the evidence evaluated, and that evidence failed it. This is a finding about the evaluated evidence, relative to the evidence set and the appraisal context in which it was evaluated. It is not a property of the world, and a different evidence set may support the claim. Where the assessment ran and the evidence evaluated lacks what the claim needs, the claim was evaluated and the finding is unsupported.

Refuted:

The evidence supports the negation of the claim. Absence of a record refutes a claim of occurrence only where rule 5 of Section 6 permits the inference from absence to non-occurrence. The report should state that finding and its basis explicitly, rather than treating it as mere absence of support.

One reporting practice applies to both questions:

Downgrade to the supportable claim:

Where the asserted claim is unsupported or not evaluated, report the strongest claim of the kind at issue that the evidence does support, or the supported claims where there is no unique strongest one, alongside the asserted claim, with its status and any finding. A claim supported at another base is a separate finding, not a downgrade; support at one base is not inferred from support at another (Section 4). "Invocation: supported. Execution: unsupported by the evidence evaluated" is actionable; "invalid" is not.

Collapsing "not evaluated" into "unsupported", or either into "refuted", destroys information a relying party needs: "we checked and it failed" and "we could not check" call for different decisions. These terms are not a result-code set. One report can say that a claim was not evaluated because a premise is unverifiable, and also identify a weaker claim that is supported.

A support assessment over the available evidence can complete while a required premise is not evaluated. The report identifies these assessments separately.

Each of these terms is relative to the evidence actually evaluated, and a report is more useful when it says what that evidence was. The evaluated set may itself be incomplete or selected. Every record in it can verify while records that would have changed the outcome were never delivered: rule 4 of Section 6, met again at the reporting layer.

A verdict applies to the evidence basis it names, and a verdict that names no basis invites exactly the widening this document is about.

Appraisal results travel. One evaluator's result may be consumed by a later composite report: that a claim was not evaluated, for instance. The later relying party then needs to distinguish an attributable appraisal from an unattributed conclusion.

Propagated or aggregated appraisal results should therefore retain enough provenance to identify:

Without that, "not evaluated" degrades across aggregation into somebody's unattributed conclusion. That is the aggregation limit of Section 6 arriving at the reporting layer. Where evaluated evidence or received appraisals support incompatible findings, the report identifies the conflict and any premise used to resolve it.

This is a consideration for reporting designs. It is not a record format and not a general retention requirement. The provenance of an appraisal identifies who is reported to have appraised what, under which context. By itself it establishes neither that the appraisal ran as reported nor that the underlying event occurred.

A reporting design that preserves these distinctions makes inflation visible: every summary that would erase one of them is a place where a stronger claim would otherwise silently replace a weaker one.

12. Applicability beyond AI agents

Nothing in this discipline is stated in terms specific to AI agents. The same bounding questions apply to:

Agent systems sharpened the problem. Actions are initiated by software whose records are produced mostly at the acting side, and delegation chains multiply the boundaries across which claims travel and inflate. But the rules in Section 6 are stated as properties of evidence rather than of agents.

One domain is worked here: the long-term preservation of records, where no agent acts and no payment is made. A preserving repository holds a record deposited by its creator. Fixity evidence can support a claim that deposited bytes have not changed relative to an accepted baseline. Identity and provenance evidence may support an authenticity assessment under stated assumptions. Archival diplomatics drew this distinction earlier and for a different substrate. It separates a record's reliability at the point of creation, the accuracy of its content, and its authenticity [INTERPARES] [DURANTI1998]. Neither finding, by itself, establishes that the record was reliable when created or that its content was accurate. Preservation does not convert custody integrity into truth of the original content. That is an instance of the claim ceiling of Section 5.

A gap in a deposited series supports no conclusion about what was never deposited, unless the deposit regime was declared, enforced, and observable to a party outside the depositor (rule 5). Whether a preserved record can still be evaluated depends on the record, the available representation and provenance information, and the evaluator's knowledge and access. This illustrates the evaluator- and time-relative unverifiability of Section 11. The Open Archival Information System (OAIS) reference model addresses continued understandability for an identified Designated Community [ISO14721].

Appendix B gives three public findings that state the same bounds outside agent systems. Applicability to the other domains listed above is asserted on the same grounds and is not worked case by case here.

14. Security considerations

This document specifies no new protocol mechanisms. Its subject is the prevention of a class of security failures that occur at the semantic layer: relying parties accepting claims stronger than the evidence supports.

Cautions about the discipline itself, including its misuse:

15. IANA considerations

This document has no IANA actions.

16. Informative References

[ANDERSON72]
Anderson, J. P., "Computer Security Technology Planning Study", ESD-TR-73-51, Vol. II, HQ Electronic Systems Division (AFSC), , <https://csrc.nist.gov/files/pubs/conference/1998/10/08/proceedings-of-the-21st-nissc-1998/final/docs/early-cs-papers/ande72.pdf>.
[AP-ECKEL]
Eckel, C., "Re: DRAFT minutes from the AGENTPROTO BoF", message to the agentproto@ietf.org mailing list, , <https://mailarchive.ietf.org/arch/msg/agentproto/dZgLXO2xr3tr8pF0yejOZhj49-0>.
[AP-FARRELL-ABC]
Farrell, S., "Re: DRAFT minutes from the AGENTPROTO BoF", message to the agentproto@ietf.org mailing list, , <https://mailarchive.ietf.org/arch/msg/agentproto/wrQqZW9Dh3Yj6N7gV9RcfN5R8Kk>.
[AP-FARRELL-CHEAT]
Farrell, S., "Re: DRAFT minutes from the AGENTPROTO BoF", message to the agentproto@ietf.org mailing list, , <https://mailarchive.ietf.org/arch/msg/agentproto/zDQjiJvUhMiv5EX2Jpgbdacs3eo>.
[AP-JIANG-CONTINUITY]
Jiang, Y., "Re: draft-bu-agentproto-security-principal-binding-06 published: focused review request", message to the agent2agent mailing list, , <https://mailarchive.ietf.org/arch/msg/agent2agent/WtvC4rdcZ4dNxjpzIm1FM0y4-0o/>.
[AP-LIU-DPOP]
Liu, D., "Re: DRAFT minutes from the AGENTPROTO BoF", message to the agentproto@ietf.org mailing list, , <https://mailarchive.ietf.org/arch/msg/agentproto/HdNkJZ48W0sg5JxK2gFkYkAK_Bs>.
[AP-MORRISON-RULES]
Morrison, B., "Re: New I-D : A Dimensional Model for Characterizing AI Agent Protocol Proposals and Their Substrates", message to the agent2agent mailing list, , <https://mailarchive.ietf.org/arch/msg/agent2agent/bRdOV5Ou66hFkNd5kt7Wd4KqHAg/>.
[AP-SAMMARTANO]
Sammartano, S., "Re: DRAFT minutes from the AGENTPROTO BoF", message to the agentproto@ietf.org mailing list, , <https://mailarchive.ietf.org/arch/msg/agentproto/ike9opEcHLoS6QDdkJ0SKzSA96Y>.
[AP-SAMMARTANO-FRAMING]
Sammartano, S., "Re: DRAFT minutes from the AGENTPROTO BoF", message to the agentproto@ietf.org mailing list, , <https://mailarchive.ietf.org/arch/msg/agentproto/ALrRxQEqT6VmBd6jBTuIKETcno8>.
[AP-SCHROCK-RETENTION]
Schrock, I., "Re: New I-D : A Dimensional Model for Characterizing AI Agent Protocol Proposals and Their Substrates", message to the agent2agent mailing list, , <https://mailarchive.ietf.org/arch/msg/agent2agent/kWvHTqephUemSNjbGQWGmoZkYSU/>.
[AP-SERGEEV-GATE]
Sergeev, M., "Re: DRAFT minutes from the AGENTPROTO BoF", message to the agentproto@ietf.org mailing list, , <https://mailarchive.ietf.org/arch/msg/agentproto/ORdiqDGi8Fno6ROXHmSljQFmEOY/>.
[CCSDS652-1]
Consultative Committee for Space Data Systems, "Requirements for Bodies Providing Audit and Certification of Candidate Trustworthy Digital Repositories", CCSDS 652.1-M-3, Issue 3, , <https://ccsds.org/Pubs/652x1m3.pdf>.
[DID-CORE]
World Wide Web Consortium, "Decentralized Identifiers (DIDs) v1.0", W3C Recommendation, Sections 9.2 and 9.11, , <https://www.w3.org/TR/2022/REC-did-core-20220719/>.
[DURANTI1998]
Duranti, L., "Diplomatics: New Uses for an Old Science", Scarecrow Press, .
[EPA-VW]
United States Environmental Protection Agency, "Learn About Volkswagen Violations", agency overview of the Clean Air Act violations, including the Notice of Violation of 18 September 2015, , <https://www.epa.gov/vw/learn-about-volkswagen-violations>.
[EPA-VW-NOV]
United States Environmental Protection Agency, "Notice of Violation", letter to Volkswagen AG, Audi AG, and Volkswagen Group of America, Inc., , <https://www.epa.gov/sites/default/files/2015-10/documents/vw-nov-caa-09-18-15.pdf>.
[FRE902]
United States, "Federal Rules of Evidence, Rule 902(13) and Rule 902(14), with Advisory Committee Notes", , <https://www.govinfo.gov/content/pkg/USCODE-2024-title28/html/USCODE-2024-title28-app-federalru-dup2-rule902.htm>.
[GRENFELL2]
Grenfell Tower Inquiry, "Grenfell Tower Inquiry: Phase 2 Report, Volume 1", HC 19-I, Part 1, Chapter 2, paragraphs 2.122-2.123, , <https://assets.publishing.service.gov.uk/media/66d817aa701781e1b341dbd3/CCS0923434692-004_GTI_Phase_2_Volume_1_BOOKMARKED.pdf>.
[GSN]
SCSC Assurance Case Working Group, "Goal Structuring Notation Community Standard, Version 3", SCSC 141C, , <https://scsc.uk/scsc-141c>.
[HAMILTON]
Court of Appeal (Criminal Division), England and Wales, "Hamilton & Ors v Post Office Ltd", [2021] EWCA Crim 577, paragraphs 136-137, , <https://caselaw.nationalarchives.gov.uk/ewca/crim/2021/577>.
[I-D.abak-agent-control-delivery-evidence]
Abak, A. T., "Evidence Requirements for Agent Control Delivery and Outcome Reconciliation", Work in Progress, Internet-Draft, draft-abak-agent-control-delivery-evidence-01, , <https://www.ietf.org/archive/id/draft-abak-agent-control-delivery-evidence-01.html>.
[I-D.bradleyb-audit-decision-records]
B, B., "Signed Decision Records for Agent Authorization: Disclosures, Entry Emission, and Ordering Evidence", Work in Progress, Internet-Draft, draft-bradleyb-audit-decision-records-00, , <https://datatracker.ietf.org/doc/html/draft-bradleyb-audit-decision-records-00>.
[I-D.bu-agentproto-security-principal-binding]
Bu, S., "Security Principal and Verifier Binding for Agent Communication Protocols", Work in Progress, Internet-Draft, draft-bu-agentproto-security-principal-binding-07, , <https://datatracker.ietf.org/doc/html/draft-bu-agentproto-security-principal-binding-07>.
[I-D.fengfar-led]
Farrell, S. and C. Feng, "Dealing with LLMs in IETF Discussions", Work in Progress, Internet-Draft, draft-fengfar-led-01, , <https://www.ietf.org/archive/id/draft-fengfar-led-01.html>.
[I-D.ietf-wimse-arch]
Salowey, J., Rosomakho, Y., and H. Tschofenig, "Workload Identity in a Multi System Environment (WIMSE) Architecture", Work in Progress, Internet-Draft, draft-ietf-wimse-arch-08, , <https://www.ietf.org/archive/id/draft-ietf-wimse-arch-08.html>.
[I-D.kuehlewind-audit-architecture]
Kuehlewind, M. and H. Birkholz, "An Architecture for Auditing Agent Delegation and Interactions", Work in Progress, Internet-Draft, draft-kuehlewind-audit-architecture-01, , <https://datatracker.ietf.org/doc/html/draft-kuehlewind-audit-architecture-01>.
[I-D.schrock-ep-outcome-binding]
Schrock, I., "Outcome Binding for Authorized Actions and Independently Observed Effects", Work in Progress, Internet-Draft, draft-schrock-ep-outcome-binding-00, , <https://datatracker.ietf.org/doc/html/draft-schrock-ep-outcome-binding-00>.
[I-D.sergeev-wexp-core]
Sergeev, M. and V. Ikher, "The Witnessed Execution Protocol (WEXP): Core Specification", Work in Progress, Internet-Draft, draft-sergeev-wexp-core-01, , <https://datatracker.ietf.org/doc/html/draft-sergeev-wexp-core-01>.
[I-D.wadkins-agentproto-action-determinability]
Wadkins, D. L., "Independent Determinability of Agent Actions", Work in Progress, Internet-Draft, draft-wadkins-agentproto-action-determinability-00, , <https://datatracker.ietf.org/doc/html/draft-wadkins-agentproto-action-determinability-00>.
[I-D.yossif-enrollment-problem]
Yossif, M. K., "Problem Statement: Enrollment and Key-Binding Assumptions in Execution Authority Evidence", Work in Progress, Internet-Draft, draft-yossif-enrollment-problem-00, , <https://datatracker.ietf.org/doc/html/draft-yossif-enrollment-problem-00>.
[IN-TOTO]
Torres-Arias, S., Afzali, H., Kuppusamy, T. K., Curtmola, R., and J. Cappos, "in-toto: Providing farm-to-table guarantees for bits and bytes", 28th USENIX Security Symposium, , <https://www.usenix.org/system/files/sec19-torres-arias.pdf>.
[INTERPARES]
Duranti, L., "The Long-term Preservation of Authentic Electronic Records: Findings of the InterPARES Project", , <https://www.interpares.org/book/index.cfm>.
[ISO14721]
International Organization for Standardization, "Space Data System Practices -- Reference model for an open archival information system (OAIS)", also issued as CCSDS 650.0-M-3, December 2024, ISO 14721:2025, Edition 3, .
[ISO15026-2]
ISO/IEC/IEEE, "Systems and software engineering -- Systems and software assurance -- Part 2: Assurance case", ISO/IEC/IEEE 15026-2:2022, .
[ISO16363]
International Organization for Standardization, "Space data and information transfer systems -- Audit and certification of trustworthy digital repositories", ISO 16363:2025, Edition 2, .
[PROV-DM]
World Wide Web Consortium, "PROV-DM: The PROV Data Model, W3C Recommendation", , <https://www.w3.org/TR/prov-dm/>.
[RFC3161]
Adams, C., Cain, P., Pinkas, D., and R. Zuccherato, "Internet X.509 Public Key Infrastructure Time-Stamp Protocol (TSP)", RFC 3161, DOI 10.17487/RFC3161, , <https://www.rfc-editor.org/rfc/rfc3161>.
[RFC3227]
Brezinski, D. and T. Killalea, "Guidelines for Evidence Collection and Archiving", BCP 55, RFC 3227, DOI 10.17487/RFC3227, , <https://www.rfc-editor.org/rfc/rfc3227>.
[RFC3552]
Rescorla, E. and B. Korver, "Guidelines for Writing RFC Text on Security Considerations", BCP 72, RFC 3552, DOI 10.17487/RFC3552, , <https://www.rfc-editor.org/rfc/rfc3552>.
[RFC4998]
Gondrom, T., Brandner, R., and U. Pordesch, "Evidence Record Syntax (ERS)", RFC 4998, DOI 10.17487/RFC4998, , <https://www.rfc-editor.org/rfc/rfc4998>.
[RFC5848]
Kelsey, J., Callas, J., and A. Clemm, "Signed Syslog Messages", RFC 5848, DOI 10.17487/RFC5848, , <https://www.rfc-editor.org/rfc/rfc5848>.
[RFC6376]
Crocker, D., Ed., Hansen, T., Ed., and M. Kucherawy, Ed., "DomainKeys Identified Mail (DKIM) Signatures", STD 76, RFC 6376, DOI 10.17487/RFC6376, , <https://www.rfc-editor.org/rfc/rfc6376>.
[RFC6973]
Cooper, A., Tschofenig, H., Aboba, B., Peterson, J., Morris, J., Hansen, M., and R. Smith, "Privacy Considerations for Internet Protocols", RFC 6973, DOI 10.17487/RFC6973, , <https://www.rfc-editor.org/rfc/rfc6973>.
[RFC9162]
Laurie, B., Messeri, E., and R. Stradling, "Certificate Transparency Version 2.0", RFC 9162, DOI 10.17487/RFC9162, , <https://www.rfc-editor.org/rfc/rfc9162>.
[RFC9334]
Birkholz, H., Thaler, D., Richardson, M., Smith, N., and W. Pan, "Remote ATtestation procedureS (RATS) Architecture", RFC 9334, DOI 10.17487/RFC9334, , <https://www.rfc-editor.org/rfc/rfc9334>.
[RFC9449]
Fett, D., Campbell, B., Bradley, J., Lodderstedt, T., Jones, M., and D. Waite, "OAuth 2.0 Demonstrating Proof of Possession (DPoP)", RFC 9449, DOI 10.17487/RFC9449, , <https://www.rfc-editor.org/rfc/rfc9449>.
[RFC9591]
Connolly, D., Komlo, C., Goldberg, I., and C. A. Wood, "The Flexible Round-Optimized Schnorr Threshold (FROST) Protocol for Two-Round Schnorr Signatures", RFC 9591, DOI 10.17487/RFC9591, , <https://www.rfc-editor.org/rfc/rfc9591>.
[RFC9711]
Lundblade, L., Mandyam, G., O'Donoghue, J., and C. Wallace, "The Entity Attestation Token (EAT)", RFC 9711, DOI 10.17487/RFC9711, , <https://www.rfc-editor.org/rfc/rfc9711>.
[RFC9942]
Steele, O., Birkholz, H., Delignat-Lavaud, A., and C. Fournet, "CBOR Object Signing and Encryption (COSE) Receipts", RFC 9942, DOI 10.17487/RFC9942, , <https://www.rfc-editor.org/rfc/rfc9942>.
[RFC9943]
Birkholz, H., Delignat-Lavaud, A., Fournet, C., Deshpande, Y., and S. Lasker, "An Architecture for Trustworthy and Transparent Digital Supply Chains", RFC 9943, DOI 10.17487/RFC9943, , <https://www.rfc-editor.org/rfc/rfc9943>.
[SCHNEIER-KELSEY]
Schneier, B. and J. Kelsey, "Secure Audit Logs to Support Computer Forensics", ACM Transactions on Information and System Security, Vol. 2, No. 2, pp. 159-176, DOI 10.1145/317087.317089, , <https://www.schneier.com/wp-content/uploads/2016/02/paper-auditlogs.pdf>.
[VC-DATA-MODEL]
World Wide Web Consortium, "Verifiable Credentials Data Model v2.0", W3C Recommendation, Section 1.1, , <https://www.w3.org/TR/2025/REC-vc-data-model-2.0-20250515/>.
[WCC-CORE]
Sergeev, M. A., "Witnessability Conceptual Core 1.0", DOI 10.5281/zenodo.21865251, , <https://doi.org/10.5281/zenodo.21865251>.
[WITMODEL]
Sergeev, M. A. and V. Ikher, "Toward a Witnessability Model for AI and Software Execution Systems: A Boundary-Based Framework for Classifying Execution Evidence", Version 1.1, bridge revision; also at DOI 10.2139/ssrn.6994720, , <https://doi.org/10.5281/zenodo.21970802>.

Appendix A. The (a,b,c) scenario worked through

In the agentproto mailing-list discussion of July 2026, in the thread on the draft minutes of the AGENTPROTO BoF, Stephen Farrell posed a gating scenario [AP-FARRELL-ABC]. In his words:

He emphasized that "the scenario I posited is one where 'b' creates 'c' and so is in a fine place to cheat" [AP-FARRELL-CHEAT].

The list discussion produced concrete mitigations, collected in a reply by Shawn Sammartano [AP-SAMMARTANO], who credited most of them to existing mechanisms:

The author proposed an enforcement/observation framing in response to that discussion [AP-SERGEEV-GATE]. The control-topology test of Section 8 refines that framing by distinguishing adversary-resistant detection coverage from a particular finding made from evidence already obtained. Farrell posed the scenario. This document does not attribute the generalized test to him. Applied to the scenario:

This is narrower than solving credential custody. It makes the assurance claim and its trust boundary reviewable before a mechanism is chosen, which is what a chartering or design review needs.

Appendix B. Three public findings

Three public findings state the same bounds outside agent systems, in the words of the bodies that made them.

In the English prosecutions arising from the Post Office Horizon accounting system, the Court of Appeal found that the prosecutor "treated what was no more than a shortfall shown by an unreliable accounting system as an incontrovertible loss", and that defendants were convicted "on the basis that the Horizon data must be correct, and cash must therefore be missing, when in fact there could be no confidence as to that foundation" [HAMILTON]. The load-bearing premise was the reliability of that system. The court recorded that the prosecutor was under a duty to investigate subpostmasters' claims that there were problems with it, and it found "failures of investigation and disclosure". On the reading taken here, that is the position rule 9 of Section 6 names. The claim rested on a premise the relying party was required to test and did not adequately test. The judgment does not address what a bounded report of the same evidence would have said, and it is not cited for that part of the rule.

In the inquiry into the Grenfell Tower fire, the panel examined the performance criteria applied in large-scale fire tests of external wall systems. Failure to meet those criteria would show a system unlikely to comply with the applicable requirement. But "the converse was not necessarily true": a system might meet the criteria and yet fail to comply with that requirement [GRENFELL2]. A bounded test result can tell against a claim in one direction without establishing it in the other.

In the third case the test results were genuine outputs of a genuine test. Vehicles were certified against emissions standards on the strength of those results [EPA-VW-NOV]. Software detected the test condition and enabled full emissions controls only then, so the vehicles "meet emissions standards in the laboratory or testing station, but during normal operation" emitted nitrogen oxides at levels up to forty times the standard [EPA-VW]. The result established the behavior of the vehicle under the conditions the test applied, which was all it had ever observed. The certificate carried it to a claim about the vehicle in use (rule 11 of Section 6).

These authorities reached findings within their own legal and regulatory settings. The findings are their own, and their connection to the rules stated here is this document's reading.

Appendix C. Changes from -00

This appendix is to be removed before publication as an RFC.

Acknowledgments

The scenario of Appendix A was posed by Stephen Farrell [AP-FARRELL-ABC]. The mitigation list and the plain statement of the key-custody limit were collected by Shawn Sammartano [AP-SAMMARTANO], who also confirmed the enforcement/observation framing used here [AP-SAMMARTANO-FRAMING]; Charles Eckel proposed that 'a' authorize 'c' to do something specific on 'a's behalf, rather than expose a long-term credential [AP-ECKEL]. Dapeng Liu proposed sender-constrained tokens (DPoP) so that 'b' could not present a token bound to 'c's key [AP-LIU-DPOP]; Section 8 discusses what they do and do not prevent. Yuning Jiang raised the continuity of an identity across restarts, instance changes, and key rotation, which is how the question behind rule 10 of Section 6 entered this work [AP-JIANG-CONTINUITY]. The distinction itself is older and is drawn in [DID-CORE]. Vladimir Ikher is the author's co-author on the earlier work [WITMODEL] from which the bases of Section 4 and the ceiling of Section 5 are restated here. This document also benefited from the broader agentproto and agent2agent mailing-list discussions of 2026 on evidence, delegation, and audit. The author thanks their participants without implying that any of them endorses this document.

Douglas Wadkins reviewed a pre-submission draft; the statement of scope in Section 1.1, the additions to Section 11, and the qualification in Section 12 are responses to his comments.

Document development disclosure

This document was drafted with substantial assistance from generative AI tools, working from the author's prior publications, specifications, and mailing-list correspondence, at the author's direction. It also incorporates material arising from adversarial review conducted with such tools: the wording of several rules, and the separation of analytical justification from documented practical confirmation stated in Section 6, were reached that way. In this version, the treatment of independence (Sections 2, 3, 8.1, and 9), rules 12 and 13 of Section 6, and the reporting terms of Section 11 were revised or added the same way. The author reviewed the resulting text and is responsible for its content, citations, and attributions. The use of large language models (LLMs) in IETF discussions is the subject of [I-D.fengfar-led], which explicitly excludes Internet-Draft and RFC text from its scope. This disclosure is voluntary.

The author also has an implementation interest in this area: the WEXP appraisal layer [I-D.sergeev-wexp-core] is the author's own work. This document is usable independently of it. No organizational independence between the two works is claimed.

Author's Address

Mikhail Sergeev
Independent Researcher