Internet-Draft Substrate Provenance Grammar August 2026
Morrison Expires 10 February 2027 [Page]
Workgroup:
Network Working Group
Internet-Draft:
draft-morrison-substrate-provenance-grammar-01
Published:
Intended Status:
Informational
Expires:
Author:
B. Morrison
Alter Meridian Pty Ltd

Substrate-Provenance Annotation Grammar for Large-Language-Model Output

Abstract

This memo specifies a wire-level annotation grammar by which a large-language-model output may carry, at emission and at the granularity of an individual assertion, a provenance label drawn from a closed enumerated vocabulary of substrate-class identifiers. The memo defines the closed vocabulary, the per-assertion attachment form, the admissibility discipline a relying party MAY apply to the labels, and two terminal output states, UNVERIFIED-INFERENCE and DECAYED-TO-UNCERTAINTY, equal-rank with assertion and denial. The memo does not specify what an inference system MUST do; it specifies the wire grammar by which a relying party may inspect what the inference system DID with respect to the substrates it consulted. The memo is Informational.

Status of This Memo

This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79.

Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/.

Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress."

This Internet-Draft will expire on 10 February 2027.

Table of Contents

1. Introduction

When a large-language-model output is consumed by another agent, by a downstream automation, or by a relying party with action authority, the consuming entity has no wire-level mechanism to distinguish three structurally different cases: (i) the model emitted the assertion from training-corpus-resident pattern without consulting any external substrate; (ii) the model emitted the assertion after consulting an external substrate whose state corroborated the assertion within an admissibility window; (iii) the model emitted the assertion after consulting an external substrate whose state did not corroborate the assertion, and the model proceeded anyway.

Existing approaches to this problem operate at the prose layer: post-hoc citation insertion by a separate retrieval orchestrator, natural-language hedge phrasing ("I believe", "it appears", "according to"), per-paragraph confidence scores rendered as adjectives, or refusal to answer. All four are surface-form-parseable rather than structurally distinguishable; all four can be defeated by a model trained to substitute hedge phrasing for substrate consultation.

This memo specifies a wire-level grammar by which the inference system declares, at the granularity of an individual assertion within its output, which substrate-class, if any, corroborated the assertion at emission. The grammar is closed-vocabulary, finite, and version-anchored; the consuming entity parses the annotation without interpreting prose. Two terminal annotations, UNVERIFIED-INFERENCE and DECAYED-TO-UNCERTAINTY, are equal-rank with assertion and denial. They are a structurally distinguishable output state, not a confidence score and not a hedge phrase.

2. Conventions and Definitions

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here.

The following terms are defined for the purposes of this document:

3. The Closed Substrate-Class Vocabulary

The vocabulary defined by this memo is closed, finite, and version-anchored. An implementation parsing a provenance annotation MUST recognise the annotation if and only if the substrate-class identifier appears in the version of the vocabulary the implementation has loaded. Unknown substrate-class identifiers MUST NOT be silently interpreted; an implementation that encounters one MUST treat the annotation as if it were UNVERIFIED-INFERENCE (Section 4).

The version-anchor scheme used by this memo is a dotted major.minor pair appearing as a leading element of the substrate-class identifier. The vocabulary defined in this revision of the memo carries the version anchor 1.0. Future revisions of this memo MAY add substrate-classes; addition is a minor-version bump. Future revisions MUST NOT remove substrate-classes without a major-version bump.

The version-1.0 closed substrate-class vocabulary:

The eight identifiers above constitute the entirety of the version-1.0 closed vocabulary.

4. Terminal Annotations

Two annotation values are terminal: they signal a structurally distinguishable output state of the inference system, equal-rank with assertion and denial, rather than corroboration by any substrate.

Both terminal annotations are first-class output tokens. A relying party parsing the wire output observes them in the same structural slot in which substrate-class identifiers appear; the parser applies the terminal-annotation disposition without inspecting prose. The two terminal annotations are structurally distinct from refusal to answer, from explicit denial, and from the absence of annotation.

5. Attachment Form

A provenance annotation is attached to an individual assertion in the inference system's output. This memo specifies the abstract attachment relationship; the concrete wire encoding is the implementation's choice and is not normative herein.

The annotation form is the tuple:

(assertion-span, substrate-class-identifier, observation-id?, ts?)

where:

Two concrete encodings are illustrative and not normative:

JSON-structured-output encoding per [RFC8259]:

{
  "assertion": "the file CHANGELOG.md contains an entry dated 2026-05-25",
  "provenance": {
    "substrate_class": "substrate.code.read",
    "observation_id": "sha256:e3b0c4...",
    "ts": "2026-05-28T11:40:00Z"
  }
}

In-line bracketed annotation, for free-text outputs:

The file CHANGELOG.md contains an entry dated 2026-05-25.
[substrate.code.read; observation-id=sha256:e3b0c4...;
 ts=2026-05-28T11:40:00Z]

6. Admissibility Discipline

A relying party MAY apply a cardinality-thresholded admissibility discipline to inference system output annotated under this grammar. The discipline is parameterised by:

Reference values:

The relying party applies the discipline by counting, for each assertion in the inference system's output, the distinct substrate-class identifiers appearing in the assertion's provenance annotations whose ts field is within W. Assertions not meeting the cardinality floor are not admitted; assertions annotated with unverified-inference or decayed-to-uncertainty are not admitted by virtue of the terminal annotation itself.

A provenance annotation MUST NOT be counted toward the k = 3 floor unless the relying party has independently re-verified it, per the Annotation Fabrication discussion in Security Considerations. Independent re-verification requires an observation-id; an annotation lacking an observation-id MUST NOT be counted toward the k = 3 floor. This requirement does not apply to the k = 2 floor, where an annotation remains admissible under the counting procedure above without independent re-verification.

Independent re-verification is only meaningful where re-querying the named substrate can CONFIRM the earlier observation rather than merely produce a second, unrelated one. A relying party MUST apply that test per annotation before counting it toward the k = 3 floor, and MUST NOT treat a later reading as confirming an earlier one where the substrate's value may legitimately have changed in between, or where the observed condition no longer exists to be queried.

substrate.do.sse-count, whose subscriber count legitimately differs on every observation, and substrate.unix.peercred, which is unavailable once the connection it described has closed, fail that test in every case. substrate.mcp.brief names state owned by a third party and may fail it depending on whether that state is stable across the interval. The test is stated as a property rather than as a list because the vocabulary is closed but the re-queryability of a given observation is not a fixed attribute of its class.

An annotation that fails the test remains admissible toward k = 2. This is a deliberate consequence rather than a defect in the vocabulary, and it fails closed, but an implementer counting toward the k = 3 floor needs to know that some annotations can never contribute to it.

This memo does not specify what the relying party MUST do with an inadmissible assertion. Common dispositions include: discarding the assertion silently, surfacing it to a human reviewer, requesting re-emission from the inference system, or substituting an explicit refusal in the relying party's own output to the next consuming entity. Each disposition is the relying party's own policy choice and is not constrained by this memo.

7. Why the Vocabulary Is Closed

A reader may ask why the substrate-class vocabulary is closed rather than open-extensible.

An open-extensible vocabulary would permit any inference system to introduce a new substrate-class identifier and emit corroboration annotations under it. A relying party encountering an unrecognised identifier would face a choice: trust the new identifier on its face, refuse the assertion, or treat the identifier as equivalent to a fallback known identifier. Each choice is structurally inferior to the closed-vocabulary posture of this memo:

The closed-vocabulary posture treats unrecognised identifiers as unverified-inference (Section 4). This preserves the admissibility discipline under vocabulary drift while leaving the relying party free to upgrade its parser to a newer vocabulary version.

8. Relation to Prior Art

The contribution of this memo is the joint articulation of a closed substrate-class vocabulary, per-assertion attachment, and two terminal annotations as a single wire-level grammar. Each adjacent prior-art family is structurally distinct from this grammar in at least one of the three components:

Retrieval-Augmented Generation [RAG] performs retrieval against an external corpus and conditions generation on the retrieved context. RAG is an inference-system architecture; this memo specifies an output-side annotation grammar. A RAG-architected inference system MAY emit annotations under this grammar; a non-RAG inference system MAY emit annotations under this grammar. The grammar is orthogonal to the architecture.

Constitutional AI [CAI] and self-critique architectures apply a second model pass to evaluate the first pass's output against a specification. The output of such a system is not annotated at the granularity of an individual assertion against an external substrate-class; it is annotated, if at all, with a critique-pass verdict against a specification authored by the model vendor. Annotation against a vendor-authored specification is structurally distinct from annotation against a substrate observable to a relying party.

Multi-agent debate [DEBATE] and related multi-pass deliberation architectures produce a single consensus output from multiple agent passes. The output's provenance is the agents' agreement process, not a substrate observable to a relying party.

Hidden-state probing [PROBE] inspects the inference system's internal activations to estimate the system's own confidence in its output. Confidence is a property of the inference system; substrate-class corroboration is a property of the relying party's own observation surface. The two are categorically different information sources.

Cryptographically-anchored append-only logs (Certificate Transparency [RFC6962], trusted timestamping per [RFC3161]) are candidate corroborating substrates under this grammar, and each is a substrate-class a future revision of the vocabulary MAY add. A chained log considered in isolation is not a per-assertion annotation grammar.

9. IANA Considerations

This memo requires no IANA actions in its present revision. A future revision may request establishment of an IANA registry for substrate-class identifiers, governed by the closed-vocabulary discipline of Sections 3 and 7.

10. Security Considerations

The grammar specified by this memo surfaces three classes of attack absent from prose-only or hedge-phrasing approaches. The mitigations described below are operational rather than wire-level; this memo specifies the grammar only, and an implementation's operational posture is its own.

10.1. Annotation Fabrication

An inference system may emit a provenance annotation citing a substrate-class corroboration that did not in fact occur. The grammar specified by this memo provides no cryptographic binding between the annotation and any observed substrate state. A relying party MUST NOT treat the annotation as evidence of corroboration; the annotation is a declaration of the inference system's claim about its own behaviour, which the relying party MAY independently verify by re-querying the named substrate with the optional observation-id. Cryptographic binding of annotations to observed substrate state is out of scope for this memo and is the subject of separate work.

Admissibility Discipline requires this independent re-verification before a provenance annotation may be counted toward the k = 3 floor; the MAY above is the relying party's general option to verify any annotation, and is superseded, for that floor, by the MUST stated there.

10.2. Vocabulary Drift

An inference system implemented against a newer version of the vocabulary may emit substrate-class identifiers not present in a relying party's older vocabulary. Per Section 3, the relying party treats unrecognised identifiers as unverified-inference. This is a fail-closed posture and is correct. A relying party operating at a substantially older vocabulary version SHOULD upgrade its parser to the current published version of this memo.

10.3. Substrate Capture

An adversary controlling a substrate identified in the vocabulary may engineer the substrate's state to corroborate assertions of the adversary's choice. The admissibility discipline of Section 6, with k = 2 or k = 3, mitigates this by requiring corroboration from substrate-class-distinct sources before admission. An adversary controlling all k substrate-classes can defeat the discipline; selection of independent substrate-classes is the relying party's operational responsibility and is not specified by this memo.

11. Privacy Considerations

Provenance annotations expose to consumers of the inference system's output the categories of substrate the inference system consulted. In typical deployments, the substrate-class identifier is a category not a record; the per-record observation-id, if emitted, may carry the privacy properties of the underlying substrate (a filesystem path, a content hash, a peer credential identifier).

A relying party emitting annotations under this grammar to a further downstream consumer SHOULD apply the same identity-binding considerations articulated in [SUBOBS] for substrate observables: pseudonymous tier observations need not carry strong identifiers; identity-bound tier observations carry the identity-binding strength of the underlying substrate.

The grammar specified by this memo does not require, and does not recommend, attachment of identifiers tying the inference system's output to a particular human end-user. Such attachments, if made, are outside the scope of this memo.

12. References

12.1. Normative References

[RFC2119]
Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels", BCP 14, RFC 2119, DOI 10.17487/RFC2119, , <https://www.rfc-editor.org/info/rfc2119>.
[RFC8174]
Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174, , <https://www.rfc-editor.org/info/rfc8174>.
[RFC8259]
Bray, T., Ed., "The JavaScript Object Notation (JSON) Data Interchange Format", STD 90, RFC 8259, DOI 10.17487/RFC8259, , <https://www.rfc-editor.org/info/rfc8259>.
[SUBOBS]
Morrison, B., "Substrate-Observation as an Alternative to Envelope Coordination for Concurrent Sessions", , <https://datatracker.ietf.org/doc/draft-morrison-substrate-observation/>.
[MCPDNS]
Morrison, B., "Discovery of Model Context Protocol Servers via DNS TXT Records", , <https://datatracker.ietf.org/doc/draft-morrison-mcp-dns-discovery/>.
[IDCOMMITS]
Morrison, B., "Identity-Attributed Git Commits via Tier-Structured Trailers", , <https://datatracker.ietf.org/doc/draft-morrison-identity-attributed-commits/>.
[IDPRONOUNS]
Morrison, B., "Identity Pronouns: A Reference-Axis Extension to ~handle Identity Systems", , <https://datatracker.ietf.org/doc/draft-morrison-identity-pronouns/>.

12.2. Informative References

[RFC6962]
Laurie, B., Langley, A., and E. Kasper, "Certificate Transparency", RFC 6962, DOI 10.17487/RFC6962, , <https://www.rfc-editor.org/info/rfc6962>.
[RFC3161]
Adams, C., Cain, P., Pinkas, D., and R. Zuccherato, "Internet X.509 Public Key Infrastructure Time-Stamp Protocol (TSP)", RFC 3161, DOI 10.17487/RFC3161, , <https://www.rfc-editor.org/info/rfc3161>.
[POSIX]
"IEEE Std 1003.1-2017, Standard for Information Technology - Portable Operating System Interface (POSIX) Base Specifications", , <https://pubs.opengroup.org/onlinepubs/9699919799/>.
[RAG]
Lewis, P., Perez, E., and A. Piktus, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", , <https://arxiv.org/abs/2005.11401>.
[CAI]
Bai, Y., "Constitutional AI: Harmlessness from AI Feedback", , <https://arxiv.org/abs/2212.08073>.
[DEBATE]
Irving, G., "AI Safety via Debate", , <https://arxiv.org/abs/1805.00899>.
[PROBE]
Azaria, A. and T. Mitchell, "The Internal State of an LLM Knows When It's Lying", , <https://arxiv.org/abs/2304.13734>.

Acknowledgements

This memo articulates a grammar layered above the substrate-observation primitive of [SUBOBS], the identity substrate of [IDPRONOUNS] and [IDCOMMITS], and the discovery mechanism of [MCPDNS]. Its development is the joint product of deployed agentic-system experience and structural analysis of adjacent prior art.

IPR Posture

The applicant of any patent rights that may be construed to read on the grammar specified by this memo will file an IPR disclosure under the IETF's standard procedures. The applicant's intended licensing posture for any such rights is royalty-free with defensive-termination, consistent with the applicant's published IPR disclosures on companion memos in the morrison-* family.

Contributors

Author's Address

Blake Morrison
Alter Meridian Pty Ltd