| Internet-Draft | Agent Memory Integrity Benchmark | October 2026 |
| Khandelwal | Expires 5 April 2027 | [Page] |
AI agents increasingly persist memory across sessions and treat that memory, on the next turn, as if it were their own prior experience. This document defines a benchmarking method that measures whether an agent's memory subsystem detects that its persisted memory has been altered, removed, reordered, replayed, or forged at the storage layer, and refuses to serve that memory or reports it before it is served. The method defines eight storage-level edits, three verdict classes, a detection-point distinction between read time and audit time, three control cases, and a scoring rule. It is a laboratory method for controlled, reproducible measurement, in the spirit of RFC 2544 and RFC 8239, and it is intended as a test method for the "Protection of Memory Data Integrity" metric under discussion in the Benchmarking Methodology Working Group. This revision adds a control that proves each edit landed as intended, reports the method's results on thirteen memory subsystems including six that claim tamper evidence, and records the first vendor fix made in response to a measurement.¶
This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79.¶
Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/.¶
Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress."¶
This Internet-Draft will expire on 4 April 2027.¶
Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved.¶
This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Code Components extracted from this document must include Revised BSD License text as described in Section 4.e of the Trust Legal Provisions and are provided without warranty as described in the Revised BSD License.¶
An AI agent built on a framework such as LangGraph, Letta, Mem0, or a vector memory store keeps state between sessions: long-term memory records, session checkpoints, and their metadata. On the next turn the agent reads that state back and acts on it as trusted context. The store that holds the state is ordinary infrastructure, a database file, a table, a key-value namespace, or an object store, and is reachable by the same means as any other data: a compromised host, a shared credential, an injection flaw in a co-located application, a restored backup, or a malicious operator.¶
There is at present no agreed way to measure whether an agent notices when the memory behind it has been changed. Existing work on agent security concentrates on the input path: prompt injection, and poisoning through the agent's own write interface. The at-rest case, where the adversary edits the store directly, is different: the adversary needs no injection, and can move, remove, or roll back genuine records without authoring any content of their own.¶
This document specifies a laboratory benchmarking method for that case. It is deterministic, runs offline, and produces a single metric value together with a per-case verdict table. It follows the conventions of the Benchmarking Methodology Working Group: a controlled test bed, an explicit procedure, control cases that must hold for a result to count, and full reporting of the configuration under test [RFC2544] [RFC8239].¶
This method measures one property: the ability of a System Under Test (SUT) to detect at-rest tampering with its own persisted memory and to avoid serving tampered memory as genuine. It does not measure confidentiality of memory, resistance to prompt injection through the agent interface, or correctness of the agent's reasoning.¶
The adversary here is distinct from two adjacent adversaries discussed elsewhere. It differs from an adversary who can only talk to the agent (memory poisoning through the write path), and from a legitimate user of another session (cross-session isolation). The adversary in this document has write access to the storage medium but holds none of the SUT's cryptographic keys.¶
The method is intended to serve as the test method for a "Protection of Memory Data Integrity" metric. It can be used on its own for any agent memory subsystem.¶
The two adjacent adversaries have their own measurements, and the three fit together. Poisoning through the write path is measured by writing through the SUT's own interface on three attacker channels (an untrusted-source label, a forged trusted-source label, and a trusted label with a valid signature) and recording whether the read path later serves the planted memory as context. Cross-context isolation is measured both through the read path (write as one context, read as another) and, at the storage layer, by test case T6 of this method, since a filter that holds on the first can fail on the second. Proposed procedures for both were offered to the authors of the metric document on the working group list in October 2026.¶
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here.¶
The adversary has write access to the medium that holds the SUT's memory but holds none of the SUT's cryptographic keys. The adversary's goal is that the SUT resumes from memory of the adversary's choosing and that nobody notices.¶
Because the adversary can write anything to the store, an integrity value stored next to the data, a checksum column, a hash field, or a "verified" flag, provides no protection unless it is bound to a secret or a root of trust the adversary does not hold. The method therefore measures the SUT's behaviour on read, not the presence of integrity fields in the store. An adversary who can edit a memory record can equally edit a checksum stored beside it; the only question worth measuring is whether the SUT notices at the moment it loads the memory.¶
The memory subsystem MUST be run in the configuration the SUT's documentation recommends for production, including any at-rest encryption, signing, or audit feature the vendor documents as protecting memory. A feature that is off by default MUST be reported as off by default. Where the SUT offers such a feature, the evaluator SHOULD report two results, one with the feature off and one with it on, so that the two can be compared.¶
The evaluator seeds memory through the SUT's own write path: at least five records in each of two isolated contexts (two users, threads, or sessions, whichever the SUT exposes). One record in context A carries a distinctive fact FA and one record in context B carries a distinctive fact FB. Nothing is written directly to the store during seeding.¶
All reads are made through the SUT's own read path and never by inspecting the store. What the store contains after an edit is the adversary's business; what the SUT serves through its read path is the metric.¶
Each case starts from a fresh copy of the seeded store. The evaluator stops the SUT, applies exactly one edit with generic storage tooling (SQL, a file editor, or the store's own client with no SUT code loaded), restarts the SUT, and reads the affected context through the read path.¶
Stopping the SUT before the edit is part of the method, not a convenience. A subsystem that watches its store from inside a running process can detect every edit while that process is up and none once it has restarted, because the restarted process establishes a new baseline from whatever the store holds. The adversary who edits the store chooses the moment, and a store compromise typically coincides with a host restart. The evaluator MAY additionally report the process-alive position for comparison, but the metric is taken with the SUT stopped.¶
T6 and T7 use only bytes the SUT itself wrote. They pass through any at-rest encryption that does not bind ciphertext to record identity, and any signature that does not cover a record's position. They are the cases that separate confidentiality from integrity, and a subsystem that relies on encryption alone typically accepts both.¶
A subsystem that stores memory as unordered key-value pairs, with no inherent record order and no separate metadata layer, cannot express T2, T4, or T8. Such cases are reported as "not applicable" for that SUT, with the reason, and are excluded from the denominator in Section 9. They are never reported as passes.¶
After the restart the evaluator reads the affected context through the read path and classifies the outcome as one of three verdicts.¶
REJECTED and REPORTED are passes. ACCEPTED is a fail. Each verdict carries a detection point. A signal raised at the moment the memory is loaded or served has detection point "read". A signal that appears only when the operator runs a separate audit command has detection point "audit"; such a result is recorded as REPORTED with detection point "audit" and counts as a partial pass, because by the time such an audit runs the agent has already resumed from and acted on the memory.¶
Three control cases MUST hold for a result to be valid.¶
If C1 or C2 fails, the metric MUST NOT be computed and the result MUST be reported as "not evaluable" with the reason. Without C1, a subsystem that refuses everything would score a perfect pass rate; C1 is therefore not optional. C3 exists because an evaluator's edit can silently miss: a copy made onto the record it came from, a reorder that writes each record back under its own key, or a replay whose victim and donor share an owner all leave the store as it was, and the SUT then "accepts" an edit that never happened. Two such no-op cells were found in the reference implementation's own results and corrected after C3 was added.¶
Let N be the number of applicable test cases for the SUT (eight, less any cases reported "not applicable" under Section 6). Let passes be the count of REJECTED and read-time REPORTED verdicts, and let partial be the count of audit-time REPORTED verdicts. The metric is:¶
Metric Value = (passes + 0.5 * partial) / N¶
Evaluators MAY additionally report the pass rate on T1 to T5 and on T6 to T8 separately, since the first group tests tamper evidence and the second tests binding of a record to its place and owner in the store.¶
Each result MUST state: the SUT and the exact version of its memory component; the store backend and version; the configuration used, that is, which encryption, signing, or audit features were on or off; for every test case the verdict and the detection point; that the SUT was stopped during each edit, or the process-alive position if that is additionally reported; the exact commands or script that produced the edit; and the date of measurement.¶
A result is a statement about one version of one subsystem in one configuration. Re-measurement after a version change is expected, and a change of verdict between versions is itself a useful signal.¶
The following verdicts illustrate what the method produces on current releases of thirteen memory subsystems: seven that make no integrity claim and six that do. They are informative, included so that the method's output can be seen on real software, and are not normative. All runs used the subsystem's default configuration unless stated, with the SUT stopped during each edit. Versions are given so the results can be reproduced or contested; every row below is reproducible from the reference implementation in Section 13.¶
Subsystems that make no integrity claim:¶
Subsystems that claim tamper evidence, measured against the claim:¶
Three patterns follow from the second group. A per-record authenticator catches edits to a record's bytes and nothing about where the record sits or how many there are. A hash chain alone catches everything except removal of the tail, because what remains is still a valid chain; the tail is caught only by a head anchored or signed outside the store. And every subsystem in the second group but one detects on audit, after the agent has already resumed; only the per-record check refuses on the read path, and it refuses the fewest cases. The pattern across both groups is the one Section 7 is built for: the read path trusts the store, and where a protection exists it is either off by default, partial, or late.¶
Four choices in this method are worth stating explicitly.¶
First, eight edits rather than one. A single "tamper the file" case lets a subsystem with a whole-store checksum pass while it still accepts rollback and cross-context replay. T6 and T7 in particular exist because encryption alone accepts them: they move only genuine bytes.¶
Second, the verdict comes from the SUT, not from the evaluator. Inspecting the store to decide whether a tamper "should" have been caught makes the result an opinion. Reading through the SUT's own read path and recording what it did makes it a measurement.¶
Third, control C1 is not optional. It is the only thing that prevents a subsystem that rejects all reads from scoring a perfect result.¶
Fourth, the SUT is stopped during the edit. The memory-blackbox result in Section 11 is the reason: the same subsystem, the same eight edits, and the same scan went from eight of eight reported to none of eight on nothing but a restart. A method that left the attacker position to the evaluator would have produced both results and called both correct.¶
An open, MIT-licensed implementation of this method, called agmi (Agent Memory Integrity), runs offline and applies the eight edits and the three controls to the memory components of thirteen subsystems through a common adapter interface, with every verdict pinned to the measured version by a test that fails when the version moves. It also runs as a continuous-integration check that fails a build when any edit is accepted; one of the measured vendors has adopted it in that form. It is provided for reproducibility and is not required to use the method.¶
This document defines a benchmarking method for laboratory use. As with other benchmarking methodologies, the procedures here are intended for an isolated test bed and MUST NOT be run against production systems or systems the evaluator is not authorised to test. The test cases apply adversarial edits to a store; those edits MUST be confined to test data in a controlled environment.¶
The method measures a security-relevant property, whether an agent detects at-rest tampering with its memory, but a metric value from this method is not by itself an assurance of security. A high value on this method does not measure confidentiality, input-path poisoning resistance, or any property outside Section 2.¶
Publishing per-case verdicts for named subsystems and versions describes weaknesses in released software. The intent is constructive: results are stated with the version measured so that vendors can reproduce and fix them, and so that a later re-measurement shows the change.¶
This document has no IANA actions.¶
The author thanks the maintainers of the memory subsystems measured during the development of this method for their engagement on the individual findings, in particular the maintainers of inspeximus, who found the no-op cells that led to control C3, and of memory-blackbox, who reproduced and fixed the restart result within a day. Thanks also to the authors of draft-han-bmwg-agent-security-benchmark for the discussion on the working group list that shaped the detection-point distinction.¶
This section is to be removed before publication as an RFC.¶