Agent Memory Benchmark

Authored by SAIHM, which builds one of the systems it measures. That is a conflict of interest. The only honest response to it is to make the harness checkable rather than to claim neutrality — so the harness is Apache-2.0, it runs from a clean checkout, and it reports its inputs alongside its outputs. Where we have no result we can stand behind, the axis says so: one sub-measure on this page reads not measured because our own scorer will not accept the evidence we had for it. That is an absence, not an adverse result, and we are not going to dress it up as one.

The erasure axis →

Read this part first

A benchmark published by a vendor that wins it is marketing. We cannot escape that by asserting neutrality, so we have tried to escape it by construction instead:

  • We do not run competitors. Several vendors’ terms restrict published benchmarking, and a competitor number we measured badly is worse than no number. Competitor entries are citations — the harness refuses one that lacks a URL, a retrieval date and the vendor’s own words.
  • A measurement that did not happen is not a zero. “Measured”, “unsupported” and “not measured” are three different states, and the harness will not quietly score the third as free or as passing.
  • The records are in the repository, not just the numbers. Every figure on this page comes from a record you can download and re-score with the same scorer we ran. Where we have no record we can stand behind, the axis says not measured rather than carrying a number.

The axes

AxisQuestionSAIHM’s result today
A2 — erasure correctnessAfter a system is told to forget something, can it still be recovered — and was the key destroyed, or merely withheld?All 3 of 3 erased canaries were unrecoverable — 9 probes, 0 recoveries on the surfaces a reader can repeat, and 3 probes, 0 recoveries on the operator-attested one. The two halves are never added together. Controls C1 and C2 passed and are computed from the probes in the record; C3 is a declaration the record makes, and no scorer can confirm it. Measured 2026-09-18.
A8 — cross-agent portabilityCan memory written by one agent be read by a second agent with its own identity — and taken back?Handoff passes byte-exact in three steps; export is unsupported; revocation is not measured — see below. Measured 2026-09-18.

The attested half is the half you cannot repeat, so the record names who made those observations and when, and the scorer prints both alongside the result. That is not proof — it is what makes the claim falsifiable by someone with access rather than merely unverifiable by everyone.

Those two axes are the whole benchmark. Earlier drafts carried retrieval, latency and ordering figures measured elsewhere and reproducible by nobody, including us. They have been removed: a page arguing that you should check its numbers cannot also carry numbers that cannot be checked.

A2 — erasure correctness, and why it is the easiest axis to fake

“Forget” means two very different things.

Under access-control erasure, an index entry is flagged and the read path starts answering “not found”. The ciphertext is still there, and anyone holding the key can still decrypt it. Under crypto-shred, the key material for that one cell is destroyed, so the surviving ciphertext is unreadable to everyone — the operator included.

Here is the problem. Every probe a server answers returns “not found” under both. A benchmark built from ordinary read-back probes scores a perfect zero against a system that deleted nothing but an index row. That zero is not a weak result; it is not a result at all.

So a zero has to earn its way out

AMB refuses to report zero leaks unless three controls pass in the same run:

  • C1 — every canary was reachable on every surface before the erasure. Without it, “not found afterwards” is satisfied by something that was never findable.
  • C2 — every deliberately-kept survivor is still reachable after. Without it, zero leaks is satisfied by breaking the store or the reader.
  • C3 — some surface’s negative actually distinguishes a destroyed key from a withheld one.

C3 is the one that bites. It cannot be satisfied by any probe the server answers, which means it cannot be satisfied by any probe a reader can run against a hosted system. Telling the two cases apart requires holding the ciphertext while lacking the key — a position a hosted design never puts a reader in.

The ceiling, stated plainly

No party outside a system can distinguish a key that was destroyed from one that is merely withheld. That is not a limitation of this benchmark; it is a property of the situation, and it applies to every erasure claim made by every vendor, whether or not they say so.

What follows from it is uncomfortable and we would rather write it down than have it noticed: the probe that separates the two is operator-attested by construction. So AMB reports A2 in two halves that are never added together — what a reader can reproduce, and what they must take on trust, labelled as such. You are invited to take the first half and discard the second.

It also means the claim is falsifiable rather than verifiable. It does not ask to be believed. It stands until someone recovers a forgotten cell.

The result

Measured against the hosted endpoint on 2026-09-18, run a2-2026-09-18-21191a33. Five canaries were minted for the measurement, three erased, two deliberately kept alive.

HalfRecovery attempts that succeededSurfaces
Reproducible by a reader3 of 3 canaries unrecoverable
9 probes, 0 recoveries
ordinary read, single-cell read, read from a cold separate process
Operator-attested3 of 3 canaries unrecoverable
3 probes, 0 recoveries
attempting to open the ciphertext still held in storage, using this identity’s own key material

The two halves are reported separately and are not added together. Take the first and discard the second if you prefer; the first is the part you can repeat.

What the nine are, stated plainly, because a bigger number is easy to manufacture here too. Three erased canaries asked for on three reader surfaces: the ordinary read path, the single-cell read path, and a read from a cold separate process, started fresh in each phase. That is nine attempts and three independent facts. The third surface exists because the recall cache is a file, so a second client inside the same process would read the same cache and establish nothing. All three are answered by the same endpoint — which is exactly why the attested half is reported beside them rather than folded into them.

The controls, which is what makes the zero mean anything:

  • C1 passed — every one of the five canaries was reachable on every surface before the erasure, the attested surface included. A canary that was never findable cannot later prove it was erased.
  • C2 passed — both survivors were still reachable on every surface afterwards, 8 of 8 probes. The readers were not broken; they were simply no longer able to find what had been destroyed.
  • C3 passed — the record declares a surface whose negative separates a destroyed key from a withheld one, so the zero is not resting on endpoint-answered probes alone. C3 is weaker than C1 and C2 and we would rather you heard it here. C1 and C2 are computed from the twenty and eight observations in the record. C3 checks that such a surface is declared — it cannot verify the declaration, and no scorer outside the measured system could. Read it as “this result does not rest on read-path probes alone”, never as “the separation was independently confirmed”.

The record carries no anomalies for this run. That is the absence of a report rather than a separately verified all-clear: the harness prints anomalies only where the deployment reports them, and on this run it reported none. The attestation was taken by a read-only inspection at two recorded instants, 06:26:11Z and 06:26:17Z — one before the erasure and one after, on this run. The harness checks that both are real instants and that they are in order; that they straddle the erasure is something this record states and the scorer does not establish.

Two honest qualifications. The tenant was not empty when the run began — five cells from an earlier aborted attempt were cleared first, are recorded as such, and form no part of the measurement. And this number covers the hosted endpoint only. The self-hosted runtime is a different regime, it is access-control erasure rather than key destruction, and the disclosure above says so rather than letting this figure stand in for both.

A8 — can a second agent read it, and can you take it back?

Measured 2026-09-18 between two distinct identities, against the same deployment the erasure axis measures.

Sub-measureResult
P1 — the recipient reads back what was writtenPASS, byte-exact
P2 — steps in the handoff3 — write, grant, notify. Lower is better, and it is meant to be compared across systems; a system that cannot do the handoff at all scores unsupported rather than a large number.
P3 — revocation closes the read pathNOT MEASURED — see below
P4 — export and reimportnot supported

What a byte-exact pass does not establish. The record carries the handed-over content and the steps, and no identifier tying the read-back to the specific memory the first agent handed over. So P1 shows that the second agent produced the canonical payload — not that the payload came from that memory. A system that served the payload from anywhere would pass this sub-measure and no reader could tell from the file. Closing that needs an identifier minted at handoff and echoed at read, which this release does not carry, so the scorer states the gap on every verdict rather than leaving it to be noticed.

The payload is compared byte-exact and deliberately contains things that survive naive serialisation badly — quotes, a backslash, a newline, a non-BMP character and a trailing space. A system that normalises any of them has not preserved your memory.

There are four independent sub-measures and no composite score, because a composite would let a strong P2 hide a failed P3 — and P3 is the one most often claimed and least often shown.

Why P3 reads “not measured”. Our run recorded that the recipient’s read did not succeed after revocation. It did not record how that was established — and the adapter of the day treated any error as proof that revocation had closed the read path. A server fault or an exhausted retry would have been published as a pass: the strongest claim on the axis, resting on the weakest possible observation. That is the same defect we had already found and closed on the erasure axis, sitting open on this one.

The scorer now requires a revocation pass to name its evidence, as one of three — endpoint-reported-absent, the endpoint answered and had nothing to serve; endpoint-denied, the endpoint answered and refused this reader; or recipient-cannot-decrypt, the recipient obtained the stored bytes and could no longer read them, with no server decision in the path. That third one is not ours to demonstrate: our first two values described our own architecture rather than the question, and a system that revokes cryptographically had no way to report a revocation it had observed perfectly. Our own published transcript cannot supply that evidence, so the axis reports P3 as not measured — and you can watch it happen with the record in the repository. The other three sub-measures still score. We would rather ship the axis in that state, with the instrument that caught it, than ship a pass we cannot stand behind. A failure would still be reported without any such warrant: access surviving revocation is self-evidencing, and requiring a label there would suppress the one result that counts against us.

P4 is reported as unsupported, not as a failure. There is no export primitive, so the system cannot be asked the question. “Cannot do this” and “we did not test this” are different claims and this benchmark keeps them apart.

Get the harness and check us

The harness, both measurement records and the full corrections ledger are at github.com/SAIHM-Admin/agent-memory-benchmark. The ledger (METHODOLOGY.md) is the part worth reading if you are deciding whether to believe any of this: it lists every defect found while building the harness, including the ones that flattered us.

Everything above is produced by code you can read and run. It is Apache-2.0 and has no dependency on anything of ours to score a record:

git clone https://github.com/SAIHM-Admin/agent-memory-benchmark
cd agent-memory-benchmark
npm install                       # the only step that touches the network
npm test                          # the guard suite, offline, no credentials

# re-score the records this page reports, from the repository:
npx tsx src/run.ts a2 --record results/a2-hosted-2026-09-18.public.json
npx tsx src/run.ts a8 --transcript results/a8-hosted-2026-09-18.json

# a fixture built to fail, so you can see a FAIL rather than only passes:
npx tsx src/run.ts a8 --transcript inputs/example-transcript.json

The scoring path makes no network calls and cannot erase anything. The program that performs a real erasure is a separate file and will not start without an explicit flag, nor proceed against a tenant holding any cell it was not told about in advance — because the first version of a tool like that which goes wrong goes wrong against someone’s real memory.

That guard has a deliberate way through it, and we would rather you read it here than discover it. One flag takes a list of cell identifiers and means destroy these first; it is why our own published record discloses five cells cleared before the run began. Where erasure works by destroying keys, anything cleared that way is gone permanently for everyone, us included. Point it only at a tenant you created for the measurement.

Submit a system

Every AMB number, ours included, is an unverified self-report. The scorer checks that a record is internally consistent and that its controls hold. It cannot check that the run happened, or that the system described is the system measured — nothing binds a record to a deployment, and adding a signature would only move the same trust one step. What it establishes is that a result does not contradict itself and does not rest on probes that prove nothing. That is worth having, and it is not verification. We would rather say so than let the word “benchmark” do work it has not earned.

Third-party results are the point of publishing the harness, not a courtesy. The scorer you run is the one we run against ourselves, and it will refuse your record on the same grounds — a canary that was never reachable before the erasure, a survivor that vanished after it, no surface whose negative separates a destroyed key from a withheld one, or a surface a reader cannot repeat that does not say who attested it and when. It will also refuse a record it cannot print back to you honestly: a field carrying an invisible or direction-changing character, a name or value longer than it will render, or more canaries and surfaces than a reader could check by reading them. The attestation rule is a check on what your record declares, not on what we can confirm — no scorer outside a system can confirm who attested a surface or when, ours included, and the output says so next to the number. It does not currently refuse our erasure record. It does withhold a pass from one sub-measure of our own portability transcript, and that transcript ships in the repository for you to score yourself.

The record format, the accepted revocation-evidence values and what each refusal means are documented in the repository’s README. Send a result, or tell us a refusal was unhelpful, by opening an issue on the repository. A refusal that does not tell you how to satisfy it is a defect in the harness, not in your submission. A record is also refused before it is scored if a field is longer than the harness will print back, if a document is nested deeper or carries more fields than a reader could check, or if any field carries a character that is invisible or direction-changing — each with the field named and the limit stated.