AI agents
Identity, Privacy, Retention, and Deletion
Govern agent state and memory across identities, purposes, access boundaries, retention, correction, and deletion.
By the end you can
- Define agent memory governance as an operational contract rather than a capability label
- Contrast Device-level memory with Account-scoped memory in “A shared household assistant attributed one person’s preference to another”
- Trace “A deleted fact can continue influencing behavior through derived artifacts” through a concrete execution path
- Produce “Create a memory governance matrix” with evidence for “Cross-user and cross-tenant retrieval is blocked by design”
Example
The United States sued Amazon over Alexa, and the complaint is about identity and deletion
On 31 May 2023 the United States sued Amazon over Alexa in the Western District of Washington. The complaint describes the shared-household failure as a design rather than an accident. More than 800,000 children under 13 had their own Alexa profiles. Each was linked to a parent's profile on shared household devices. The identity boundary was drawn around the device and the account that owned it. Not around the person speaking to it.
The second half of the complaint is about deletion. Until mid-2019, the United States alleges, Amazon honoured a parent's deletion request against the audio and kept the writing: “Specifically, when users requested deletion of their voice recordings, or when parents used the Alexa Privacy Hub or the Alexa App to delete their children’s voice recordings, Alexa deleted the voice recordings but retained written transcripts of those recordings in residual data stores.” Those transcripts went on being used for product improvement.
The case was terminated on 20 July 2023 and settled for a $25,000,000 civil penalty. Both halves of the lesson sit in one filing. An identity boundary set by hardware, and a deletion that reached the record the user pointed at and stopped there.
- Decision at stake: Govern agent state and memory across identities, purposes, access boundaries, retention, correction, and deletion. The complaint says those decisions were effectively taken once, at the device.
- Hidden assumption: A shared device is a reliable proxy for one person's identity. That assumption is what links the profiles of more than 800,000 children under 13 to a parent's profile on the same household hardware.
- Primary control question: A deleted fact can continue influencing behavior through derived artifacts. Here it was written transcripts, kept in residual data stores after the voice recordings themselves were deleted, and still used for product improvement.
- Evidence to collect: Cross-user and cross-tenant retrieval is blocked by design, and a deletion request demonstrably reaches every store that holds a copy — residual data stores included — not only the record the user pointed at.
A federal order already spells out what "partitioned" and "deleted" have to mean
Agent identity can refer to end users, accounts, devices, organizations, service principals, or delegated actors. Memory has to be partitioned by the identity and the purpose that justify collecting it and reading it back. Partitioned means the agent cannot read across the wall. What it learned for one account is unavailable while it works for another.
Deletion should cover primary records and derived indexes, summaries, caches, and embeddings according to documented timelines. That is not a design aspiration. It is the operative part of an order in force. The Federal Trade Commission's order against Everalbum, issued 6 May 2021, put three clocks on one request: photos and videos of deactivated users within 30 days, all non-consented Face Embeddings within 90 days, and any “Affected Work Product” within 90 days. The order defines that last term itself: ““Affected Work Product” means any models or algorithms developed in whole or in part using Biometric Information Respondent collected from Users of the “Ever” mobile application.” The Commission's public summary lined the three items up: “Part III of the proposed order requires Respondent to delete (A) photos and videos of Ever app Users who requested deactivation of their accounts, (B) face recognition data that it created without obtaining Users’ affirmative express consent, and (C) models and algorithms it developed in whole or in part using images from Users’ photos.”
Each deletion had to be confirmed to the Commission in a statement sworn under penalty of perjury. The order also carries an express carve-out. Retention is permitted where a law, regulation, court order or the safeguarding of evidence in pending litigation requires it. Embeddings named, models named, days counted, exception written down. That is what a retention policy looks like when someone else drafts it for you.
Storage without a named identity and purpose behind it cannot be walled off from the next account or removed on any documented timeline.
Case
GDPR Article 17(1), and a Californian rule that names the exception out loud
Two independent legal systems write the obligation down. Article 17(1) of the GDPR gives a person the right to erasure without undue delay. It applies when the data are no longer necessary for the purposes they were collected for. California states a comparable right and then, unlike the GDPR, specifies the mechanism. The specification concedes that copies persist.
California's Civil Code § 1798.105(c)(1) requires a business to delete the consumer's personal information from its records. It must also notify service providers, contractors and third parties to do the same. The implementing regulation, § 7022(b)(1), then allows three ways of complying: “Permanently and completely erasing the personal information from its existing systems except archived or backup systems, deidentifying the personal information, or aggregating the consumer information.”
Read the middle of that sentence: except archived or backup systems. § 7022(d) makes the exception explicit. Compliance for data held on archived or backup systems can be delayed until that system is restored to an active system, or is next accessed or used. The statute behind the regulation was added by AB 375, chaptered on 28 June 2018. The regulation's current text takes effect on 1 January 2026.
Deletion is a system property here, not a database command. Erasing the row is the easy half. Proving the agent no longer behaves as though it had read it is the other half. And the regulator has already written into the rule that some copies are still sitting somewhere.
Visual
The five boundaries already have numbers in a published framework
Identity boundary, Purpose boundary, Access policy, Retention policy, and Correction and deletion are decisions taken before personalization. Not after it. They are also not this lesson's invention. Each one has an identifier in the NIST Privacy Framework, Version 1.0, published on 16 January 2020.
The identity boundary is ID.IM-P3: the categories of individuals whose data are processed are inventoried. The purpose boundary is ID.IM-P5, which the Core states in a single line: “ID.IM-P5: The purposes for the data actions are inventoried.” Correction is CT.DM-P3, data elements can be accessed for alteration. Deletion is CT.DM-P4, data elements can be accessed for deletion. The retention schedule is CT.DM-P5, data are destroyed according to policy.
Inventoried is the load-bearing word in both identity subcategories. An inventory is a list you can be handed and audited against. It is what the first two boundaries produce, and it is why they cannot be reconstructed later. The first boundary says whose data this is. The second says what it was gathered for. Neither can be added after the agent has already learned from everything. And a reader who has to argue the point with a procurement team has the citations to argue it with: ID.IM-P3, ID.IM-P5, CT.DM-P3, CT.DM-P4 and CT.DM-P5.
- 1
Identity boundary
Defines whose data and authority a task represents.
- 2
Purpose boundary
States why information may be stored and later retrieved.
- 3
Access policy
Controls which agents, tools, people, and tenants may use it.
- 4
Retention policy
Sets expiry, archival, legal hold, and review schedules.
- 5
Correction and deletion
Propagates changes through derived stores and behavioral state.
Analogy
A bank vault with separate safe-deposit boxes, and a court order about what the bank remembers
Memory partitions work like safe-deposit boxes tied to identities, access rights, and retention rules. One customer's key should not open another box. Emptying a box removes what was inside it. It leaves the summaries, embeddings, and routing rules that box already shaped still acting on the agent.
That last clause is not a metaphor. It is § III(C) of the stipulated order that settled the Alexa case. The order bars Amazon from using deleted Alexa App geolocation information, voice information and children's personal information to create or improve any Data Product — “any model, derived data, or other tool developed using” that information. Then it writes the gap into the remedy: “For the avoidance of doubt, the foregoing provision does not limit Defendant’s use of any Data Product created prior to the processing of the deletion of Alexa App Geolocation Information, Voice Information, or Children’s Personal Information as provided for in Section I, IV(A) and IV(B).”
§ I of the same order sets a 90-day deletion for inactive child profiles. The judgment is a $25,000,000 civil penalty. Judge Tana Lin signed it, and it was entered on 19 July 2023. So the box is emptied on a schedule, the models built from its contents stay in service, and a federal court has said so in writing.
Memory governance must cover both stored data and its continuing influence on decisions.
Comparison
Scope holds where data is read, not in the prompt
Device-level memory, Account-scoped memory, and Task-scoped state answer the question of whose data this is. Each one answers it more precisely than the last.
Device-level memory follows a shared device or installation. It is convenient, it is ambiguous about identity, and it carries the highest disclosure risk. The Alexa complaint is what that risk looks like once it is written down as an allegation: the profiles of more than 800,000 children under 13, linked to a parent's profile on shared household devices. Account-scoped memory ties data to an authenticated account or tenant. That is a clearer partition. It is still exposed to shared accounts, and it depends on delegated roles being modelled properly. Task-scoped state exists only for one bounded workflow: strong minimization, limited continuity, and the safe default when you are not sure which of the other two you have earned.
A scope is only real if cross-user and cross-tenant retrieval is blocked by design rather than by convention. The boundary has to be enforced where the data is read, not left to the prompt. Deletion is the harder half of the same question, and the two Alexa filings show why. An identity boundary drawn at the device made the retention question unanswerable per person. The retention failure then survived into the transcripts and the models.
Device-level memory
Data follows a shared device or installation.
- Convenient
- Identity ambiguity
- High disclosure risk
Account-scoped memory
Data belongs to an authenticated account or tenant.
- Clearer partition
- Shared-account issues
- Needs delegated roles
Task-scoped state
Information exists only for one bounded workflow.
- Strong minimization
- Limited continuity
- Safe default
Key idea
A deleted fact can continue influencing behavior through derived artifacts
Removing the original message may leave summaries, embeddings, profiles, cached contexts, or learned routing rules behind. The user sees deletion. The agent still acts on the information.
How much of it is recoverable is not a matter of opinion. It has been measured on a deployed model. In 2021 Carlini and eleven co-authors attacked GPT-2 as a black box and reported the result in one line: “In particular, among 600,000 (honestly) generated samples, our attacks find that at least 604 (or 0.1%) contain memorized text.”
The route to that figure is the part to hold on to. They generated 600,000 samples, hand-inspected 1,800 candidates, and confirmed 604 unique memorized training examples. The aggregate true positive rate was 33.5%, and 67% for their best attack configuration. Among what came back was one individual's full name, physical address, email address, phone number and fax number. It came out of the deployed model itself. Not out of a database anyone still had a delete button for.
Maintain derivation lineage and verify behavioral removal, not only record deletion.
A deletion that stops at the original record leaves the user believing something is gone while the agent keeps acting on it.
Steps
Create a memory governance matrix, and write the limit of the deletion test into it
A governance matrix makes the identity-by-purpose grid visible, and it should be built against one real workflow. Map identities: users, delegates, organizations, devices, services, and shared contexts. Map purposes: tie every memory class to a permitted future use. Set access scopes across tenant, role, tool, and agent. Define retention, with expiry, correction, legal hold, and archival behavior. Then test deletion end to end. Remove a record and verify that indexes, summaries, caches, and outputs no longer use it. One row per kind of user, one column per purpose, and a rule in every cell that somebody has to sign.
The last step has a limit, and the limit belongs in the cell rather than in a footnote. The same model parameters can be reached from different datasets. So an entity can satisfy the standard approximate-unlearning definition without changing the model at all. Thudi and three co-authors showed that in 2022, and their abstract states the consequence flatly: “Our results show that even for a given training trajectory one cannot formally prove the absence of certain data points used during training.”
So a cell cannot honestly say the data is gone from the model. What it can say is which specific deletion algorithm was run, and where its audit trail lives. Those four authors argue that an algorithm designed for external scrutiny is the only auditable claim available. Filling the cells is what surfaces that awkwardness. Where the rows stay separate, and no cell lets one row read another's data, the matrix is doing its job.
- 1
Map identities
List users, delegates, organizations, devices, services, and shared contexts.
- 2
Map purposes
Tie every memory class to a permitted future use.
- 3
Set access scopes
Define tenant, role, tool, and agent boundaries.
- 4
Define retention
Choose expiry, correction, legal hold, and archival behavior.
- 5
Test deletion end to end
Remove a record and verify indexes, summaries, caches, and outputs no longer use it.
The cost of deletion is fixed before training, not after the first request
Privacy controls belong in the memory architecture before personalization begins. Retrofitting identity and deletion after deployment is difficult and incomplete. The difficulty has been measured.
Sharding the training data is one way to pay for deletion in advance. SISA does exactly that: each point influences only one constituent model, so forgetting a point costs one shard rather than everything. Bourtoule and seven co-authors published it in 2021, and their abstract puts a number on it: “Under no distributional assumptions, for simple learning tasks, we observe that SISA training improves time to unlearn points from the Purchase dataset by 4.63×, and 2.45× for the SVHN dataset, over retraining from scratch.” For ImageNet classification the improvement was 1.36×.
That 4.63× is not a deletion feature anyone shipped after a complaint arrived. It is the return on a partitioning decision taken before a single point was trained on. It is what "privacy controls belong in the memory architecture before personalization begins" looks like in numbers. Whoever answers a deletion request is the person who has to chase the copies. Give that person a way to find everything the deleted fact was used to build. Governance holds when the system itself enforces the separation between users and tenants, and when deletion reaches the derived copies as well as the original.
Personalization built first and governed later hands you a memory you cannot fully answer for when someone asks what you know about them.
Key takeaways
- Agent identity can refer to end users, accounts, devices, organizations, service principals, or delegated actors — and a device is a proxy for none of them. The Alexa complaint alleges more than 800,000 children under 13 had profiles linked to a parent's profile on shared household devices.
- Deletion should cover primary records and derived indexes, summaries, caches, and embeddings according to documented timelines. The FTC's Everalbum order, issued 6 May 2021, sets 30 days for photos and videos, and 90 days for non-consented Face Embeddings and for the models trained on them. Each one is sworn under penalty of perjury.
- The identity and purpose boundaries are inventory items with numbers: ID.IM-P3 and ID.IM-P5 in the NIST Privacy Framework, Version 1.0, published 16 January 2020. Correction, deletion and destruction-on-schedule are CT.DM-P3, CT.DM-P4 and CT.DM-P5.
- A record can be deleted while what was built from it stays in service. § III(C) of the Alexa stipulated order expressly does not limit use of any Data Product created before the deletion was processed. § 7022(d) of the California regulations delays compliance for archived or backup systems until that system is next restored, accessed or used.
- The residue is measurable, not rhetorical. Carlini and eleven co-authors confirmed 604 memorized training examples among 600,000 generated GPT-2 samples, including one person's full name, physical address, email address, phone number and fax number.
- Architecture decided before training is what makes deletion affordable and auditable. SISA cut unlearning time by 4.63× on Purchase and 2.45× on SVHN. Thudi and three co-authors showed that the absence of a data point cannot be formally proved even from the full training trajectory.