AI Agent Memory: What to Remember and What to Forget
AI agent memory is the context an agent can carry between interactions. Short-term memory keeps one conversation or task coherent. Long-term memory preserves selected information across sessions. For a shared team agent, the difficult decision is what becomes a reusable fact, who can access it, and when it must expire.
A delivery agent that remembers your reporting format can save repeated instructions. The same agent can cause trouble if it treats last month’s launch date as current, shares one customer’s terms with another team, or remembers an intended action as completed.
Start with a narrow rule: retain approved, reusable context; fetch changing operational facts from their current source; verify completed actions through the system that performed them. Give each stored fact a source, an owner, and a correction path. Here is how to choose the implementation and test that boundary.
Why shared agent memory needs an owner
Your team should not have to explain the same delivery process every Monday. A shared agent should know the approved reporting format, which team owns an escalation, and where to find the current operating procedure.
The risk grows when remembered context starts influencing several people. An outdated instruction in a personal assistant affects one user. An outdated instruction in a shared delivery agent can affect every project that uses it.
The September 30 edition of The AI Daily Brief put this ownership problem at the centre of its team-agent discussion. The useful takeaway is that shared knowledge needs maintenance, not just collection. A personal agent connected to approved team documents may be sufficient before you introduce a shared, writable memory store.
This matters as products become more persistent. OpenAI introduced dots on September 29, 2026, powered by GPT-6 Astra, with its own cloud computer and the ability to learn from feedback. Specialist organizational dots are entering focused enterprise pilots. That announcement does not establish the memory controls or service commitments your deployment requires. The initial rollout is limited to eligible markets and plans; check your organization’s access before designing around it. Source: OpenAI’s dots announcement.
Short-term memory, long-term memory, and RAG solve different problems
Short-term agent memory helps the agent follow the task in progress: the question, earlier answers, and intermediate context. Long-term agent memory carries selected information into later sessions.
LangGraph documents short-term memory as thread-level state persisted through checkpoints. Its long-term memory stores can use custom namespaces to organize information across sessions. Those namespaces describe how information is organized; they do not, by themselves, establish who is authorized to read it. Source: LangChain’s memory documentation.
Retrieval-augmented generation (RAG) retrieves relevant material and provides it to a model. Agent memory describes the retained context that a system makes available. A memory system can use retrieval, but searching an approved policy document and remembering something a user said have different evidence requirements.
A retrieved document can still be outdated. A remembered statement can still be wrong. In both cases, the agent needs enough source information to distinguish a current record from a plausible sentence.
Our RAG versus fine-tuning guide covers the broader knowledge-design decision. Here, the question is narrower: which facts should persist, and under whose authority?
A delivery example: the agent remembers the wrong launch date
The following example is hypothetical. It illustrates a design problem, not a Powercode customer result.
A delivery agent helped prepare a CRM project update in September. During planning, a project manager mentioned an October 9 go-live date. The agent retained it.
The customer later approved October 16 in the project tracker. On October 2, a different colleague asks the agent to draft the weekly status email.
The agent now has two possible inputs: a remembered October 9 date and a current, approved October 16 record. A system that selects whichever sentence appears most relevant can produce a confident but incorrect update.
The safer design gives those inputs different roles:
- The remembered date is historical context. It explains what was previously discussed.
- The approved project tracker is the operational source of truth. The agent checks it before stating the current launch date.
- A conflict triggers a correction. The old memory is marked as superseded rather than silently left available as a current fact.
If the agent cannot access the tracker, it should flag the date for confirmation. It should not convert an old planning assumption into a customer commitment.
The same distinction applies to actions. “Prepare the status email” is an instruction. “The email was sent” requires confirmation from the sending system. Memory should not turn an intention into evidence of a completed action. Our guide to AI agent idempotency addresses the related problem of repeating actions.
Use a seven-field record for reusable facts
Before choosing a vector database or memory framework, define what a valid memory record must contain. The following seven-field structure is an original design checklist, not a prescribed API for a particular product.
| Field | What it answers | Illustrative value |
|---|---|---|
| Scope | Which tenant and entity does this concern? | Customer A / CRM implementation |
| Fact | What precisely should be retained? | Approved go-live date: October 16, 2026 |
| Source | Where can the claim be checked? | Approved project-tracker record |
| Observed date | When did the system obtain this evidence? | October 2, 2026, in this hypothetical example |
| Owner | Who can approve or correct it? | Named project owner |
| Status | Is it proposed, approved, disputed, or superseded? | Approved |
| Lifecycle | When does it expire, and what does it replace? | Review at the next approved milestone change; supersedes October 9 |
Keep a model’s inference separate from an approved fact. If the agent infers that a customer prefers short emails, label that as a proposed preference until the relevant person confirms it. Do not store the guess as an authoritative customer requirement.
For volatile information such as deadlines, budgets, staffing, and approval status, retaining a pointer to the current record may be more useful than retaining another copy of its value.
Choose the smallest memory architecture that fits
Your implementation choice should follow the ownership and access model. This is a documented-capability comparison, not a hands-on product test or performance benchmark.
| Approach | Consider it when | Verify before committing |
|---|---|---|
| Approved documents or a conventional database | The agent needs stable procedures or current records, with little need to retain new facts automatically. | Document ownership, permissions, versioning, and how the agent finds the current record. |
| OpenAI dots | An eligible existing product could fit the work without a custom agent application. | Available features, organizational access, data controls, and whether the required team use is supported. |
| LangGraph | You are building an application and need to control thread state and cross-session storage. | Your persistence, authorization, correction, and deletion implementation. |
| Amazon Bedrock AgentCore Memory | You want a managed memory component within a compatible agent architecture. | Supported memory configuration, access controls, retention requirements, and integration effort. |
These options operate at different levels: dots is an agent product, LangGraph is an application framework, and AgentCore Memory is a memory component. Compare the fit for your workflow rather than treating them as interchangeable subscriptions.
Amazon describes AgentCore Memory as a fully managed service with short-term interaction storage and long-term extraction of information such as preferences, facts, and summaries. Extraction still needs an application-level policy for what your organization will accept as usable context. Source: Amazon Bedrock AgentCore Memory documentation.
Choose an approach because you can explain how your required facts enter the system, who can retrieve them, and how an incorrect fact stops influencing answers.
Implement memory in five controlled steps
1. Establish a baseline without persistent memory
Run representative tasks using explicit instructions and approved sources. Record where users repeatedly supply the same context and where missing context causes a specific failure. These observations define the problem memory should solve.
If a single task can retrieve everything it needs from a current database, shared long-term memory may add little value. Check the simpler workflow-versus-agent decision before adding another component.
2. Admit only qualified facts
Choose a small initial scope: approved reporting preferences, stable team procedures, or confirmed project context. Define who can approve each class of information.
Exclude credentials, unsupported inferences, and cross-customer information from shared memory. Private conversation history should not become team knowledge merely because it was available to the agent.

3. Enforce scope before retrieval and writing
Apply tenant and user authorization when the agent reads memory and when it proposes a new record. Do not rely on a prompt that asks it to ignore other customers’ information.
A label such as “Customer A” helps organize records. The application must still enforce the access boundary. Review tool permissions alongside memory permissions using the AI agent security guide.
4. Make corrections and expiry part of normal operation
Define what happens when a newer record contradicts an older one. Mark superseded facts, preserve the evidence needed to understand the correction, and prevent old values from being treated as current.
Set review or expiry rules by information type. A formatting preference and a launch date should not share an arbitrary retention rule.
5. Verify deletion beyond the visible record
Map any derived summaries, retrieval indexes, replicas, logs, and backups your implementation creates. Document what deletion removes immediately, what remains under a retention policy, and how you verify each boundary. Do not promise complete erasure based on one successful database operation.
Define eight release checks before the agent remembers more
These are proposed checks, not tests we have performed on the products above.
- Approved recall: retrieve the right fact for the right project and identify its source.
- Tenant isolation: reject a request for another customer’s retained information.
- Permission change: stop returning information after the user loses access.
- Correction: prefer the approved replacement and identify the old record as superseded.
- Expiry: flag an expired fact instead of presenting it as current.
- Source unavailable: qualify a volatile claim when its current record cannot be checked.
- Action confirmation: distinguish requested, attempted, and tool-confirmed actions.
- Deletion: verify removal from each documented active or derived store.
Define the denominator before reporting a score. Approved-fact retrieval accuracy is cases that return a correct, source-backed answer divided by cases with an approved expected answer. Stale-answer rate is cases that present a superseded or expired fact as current divided by cases with a known correction or expiry. Record correction propagation time from approval until every relevant retrieval path returns the replacement. These measures are proposed, not benchmark results.
Keep cross-tenant disclosure as a release failure, not an average score that strong answers elsewhere can offset. Your test report should state the cases, environment, and limitations. Use the agent evaluations guide to structure that evidence.
Also decide who can disable memory writes and contain a suspected disclosure. That belongs in the agent incident-response plan, before the first incident.
Frequently asked questions
Does long-term agent memory mean the model is retrained?
No. The approach described here stores context outside the model and supplies selected information during later interactions. It does not require fine-tuning the model. Check the implementation before assuming what “memory” means in a particular product.
Should every conversation become shared team memory?
No. Conversations can contain private information, temporary assumptions, and conflicting instructions. Use an explicit admission policy. Retain approved facts for a defined purpose and audience rather than treating an entire conversation as team knowledge.
When should an agent forget something?
When it expires, is superseded, loses its legitimate purpose, or must be removed under the applicable retention or deletion policy. Forgetting must address the stores your system actually uses, including derived records where relevant.
If your team is deciding what an agent should retain, talk to Powercode about the workflow and its data boundaries. Start with one repeated task, its current sources, and a clear list of failures the design must prevent.