Memory hook: Context is selected evidence, not a transcript dump.
Must remember
- Use clear role/task instructions, examples, output schemas, dynamic context and defensive boundaries. Version prompts and evaluate changes across representative tasks and adversarial inputs.
- Short-term context supports the current session; long-term stores preserve selected facts or semantic memories. Apply authorization, provenance, expiry/deletion and tenant isolation to every read/write.
- Context accumulation eventually exceeds a window. Retrieval, summarization and compaction can help, but preserve exact identifiers, commitments and authoritative source references needed for continuity.
- Sliding-window amnesia loses older facts; summary drift distorts meaning; vector-only recall can miss exact entities or chronology. Combine structured state with semantic retrieval where appropriate.
- Multi-agent RAG requires suitable chunking, embedding versions, metadata/security filters and retrieval precision. Search/MCP-accessible sources must preserve the requesting identity’s access rights.
- Fine-tuning is a separate controlled lifecycle with curated data, evaluation and retraining frequency. Use it for measured behavior needs rather than as a replacement for current factual retrieval.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| An agent repeatedly forgets a customer identifier | Store exact structured session state instead of relying only on semantic memory. |
| Agents need current policy evidence | Authorized retrieval with source versions and relevance evaluation. |
Traps
- Summarization can discard a constraint that matters later.
- Shared vector memory can leak data unless access is enforced before retrieval.
Active recall
1. What is summary drift?
Progressive distortion or omission as context is repeatedly summarized.
2. Why retain provenance in memory?
To verify, update, revoke or explain a remembered fact.
3. How choose chunk size?
Evaluate retrieval relevance, context completeness, latency and token cost on real tasks.
4. What should be compacted cautiously?
Exact IDs, permissions, pending approvals, commitments and safety-critical constraints.
5. When is structured state preferable to embeddings?
For exact values, workflow status, ownership and deterministic lookup.