Memory hook: Follow the data into every copy, model and tool.
Must remember
Object storage, block volumes, database records and ephemeral storage have different access and deletion behavior. Discover structured, semistructured and unstructured data, classify it, map flows and apply policy at every destination. Replicas, caches, logs, backups and model-training datasets can preserve sensitive copies after a primary record is deleted.
Choose encryption and key ownership according to separation, recovery and legal requirements. Customer-managed keys provide control but also create availability duties. External key control can add another dependency. Tokenization, masking, rights management and DLP solve different problems: rights management can restrict permitted use of protected documents; DLP observes/enforces data movement where it has visibility.
Application security requires a secure lifecycle, threat modeling, dependency controls, test coverage and API authorization. Gateways provide centralized policy, but backends still need object-level authorization. WAFs, database activity monitoring and sandboxing supplement correct application design. Container orchestration needs control-plane, image, workload-identity and network protections.
For AI/ML, validate dataset origin, permissions and integrity. Minimize sensitive training data; consider memorization, inference and model-extraction risks. Protect model artifacts and evaluation data as assets. Prompt injection can arrive through retrieved documents and tool outputs, not only the user prompt. Grounding/RAG improves access to information but must enforce document permissions at retrieval time.
Constrain agent tools to specific allowed actions, use short-lived identities, require approvals where appropriate and log actions without exposing secrets. Evaluate both task quality and safety under adversarial inputs. Automated threat detection can produce false positives and drift; retain accountable human review and measurable performance.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Sensitive documents used in RAG | Permission-aware retrieval, minimized context and protected logs. |
| Need to control document usage after sharing | Appropriate information-rights management alongside access controls. |
| AI tool can modify resources | Scoped identity, allowed operations and monitored approval boundaries. |
Traps
- A retrieval filter is only effective if it cannot be bypassed by another query path.
- Encryption does not remove obligations for copies and derived datasets.
Active recall
1. Why inventory derived datasets?
They can preserve protected information and create separate use/retention obligations.
2. What does a customer-managed key add operationally?
Key access, availability, rotation and recovery responsibilities.
3. Does a WAF fix broken object-level authorization?
No. The application must enforce the user’s permission for each object/action.
4. Can prompt injection come from a document?
Yes. Retrieved or tool-provided content is an untrusted instruction source.
5. Why monitor AI security tools after deployment?
Performance and data distributions change; false positives/negatives require ongoing evaluation.