Memory hook: Measure groundedness and task success alongside latency, cost and harm.
Must remember
- Keep representative evaluation datasets separate from tuning examples. Include adversarial, no-answer, multi-turn and subgroup cases. Use rubrics for relevance, correctness, groundedness, safety and action success; calibrate model judges with human review.
- Inspect retrieval/index health and ingestion freshness separately from model performance. Drift can affect input distributions, tool contracts or knowledge availability. A fluent answer with a real citation may still be unsupported by that source.
- Configure supported safety filters/guardrails, risk detection and moderation. Test both direct injection and indirect instructions in retrieved documents, images and tool outputs. Restrict agent tools/credentials and enforce oversight outside the model.
- Use managed identity, least-privilege roles, private networking where appropriate and controlled secret access. Redact sensitive telemetry; record provenance, approvals, model/prompt/tool versions and correlation IDs for investigation.
- Trace each retrieval, model and tool step. Distinguish quota/rate limiting, invalid inputs, failed authorisation, network timeouts and model-quality failures. Retry only appropriate failures with bounded backoff; reconcile ambiguous external actions by operation ID.
- Evaluate image/audio safety and accessibility as well as text. Human escalation, documented limitations and a feedback process remain necessary for high-consequence use. No model-generated confidence value is an independent guarantee.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Quality drops after a deployment | Compare fixed evaluation cases and component versions. |
| Agent repeats a tool indefinitely | Enforce limits and inspect state/termination conditions. |
| Telemetry contains personal prompts | Apply minimisation, redaction, access and retention policy. |
Traps
- A judge model has its own biases.
- A passing safety filter is not proof of factual correctness.
- More logging can increase privacy exposure.
Active recall
1. What makes an evaluation reproducible?
Versioned data, model, prompt, index, tool contracts and scoring criteria.
2. Why include no-answer examples?
To verify abstention instead of unsupported invention.
3. What should be traced for a slow agent?
Model, retrieval and tool spans, retries and token use.
4. Why can a correct source citation still be misleading?
The cited passage may not support the actual claim.
5. Where should consequential action approval occur?
In trusted application/workflow controls before tool execution.