certslothcertsloth
AI-103/Topic 05

Azure / Associate

AI Evaluation, Safety and Observability

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Measure groundedness and task success alongside latency, cost and harm.

Must remember

  • Keep representative evaluation datasets separate from tuning examples. Include adversarial, no-answer, multi-turn and subgroup cases. Use rubrics for relevance, correctness, groundedness, safety and action success; calibrate model judges with human review.
  • Inspect retrieval/index health and ingestion freshness separately from model performance. Drift can affect input distributions, tool contracts or knowledge availability. A fluent answer with a real citation may still be unsupported by that source.
  • Configure supported safety filters/guardrails, risk detection and moderation. Test both direct injection and indirect instructions in retrieved documents, images and tool outputs. Restrict agent tools/credentials and enforce oversight outside the model.
  • Use managed identity, least-privilege roles, private networking where appropriate and controlled secret access. Redact sensitive telemetry; record provenance, approvals, model/prompt/tool versions and correlation IDs for investigation.
  • Trace each retrieval, model and tool step. Distinguish quota/rate limiting, invalid inputs, failed authorisation, network timeouts and model-quality failures. Retry only appropriate failures with bounded backoff; reconcile ambiguous external actions by operation ID.
  • Evaluate image/audio safety and accessibility as well as text. Human escalation, documented limitations and a feedback process remain necessary for high-consequence use. No model-generated confidence value is an independent guarantee.

Choose under exam pressure

Requirement Choice and reason
Quality drops after a deployment Compare fixed evaluation cases and component versions.
Agent repeats a tool indefinitely Enforce limits and inspect state/termination conditions.
Telemetry contains personal prompts Apply minimisation, redaction, access and retention policy.

Traps

  • A judge model has its own biases.
  • A passing safety filter is not proof of factual correctness.
  • More logging can increase privacy exposure.

Active recall

1. What makes an evaluation reproducible?

Versioned data, model, prompt, index, tool contracts and scoring criteria.

2. Why include no-answer examples?

To verify abstention instead of unsupported invention.

3. What should be traced for a slow agent?

Model, retrieval and tool spans, retries and token use.

4. Why can a correct source citation still be misleading?

The cited passage may not support the actual claim.

5. Where should consequential action approval occur?

In trusted application/workflow controls before tool execution.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.