certslothcertsloth
MLA-C02/Topic 06

AWS / Associate

Evaluation and Responsible AI

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Measure the answer, the experience and the harm; averages can hide who fails.

Must remember

  • Use representative held-out tasks, stable baselines and subgroup analysis. Compare quality, robustness, safety, latency, cost and business outcomes. A benchmark unrelated to the real workflow is weak evidence of suitability.
  • BLEU emphasises n-gram precision against references and is associated with translation. ROUGE uses overlap/recall-oriented measures often applied to summaries. BERTScore uses contextual representations for semantic similarity. None alone proves factual correctness or useful business outcomes.
  • LLM-as-a-judge can scale evaluation but introduces judge bias, model/version sensitivity and prompt dependence. Calibrate against human judgement and inspect disagreements. Bedrock evaluation capabilities support supported automated and human evaluation approaches.
  • Evaluate RAG retrieval separately from answer groundedness and relevance. Evaluate an agent's tool use, action validity and completed task, not just its final wording. Include adversarial inputs and safe refusal/escalation tests.
  • Responsible AI includes fairness, robustness, safety, privacy, transparency, explainability, accountability and veracity. Representative data, label review, human audits and subgroup metrics help find harmful differences hidden by aggregate accuracy.
  • A transparent system exposes relevant workings and limitations; an explanation helps a person understand a particular result or behaviour. Model cards document intended use, evidence, limitations and risk. An open-source model is not automatically interpretable, safe or appropriately licensed.
  • Consider intellectual-property rights, deceptive or biased outputs, environmental impact and user trust. Human-centred design needs clear AI disclosure, feedback and contestability where appropriate. Guardrails reduce specified risks but cannot certify that every response is harmless or true.

Choose under exam pressure

Requirement Choice and reason
Summary wording differs but meaning is similar Use semantic and human evaluation alongside overlap metrics.
High overall accuracy but poor outcomes for one group Subgroup analysis and fairness investigation.
High-consequence decision Human oversight, documented limits and a suitable error/appeal process.

Traps

  • Fairness has multiple definitions and trade-offs.
  • A hallucination score is evidence, not a guarantee.
  • Removing all sensitive columns does not necessarily remove proxy bias.

Active recall

1. Why can accuracy conceal unfairness?

Large groups dominate averages; inspect meaningful subgroups and error costs.

2. Can BLEU verify a factual claim?

No. Reference overlap does not independently establish truth.

3. What is the danger of using only one model as a judge?

Its own biases and blind spots can systematically favour poor answers.

4. What belongs in a model card?

Purpose, training/evaluation context, limitations, risk considerations and operating assumptions.

5. Why evaluate retrieval before generation?

A grounded answer requires relevant authorised evidence; missing evidence is a different failure from misuse of good evidence.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.