Memory hook: Test the answer and the path taken.
Must remember
- Build golden datasets with expected outcomes, authorized tools and edge cases. Evaluate retrieval, final response, tool selection/arguments, trajectory, safety and successful completion.
- ADK evaluation sets, managed generative evaluation and custom raters address different checks. Calibrate automated judges with human review; avoid evaluating only easy happy paths.
- Separate development evaluations from continuous production sampling. Redact sensitive data, preserve reproducible versions and compare model/prompt/tool changes against a fixed baseline.
- Agent Runtime offers managed agent hosting; Cloud Run suits supported stateless/container workloads; GKE provides Kubernetes control. Compare execution duration, state, networking, scaling and cost.
- Trace model calls, retrieval and tool spans with correlation IDs. Diagnose reasoning loops, slow tools, failed permissions, poor retrieval and unavailable dependencies as distinct problems.
- Roll out gradually with quality, latency, cost and safety gates. Keep a known-good configuration and artifact for rollback; monitor token usage and concurrency as well as request count.
Review details
An evaluation matrix should include answer correctness, retrieved evidence, tool selection, argument validity, action order, authorization and actual completion. A test can fail even if its final text is correct—for example, if an agent read another tenant's document or executed a duplicate payment.
Use fixed regression cases plus fresh production/adversarial samples, keeping expected evidence separate from model training examples. Calibrate autoraters, inspect false passes/failures, and record model/prompt/retrieval/tool/policy versions. When latency rises, inspect spans and repeated loops before buying a larger model or adding more agents.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| An agent answers correctly but calls an unauthorized tool | Fail the evaluation; outcome alone is insufficient. |
| Latency rises after adding a tool | Inspect tool spans and retry behavior before changing the model. |
Traps
- A golden dataset that mirrors training examples overstates quality.
- A healthy container does not prove an agent is completing tasks correctly.
Active recall
1. What is trajectory evaluation?
Checking the sequence and correctness of agent/tool actions, not only the final text.
2. Why version evaluation inputs?
To make before/after comparisons reproducible.
3. How identify a runaway loop?
Repeated spans/actions with rising tokens/time and no progress toward completion.
4. When choose GKE?
When Kubernetes-specific runtime, networking or workload control is required.
5. What should stop a rollout?
A breached quality, safety, reliability or cost threshold defined before release.