Memory hook: Outcome, trajectory, quality, cost.
Must remember
- Evaluate memory retrieval, knowledge relevance, prompt behavior, tool choice/arguments and final task completion separately. A correct final answer can hide an unauthorized or wasteful trajectory.
- Use human review, built-in/custom metrics and calibrated LLM judges with golden and adversarial datasets. Validate synthetic test data and avoid judging only examples used during development.
- Trace observable model, agent and tool spans with correlation IDs. Foundry tracing and structured logs should show execution, timing, tokens, errors and approved replay context without indiscriminate sensitive-data capture.
- Monitor agent health, coordination failures, drift, quality regressions and service availability. Define SLOs for successful task completion and latency, with runbooks for recurring failure patterns.
- Optimize bottlenecks through model selection, bounded parallelism, retrieval, caching and prompt length. Respect rate limits; track tokens, tool calls, quotas, allocations and chargeback at useful ownership boundaries.
- Use feedback loops and controlled A/B/canary comparisons. An automated improvement loop must still pass quality, safety and cost gates before changing production.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A workflow succeeds but costs ten times more | Inspect repeated tool/model calls, context growth and runaway coordination. |
| Quality drops after a prompt release | Compare versioned evaluation results and traces, then roll back if gates fail. |
Traps
- Low latency is not success if the agent silently skips required work.
- An LLM judge’s score is not a substitute for calibrated evidence.
Active recall
1. What is trajectory quality?
Whether the agent used appropriate authorized steps/tools to complete the task.
2. Why correlate traces across agents?
To reconstruct dependencies and locate the source of a distributed failure.
3. What should a cost budget constrain?
Relevant tokens, tools, model capacity and workflow execution, with an enforcement response.
4. How detect behavioral drift?
Repeated evaluation against stable criteria plus production feedback and telemetry.
5. Why retain a human review channel?
Ambiguous or high-impact cases need judgment beyond automated scoring.