Memory hook: Data changes; behavior changes; respond with evidence.
Must remember
- Data drift changes input distributions; concept drift changes the relationship between inputs and desired outputs; training-serving skew is a mismatch between training and production features. They require different investigations.
- Model monitoring and continuous evaluation track quality, distributions and feature attribution where supported. Ground-truth labels may arrive late, so use leading indicators without confusing them with proven accuracy.
- Explainability helps understand predictions, but an explanation is not a causal proof. Evaluate performance and fairness across relevant subgroups and document limitations.
- Protect data and model access with identity, encryption, network controls and logging. Guard against prompt injection, sensitive-data disclosure and malicious tool use; Model Armor and safety filters supplement authorization and application validation.
- For generative systems, evaluate retrieval relevance, groundedness, response usefulness, safety, latency and cost. Trace failures to retrieval, prompt, model, tool or policy rather than retraining indiscriminately.
- Set alert thresholds with a response owner and compare model versions. Retrain, roll back, change retrieval or restrict functionality according to the identified cause.
Review details
Feature attribution drift is a change in how features influence predictions; it differs from a change in input distributions. Model monitoring support depends on model/task/version, so do not assume every custom or generative model exposes identical signals. Keep per-version baselines and stratify important user groups.
For delayed labels, monitor schema, missing features, data distributions, latency and user signals as leading indicators, while scheduling true outcome evaluation when labels arrive. Fixing retrieval freshness or a broken feature transform can be the correct response; automatically retraining on every alert may amplify the problem.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Input distribution shifts but labels are delayed | Investigate drift and operational impact while collecting outcome evidence. |
| A RAG answer cites the wrong passage | Inspect retrieval and grounding before assuming the model needs fine-tuning. |
Traps
- Drift is a warning signal, not automatic proof that accuracy has fallen.
- Safety filters cannot authorize a tool action on behalf of a user.
Active recall
1. How does concept drift differ from data drift?
Concept drift changes the predictive relationship; data drift changes the input distribution.
2. Why evaluate subgroups?
Overall metrics can conceal poor or harmful performance for particular populations.
3. What is training-serving skew?
Different feature generation or availability between training and production.
4. What should an alert include?
The affected model/version, evidence, impact and an actionable response path.
5. Why monitor token consumption?
It affects cost, latency and sometimes runaway workflow behavior.