Memory hook: Metrics show the trend, logs explain events, traces follow the journey.
Must remember
- Observability is the ability to understand internal system behaviour using emitted evidence. Monitoring detects conditions you have chosen to watch; useful telemetry also helps investigate unexpected problems.
- Metrics are numeric measurements over time, suited to rates, resource use and service health. Logs record events with context. Distributed traces connect spans across services to show where a request spent time or failed. Correlation identifiers help connect the signals.
- Prometheus collects and queries time-series metrics. A common model is to scrape HTTP metrics endpoints using service discovery. An exporter exposes metrics for a component that does not provide them directly. PromQL is the query language.
- Alertmanager groups, deduplicates and routes alerts, and supports silencing/inhibition. A dashboard displays evidence; it is not a substitute for an actionable alert with ownership and a response plan.
| Metric concept | Example and exam distinction |
|---|---|
| Counter | Total completed requests; normally rises, with resets possible |
| Gauge | Current queue depth or memory usage; can rise and fall |
| Histogram | Distribution of observations such as request duration |
| Labels | Dimensions such as service or status class; every combination can create another time series |
- Use a counter's rate over an interval for requests per second, rather than treating its cumulative total as a current rate. Avoid unbounded labels such as user IDs or raw request IDs; high cardinality increases storage and processing work.
- OpenTelemetry supplies vendor-neutral instrumentation, telemetry APIs/SDKs and collection/export components. Its Collector can receive, process and export signals. It is not itself the complete long-term storage and dashboard backend.
- Jaeger is associated with distributed tracing. Fluentd and Fluent Bit collect and forward logs or telemetry. Recognise the tool's job before matching it to a scenario.
- Latency, traffic, errors and saturation are useful service health signals. Latency percentiles reveal slow-tail behaviour that an average can hide. Measuring only CPU can miss a user-visible failure caused by another service.
- An SLI is a measured indicator; an SLO is the target for it; an SLA is a commitment with contractual implications. An error budget expresses tolerated unreliability against an SLO over its chosen window. A 99.9% request-success SLO permits 0.1% unsuccessful eligible requests under that definition.
- Kubernetes resource metrics, such as those often exposed through Metrics Server for
kubectl topand autoscaling, do not constitute a complete historical observability platform. Cost visibility also needs resource attribution, requests versus usage, retention and billing awareness.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Alert when a service's error rate remains high | Metrics and a meaningful alert rule |
| Find which backend made one distributed request slow | A trace with spans and propagated context |
| Read the exception and application context around a failure | Relevant logs, correlated with the request |
| Standardise instrumentation while retaining backend choice | OpenTelemetry |
| Reduce duplicate incident notifications | Alert grouping and deduplication through Alertmanager |
Traps
- Collecting more logs is not automatically better; noisy or sensitive logs create cost and security problems.
- Average latency can hide a poor experience for the slowest requests.
- A trace is not necessarily a record of every request; sampling affects what is retained.
- Prometheus metrics are not a substitute for a complete, exact per-request financial ledger.
Active recall
1. Checkout takes five seconds across four services. Which signal best identifies the slow segment?
A distributed trace whose spans follow the request across the services. Metrics show aggregate trends, while logs can add detailed context for the slow span.
2. A value records the total number of completed requests since startup. Counter or gauge?
Counter. It normally increases and can reset after restart. Calculate a rate over time when you need current throughput; a queue's current depth would be a gauge.
3. Why is a unique request ID usually a poor Prometheus label?
It produces extremely high cardinality because each distinct label combination creates a separate time series. Put request-level detail in suitable logs or traces instead.
4. Does adding the OpenTelemetry Collector automatically provide a permanent searchable telemetry database?
No. The Collector receives, processes and exports telemetry. A configured backend supplies the required retention, query and visualisation capabilities.
5. A team targets 99.9% successful eligible requests. Which term describes the target and which describes its measurement?
The target is an SLO; the measured success ratio is an SLI. The remaining 0.1% represents the request-based error budget under that definition and measurement window.