certslothcertsloth
KCNA/Topic 09

CNCF / Associate

Observability, metrics, logs and traces

4 min read5 recall promptsReviewed 2026-10-10

Memory hook: Metrics show the trend, logs explain events, traces follow the journey.

Must remember

  • Observability is the ability to understand internal system behaviour using emitted evidence. Monitoring detects conditions you have chosen to watch; useful telemetry also helps investigate unexpected problems.
  • Metrics are numeric measurements over time, suited to rates, resource use and service health. Logs record events with context. Distributed traces connect spans across services to show where a request spent time or failed. Correlation identifiers help connect the signals.
  • Prometheus collects and queries time-series metrics. A common model is to scrape HTTP metrics endpoints using service discovery. An exporter exposes metrics for a component that does not provide them directly. PromQL is the query language.
  • Alertmanager groups, deduplicates and routes alerts, and supports silencing/inhibition. A dashboard displays evidence; it is not a substitute for an actionable alert with ownership and a response plan.
Metric concept Example and exam distinction
Counter Total completed requests; normally rises, with resets possible
Gauge Current queue depth or memory usage; can rise and fall
Histogram Distribution of observations such as request duration
Labels Dimensions such as service or status class; every combination can create another time series
  • Use a counter's rate over an interval for requests per second, rather than treating its cumulative total as a current rate. Avoid unbounded labels such as user IDs or raw request IDs; high cardinality increases storage and processing work.
  • OpenTelemetry supplies vendor-neutral instrumentation, telemetry APIs/SDKs and collection/export components. Its Collector can receive, process and export signals. It is not itself the complete long-term storage and dashboard backend.
  • Jaeger is associated with distributed tracing. Fluentd and Fluent Bit collect and forward logs or telemetry. Recognise the tool's job before matching it to a scenario.
  • Latency, traffic, errors and saturation are useful service health signals. Latency percentiles reveal slow-tail behaviour that an average can hide. Measuring only CPU can miss a user-visible failure caused by another service.
  • An SLI is a measured indicator; an SLO is the target for it; an SLA is a commitment with contractual implications. An error budget expresses tolerated unreliability against an SLO over its chosen window. A 99.9% request-success SLO permits 0.1% unsuccessful eligible requests under that definition.
  • Kubernetes resource metrics, such as those often exposed through Metrics Server for kubectl top and autoscaling, do not constitute a complete historical observability platform. Cost visibility also needs resource attribution, requests versus usage, retention and billing awareness.

Choose under exam pressure

Requirement Choice and reason
Alert when a service's error rate remains high Metrics and a meaningful alert rule
Find which backend made one distributed request slow A trace with spans and propagated context
Read the exception and application context around a failure Relevant logs, correlated with the request
Standardise instrumentation while retaining backend choice OpenTelemetry
Reduce duplicate incident notifications Alert grouping and deduplication through Alertmanager

Traps

  • Collecting more logs is not automatically better; noisy or sensitive logs create cost and security problems.
  • Average latency can hide a poor experience for the slowest requests.
  • A trace is not necessarily a record of every request; sampling affects what is retained.
  • Prometheus metrics are not a substitute for a complete, exact per-request financial ledger.

Active recall

1. Checkout takes five seconds across four services. Which signal best identifies the slow segment?

A distributed trace whose spans follow the request across the services. Metrics show aggregate trends, while logs can add detailed context for the slow span.

2. A value records the total number of completed requests since startup. Counter or gauge?

Counter. It normally increases and can reset after restart. Calculate a rate over time when you need current throughput; a queue's current depth would be a gauge.

3. Why is a unique request ID usually a poor Prometheus label?

It produces extremely high cardinality because each distinct label combination creates a separate time series. Put request-level detail in suitable logs or traces instead.

4. Does adding the OpenTelemetry Collector automatically provide a permanent searchable telemetry database?

No. The Collector receives, processes and exports telemetry. A configured backend supplies the required retention, query and visualisation capabilities.

5. A team targets 99.9% successful eligible requests. Which term describes the target and which describes its measurement?

The target is an SLO; the measured success ratio is an SLI. The remaining 0.1% represents the request-based error budget under that definition and measurement window.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.