Memory hook: Metrics point; logs explain; traces connect; profiles locate cost.
Must remember
Collect platform and application telemetry with the Ops Agent, OpenTelemetry and supported integrations. Managed Service for Prometheus handles Prometheus-style metrics; Cloud Monitoring supports dashboards, alerts and SLO observation. Synthetic checks exercise a user journey externally. A server reporting healthy does not prove login and checkout work.
Use structured logs with severity, service, deployment version and trace correlation. Logs Explorer queries events; routing sinks send selected entries to destinations such as BigQuery, Pub/Sub or Cloud Storage. Exclusions/sampling reduce cost but can remove evidence. Protect audit/security logs and redact sensitive fields before export. Multi-project logging and metrics scopes need explicit IAM and ownership.
Distributed traces connect spans across services. Follow the critical path: downstream timeouts, repeated retries, lock contention or network hops may dominate a request. Profiles identify CPU/memory hot spots; high average CPU does not reveal which code path wastes it. Correlate a regression with a deployment before assuming infrastructure is undersized.
Alert on actionable user impact and rapid error-budget consumption. Route incidents to an owner with a runbook; avoid paging on every harmless transient. Stabilize by rollback, traffic drain or capacity increase as evidence supports, then investigate root causes. AI-assisted analysis is a hypothesis source, not authority to change production without verification.
FinOps joins engineering, finance and product decisions. Attribute costs by project/labels, remove idle capacity, rightsize requests, tune logs and data transfer, and select commitments for stable demand. Spot capacity suits interruption-tolerant work. Dynamic Workload Scheduler and reservations address specific scheduling/capacity needs. Compare cost per useful transaction/job, not only the cheapest VM hour.
Review details
Use the four golden signals—latency, traffic, errors and saturation—plus service-specific indicators. High-cardinality metric labels such as unbounded user IDs can increase cost and make analysis harder. Synthetic checks test an external path; they complement, rather than replace, real-user telemetry.
If observability is missing, trace the telemetry path itself: agent/exporter → credentials/network → ingestion → filter/exclusion → sink/destination → query/permissions. A logs-based metric does not automatically become a useful alert. FinOps recommendations are evidence to investigate; rare peaks and failure headroom may not be visible in a short observation window.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Requests slow across many services | Correlated distributed traces. |
| High CPU inside one process | A profiler plus application context. |
| Growing observability bill | Review retention, volume, cardinality, exclusions and sampling without losing required evidence. |
Traps
- An alert without an owner/runbook often creates noise.
- A recommendation is not evidence that a workload can tolerate its proposed reduction.
Active recall
1. What does a span represent?
A timed operation within a distributed trace.
2. Why add deployment version to telemetry?
To correlate regressions with a particular release.
3. Why can high-cardinality labels be costly?
They multiply distinct metric time series and operational overhead.
4. What is an actionable alert?
One indicating a meaningful condition with a responsible responder and useful next steps.
5. What is a useful FinOps unit metric?
Cost per successful transaction, customer, job or another meaningful business output.