certslothcertsloth
SCS-C03/Topic 01

AWS / Specialty

Monitoring & Audit

5 min read5 recall promptsReviewed 2026-10-10

Memory hook: Metrics show symptoms, logs explain events, traces follow requests, and audit records identify changes.

Must remember

Separate the evidence questions

  • CloudWatch: how is the workload behaving? Use metrics, logs, dashboards and alarms for operational evidence. CloudTrail: which identity called which API, against what resource and when? AWS Config: what resource configuration existed, and did an evaluated rule consider it compliant?
  • These sources complement each other. A slow API can require a latency alarm, application logs and a trace; identifying an administrator's change requires audit evidence. Check that the needed events, resources and retention were actually configured before promising historical answers.
  • AWS X-Ray follows instrumented requests across services and downstream calls. Its trace map helps find latency, errors and bottlenecks. Traces are not a replacement for every application log or API audit event; instrumentation and sampling affect visibility.

Metrics and logs

  • A metric is a numeric time series identified by namespace, name and dimensions. Choose meaningful statistics and evaluation periods: average latency can conceal slow tail requests, while a total error count without request volume may mislead.
  • Alarms evaluate metric conditions; actions notify or invoke supported responses. Composite alarms combine alarm states to reduce noisy paging. Treat missing data intentionally rather than assuming missing means healthy. Metric streams continuously deliver selected metric updates to downstream consumers.
  • Logs Insights queries log events. Metric filters count matching events into metrics. Subscriptions forward matching logs to supported destinations; export writes log data to S3 for a different processing workflow. These are distinct mechanisms.
  • The CloudWatch agent collects additional guest-OS and application signals. Standard EC2 metrics do not automatically reveal every filesystem or memory measurement.
  • Container Insights and Lambda Insights add workload-specific visibility; Contributor Insights identifies prominent contributors; Application Insights helps correlate application problems. Their scope and collection costs need deliberate configuration.

Events, audit and compliance

  • EventBridge rules match events on buses and deliver them to targets. Scheduler invokes targets on time-based schedules. An archive retains selected events for replay; a replay can repeat a business action, so consumers still need idempotency.
  • A successfully accepted event can fail to match a rule or fail later delivery. Check source/detail pattern, bus, target configuration, resource permissions and retry/dead-letter behavior separately.
  • CloudTrail management events describe control-plane activity; selected data events provide supported resource-level activity. CloudTrail Insights detects unusual supported API activity. A rule reacting to an API call via EventBridge needs the appropriate event path and coverage.
  • Config recording tracks selected resource configurations; rules evaluate compliance. Notifications and remediation are separately configured. A remediation role can mutate resources, so evaluation should not be confused with automatic repair.

Additional published-scope tools

  • Amazon Managed Service for Prometheus stores and queries compatible operational metrics, especially for container workloads using PromQL. Amazon Managed Grafana visualizes metrics, logs and traces from multiple sources. The dashboard layer is different from the metric storage/query layer.
  • AWS Health Dashboard reports AWS service events and account-relevant impacts. Combine it with workload telemetry: a healthy AWS status does not prove your application or configuration is healthy.
  • Keep logs and traces useful: redact sensitive data, set retention, scope collection and correlate request identifiers. Broad logging can create both sensitive-data exposure and substantial ingestion charges.

Choose under exam pressure

Clue in the requirement Choose or investigate
Alert on errors or latency CloudWatch metric alarm with an action
Determine who deleted a database CloudTrail audit events
Review a resource's historical configuration compliance Config history and rule evaluations
Find the slow downstream call in a distributed request X-Ray tracing
Query container metrics using PromQL Managed Service for Prometheus
Visualize several telemetry sources together Managed Grafana
Trigger work for matching application events EventBridge rule and target
Investigate a relevant AWS service disruption AWS Health Dashboard

Traps

  • An alarm with no action is not an email subscription; rule creation alone does not establish every permission needed for delivery.
  • A trace sample or a log metric is not a complete security audit record.
  • A Config rule can report a problem without correcting it. Replaying an event can repeat side effects rather than merely replaying a picture of history.

Active recall

1. Users report slow checkout, but CPU is normal. Which evidence can identify the slow downstream service?

Use distributed tracing such as X-Ray with suitable instrumentation. CPU describes one resource; a trace follows the request through network, application and database dependencies.

2. An auditor asks who changed a security group, then whether it was noncompliant. Is one CloudWatch alarm sufficient?

No. CloudTrail identifies the caller and API change; appropriately recorded Config history and rule evaluations answer configuration/compliance questions. An alarm only evaluates its configured signal.

3. EventBridge accepted an event, but Lambda has no new log. What should you check before broadening permissions?

Verify the bus and event pattern first, then the target and invocation permissions/delivery failures. A valid event can be accepted without matching the intended rule.

4. A Kubernetes team wants PromQL queries and shared dashboards. Which services serve the different roles?

Managed Service for Prometheus provides compatible metric ingestion, storage and queries; Managed Grafana provides visualizations across that and other data sources. A dashboard alone does not collect every needed metric.

5. Replaying last week's order events charged customers again. What architecture property was missing?

Durable idempotent processing keyed to the business operation. Archives and replay help recovery or analysis, but repeated delivery must not automatically repeat a payment.

Terraform anchor: References order resource creation; an event-pattern precondition checks a declared assumption, not end-to-end runtime delivery.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.