Memory hook: Healthy, authorized, reachable, reversible.
Must remember
- Cloud Run can build from source or deploy a container. Configure the listening port, runtime identity, ingress, invoker permissions, concurrency, resource limits and minimum/maximum instances.
- Eventarc or Pub/Sub invocation needs both a correctly configured receiver and authorized delivery identity. A deployed service is not necessarily publicly callable.
- Cloud Run revisions support traffic migration and rollback. Set maximum instances with downstream connection limits in mind; retries can multiply traffic during an outage.
- GKE Deployments manage replicas and rolling updates; Services provide stable access; Gateway/Ingress expose appropriate traffic. Resource requests influence scheduling and autoscaling.
- Readiness controls traffic eligibility; liveness restarts an unhealthy container; startup probes protect slow initialization. An aggressive liveness probe can turn dependency slowness into a restart storm.
- Horizontal Pod Autoscaler adjusts replicas from supported metrics; node capacity must also exist. Use controlled rollout checks, graceful termination and durable external state.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A process starts slowly but then works | A suitable startup probe rather than repeated liveness failures. |
| Reduce deployment risk | Shift a small traffic share, compare errors/latency, then promote or roll back. |
Traps
- A ready pod and a healthy external load balancer are separate checks.
- Increasing replicas can overload a database that cannot accept more connections.
Active recall
1. What does readiness failure do?
Removes a pod from normal service traffic eligibility without necessarily restarting it.
2. Why cap Cloud Run instances?
To control cost and protect limited downstream capacity, while accepting possible queuing/rejection tradeoffs.
3. What makes a Pub/Sub push fail with authorization errors?
Missing invoker permission, wrong delivery identity or token audience configuration.
4. Why set CPU requests for metric-based HPA?
Utilization-based scaling depends on the requested resource baseline.
5. What must be retained for rollback?
A known-good artifact/revision and compatible configuration and data behavior.