Reviewed 10 October 2026 against the current CNCF curriculum and Linux Foundation exam page. This is the four-domain scope: Kubernetes fundamentals 44%, container orchestration 28%, cloud native application delivery 16%, and cloud native architecture 12%. Older revision material can show a different five-domain outline. KCNA is a conceptual multiple-choice exam; this guide does not imply a particular hands-on exam runtime version.
Memory hook: Know what runs, who decides, how traffic and data move, how changes arrive, and how you know the result is healthy.
Kubernetes fundamentals — 44%
Architecture: the API server receives requests; etcd persists Kubernetes object state; controllers reconcile desired and observed state; the scheduler chooses a node; the kubelet works with the runtime to execute assigned Pods. A node can be physical or virtual. Backing up etcd is not a substitute for backing up application volume data.
Objects: a Pod groups closely coupled containers with a shared network identity and declared volumes. A Deployment manages stateless replicas and rollouts, a StatefulSet supports stable per-replica identity, a DaemonSet covers eligible nodes, a Job runs to completion, and a CronJob creates scheduled Jobs. StatefulSet does not implement the database's replication or backups for you. Labels group resources; annotations carry additional metadata. A namespace organises objects and policy scope without automatically isolating traffic.
Administration: kubeconfig contexts select a cluster, identity and default namespace. Remember get for lists, describe for details/events, logs for process output, explain for fields, and auth can-i for permission checks. Declarative manifests describe intended state through apiVersion, kind, metadata and spec; status reports observed state. ConfigMaps hold ordinary settings; Secrets hold sensitive values. Base64 is not encryption. Changing an environment-variable source does not update an already running process automatically.
Scheduling: requests influence placement; limits constrain runtime consumption. CPU over a limit is normally throttled; memory exhaustion can kill a container. 500m means half a CPU. Node selectors and affinity express placement; anti-affinity and topology spread can improve failure-domain distribution. Tolerations permit placement despite a taint but do not guarantee it. HPA changes replicas, VPA changes resource requests according to configuration, and node autoscaling changes node capacity. ResourceQuota limits a namespace's aggregate use; LimitRange sets defaults or bounds.
Containers: an image is the artifact, a registry distributes it, and a container is a running instance. Containers ordinarily share a host kernel. Linux namespaces isolate views; cgroups account for and constrain resources. Tags can move; digests identify content. OCI defines interoperable specifications; CRI runs, CNI connects, CSI stores. containerd and CRI-O are runtimes. A Docker-built compatible image does not require Docker Engine as the Kubernetes runtime. Use minimal patched images, avoid embedded secrets, and separate build tools from the final image with multi-stage builds.
Container orchestration — 28%
Networking: clients should normally use Service discovery rather than memorised Pod IPs. ClusterIP is internal, NodePort exposes a node port, LoadBalancer requests a supported external integration, and ExternalName returns a DNS alias. A headless Service exposes endpoint discovery without the normal virtual ClusterIP. Service selectors must match the intended Pods, and targetPort must match the application's listener. Ingress and Gateway API require implementations; an object by itself cannot provide the traffic path.
Security: authentication establishes identity, authorisation grants allowed actions, and admission validates or mutates a proposed object change. Use workload ServiceAccounts and least-privilege RBAC. A RoleBinding referencing a ClusterRole still grants applicable access within the binding's namespace; a ClusterRoleBinding grants across the cluster. NetworkPolicy requires an enforcing plugin. If both ends are isolated, both source egress and destination ingress must allow the connection. Policy does not encrypt traffic; mutual TLS addresses a different requirement. Harden workloads with appropriate security contexts and admission rules.
Troubleshooting: follow the failing layer. Pending suggests placement or setup evidence; ImagePullBackOff suggests image retrieval; CrashLoopBackOff suggests repeated failure, with prior logs and exit reasons needed for diagnosis. Running does not imply Ready. Check context, namespace, events and recent changes before guessing. For a failed Service, inspect DNS, selectors, EndpointSlices, readiness, target ports, listeners and policy.
Storage: emptyDir survives a container restart within a Pod but not removal of that Pod from its node. A PVC requests storage; a PV represents it; a StorageClass and provisioner can supply it dynamically. RWO means one node, not necessarily one Pod; RWOP is the stricter single-Pod mode for supported CSI storage. RWX needs a backend that supports sharing. Retain and Delete reclaim policies have different data consequences. Storage topology and delayed binding can affect placement. Persistent storage still needs backup and tested restoration.
Cloud native application delivery — 16%
Release pipeline: commit, build, test and scan, publish an immutable artifact, deploy, then verify. CI integrates and checks changes. Continuous delivery keeps them releasable and may retain manual approval. Continuous deployment automatically releases qualifying changes. Build once and promote the same reviewed artifact when possible.
GitOps: keep desired state declarative and versioned, let agents pull it, and reconcile continuously. Argo CD and Flux are examples. A one-off kubectl apply command is not itself a continuous GitOps system. Helm packages templates and values as charts; Kustomize applies overlays and patches to manifests.
Release choices: rolling updates gradually replace replicas, blue-green prepares another environment before switching traffic, and canary limits initial exposure while measuring results. Surge capacity, readiness and rollback compatibility matter. Reverting an image does not undo database changes or external side effects. A software bill of materials inventories components; it does not prove an artifact is vulnerability-free.
Debugging: readiness controls traffic eligibility; liveness detects when a container needs restarting; startup probes protect slow initialisation before the other checks take over. Use logs --previous for the last terminated container and -c to select the intended container. Prefer evidence over deleting Pods until a symptom temporarily disappears.
Cloud native architecture — 12%
Observability: metrics show numeric trends, logs record events, and traces connect a request's spans across services. Prometheus collects time-series metrics and uses PromQL; Alertmanager groups and routes alerts. Counters accumulate, gauges rise and fall, and histograms describe distributions. Use a counter's rate for throughput and avoid unbounded metric labels. OpenTelemetry instruments and transports telemetry; a backend supplies long-term analysis and storage. An SLI is a measurement, an SLO its target, and an SLA a contractual commitment. Latency, traffic, errors and saturation help describe service health.
Principles: cloud native applies on premises and at the edge as well as in public clouds. Favour automation, explicit desired state, resilience and manageable change. Microservices permit independent deployment but add distributed-system complexity. Immutable delivery replaces versioned components rather than relying on unrecorded manual edits. Elasticity adjusts capacity with demand. Serverless shifts infrastructure operation to a platform without removing responsibilities for code, identity or data.
Ecosystem and community: choose tools by their job: Kubernetes orchestrates; Harbor stores images; etcd stores cluster state; Envoy proxies traffic; meshes add traffic features; Fluentd/Fluent Bit collect telemetry; Open Policy Agent evaluates policy. CNCF is a neutral home for projects. Sandbox, Incubating and Graduated indicate project maturity processes, not a guarantee for every installation. Follow governance, contribution instructions, code of conduct and licence requirements. Documentation, translations, testing and reproducible bug reports are valuable contributions alongside code.
Last-minute traps
- A namespace is not automatically a security or network boundary.
- Readiness failure does not by itself restart a container.
- A toleration does not reserve or force a node placement.
- More replicas do not create more node capacity or repair a broken application.
- RWO is not the same promise as RWOP.
- A mutable image tag is not a content identity.
- A Git repository is not a safe place for plain credentials.
- A persistent disk, successful rollout or Graduated project still needs operational verification.
Final active recall
1. A Deployment has the correct desired replica count, but new Pods remain Pending. What evidence and scaling distinction matter?
Inspect events, requests, available node resources, placement constraints and PVC status. HPA changes the desired replica count; node autoscaling changes the capacity that may be needed to place those replicas. A matching toleration alone does not guarantee placement.
2. A healthy-looking Pod receives no traffic through its Service. What is your shortest useful investigation path?
Confirm the context and namespace, then check the Service selector, matching Pod labels, EndpointSlices, readiness, Service target port and application listener. Continue through DNS and NetworkPolicy as the evidence requires. Running is not equivalent to Ready.
3. A team wants reproducible releases and automatic correction of manual cluster edits. Which artifact reference and delivery model fit?
Pin reviewed image digests and keep desired configuration in a reviewed repository. Use a GitOps reconciler such as Argo CD or Flux to pull and continuously reconcile the declared state. Keep credentials protected outside plain repository content.
4. The previous release is restored after an incident, but customer data remains corrupted. Why is that possible?
Application rollout reversal does not automatically reverse database writes or schema changes. Persistence preserves the current data, including damage. Recovery needs an appropriate backup or data-repair strategy and verified restoration.
5. A request is slow across several services, and the team wants to explain the delay. Which telemetry and community assumptions should they avoid?
Use a distributed trace to locate the slow spans and correlate it with metrics and logs. OpenTelemetry can produce and move those signals but needs a suitable analysis backend. Choosing a CNCF Graduated project does not remove the need to configure, instrument and operate it correctly.
Sources and scope
- Current CNCF KCNA curriculum
- CNCF exam domains
- Linux Foundation KCNA information
- Kubernetes concepts
- OpenGitOps principles
- CNCF glossary
The curriculum mapping is attributed to CNCF under CC BY 4.0. The explanations and recall questions are original study material. The official outline names broad competencies rather than an exhaustive list of individual commands; use the linked topic references to deepen areas you cannot yet explain confidently.
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Cluster components and desired state
Memory hook: The API receives, etcd remembers, controllers reconcile, the scheduler places, the kubelet runs.
Must remember
- Kubernetes orchestrates containerised applications. You declare the desired state; control loops repeatedly compare it with the observed state and act on the difference. This supports recovery without a person restarting every failed application.
- A cluster combines a control plane with worker nodes. A node can be a virtual or physical machine. Highly available control planes avoid making one machine the only route to cluster management.
- kube-apiserver exposes the Kubernetes API and is the entry point for clients and components. etcd stores cluster configuration and state. Backing up etcd protects Kubernetes object data; it does not automatically back up files inside application volumes.
- kube-controller-manager runs reconciliation controllers, such as the controller that maintains a requested replica count. kube-scheduler selects a suitable node for an unscheduled Pod. It does not start that Pod's containers itself.
- kubelet runs on each node and works with a container runtime to keep assigned Pods running. kube-proxy, or a networking implementation that replaces its role, supplies Service traffic forwarding. A cloud controller integrates supported cloud infrastructure.
- A Pod is the smallest deployable Kubernetes unit: one or more closely coupled containers with a shared network identity and declared volumes. Containers in one Pod reach each other through
localhost; separate Pods have separate identities. - An object manifest normally has apiVersion, kind, metadata and spec.
specexpresses intent;statusreports observed conditions. YAML is a representation of API objects, not a different control plane. - Labels identify and group objects; selectors find matching groups. Annotations carry extra descriptive metadata. Namespaces organise namespaced objects; nodes and PersistentVolumes are examples of cluster-scoped objects.
Read the story behind a Deployment: submit its manifest → controllers create the required child objects → scheduler assigns Pods → kubelets start containers → readiness determines whether application endpoints should receive traffic. A controller may replace a failed Pod with a new identity rather than repair the old Pod in place.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Maintain three application replicas after one Pod disappears | A workload controller, such as a Deployment; a standalone Pod does not provide that replica-management loop |
| Identify where cluster object state is persisted | etcd; an image registry stores images instead |
| Place a Pod on an eligible node | Scheduler; kubelet performs execution after assignment |
| Group all frontend objects for selection | A label such as app=frontend; an annotation is not the normal grouping mechanism |
| Add a new API resource with its own reconciliation logic | A CustomResourceDefinition plus a controller; an operator packages application-specific operational knowledge |
Traps
- Kubernetes does not provide a complete application database, source repository or build pipeline just because it runs containers.
- A namespace is a scope for organisation and policy, not an automatic network isolation boundary.
- Pod IPs and container writable layers are poor places to keep an application's durable identity or only copy of data.
- Losing API availability affects management and reconciliation; it does not mean every already running container instantly stops.
02 · Containers, images and runtime interfaces
Memory hook: Build an image, store it in a registry, run it as a container, group it in a Pod.
Must remember
- An image packages application files, dependencies and startup metadata. A container is a running instance with a writable layer and runtime isolation. A registry stores and distributes image artifacts; it does not schedule applications.
- Containers ordinarily share the host kernel. Virtual machines have a guest operating system and kernel. Containers are lightweight process isolation, not a guarantee of the same isolation boundary as separate VMs.
- Linux namespaces give processes isolated views of resources such as process IDs and networking. Control groups, or cgroups, account for and constrain resource use. Linux namespaces and Kubernetes namespaces are different concepts.
- Image layers enable reuse and caching. A tag such as
stablecan point to a different image later; a digest identifies specific content. Pin digests when an exact, reproducible artifact matters. - A Dockerfile describes an image build. A multi-stage build separates compilation tools from the final runtime image. Smaller images can reduce download time and attack surface; they still need patching and vulnerability review.
- The Open Container Initiative, OCI, defines interoperable image, runtime and distribution specifications. Kubernetes' Container Runtime Interface, CRI, lets the kubelet talk to runtimes such as containerd and CRI-O. Image-building tools and node runtimes serve different purposes.
- Kubernetes also uses CNI for container networking integrations and CSI for storage integrations. Remember: CRI runs; CNI connects; CSI stores. These interfaces make implementations replaceable without making their capabilities identical.
- Treat container filesystems as replaceable. Put runtime settings in configuration, keep credentials out of image layers, run as a non-root user where possible, drop unneeded privileges and rebuild patched images instead of manually editing every running container.
- An ordinary init container completes setup before the main application starts. A sidecar supports the application while it runs, for example with a proxy or telemetry agent. Use separate Pods when components need independent scaling or lifecycles.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Guarantee the deployed artifact matches the reviewed artifact | An image digest, supported by signing/provenance checks where required |
| Give a process its own view of networking and process IDs | Linux namespace isolation |
| Restrict a container's resource consumption | Resource controls implemented using mechanisms such as cgroups |
| Run OCI images on a Kubernetes node | A CRI-compatible runtime such as containerd or CRI-O |
| Produce a smaller production image after compiling code | A multi-stage build with a minimal final runtime stage |
Traps
- An image built with Docker can run through a compatible runtime without Docker Engine being the Kubernetes runtime.
- Deleting a secret in a later image layer does not reliably remove it from earlier layers or build history.
latestis a tag, not a promise that an already running Pod continuously updates itself.- A container image must support the target operating system and CPU architecture; packaging does not eliminate platform compatibility requirements.
03 · kubectl, configuration and access
Memory hook: Context selects the destination; authentication proves identity; authorisation grants actions; admission checks the change.
Must remember
- kubectl is a client of the Kubernetes API. A kubeconfig context combines a cluster, credentials and a default namespace. Check the context before acting; an accurate command aimed at the wrong cluster is still wrong.
getlists objects,describeadds object details and events, andexplaindescribes API fields.-n team-achooses a namespace;-Alists across namespaces when authorised.-o yamlor-o jsonexposes object structure.- Imperative commands describe an immediate operation, such as creating an object. Declarative management records the desired configuration in manifests and applies it repeatedly.
kubectl apply -f file.yamlsubmits desired configuration; it does not establish a continuously running GitOps controller. - A ConfigMap carries non-confidential settings. A Secret carries sensitive values and can be mounted or injected into a Pod. Base64 in a Secret manifest is encoding, not encryption; encryption at rest, restricted access and safe handling are separate controls.
- Mounted configuration can update over time, but an application must notice and reload it. Values injected as environment variables do not automatically change in an existing process; replacing Pods is often the intended rollout mechanism.
- Authentication determines who makes an API request. Authorisation, commonly RBAC, determines whether that identity may perform the requested verb on the resource. Admission can validate or mutate an authorised object-creation or modification request before persistence.
- Role and RoleBinding usually express permissions within a namespace. ClusterRole defines reusable or cluster-scoped permissions. A RoleBinding can reference a ClusterRole but grants applicable permissions only in the binding's namespace; a ClusterRoleBinding grants across the cluster.
- ServiceAccounts are workload identities. Prefer short-lived, scoped credentials and omit API token mounting when an application does not need Kubernetes API access. Grant least privilege; a namespace does not justify giving every application cluster-admin.
- Pod Security Standards offer Privileged, Baseline and Restricted profiles. Pod Security Admission can enforce, warn or audit at namespace scope. Application identity, network policy and container hardening protect different layers.
Read-only command recognition:
kubectl config current-context
kubectl get pods -n team-a -o wide
kubectl describe pod app -n team-a
kubectl explain deployment.spec
kubectl auth can-i get secrets -n team-a
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Keep a non-sensitive API endpoint outside the image | ConfigMap; rebuilds should not be necessary for each environment setting |
| Allow a workload to read only selected resources in its namespace | A ServiceAccount with narrowly scoped RBAC |
| Check whether the current identity may list Pods | kubectl auth can-i list pods in the intended namespace |
| Reject privileged workloads during admission | Appropriate Pod Security Admission enforcement or another suitable admission policy |
| Identify why a user receives Forbidden after logging in | Investigate authorisation, bindings, verbs and resource scope |
Traps
- Possessing a valid identity does not imply permission to perform every action.
- A Secret object is not automatically encrypted merely because its YAML looks unreadable.
- Listing or watching Secrets can expose secret data; “read-only” is not necessarily harmless.
- ConfigMap or Secret changes do not rebuild an image or automatically restart every consumer.
04 · Workloads, scheduling and scaling
Memory hook: Deployment repeats; StatefulSet identifies; DaemonSet covers nodes; Job finishes; CronJob schedules runs.
Must remember
| Workload | What it manages | Recognise this requirement |
|---|---|---|
| Deployment | ReplicaSets and replaceable Pods | Stateless services, rolling updates and replica count |
| StatefulSet | Pods with stable identities and ordered lifecycle features | Stable names and per-replica persistent storage |
| DaemonSet | A Pod on each eligible node | Node log collectors, network agents or monitoring agents |
| Job | Work that runs to completion | A batch calculation or one-off migration |
| CronJob | Jobs created on a schedule | Regular reports or scheduled maintenance work |
- A ReplicaSet maintains the requested replica count; a Deployment adds rollout management. A StatefulSet does not make database replication, backups or application consistency automatic.
- The scheduler filters and scores nodes against a Pod's requirements. Requests express resources needed for scheduling. Limits constrain runtime usage. For CPU,
500mis half a CPU; memory units such asMiandGiare binary quantities. - CPU over a limit is normally throttled. Exceeding available or permitted memory can lead to an out-of-memory kill. A large request can keep a Pod Pending even while actual node usage appears low.
- nodeSelector matches node labels. Node affinity supports richer required or preferred placement rules. Pod affinity/anti-affinity considers other Pods. Topology spread constraints help distribute replicas across failure domains.
- A taint discourages or prevents placement on a node. A matching toleration allows a Pod to tolerate that taint; it does not guarantee the Pod will be scheduled there. Other resource and placement checks still apply.
- Horizontal Pod Autoscaler, HPA, adjusts replica count from observed metrics. Vertical Pod Autoscaler, VPA, adjusts resource requests according to its configuration. Node autoscaling changes the available node capacity. None of these replaces a functioning application design.
- ResourceQuota limits aggregate resource use in a namespace. LimitRange can set per-object defaults and bounds. They do not grant permissions to users.
- Multiple replicas improve availability only when placement, capacity, readiness and dependent services support it. A PodDisruptionBudget limits certain voluntary disruptions; it does not prevent every hardware failure or guarantee recovery capacity.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Run one node telemetry agent on every eligible node | DaemonSet |
| Add more web replicas when measured demand rises | HPA with a suitable metrics source |
| Add capacity because Pods cannot fit on existing nodes | Node autoscaling or capacity changes; more replicas alone cannot create node resources |
| Keep replicas away from a single failure domain | Suitable anti-affinity or topology spread configuration |
| Run a report every night and record completion | CronJob creating Jobs |
Traps
- A toleration removes a placement obstacle; it does not attract or reserve a node for the Pod.
- Requests influence scheduling even when current measured usage is low.
- A Deployment needs usable labels and selectors to manage the intended Pod set.
- Increasing replicas does not fix a broken image, incorrect configuration or an unavailable external database.
05 · Services, DNS and network security
Memory hook: Pod IPs change; Services give a stable destination; DNS gives a name; policy controls the path.
Must remember
- Each Pod has a network identity. Containers inside one Pod share its network namespace and port space, so they can communicate through
localhostbut cannot both bind the same address and port. - A network implementation provides Pod connectivity, commonly through CNI plugins. CoreDNS commonly provides cluster DNS. The mechanism routing Service traffic may be kube-proxy or another implementation.
- A Service describes a stable way to reach a changing set of backends. Label selectors usually identify its Pods, and EndpointSlices record backend addresses and readiness information. A Service does not create those application Pods.
| Service type | Recognition cue |
|---|---|
| ClusterIP | Stable address inside the cluster; the default type |
| NodePort | A port exposed on nodes, subject to network reachability and firewalls |
| LoadBalancer | An external load balancer supplied by an available integration |
| ExternalName | DNS alias to another name; no ordinary Pod proxying |
- A headless Service uses
clusterIP: Noneand allows discovery of endpoint addresses rather than providing the normal virtual ClusterIP. It is useful when clients need individual backend identities. service.namespace.svc.<cluster-domain>is the normal fully qualified Service name pattern. The cluster domain is oftencluster.local, but it is configurable.portis the Service-facing port;targetPortis the backend application port.- Ingress expresses HTTP/HTTPS routing such as host and path rules. It requires a controller. Gateway API provides more expressive, role-oriented networking APIs and also needs an implementation. Neither an API manifest nor a Service type guarantees external infrastructure exists.
- A NetworkPolicy selects Pods and permits particular ingress or egress traffic. Enforcement requires a supporting network plugin. Policies are additive allow rules: when both ends are isolated, the source egress and destination ingress rules must permit the connection.
- Without applicable isolation policies, standard Kubernetes network policy behaviour is permissive. Namespace separation alone does not block packets. A default-deny policy often needs explicit allowances for DNS and application dependencies.
- A service mesh adds application traffic features such as identity-based mutual TLS, traffic management and telemetry. It can use proxies such as Envoy. It complements basic networking; it does not make RBAC or application authorisation unnecessary.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Reach a backend despite replacement Pod IPs | Service and cluster DNS |
Route shop.example and /api to different applications |
Ingress or an appropriate Gateway API implementation |
| Block arbitrary Pod-to-Pod access | NetworkPolicy plus an enforcing plugin |
| Discover each stateful backend directly | A headless Service and suitable DNS discovery |
| Encrypt and authenticate service-to-service connections | Mutual TLS, often provided through a service mesh |
Traps
- A ClusterIP Service is not automatically reachable from the public internet.
- A Service with the wrong selector can exist while having no usable application endpoints.
- NetworkPolicy controls connectivity; it does not by itself encrypt packets or decide Kubernetes API permissions.
- A NetworkPolicy object has no useful enforcement effect if the installed network implementation does not support it.
06 · Volumes, claims and persistence
Memory hook: A volume is mounted, a claim asks, a PV supplies, a StorageClass provisions.
Must remember
- The writable layer of a container is disposable. A volume exposes storage to containers in a Pod. Containers share the volume only when each mounts it; their independent root filesystems are not automatically shared.
- emptyDir starts empty when a Pod is assigned to a node. It survives individual container restarts within that Pod but is removed when the Pod is removed from the node. Use it for scratch data or cooperation between containers, not the only copy of important records.
- A PersistentVolume, PV, represents a storage resource in the cluster. A PersistentVolumeClaim, PVC, is a namespaced request for storage with requirements such as capacity and access mode. A Pod mounts a claim; the claim binds to suitable storage.
- Static provisioning supplies an existing PV. Dynamic provisioning uses a StorageClass and provisioner to create backing storage for a suitable claim. CSI provides an integration standard between Kubernetes and storage systems.
- A PV is cluster-scoped; a PVC is namespaced. The workload and its claim normally belong to the same namespace. Persistent storage has a lifecycle separate from an individual Pod, but its survival still depends on deletion and reclaim settings.
| Access mode | Meaning to remember |
|---|---|
| ReadWriteOnce, RWO | Read-write mounting from one node; several Pods on that node may use it |
| ReadOnlyMany, ROX | Read-only mounting from multiple nodes |
| ReadWriteMany, RWX | Read-write mounting from multiple nodes |
| ReadWriteOncePod, RWOP | Read-write mounting by one Pod, with supported CSI storage |
- Access modes depend on driver and backend capabilities. Choosing RWX in YAML does not turn a single-node block disk into a distributed shared filesystem.
- A Retain reclaim policy keeps released storage for deliberate recovery or cleanup. Delete removes the associated storage through its supported provisioner when reclamation occurs. Know which data may disappear before deleting a claim.
- Storage topology matters: a disk can be restricted to a zone or node. A StorageClass with WaitForFirstConsumer delays binding or provisioning until scheduling information is available, helping avoid incompatible placement.
- Block storage provides devices, file storage provides filesystem access, and object storage exposes objects through an API. A cloud object bucket is not automatically a normal POSIX filesystem volume.
- Persistence is not a backup strategy. Protect against application corruption, accidental deletion and infrastructure loss with appropriate backup, restore testing and storage durability choices.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Share disposable scratch files between containers in one Pod | A mounted emptyDir |
| Retain application data through Pod replacement | PVC-backed persistent storage with appropriate lifecycle settings |
| Create storage automatically when a claim is submitted | A suitable StorageClass and dynamic provisioner |
| Allow different nodes to write to one supported shared filesystem | RWX-capable storage, not merely an RWX declaration |
| Coordinate storage allocation with workload placement | WaitForFirstConsumer where supported and appropriate |
Traps
- RWO means one node, not necessarily one Pod.
- Deleting a Pod and deleting its PVC are different operations with different data consequences.
- A Pending PVC can indicate missing capacity, an incompatible class or access mode, or delayed binding; it is not always a broken application image.
- A persistent disk can preserve corrupted data just as persistently as correct data.
07 · Troubleshooting and application debugging
Memory hook: Find the failing layer: placement, image, process, readiness, Service, network or dependency.
Must remember
- Start with scope and evidence: current context, namespace, affected objects, recent changes and the intended outcome.
kubectl getgives the overview;describeand events explain lifecycle failures; logs explain application behaviour. - A Pending Pod has not completed setup. Events may reveal insufficient resources, incompatible node selection, a missing volume or other scheduling/setup problems. Do not assume that an application exception causes every Pending state.
- ImagePullBackOff points to repeated image-pull failures: check the image name/tag/digest, registry access, credentials and connectivity. CrashLoopBackOff means a container repeatedly exits or fails and restart attempts are backing off; it describes a symptom, not a unique root cause.
- For crashes, inspect the current and previous container logs, exit reason, command, configuration, Secrets and resource limits.
OOMKilledmakes memory behaviour a priority. In multi-container Pods, select the correct container with-c. - Readiness asks whether a container should receive application traffic. Liveness asks whether a running container needs restarting. Startup gives a slow-starting application time to initialise before liveness and readiness checks take over.
- A Pod can be Running without being Ready. A failed readiness check can remove its endpoint from normal Service traffic without restarting the container. Repeated liveness or startup failures can trigger container restart according to the configured behaviour.
- For connectivity, trace name resolution → Service → EndpointSlices → Pod readiness → target port → process listener → policy and dependencies. A mismatched Service selector or wrong target port can look like an application outage.
- For node-level issues, check node readiness, pressure conditions, resource availability and the kubelet/runtime/networking layer. A central application log alone cannot explain every infrastructure failure.
- kubectl exec runs a command in an existing container; kubectl debug can introduce a diagnostic environment when supported and authorised. Minimal images may lack shells or debugging tools. Debug actions require appropriate permissions and can change or expose a workload.
- Change one plausible cause at a time, observe the result and verify user-facing recovery. Repeatedly deleting Pods can erase useful evidence while a controller reproduces the same broken specification.
Read-only command recognition, using example names:
kubectl get pods -n demo -o wide
kubectl describe pod web -n demo
kubectl logs web -n demo -c app --previous
kubectl get events -n demo --sort-by=.metadata.creationTimestamp
kubectl get endpointslices -n demo -l kubernetes.io/service-name=web
Choose under exam pressure
| Symptom | First useful evidence |
|---|---|
| Pod is Pending | Events, resource requests, node eligibility and claims |
| Container cannot download its image | Image reference, registry connectivity and pull credentials |
| Container restarts repeatedly | Previous logs, exit reason, configuration and probe behaviour |
| Service exists but has no ready backends | Selector, EndpointSlices and readiness conditions |
| Slow startup causes repeated restarts | Startup probe and probe timing, then the actual startup dependency |
Traps
- A readiness failure is not the same action as a liveness failure.
- Exit status and events are evidence; status labels alone do not establish a root cause.
- The relevant error may be in a previous container instance or another container in the Pod.
- A liveness check that fails whenever an external dependency is temporarily unavailable can create unnecessary restart storms.
08 · Application delivery, CI/CD and GitOps
Memory hook: CI proves the change; delivery prepares it; deployment releases it; GitOps keeps reality aligned.
Must remember
- Continuous integration, CI, merges and validates changes frequently through automated builds and tests. Continuous delivery keeps a validated release ready, often with a human approval before production. Continuous deployment automatically releases qualifying changes to production.
- A typical path is commit → build → test/scan → immutable artifact → deploy → verify. Build once and promote the reviewed artifact when possible. Rebuilding separately for every environment risks deploying content different from what was tested.
- GitOps uses declarative desired state kept in version control, with agents pulling that state and continuously reconciling the live system. Git history supports review and traceability; the reconciler detects and corrects drift.
- Argo CD and Flux are recognisable GitOps delivery projects. A CI script that runs
kubectl applyis automation, but it does not by itself provide continuous pull-based reconciliation. - Helm packages applications as charts with templates and values, and tracks installed releases. Kustomize composes and patches Kubernetes manifests using bases and overlays without requiring a general template language. Either can supply desired manifests to a delivery process.
| Strategy | What changes | Main trade-off |
|---|---|---|
| Recreate | Stop the old version, then start the new | Simple, but may cause downtime |
| Rolling update | Replace instances gradually | Old and new versions may coexist |
| Blue-green | Prepare a second environment, then switch traffic | Fast switch or reversal, but extra capacity |
| Canary | Send a limited share of users or traffic to the new version | Lower initial exposure, but needs measurement and traffic control |
- Kubernetes Deployments support rolling updates and recreates. A meaningful canary requires a deliberate way to control exposure and judge results; simply creating a new Pod is not a complete release strategy.
maxSurgecontrols temporary extra replicas during a rolling update;maxUnavailablecontrols permitted unavailability. Readiness helps prevent premature traffic to new Pods. Resource headroom matters when old and new replicas overlap.- Rollback restores a prior application configuration or artifact. It does not automatically reverse database schema migrations, external messages or user-visible data changes. Design compatibility and recovery before releasing.
- Secure the supply chain with reviewed dependencies, trusted registries, image scanning, provenance/signature checks and least-privilege pipeline credentials. SBOMs describe software components; they are not proof that a release has no vulnerabilities.
- Deployment success needs observation: healthy Pods are necessary but do not alone prove acceptable latency, error rate or business behaviour.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Automatically detect and repair changes made outside reviewed manifests | A GitOps reconciler |
| Install a configurable, reusable application package | Helm chart and release |
| Maintain environment-specific manifest changes over a shared base | Kustomize overlays |
| Expose only a small share of production traffic initially | Canary with explicit traffic control and success criteria |
| Switch back quickly while keeping an old environment available | Blue-green, with a compatible data strategy |
Traps
- Continuous delivery does not necessarily mean automatic production deployment.
- Git is not a secret manager; plain credentials in repository history remain exposed even after a later deletion.
- An automatic rollback cannot safely undo every stateful business operation.
- An image scan and a successful deployment are different checks; neither replaces runtime monitoring.
09 · Observability, metrics, logs and traces
Memory hook: Metrics show the trend, logs explain events, traces follow the journey.
Must remember
- Observability is the ability to understand internal system behaviour using emitted evidence. Monitoring detects conditions you have chosen to watch; useful telemetry also helps investigate unexpected problems.
- Metrics are numeric measurements over time, suited to rates, resource use and service health. Logs record events with context. Distributed traces connect spans across services to show where a request spent time or failed. Correlation identifiers help connect the signals.
- Prometheus collects and queries time-series metrics. A common model is to scrape HTTP metrics endpoints using service discovery. An exporter exposes metrics for a component that does not provide them directly. PromQL is the query language.
- Alertmanager groups, deduplicates and routes alerts, and supports silencing/inhibition. A dashboard displays evidence; it is not a substitute for an actionable alert with ownership and a response plan.
| Metric concept | Example and exam distinction |
|---|---|
| Counter | Total completed requests; normally rises, with resets possible |
| Gauge | Current queue depth or memory usage; can rise and fall |
| Histogram | Distribution of observations such as request duration |
| Labels | Dimensions such as service or status class; every combination can create another time series |
- Use a counter's rate over an interval for requests per second, rather than treating its cumulative total as a current rate. Avoid unbounded labels such as user IDs or raw request IDs; high cardinality increases storage and processing work.
- OpenTelemetry supplies vendor-neutral instrumentation, telemetry APIs/SDKs and collection/export components. Its Collector can receive, process and export signals. It is not itself the complete long-term storage and dashboard backend.
- Jaeger is associated with distributed tracing. Fluentd and Fluent Bit collect and forward logs or telemetry. Recognise the tool's job before matching it to a scenario.
- Latency, traffic, errors and saturation are useful service health signals. Latency percentiles reveal slow-tail behaviour that an average can hide. Measuring only CPU can miss a user-visible failure caused by another service.
- An SLI is a measured indicator; an SLO is the target for it; an SLA is a commitment with contractual implications. An error budget expresses tolerated unreliability against an SLO over its chosen window. A 99.9% request-success SLO permits 0.1% unsuccessful eligible requests under that definition.
- Kubernetes resource metrics, such as those often exposed through Metrics Server for
kubectl topand autoscaling, do not constitute a complete historical observability platform. Cost visibility also needs resource attribution, requests versus usage, retention and billing awareness.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Alert when a service's error rate remains high | Metrics and a meaningful alert rule |
| Find which backend made one distributed request slow | A trace with spans and propagated context |
| Read the exception and application context around a failure | Relevant logs, correlated with the request |
| Standardise instrumentation while retaining backend choice | OpenTelemetry |
| Reduce duplicate incident notifications | Alert grouping and deduplication through Alertmanager |
Traps
- Collecting more logs is not automatically better; noisy or sensitive logs create cost and security problems.
- Average latency can hide a poor experience for the slowest requests.
- A trace is not necessarily a record of every request; sampling affects what is retained.
- Prometheus metrics are not a substitute for a complete, exact per-request financial ledger.
10 · Cloud native principles, ecosystem and community
Memory hook: Design for change and failure; choose a tool by its job; contribute through an open process.
Must remember
- Cloud native describes ways to build and operate applications that adapt to dynamic environments through automation, resilience and manageable change. It is not limited to one public cloud: cloud native systems can run on premises, at the edge or across cloud providers.
- Microservices split an application into independently deployable services around clear responsibilities. They can improve independent scaling and ownership but add distributed communication, data-consistency and operational complexity. A monolith can still use containers and modern delivery practices.
- Loose coupling limits how much components must know about each other. APIs, events and queues can reduce direct dependencies, but contracts, retries and failure handling remain necessary. Repeated retries need backoff and safe/idempotent behaviour to avoid amplifying failures.
- Immutable infrastructure favours replacing a versioned component over manually patching it in place. Declarative configuration records intended state. Automation and reconciliation make repeatable changes possible; neither prevents a bad specification from being applied consistently.
- Resilience plans for partial failure using replicas, health checks, distribution, timeouts, recovery and tested dependencies. Elasticity adjusts capacity with demand. Scalability is the ability to handle growth; it does not require that every scaling decision be automatic.
- Serverless moves more provisioning and scaling responsibility to a platform and often uses event-triggered execution. Servers still exist; users retain responsibility for code, permissions, data and service limits. Knative is an ecosystem example for serverless workloads on Kubernetes.
| Need | Recognisable ecosystem examples |
|---|---|
| Container orchestration | Kubernetes |
| Node container runtime | containerd or CRI-O |
| Metrics and alerting | Prometheus |
| Telemetry instrumentation and transport | OpenTelemetry |
| Log collection and forwarding | Fluentd or Fluent Bit |
| Service traffic proxy or mesh | Envoy; Linkerd or Istio for mesh capabilities |
| Desired-state application delivery | Argo CD or Flux |
| Image registry | Harbor |
| Cluster state key-value store | etcd |
| Policy decisions | Open Policy Agent |
- CNCF is a vendor-neutral home for open source cloud native projects within the Linux Foundation. Its landscape groups projects by purpose; inclusion is not a statement that every project solves every problem or has identical maturity.
- CNCF maturity levels include Sandbox, Incubating and Graduated. Graduation reflects evidence of project maturity, governance and adoption under CNCF criteria; it does not certify that a particular deployment is secure or guaranteed to fit your requirements. Archived status is different from active graduation.
- Open source communities collaborate through public repositories, issues, pull requests, design proposals, reviews and meetings. SIGs, working groups, maintainers and contributors organise work. Read a project's contribution guide, governance and code of conduct before submitting changes.
- Useful contributions include documentation, translations, reproducible bug reports, testing, accessibility, triage and code. Follow the project's licence and contribution requirements, such as DCO sign-off or a CLA when requested. Open source does not mean “no licence conditions.”
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Select among many projects for one platform capability | Identify the category, integration needs, maturity and operational trade-offs |
| Reduce unrecorded manual production changes | Versioned configuration, reviewed delivery and reconciliation |
| Begin contributing without being an experienced programmer | Documentation fixes, testing, issue triage or reproducible reports |
| Decouple request handling from background work | Appropriate events or queues with explicit failure handling |
| Assess a project's decision-making and contribution process | Its governance, maintainer rules and contribution documentation |
Traps
- Running an unchanged application in a public cloud does not automatically give it resilient cloud native behaviour.
- Microservices, Kubernetes and a service mesh are architectural choices, not compulsory ingredients for every application.
- A Graduated project can still be misconfigured, compromised or unsuitable for a particular requirement.
- Open source permits use under its licence; it does not eliminate ownership, attribution or redistribution obligations.