Reviewed 10 October 2026 against the linked published scope. Performance exam. Curriculum weights: cluster administration 25%, workloads 15%, networking 20%, storage 10%, troubleshooting 30%.
Memory hook: Correct context, declarative fix, then verify the workload from the client.
Read the essentials, cover the answers and explain the decision aloud. Open the topic summaries below whenever a distinction is unclear.
Cluster administration and workloads
- Before a task, confirm context, namespace and target node.
kubectl config current-context,kubectl get nodesandkubectl get pods -Aestablish scope. Do not solve a correct problem in the wrong cluster. - API server exposes the control plane; etcd stores cluster data; scheduler assigns unscheduled Pods; controllers reconcile desired state. Kubelet runs workloads through the container runtime and reports node status.
- A kubeadm cluster requires suitable runtime/network preparation and a CNI. Upgrade with the supported version-skew and control-plane/node order; drain eligible workloads and verify health. An etcd snapshot is a control-plane recovery asset, not a backup of every application volume.
- Deployment manages replaceable Pods through ReplicaSets; StatefulSet supplies stable identities and storage associations; DaemonSet targets nodes; Job completes work; CronJob creates scheduled Jobs. A naked Pod has no Deployment controller to replace it.
- Resource requests influence scheduling; limits constrain runtime use. CPU excess can be throttled; memory excess can terminate a container. A Pod remains Pending when constraints or capacity cannot be satisfied.
- Node selectors/affinity attract eligible nodes; taints repel unless tolerated. A toleration permits placement but does not force it. Topology spread and anti-affinity reduce concentration when capacity permits.
- Helm templates/packages releases; Kustomize overlays configuration. Inspect rendered manifests and resource scope before applying changes. CRDs extend API types; controllers/operators implement associated reconciliation.
Networking and storage
- Pods receive network identities through the cluster network. A ClusterIP Service provides stable access to selected ready endpoints; NodePort exposes a node port; LoadBalancer relies on implementation integration. A headless Service supports direct endpoint discovery.
- Service
port,targetPort, selector and Pod readiness must agree. Inspect Service, EndpointSlices, Pod labels and DNS separately. Ingress or Gateway API objects require a compatible controller; writing the object alone does not install one. - NetworkPolicies depend on CNI enforcement. Once isolated for a direction, a Pod needs matching allow rules. Both source egress and destination ingress must permit a connection when both are restricted. DNS is easily blocked accidentally.
- PV is supplied storage, PVC a claim and StorageClass a provisioning policy. Binding depends on class, capacity, modes and topology.
WaitForFirstConsumerdelays placement-sensitive provisioning/binding. - Access mode describes supported mount access, not filesystem permissions. RWO is one node, not necessarily one Pod; RWOP is one Pod where supported. Reclaim policy Retain preserves storage for manual handling; Delete may remove the backing resource.
Troubleshooting and recovery
- Start with
get,describe, events and logs.kubectl logs --previouscan reveal a crashed container;execinspects a running one. Node/runtime logs and kubelet status matter when Pod logs never appear. - Pending: inspect scheduler events, requests, affinity/taints and PVCs. ImagePullBackOff: inspect image reference, registry connectivity and pull credentials. CrashLoopBackOff: inspect process exit, configuration and prior logs.
- Readiness removes an unready Pod from ordinary Service endpoints; liveness can restart it; startup probes postpone the other probes during initialization. A bad liveness probe can create an outage.
- Cordon prevents ordinary new scheduling; drain evicts eligible workloads subject to restrictions. A PDB limits voluntary disruptions and is not a guarantee against node failure.
- Verify after every change: rollout completion, available replicas, endpoints, client request and persistence. Save intended manifests and inspect diffs. A command returning zero is not sufficient evidence of service recovery.
Practical traps
- A toleration does not reserve a node, a Service does not start a process, and a PVC does not guarantee it can bind.
- RBAC grants combine; another RoleBinding cannot subtract permission. Check the API group, subresource and workload identity, not only an administrator's successful request.
- Follow an HTTPRoute from parent/listener to backend Service port, then endpoints. A created route is not proof it attached or forwards traffic.
- Context, namespace, filenames and persistent end state are part of correctness. Practice against the runtime version in the exam instructions.
Kubestronaut exam preparation
CKA supplies cluster administration practice for the five-exam Kubestronaut path. Rehearse using the current permitted resources, including the exam's own Quick Reference. CertSloth is revision material for preparation; it is not an allowed in-exam reference.
Final active recall
1. Pending Pod with unbound PVC: restart kubelet first?
No. Inspect PVC events, StorageClass, capacity and topology before changing the node.
2. A readiness probe fails. Does Kubernetes restart the container because of that probe?
No. Readiness affects traffic eligibility; liveness/startup failure policies can restart it.
3. What permits a tainted node but does not attract a Pod?
A matching toleration.
4. Service has no endpoints. What do you check?
Selector-to-label match, Pod readiness and EndpointSlice state.
5. Does a PDB protect against every outage?
No. It constrains supported voluntary disruptions, not arbitrary hardware failure.
Sources and further practice
- Official exam scope
- Objective-to-topic coverage map. Each full topic links to its supporting primary technical documentation.
- Kubernetes concepts
- Current candidate FAQ and environment versions
- Kubestronaut requirements
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Cluster Lifecycle, RBAC and Extensions
Memory hook: API decides; controllers reconcile; nodes execute.
Must remember
The API server validates requests and exposes cluster state; etcd persists it. The scheduler assigns unscheduled Pods to nodes; controllers reconcile desired state. Node kubelets supervise Pods through a CRI runtime; CNI provides networking and CSI storage integration.
Before kubeadm installation, verify hostnames, addresses, time, ports, container runtime/cgroups and the version-specific swap requirements. kubeadm init creates a control plane; join information adds nodes. A highly available control plane needs multiple control-plane instances, a stable API endpoint/load balancer and a quorum-safe etcd design. A load balancer alone does not replicate etcd.
Upgrade deliberately: inspect version compatibility and kubeadm's upgrade plan, upgrade control-plane nodes in the documented order, then workers. Drain a node before disruptive maintenance, accounting for PodDisruptionBudgets and local data, and uncordon afterward. Do not force a blocked drain without understanding which workload guarantee it protects. Back up etcd and practise version-appropriate recovery in a sandbox.
RBAC grants verbs on API resources. A Role is namespaced; a ClusterRole can describe broader permissions. A RoleBinding can reference a ClusterRole while granting its namespaced permissions only in the binding's namespace. A ClusterRoleBinding grants cluster-wide scope. Bind users, groups or service accounts; inspect with kubectl auth can-i using the required identity/context.
RBAC permissions are additive: there is no ordinary RBAC deny rule to cancel another grant. Core resources use apiGroups: [""]; Deployments use apps; Pod logs are the separate pods/log subresource. Test the precise verb, resource and namespace, for example kubectl auth can-i get pods --as=system:serviceaccount:team:reader -n team. Impersonation itself requires permission; testing as the administrator does not prove the workload identity can perform the operation.
Helm templates and packages charts into releases with values; inspect rendered manifests, release history and upgrades. Kustomize transforms base YAML with overlays and patches without template substitution; inspect kubectl kustomize PATH. A CRD adds an API kind; an operator includes a controller that acts on custom resources. Installing a CRD alone does not implement reconciliation.
Practical drill: explain the path from an authenticated Deployment request to running containers, naming which component handles each step.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Grant one namespace access | A RoleBinding with only the required verbs/resources. |
| Configure a packaged application | Helm values and a controlled release. |
| Patch environment-specific YAML | Kustomize overlays. |
| Automate custom-resource behavior | Install the matching operator, not just its CRD. |
Traps
- Namespaces do not by themselves enforce network isolation.
- Deleting or editing etcd data is not an ordinary application troubleshooting step.
02 · Workloads, Scheduling and Self-Healing
Memory hook: Requests place; limits constrain; probes decide.
Must remember
Deployments manage ReplicaSets and rolling changes. StatefulSets supply stable ordinal identity and per-Pod storage patterns; DaemonSets place node-local workloads; Jobs run to completion and CronJobs schedule Jobs. Pick the controller that expresses the workload instead of creating unmanaged Pods.
Use kubectl rollout status, history and undo for Deployment progress and recovery. maxSurge permits extra Pods; maxUnavailable controls how many desired replicas can be unavailable during a rolling update. A rollback restores an earlier Pod template, not external database state.
Readiness determines eligibility for normal Service traffic; liveness can restart a failing container; startup delays the other probes while an application initializes. An overly aggressive liveness probe can turn a dependency outage into a restart storm. Resource requests influence scheduling; limits constrain use, with CPU throttling and possible memory termination.
Node selectors and required affinity constrain placement; preferred affinity expresses preference. Taints repel Pods without matching tolerations; a toleration permits placement but does not select a particular node. Topology spread and anti-affinity reduce correlated failures. Admission controls, namespace quotas and LimitRanges can reject or default Pod configuration before scheduling.
HPA scales replica counts from configured metrics; utilisation-based CPU scaling requires appropriate requests and metrics availability. Node autoscaling supplies node capacity rather than changing desired application replicas. A PodDisruptionBudget limits voluntary disruption; it cannot prevent every involuntary failure.
ConfigMaps store nonsecret settings; Secrets hold sensitive values but base64 encoding is not encryption. Protect API access and at-rest storage. Environment variables are captured at container start; mounted projected data can update with qualifications, including subPath limitations. Restart or reload the application when its consumption method requires it.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Slow startup but healthy later | Use a startup probe and suitable thresholds. |
| Spread replicas across failure domains | Topology constraints/anti-affinity plus sufficient capacity. |
| Scale on measured load | HPA with correct requests and metrics. |
Traps
- A toleration is permission, not a placement guarantee.
- A readiness failure does not itself restart the container.
03 · Services, DNS and Network Policy
Memory hook: Selector finds endpoints; policy permits the path.
Must remember
Pods communicate using cluster networking supplied by the CNI. A Service gives stable discovery for a changing set of endpoints. Check labels/selectors, target ports and EndpointSlices when traffic reaches the Service but no application responds.
| Type | Use |
|---|---|
| ClusterIP | Internal stable virtual service address. |
| NodePort | A port exposed on nodes, with the Service forwarding to backends. |
| LoadBalancer | Requests an external load balancer from an available implementation. |
Headless (clusterIP: None) |
DNS discovery of endpoints rather than a normal virtual IP. |
Ingress describes HTTP/S routing, but needs an Ingress controller. TLS certificates normally come from referenced Secrets; matching host/path and backend port matters. Gateway API separates infrastructure ownership (GatewayClass/controller), listeners (Gateway) and application routing (HTTPRoute). Inspect parent references, allowed routes, listener hostnames and cross-namespace ReferenceGrants where needed. Gateway resources also require an implementation.
CoreDNS resolves service names such as api.team.svc.cluster.local, with the actual cluster domain configurable. Test name resolution separately from network reachability. Check namespace search suffixes, CoreDNS Pods/Service, resolv.conf and DNS egress policy before blaming application code.
Gateway routing drill: inspect kubectl get gateway,httproute -n team -o yaml, then trace parentRefs to the listener and backendRefs to the Service. An HTTPRoute backend port is the Service port, not an arbitrary container port. Check host/path matching and the controller's status conditions before testing a request with the intended Host header. A route that exists but never attaches to its Gateway cannot deliver the requested traffic.
NetworkPolicy is enforced only by supporting networking implementations. Policies are additive allow lists. Once a Pod is isolated for ingress or egress, applicable allowed traffic must be specified; when both ends are isolated, both source egress and destination ingress must permit the connection. An empty selector selects all Pods in that namespace. Namespace and Pod selectors in the same peer entry are AND; separate entries are alternatives.
Practical drill: draw client → DNS → Service → endpoint → container. For each arrow, identify one read-only command and one possible configuration fault.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| HTTP host/path routing | Ingress or Gateway API with a working controller. |
| Service has no backends | Inspect selectors, Pod readiness and EndpointSlices. |
| Restrict Pod communication | NetworkPolicies plus a capable CNI implementation. |
Traps
- Creating a LoadBalancer Service does not guarantee a load-balancer implementation exists.
- Default-deny egress can also block DNS.
04 · Volumes, Claims and Storage Classes
Memory hook: Claim requests; class provisions; policy reclaims.
Must remember
A Pod's container filesystem is not durable application storage. emptyDir lives with the Pod and survives container restarts, but disappears when the Pod is removed. hostPath exposes a node path and ties behavior to that node; it is not a general portable persistent-storage design.
A PersistentVolume represents storage; a PersistentVolumeClaim requests capacity and access characteristics. A StorageClass identifies a provisioner and parameters for dynamic provisioning. With WaitForFirstConsumer, binding can wait for workload placement so topology and storage availability align. An immediate binding may otherwise choose storage in the wrong zone for a constrained Pod.
Access modes describe permitted mounting patterns: ReadWriteOnce is read-write by one node (potentially several Pods there), ReadOnlyMany is many-node read access, ReadWriteMany is multi-node read/write when the storage supports it, and ReadWriteOncePod restricts access to one Pod with supported CSI storage. Modes do not transform an underlying disk into a distributed filesystem.
Reclaim policy decides what happens after the claim releases a PV. Delete generally removes dynamically provisioned backing storage through the provisioner; Retain preserves it for manual recovery/reuse. Deleting a workload is not the same as deleting its PVC. Inspect the actual PV/class rather than assuming retention.
Expansion requires class and driver support; filesystem expansion behavior depends on the volume and driver. A larger claim can request growth, but Kubernetes does not provide ordinary volume shrinking. volumeMode: Block exposes a raw block device rather than a mounted filesystem.
Diagnose Pending claims through events, matching class, size, access modes, provisioner health and topology. Diagnose mount failures through CSI/controller/node logs, attachment limits, permissions and application paths. Never delete a data-bearing claim merely to make a warning disappear.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Provision storage on claim creation | StorageClass with a working dynamic provisioner. |
| Match disk zone to Pod placement | WaitForFirstConsumer where appropriate. |
| Keep data after claim deletion | A deliberate Retain reclaim policy and recovery procedure. |
Traps
- ReadWriteOnce means one node, not universally one Pod.
- A Retain volume still needs a tested backup and recovery plan.
05 · Troubleshooting Under Time Pressure
Memory hook: Context first; events next; change one cause.
Must remember
Start with kubectl config current-context and the requested namespace. Use get, describe and sorted events to identify the failing layer before editing. Verify the target context again when switching between tasks.
| Symptom | First useful evidence |
|---|---|
| Pod Pending | Scheduling events: requests, affinity, taints, quota and PVC binding. |
| ImagePullBackOff | Image reference, registry reachability and image-pull credentials. |
| CrashLoopBackOff | kubectl logs POD --previous, exit code, command and probes. |
| Running but unavailable | Readiness, listener/targetPort, selectors and EndpointSlices. |
| Node NotReady | Node conditions, kubelet/runtime health, disk pressure and network plugin. |
| API unavailable | API endpoint, control-plane static Pods, certificates and etcd health. |
kubectl logs -c CONTAINER selects the right container; --previous reads a prior terminated instance. kubectl exec is useful when tools exist in the container; minimal images may lack a shell. Use an approved debug container or diagnostic Pod when appropriate, without assuming it has identical policy/identity to the failing workload.
On a node, systemctl status kubelet, journalctl -u kubelet and runtime tools such as crictl help when the API path is broken. Static control-plane Pod manifests are managed by the kubelet; keep backups outside its watched manifest directory to avoid accidental extra Pods.
kubectl top requires a metrics API and shows recent usage; it is not a full historical monitoring system. Compare usage with requests, limits and node allocatable capacity. Check DNS, Service endpoints, policies and routing in sequence for network failures.
Timed drill: diagnose a deliberately broken selector, an oversized request and a wrong image in a disposable cluster. State the observed cause, make the smallest correction, then verify the requested application behavior. A green Pod phase alone is not proof of success.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Application recently restarted | Previous container logs and termination status. |
| Cluster API unavailable | Node-level kubelet/runtime and control-plane diagnostics. |
| Service DNS resolves but request fails | Check endpoints and application ports, then traffic policy. |
Traps
- Blind restarts can erase evidence and fail to fix the cause.
- Editing the wrong context can complete the wrong task perfectly.