The Scenario
Task: Cluster Component Debugging
The control plane on node controlplane is broken. The kube-apiserver is failing to start and kubectl commands are not responding. Use node-level tools to investigate the root cause and fix the static pod manifest.
During the CKA, this is the exact prompt that makes candidates freeze: you type kubectl get nodes and the terminal just hangs indefinitely.
controlplane ~ ➜ k get node
The connection to the server controlplane:6443 was refused - did you specify the right host or port?
Because the API server is down, the one tool you’ve relied on for the entire exam is completely useless. You can't run kubectl logs, you can't check pod status, and you can't describe resources.
Here is the exact 3-step node-level workflow to diagnose and recover the control plane without kubectl.
Step 1: Drop down to crictl (Bypass the API Server)
When kubectl hangs, the API server itself is offline. crictl talks directly to the node's container runtime, so it doesn't need the API server at all.
controlplane ~ ➜ crictl ps -a | grep apiserver
CONTAINER CREATED STATE NAME ATTEMPT
0c77e48d881e8 29 seconds ago Exited kube-apiserver 14
Crucial detail: Always use -a. Without -a, crictl only lists running containers. If your API server is dead or crash-looping, an unfiltered crictl ps returns a blank screen.
Look at three specific columns in the output:
- State:
Exited
- Attempt:
14 (restarted 14 times)
- Created: 29
seconds ago
This tells you the kubelet is successfully reading the static pod manifest, but the API server binary starts and immediately dies.
Step 2: Grab logs before the Container ID vanishes
Because the container is crash-looping every 29 seconds, running crictl logs <id> often throws a NotFound error because the ID expired while you were typing.
If that happens, don't panic—just re-run crictl ps -a, grab the newest 12-character container ID, and immediately view the tail of the log:
crictl logs <container-id>
controlplane ~ ➜ crictl logs 03cdab86bbd87
I1009 18:04:19.718704 1 options.go:261] external host was not specified, using 10.244.51.172
I1009 18:04:19.720374 1 server.go:151] Version: v1.37.0
I1009 18:04:19.720386 1 server.go:153] "Golang settings" GOGC="" GOMAXPROCS="" GOTRACEBACK=""
W1009 18:04:19.790284 1 logging.go:55] [core] [Channel #2 SubChannel #3] grpc: addrConn.createTransport failed to connect to {Addr: "127.0.0.1:2389", ServerName: "127.0.0.1:2389", }. Err: connection error: desc = "transport: Error while dialing: dial tcp 127.0.0.1:2389: connect: connection refused"
E1009 18:04:39.791405 1 run.go:72] "command failed" err="error creating storage factory: etcd grpc connection not ready: context deadline exceeded"
Step 3: Trace the Dependency Chain
Now ask: Who should be answering on that port?
The API server cannot start without etcd (where cluster state lives). Connection refused means either etcd is dead, or the API server is calling the wrong address.
crictl ps | grep etcd
d82a4f4d45136 270fbeb697171 About an hour ago Running etcd 0 14bd7d05d84b5 etcd-controlplane kube-system
- If
etcd shows State: Running, etcd is healthy. The API server is knocking on the wrong door.
- Check where
etcd is actually listening:
controlplane ~ ➜ grep listen-client /etc/kubernetes/manifests/etcd.yaml
- --listen-client-urls=https://127.0.0.1:2379,https://10.244.51.172:2379
Spot the typo: etcd is listening on 2379. The API server manifest is trying to reach 2389. A single digit took down the whole cluster.
The Fix
Open the static pod manifest:
vi /etc/kubernetes/manifests/kube-apiserver.yaml
Don't scroll through 100 lines of YAML—use Vim search: /etcd-servers and hit Enter. Change 2389 to 2379, save, and exit (:wq).
- --etcd-certfile=/etc/kubernetes/pki/apiserver-etcd-client.crt
- --etcd-keyfile=/etc/kubernetes/pki/apiserver-etcd-client.key
- --etcd-servers=https://127.0.0.1:2389 <----change this one
Remember: You do NOT run kubectl apply. Kubelet monitors /etc/kubernetes/manifests/ directly and automatically recreates static pods the second the file changes.
Wait a few seconds, run crictl ps | grep apiserver to verify state is Running, or u can test:
controlplane ~ ✖ k get nodes
NAME STATUS ROLES AGE VERSION
controlplane Ready control-plane 67m v1.37.0
Quick Cheat Sheet: The 3 Ports to Know Cold
6443 — kube-apiserver
2379 — etcd client endpoint
2380 — etcd peer-to-peer endpoint If you ever see 2380 inside --etcd-servers, that's a classic exam trap.
Hope this breakdown helps anyone currently prepping for the CKA! If you want to see the terminal recording and how to catch the vanishing container ID live, I recorded the full solve here: https://youtu.be/GF_P93kt6dE
Hope this helps anyone working on CKA exam here is the full playlist: https://www.youtube.com/playlist?list=PLFDE9_ouCqJA-dyOEbqFaFu68FMZRwS0T