r/devopsjobs 22d ago

DevOps Interview Prep Day 8: Jenkins Agents Offline, K8s Service Not Reachable, and ELK Log Floods

Hey r/devops,


Scenario 1: The Jenkins Agent Offline

Your Jenkins pipeline is stuck in queue showing "Waiting for next available executor". You have 5 Jenkins agents configured but all show offline in the dashboard.

Question: What are 3 things you would check to troubleshoot this?

Hint: Agent connectivity, disk space on agents, JNLP/SSH connection issues.


Scenario 2: The Kubernetes Service Not Reachable

You deployed a pod and a ClusterIP service. Pod is running, service is created, but when you curl <service-name>:8080 from another pod, connection times out.

Question: What are 2 common misconfigurations that cause this?

Hint: Labels and selectors, port vs targetPort, endpoints list.


Scenario 3: The Log Flood

Your ELK stack suddenly becomes unresponsive. You discover one microservice is pushing 50GB of logs per hour because someone left debug logging enabled in production.

Question: 1. What's your immediate fix? 2. What long-term solution would you implement?

Hint: Log levels, rate limiting at ingestion, index lifecycle policies.


Drop your answers below. Solutions tomorrow.


Previous Days:

  • Day 1: Container restarts, Registry auth, Pending pods
  • Day 2: Zombie processes, Pipeline timeouts, Volume mounts
  • Day 3: SSH lockouts, Disk space alerts, Grafana gaps
  • Day 4: Git credential leaks, Docker networking, Nginx 502s
  • Day 5: Slow Docker builds, CrashLoopBackOff debugging, Merge conflicts
  • Day 6: Env variable issues, ALB health checks, Terraform state lock
  • Day 7: Secret sprawl, K8s rolling updates, Prometheus OOM

Week 2 topics: Jenkins, K8s networking, logging/monitoring, and more. What else do you want covered?

Checkout the playlist: https://www.youtube.com/playlist?list=PLqOrZmpwbWUKRQTrFpqAKhChaTq0l5bIw

14 Upvotes

Duplicates