1 What DevOps actually does
The textbook definition (“a culture that unites development with operations”) doesn't help you in an interview, because it doesn't say what you do on Monday morning. Here's a phrasing that helps:
“DevOps owns the road between the code developers write and the code running for the customer. It builds that road to be automated and repeatable, builds the environments it runs in, and makes sure that when something breaks, you see it fast and fix it fast.”
The role has four pillars. All four show up in job postings, in different orders:
🚚 Automated delivery (CI/CD)
You build the pipeline that takes the code from Git, compiles it, tests it, packages it, and puts it on servers. Tools: GitHub Actions, Jenkins, GitLab CI, Azure DevOps.
🏗️ Infrastructure as the runtime environment
You create and maintain the places where the code runs: virtual machines, Kubernetes clusters, databases, networks. Described in code, not through a UI. Tools: Terraform, Kubernetes, Docker, Ansible.
👁️ Observability
You make the system tell you its own state: metrics, logs, alerts, dashboards. Without that, the first one to find out there's a problem is the customer. Tools: Prometheus, Grafana, Datadog, ELK.
🚨 Reliability and incidents
You're on call, you take the incident, communicate it, close it, then write up what happened and what changes so it doesn't repeat. This is the part you've been doing for 20 years.
Pillar 4 you have solid: 24/7 production control at Metro, zero-downtime messaging migration at Volvo Cars, application discovery at Etihad. Pillar 3 you have partly — you've worked with Nagios and Tivoli, which are the same idea as Prometheus, but with the previous generation of tools. Pillars 1 and 2 are the ones you've touched at DISH and RapidBid, and the ones you need to be able to explain best, because that's where the interview will push.
What DevOps is NOT
- It's not “the person who has access to production.” Access is a consequence, not the role. A good DevOps engineer reduces the number of people who have to go into servers manually, including themselves.
- It's not a sysadmin with a new name. A sysadmin administers existing servers. DevOps creates them from code, destroys them, recreates them identically. If a server can't be recreated from scratch automatically, it isn't built right.
- It doesn't write the product's functionality. It writes code, but automation code: pipelines, manifests, infrastructure modules, scripts.
- It's not a separate team that receives tickets. When it becomes that, you've reinvented exactly the wall DevOps was supposed to tear down.
Why the role appeared
Historically, there were two teams with opposite goals. Development was paid to change things. Operations was paid to keep nothing changing, because any change could bring production down. The result: rare, large, risky releases, every two or three months, at night, with everyone in a meeting.
The DevOps answer is counterintuitive: you ship more often, not less. A small, automated release of twenty lines is easy to verify and easy to roll back. A release with three months of work behind it can be neither verified nor rolled back. The risk doesn't come from frequency, it comes from size.
“The paradox is that stability goes up when you ship more often, not less. A small, automated, reversible change is safer than a quarterly release where three months of changes have piled up.”
2 A DevOps engineer's day
This is what shows up consistently in role descriptions and in accounts from people in the field. It's not a typical day — it's a mix that changes daily, and getting interrupted is part of the job.
| Time slot | What happens | Why it matters |
|---|---|---|
| Morning | You check the dashboards and the overnight alerts. You look at whether any automated release failed, whether any error rate went up, whether any disk filled up. | This is the moment you catch problems before the customer does. |
| Stand-up | 15 minutes with the development team. What's blocking whom. What's shipping today. | DevOps sits inside the product team, not in a separate tower. |
| Build block | Planned work: a new pipeline, a Terraform module, a migration, a cluster upgrade, automating a manual operation. | This is where you produce long-term value. It's also the part most easily eaten up by interruptions. |
| Developer support | “My build won't pass,” “I can't connect to the staging database,” “why did the deploy fail.” Troubleshooting with someone else. | Half your reputation on the team is built here. |
| Releases | You watch releases go out to staging and production. If something doesn't look right, you stop it and roll back. | On a mature team, this is pressing a button and watching graphs. |
| Any time | Incident. Everything planned stops. | See chapter 18. |
| Periodic | Cloud cost analysis, applying security updates, permission reviews, a backup restore drill. | Work nobody asks you for, that saves you once a year. |
Don't recite the list. Pick three things and connect them: “In the morning I check what happened overnight, then I have a block of planned work — usually automation or infrastructure — and the rest of the day is split between developer requests and releases. When an incident comes up, everything planned stops.”
3 The big picture: from commit to production
If you keep one single image from this page, let it be this one. Every chapter that follows explains one box from the diagram below.
If they ask something broad (“walk me through a pipeline”), describe the flow in this order, out loud, box by box. You've got nine ready-made sentences and you can't really get stuck in the middle, because each box hands you the next one.
4 Git and the workflow
Git is where the truth lives. Everything that follows — pipeline, image, release — starts from a commit. If it's not in Git, it doesn't exist.
commit
A snapshot of the code at a given moment, with a unique identifier (a hash). This is the grain of sand everything else is built from.
branch
A parallel line of work. You branch off main, work in isolation, come back. Nobody writes directly into main.
Pull Request
The request to bring your branch into main. This is where code review happens and where the automated checks run.
merge
The actual merging. Usually blocked until tests pass and at least one colleague has approved.
tag
A label on a commit, usually a version number: v1.4.2. Releases to production are made from tags, not branches.
revert
A new commit that undoes the effect of an old one. It doesn't erase history — that matters, because history is the evidence.
Two branching models
| Trunk-based (recommended today) | GitFlow (older) | |
|---|---|---|
| What it looks like | Short branches, 1–2 days, merged into main quickly. main is always releasable. | Long branches: develop, release/*, hotfix/*, feature/*. |
| Fits when | You ship often, you have solid automated tests. | You ship rarely, in versions, you need to maintain several versions in parallel. |
| The problem | Demands testing discipline; without tests, you break main. | Long branches drift apart and merging becomes painful (“merge hell”). |
“Trunk-based, with branches of at most a day or two and a protected
main: mandatory automated checks and at least one approval. The reason is that long branches drift from the trunk and the problem shows up all at once, at merge time. GitFlow makes sense if you need to support several versions in parallel for customers.”
GitOps — the idea that ties Git to Kubernetes
In GitOps, the cluster's desired state lives in a Git repository, and an agent running in the cluster (Argo CD or Flux) continuously compares what's in Git with what's in the cluster and corrects the difference. Practical consequences:
- Nobody runs
kubectl applymanually in production anymore. You change the file, open a PR, it gets approved, the agent applies it. - If someone changes something directly on the cluster, the agent reverts it — the drift self-corrects.
- Rolling back to the previous version is a
git revert. The release history is the Git history.
5 CI/CD — the concepts
The three letters that get mixed up
Continuous integration
On every push, the code gets built and tested automatically. The goal: find out in minutes if you broke something, not in weeks.
Continuous delivery
Continuous delivery. Every commit that passes the checks is ready to go to production. Pressing the button stays with a human.
Continuous deployment
Continuous deployment. Same thing, but without the button: whatever passes the tests reaches the customer automatically.
“What's the difference between continuous delivery and continuous deployment?” Short answer: human approval. With delivery, the artifact is ready any time, but someone presses the button. With deployment, no one presses anything. Very few companies actually do true deployment in production, and it's completely normal to say so.
The stages of a pipeline, in order
| # | Stage | What it does | If it fails |
|---|---|---|---|
| 1 | Checkout | Pulls the code from the repository onto the build machine. | An access or network problem. |
| 2 | Lint / format | Checks style and obvious mistakes. Runs in seconds. | Stops immediately. It's the cheapest filter. |
| 3 | Build | Compiles, installs dependencies. | No point testing something that doesn't even build. |
| 4 | Unit tests | Tests small, isolated pieces. Minutes. | Stops. This is where you catch 80% of regressions. |
| 5 | Security scan | Vulnerable dependencies, secrets left in code, images with CVEs. | Usually stops on critical vulnerabilities. |
| 6 | Packaging | Builds the Docker image and pushes it to the registry, tagged with the commit hash. | The artifact is what will ship. It isn't rebuilt for each environment. |
| 7 | Deploy to staging | Puts the artifact in an environment as close as possible to production. | This is where you catch configuration issues. |
| 8 | Integration / e2e tests | Checks the whole system, with real databases. | Slow and flaky. That's why there are few of them, carefully chosen. |
| 9 | Approval | A human presses the button. Or an allowed time window. | The control gate for production. |
| 10 | Deploy to production | Rolling, blue-green or canary. | See below. |
| 11 | Post-deploy verification | Smoke tests plus watching the metrics for a few minutes. | If it doesn't look right, you roll back automatically. |
You build once, you ship everywhere. The same Docker image goes to dev, staging and production; only the configuration differs, injected from outside. If you rebuild the image for each environment, you haven't tested what you're shipping. It's a sentence worth remembering for the interview.
Deployment strategies
At RapidBid you did blue-green with Jenkins. It's the easiest of the three to explain, and you've actually lived it. Have a sentence ready about what the switchover looked like, and one about what you do with the database, because that's always where the next question goes: the schema has to be compatible with both versions at the moment of the switch, otherwise rolling back is no longer possible.
6 GitHub Actions, in detail
It's the automation system built into GitHub. It runs on machines that GitHub provides (a “runner”), triggered by events in the repository. It's become the default choice for most new teams, and it's what most job postings ask for.
6.1 The vocabulary, from big to small
| Term | What it is |
|---|---|
| workflow | A YAML file in .github/workflows/. A repository can have any number of them. |
| event | What triggers it: push, pull_request, schedule (cron), workflow_dispatch (manual button), release. |
| job | A group of steps that run on the same machine. Jobs run in parallel by default. |
| runner | The machine that executes. Either hosted by GitHub (ubuntu-latest), or your own (self-hosted), for access to your internal network or special hardware. |
| step | A command (run:) or a reused action (uses:). |
| action | A packaged piece of automation, yours or someone else's: actions/checkout, docker/build-push-action. |
| artifact | Files saved at the end of a job, so another job can pick them up or you can download them. |
| environment | A named environment (staging, production) that can have its own secrets, protection rules and human approvers. |
6.2 A complete workflow, commented line by line
# The name that shows up in the Actions tab name: Build and deploy # ---- WHEN it runs ---- on: push: branches: [main] # only on main paths: # and only if something relevant changed - 'src/**' - 'Dockerfile' pull_request: # and on every PR, to catch issues early workflow_dispatch: # plus a manual button in the UI # ---- Permissions: the minimum needed, explicit ---- permissions: contents: read # can read the code id-token: write # can request an OIDC token (for the cloud, no passwords) # ---- Only one run per branch; the old one gets canceled ---- concurrency: group: deploy-${{ github.ref }} cancel-in-progress: true env: IMAGE: ghcr.io/${{ github.repository }} jobs: test: runs-on: ubuntu-latest strategy: matrix: # same job, on 3 Node versions, in parallel node: [18, 20, 22] steps: - uses: actions/checkout@v4 - uses: actions/setup-node@v4 with: node-version: ${{ matrix.node }} cache: npm # keeps dependencies between runs = faster build - run: npm ci - run: npm test build: needs: test # only starts if “test” passed runs-on: ubuntu-latest outputs: tag: ${{ steps.meta.outputs.tag }} steps: - uses: actions/checkout@v4 - id: meta run: echo "tag=${GITHUB_SHA::7}" >> $GITHUB_OUTPUT - uses: docker/login-action@v3 with: registry: ghcr.io username: ${{ github.actor }} password: ${{ secrets.GITHUB_TOKEN }} - uses: docker/build-push-action@v6 with: push: true tags: ${{ env.IMAGE }}:${{ steps.meta.outputs.tag }} deploy: needs: build if: github.ref == 'refs/heads/main' # we do NOT deploy from PRs runs-on: ubuntu-latest environment: production # this is where human approval is required steps: - uses: google-github-actions/auth@v2 with: workload_identity_provider: ${{ secrets.WIF_PROVIDER }} service_account: ${{ secrets.GCP_SA }} - run: | kubectl set image deployment/api \ api=${{ env.IMAGE }}:${{ needs.build.outputs.tag }} -n prod kubectl rollout status deployment/api -n prod --timeout=180s
kubectl rollout status --timeout=180s turns the deployment into one that can actually fail. Without it, the set image command returns success instantly — it only requests the change — and the pipeline goes green even if no new pod ever started. It's exactly the kind of detail a good interviewer looks for.
6.3 Secrets and authentication
Three levels, in the order they're looked up: organization (shared across multiple repositories), repository, environment (the most restricted — they only exist in jobs that declare environment:). Environment secrets are the safest for production, because they can be tied to approvers.
What you need to know:
- Secrets are encrypted and hidden in logs, but a step that runs arbitrary code can extract them. That's why it matters which steps you hand them to.
- They're not automatically passed to workflows in other repositories (forks) — that's the protection against a malicious PR.
secrets.GITHUB_TOKENis generated automatically for every run and expires at the end. Set it to “read-only” by default at the organization level, and explicitly request more withpermissions:.
OIDC — how you get rid of keys completely
The old way: you put an AWS_SECRET_ACCESS_KEY in the repository's secrets. The key is permanent; if it leaks, it stays valid until you rotate it manually. The current, recommended way: OpenID Connect. The workflow requests a short-lived token from GitHub that describes who it is (“repository X, branch main”), and the cloud is configured to trust GitHub and hand out a temporary role in exchange for that token. No password sits anywhere anymore.
The trust condition in the cloud has to be written against immutable identifiers. If you write repo:company/payment-service and the repository gets deleted, someone can create a different repository with the exact same name and inherit the access. The safe form uses the numeric identifiers: repository_owner_id:12345:repository_id:67890. There's one more subtlety: for reusable workflows, id-token: write has to be granted in the calling workflow, not the called one.
6.4 The security of external actions
- Pin to a hash, not a tag.
uses: some/action@v3means “whatever is currently in v3” — and the author can move the tag onto new code.uses: some/action@a81bbbf...means exactly that code. You let Dependabot propose the updates. - Avoid
pull_request_targetunless you know exactly what you're doing. Unlikepull_request, it runs with the repository's secrets and with write permissions, in the context of the PR — including one that came from a stranger. - Don't put PR data directly into
run:. A PR title that contains shell characters becomes executed code. Pass the value through an environment variable and use the variable.
6.5 Reuse: two different mechanisms
Reusable workflow
A whole workflow, called from another one with uses: at the job level. It contains complete jobs, it can require approvals, it can use environments. It starts being worth it from three or four repositories doing the same thing.
Composite action
A group of steps packaged as an action, inserted inside an existing job. Good for short, repetitive sequences (“log in and set up the tool”).
6.6 Troubleshooting
- Add the repository secret
ACTIONS_STEP_DEBUG=trueand the logs become much more detailed. - A job that “doesn't start” almost always has a false
if:condition or apaths:filter that didn't match — not a bug. - “Works locally, doesn't work in CI” usually means a dependency installed on your machine but missing on the clean runner, or an environment variable that only exists on your side.
- For a slow run, look at caching first: without
actions/cacheor withoutcache:in the setup action, you reinstall everything every time.
This is the area you need to practice hands-on, not just read about. A personal repository with a workflow that builds an image and pushes it to the GitHub Container Registry gives you the right to say “I've written workflows” without overstating it. The honest phrasing, if they ask: “In my own projects I work with GitHub Actions; my production, at-scale pipeline experience is with Jenkins and Azure DevOps, and the concepts transfer directly — the stages, the single artifact, the approval gate are the same.”
7 Jenkins and Azure DevOps — what you already have
Jenkins is the classic automation system, dating back to 2011, installed on your own servers. It has a controller and agents that execute the work. The major difference from GitHub Actions: Jenkins is an application you administer yourself — with plugins, updates, backups — while Actions is a service.
7.1 Anatomy of a Jenkinsfile
pipeline { agent { label 'linux-docker' } // which agent runs this environment { // variables for the whole pipeline IMAGE = "registry.intern/api" REG = credentials('registry-creds') // secret from Jenkins Credentials } triggers { pollSCM('H/5 * * * *') } // or webhook from Git stages { stage('Build') { steps { sh 'mvn -B clean package' } } stage('Test') { steps { sh 'mvn test' } post { always { junit 'target/surefire-reports/*.xml' } } } stage('Image') { steps { sh "docker build -t $IMAGE:${GIT_COMMIT} ." } } stage('Deploy green') { steps { sh './deploy.sh green' } } stage('Switch traffic') { input { message 'Switch traffic to green?' } // human approval steps { sh './switch.sh green' } } } post { failure { mail to: 'team@company.com', subject: "Pipeline failed: ${env.JOB_NAME}" } always { cleanWs() } } }
7.2 Translating Jenkins → GitHub Actions
If they ask you “have you worked with GitHub Actions?”, the right answer isn't yes or no. It's to show you know the mapping between the two — because that's exactly what someone doing a migration does, and migration is a task that comes up often in job postings.
| Jenkins | GitHub Actions | Note |
|---|---|---|
agent { label 'x' } | runs-on: or container: | The Jenkins agent is a machine you maintain yourself; the GitHub runner is clean on every run. |
stages { stage {...} } | jobs: | Careful: Jenkins stages implicitly share the same workspace. Actions jobs don't. Dependencies have to be redesigned using needs: plus artifacts. |
environment { ... } | env: at job or step level | Nearly identical. |
credentials('id') | ${{ secrets.NAME }} | Actions adds the “environment secrets” layer, which Jenkins doesn't have natively. |
triggers { pollSCM } | on: push / on: schedule | Actions is triggered by events; it doesn't poll the repository. |
input { message } | environment: with required reviewers | Human approval moves out of the pipeline and into the environment configuration. |
sh 'command' | run: command | Identical in practice. |
| plugins (over 1800) | actions from the Marketplace | Both are someone else's code. With Actions you pin it to a hash; with Jenkins you update and hope. |
post { always { } } | if: always() on a step | Same idea, different syntax. |
In Jenkins, the “Test” stage sees the files produced by the “Build” stage, because they run in the same working directory. In GitHub Actions, every job starts on a clean machine. If you translate stage-to-job without thinking it through, the test job finds nothing. The fix: either put everything in a single job with multiple steps, or upload the result with actions/upload-artifact and download it in the next job.
7.3 Azure DevOps, in brief
The Microsoft suite: Repos (Git), Pipelines (CI/CD), Boards (tickets), Artifacts (packages), Test Plans. The pipelines are YAML, similar in shape to Actions: trigger, pool, stages, jobs, steps. What's specific: service connection (the identity the pipeline uses to enter Azure), variable group (shared variables, optionally linked to Key Vault), environment with approvals and checks — the equivalent of a production gate.
At RapidBid you worked with Azure DevOps and Jenkins with blue-green deployments. Don't downplay that just because it isn't GitHub Actions. A pipeline is a pipeline: stages, artifact, gate, rollback. A phrasing that works: “My production pipelines were on Jenkins and Azure DevOps. The concepts are the same, the syntax differs; what's specific to Actions is the isolated-jobs model and OIDC authentication, and those are the things I studied because they're what can trip you up during a migration.”
8 Docker and containers
8.1 What a container actually is
A container is an ordinary process on the host machine, which sees only a slice of the system: its own filesystem, its own network, its own process list. The isolation is done by the Linux kernel (namespaces and cgroups). It doesn't have its own kernel — it uses the host's.
8.2 Image and container
The image is the template: an immutable package with the application and everything it needs. The container is an instance running from that image. One image, any number of containers. The analogy worth remembering: the image is the class, the container is the object.
8.3 Layers and why the order matters
Every instruction in the Dockerfile creates a layer. Layers get cached and reused — but when a layer changes, everything after it gets rebuilt. Hence the practical rule: put what rarely changes first.
FROM node:20 WORKDIR /app COPY . . # any file that changed RUN npm install # invalidates this CMD ["node","server.js"]
FROM node:20 WORKDIR /app COPY package*.json ./ # rarely changes RUN npm ci COPY . . # changes often, but it's last CMD ["node","server.js"]
8.4 Multi-stage builds
You build in a large image, with a compiler and tools, then copy just the result into a small image. The final image doesn't contain the compiler, so it's smaller and has a reduced attack surface.
# --- stage 1: build --- FROM golang:1.23 AS build WORKDIR /src COPY go.mod go.sum ./ RUN go mod download COPY . . RUN CGO_ENABLED=0 go build -o /out/api ./cmd/api # --- stage 2: run --- FROM gcr.io/distroless/static:nonroot COPY --from=build /out/api /api USER nonroot:nonroot # do NOT run as root EXPOSE 8080 ENTRYPOINT ["/api"]
8.5 Security rules for images
- Don't run as root. By default, the process inside the container is root. If someone breaks out of the container, they break out as root.
- Read-only filesystem for the main process, plus volumes for whatever actually needs to be written.
- Don't put secrets in the image. Even if you delete them in a later layer, the layer where they were added still exists and can be extracted. Secrets get injected at runtime.
- Pin base image versions.
FROM node:20.11.1-alpine, notFROM node:latest— otherwise today's build differs from yesterday's. - Scan the image (Trivy, Grype) as a pipeline step.
8.6 What gets lost when the container dies
The container's filesystem is ephemeral: it disappears along with the container. Whatever needs to survive goes into a volume or an external service (database, object storage). Logs get written to standard output, not to files inside the container — the platform collects them.
At DISH you had your applications in containers, in Kubernetes. If they ask you about Docker, your anchor is practical: “I worked with images in the context of shipping to GKE — one image per service, tagged with the commit, the same image promoted across environments.” Then show that you know the three hygiene rules: multi-stage, non-root user, no secrets in layers.
9 Kubernetes — fundamentals
9.1 The problem it solves
You have fifty containers that need to run on twenty machines. Who decides which machine each one goes on? What happens when a machine dies at three in the morning? How do you route traffic to containers that come and go and change address? How do you ship a new version without stopping the service? Kubernetes is the answer to all of that, in one package.
9.2 The core idea: desired state and reconciliation
You don't give Kubernetes commands. You declare what you want the world to look like — “I want three copies of the app, version 2.1, with this much memory” — and a set of control loops constantly compares reality against your declaration and acts to close the gap. Delete a pod, another one appears. A node dies, its pods get recreated somewhere else. Not because someone reacted, but because the loop never stops.
9.3 Cluster architecture
“Kubernetes is a declarative system. I describe the desired state in a file, send it to the API server, which saves it in etcd. From there, the controllers constantly compare what I asked for against what exists and correct the difference: the scheduler picks the node, the kubelet on that node starts the container. The practical consequence is that I don't have to react when something dies — the loop does.”
10 Kubernetes — the objects
10.1 Pod
The smallest unit Kubernetes can schedule on a node. It holds one or more containers that share the same IP address, the same ports, and can share volumes — containers inside a pod can see each other on localhost. In practice, one main container, plus maybe a helper one (sidecar): a log collector, a network proxy.
A pod is disposable. It doesn't get repaired: it gets deleted and another one gets created, with a different IP. That's also why you never connect directly to a pod's IP, but to a Service instead.
10.2 Deployment
The object you work with daily. You tell it which image, how many copies and how to do the update; it creates a ReplicaSet, which creates the pods. For every new version it creates a new ReplicaSet and scales down the old one — that's also where the ability to roll back comes from.
apiVersion: apps/v1 kind: Deployment metadata: name: api namespace: prod spec: replicas: 3 # how many copies I want selector: matchLabels: { app: api } # which pods belong to me strategy: type: RollingUpdate rollingUpdate: maxSurge: 1 # I'm allowed 1 extra pod during the rollout maxUnavailable: 0 # but NONE fewer → zero downtime template: # from here down is the pod recipe metadata: labels: { app: api } spec: containers: - name: api image: registry/api:9f2c1ab # NOT “latest”: you want to know what's running ports: [{ containerPort: 8080 }] env: - name: DB_HOST valueFrom: configMapKeyRef: { name: api-config, key: db_host } - name: DB_PASS valueFrom: secretKeyRef: { name: api-secret, key: db_pass } resources: requests: { cpu: "200m", memory: "256Mi" } # for scheduling limits: { cpu: "1", memory: "512Mi" } # ceiling at runtime readinessProbe: # can I receive traffic? httpGet: { path: /ready, port: 8080 } initialDelaySeconds: 5 periodSeconds: 5 livenessProbe: # am I still alive? httpGet: { path: /healthz, port: 8080 } initialDelaySeconds: 20 periodSeconds: 10
10.3 Service — the four types
| Type | What it does | When you use it |
|---|---|---|
ClusterIP | Internal IP, visible only from inside the cluster. Default. | Communication between services. Most cases. |
NodePort | Opens the same port on every node. | Rare, in development or behind your own load balancer. |
LoadBalancer | Requests a real load balancer from the cloud provider, with a public IP. | Direct exposure. Costs money: one load balancer per service. |
ExternalName | Just a DNS alias to an external name. | So you can give the same name to a managed database outside the cluster. |
The Service finds pods through labels (selector), not by name. If you change the label on the pod and forget the selector, the Service is left with no endpoints and you get a connection error even though every pod is running fine. It's one of the most common causes of “it works on pod X, it doesn't work through the service.”
10.4 Ingress
A LoadBalancer Service per application gets expensive and hard to manage. Ingress is the layer on top: a single entry point, which routes by domain name and path to different internal services, and terminates TLS in one place. It needs an ingress controller installed in the cluster (nginx, Traefik, the provider's own) — the Ingress object by itself does nothing.
10.5 ConfigMap and Secret
Configuration doesn't go into the image. It goes into a ConfigMap (regular settings) or a Secret (passwords, keys, certificates) and gets injected into the pod, either as environment variables or as mounted files.
“What's the difference between ConfigMap and Secret?” The wrong answer is “Secret is encrypted.” Secret is just base64-encoded, which means zero protection — anyone can read it with base64 -d. What it adds on top: it's treated differently (it doesn't show up in logs, it can be encrypted at rest in etcd if you turn on encryption at rest, it can be restricted separately through RBAC). For real secrets, you use external solutions: External Secrets Operator with Vault, AWS Secrets Manager or Google Secret Manager, and Sealed Secrets if you want the secrets in Git.
10.6 The other workload types
| Object | What for | The key difference |
|---|---|---|
| Deployment | Stateless applications: APIs, websites, services. | Pods are interchangeable, they get random names. |
| StatefulSet | Databases, queues, anything with identity: PostgreSQL, Kafka, Elasticsearch. | Stable, ordered names (db-0, db-1), each with its own disk, started and stopped in order. |
| DaemonSet | One pod per node: log collector, monitoring agent, network plugin. | You don't specify the copy count — it depends on the number of nodes. |
| Job | A task that runs once and finishes: a migration, an import. | It tracks successful completion, not staying up. |
| CronJob | A Job on a schedule: nightly backup, daily report. | Regular cron syntax. |
“Deployment for anything stateless, because the pods are interchangeable and can be replaced in any order. StatefulSet when each instance has its own identity and its own disk — a database, a Kafka. The practical difference shows up when a node goes down: with a Deployment, the pod gets recreated immediately, anywhere. With a StatefulSet, Kubernetes doesn't automatically recreate the pod until it's sure the old one has died, specifically so you don't end up with two instances writing to the same disk. That means recovery is slower, on purpose.”
10.7 Storage
Three related objects: PersistentVolume (the actual disk), PersistentVolumeClaim (the application's request: “I want 20 GB, fast”) and StorageClass (the recipe that automatically creates the disk at the provider). In practice you only write the PVC; the rest resolves itself. Worth remembering: a disk in the cloud is usually tied to a zone — a pod that needs it can only be scheduled in that zone.
11 Delivery and health: the two things that get asked most
11.1 Rolling update, step by step
“Kubernetes doesn't roll back automatically to the previous version. If the new pods don't become ready, the rollout gets stuck and you're left running the old version — which is actually good behavior, because the service doesn't go down. But someone has to notice and run kubectl rollout undo. That's why my pipeline has kubectl rollout status --timeout, which turns that stuck state into a visible pipeline failure.”
# update the image kubectl set image deployment/api api=registry/api:9f2c1ab -n prod # watch — exits non-zero if it doesn't succeed within 180s kubectl rollout status deployment/api -n prod --timeout=180s # revision history kubectl rollout history deployment/api -n prod # roll back to the previous version kubectl rollout undo deployment/api -n prod # roll back to a specific revision kubectl rollout undo deployment/api --to-revision=4 -n prod # clean restart, no image change (useful after changing a Secret) kubectl rollout restart deployment/api -n prod
11.2 Probes — the three types and what they're for
If you put a database check inside livenessProbe, here's what happens the moment the database has one bad second: all the pods fail the probe at the same time, Kubernetes kills every one of them, they restart, they reconnect to the already-overloaded database, and they fail again. You've turned a temporary degradation into a full outage, with the restart loop acting as an amplifier. The rule: liveness only checks whether the process is healthy; external dependencies get checked in readiness, where the consequence is being pulled from traffic, not getting killed.
11.3 Graceful shutdown
When a pod is stopped, Kubernetes does two things in parallel: it removes the pod from the Service's list of endpoints and sends SIGTERM to the process. Because these happen in parallel, there's a window of a few hundred milliseconds where the pod is still receiving new requests but has already started shutting down. The standard fix is a preStop hook that waits a few seconds before shutdown, plus an application that handles SIGTERM properly: it stops accepting new connections, finishes what it's already working on, then exits. terminationGracePeriodSeconds (default 30) is how much patience Kubernetes gives it; after that comes SIGKILL.
12 Resources and scaling
12.1 requests and limits
| QoS class | How you get it | What it means |
|---|---|---|
Guaranteed | requests = limits, for all resources. | The last one evicted when the node is under pressure. For critical components. |
Burstable | requests < limits. | The usual case. Can temporarily exceed its reservation if the node has room. |
BestEffort | No requests and no limits. | The first one killed when the node runs out of memory. Avoid in production. |
12.2 The three kinds of scaling
HPA — horizontal
Changes the number of pods based on a metric (CPU, memory, or a custom one). The one used most often.
VPA — vertical
Changes the pods' requests/limits. Useful for figuring out the right values; in automatic mode it restarts the pods.
Cluster Autoscaler
Adds or removes nodes. Kicks in when pods are stuck Pending because there's no room left.
If the service is bottlenecked by something other than CPU — the number of connections to the database, the length of a queue, a slow external API — then CPU stays low while the service is struggling, and HPA never kicks in. Worse: if it does kick in, the new pods also request connections to that same database and make things worse. The right answer: scale on the metric that actually describes the load — queue depth, requests waiting — through custom metrics or KEDA.
12.3 Where pods land — placement
- nodeSelector / nodeAffinity — “only place this pod on GPU nodes” or “preferably in zone A”.
- podAntiAffinity — “don't put two copies of the same service on the same node”. Without this, all three replicas can end up on a single node, and one node going down means the service goes down.
- topologySpreadConstraints — spreads pods evenly across availability zones.
- taints and tolerations — the reverse mechanism: the node rejects everything except pods that explicitly “tolerate” it. This is how you reserve nodes for special workloads.
- PodDisruptionBudget — “at least two pods available at all times”. Protects you during planned operations, like node upgrades.
13 Security in the cluster
13.1 RBAC — who's allowed to do what
Four objects, two pairs. Role grants permissions within a single namespace; ClusterRole across the whole cluster. RoleBinding and ClusterRoleBinding tie the role to someone: a user, a group, or a ServiceAccount (the identity a pod runs as).
kind: Role metadata: { namespace: team-a, name: developer } rules: - apiGroups: [""] resources: [pods, pods/log, services, configmaps] verbs: [get, list, watch] # read-only - apiGroups: [apps] resources: [deployments] verbs: [get, list, watch, patch] # can change the image, cannot delete
The command you use to check: kubectl auth can-i delete pods --as=ana -n prod.
13.2 Pod Security Admission
PodSecurityPolicy was deprecated in version 1.21 and removed in 1.25. The built-in replacement is Pod Security Admission: you put a label on the namespace and the cluster rejects pods that don't meet that level.
| Level | What it allows |
|---|---|
privileged | Everything. For system components. |
baseline | Blocks the obvious privilege escalations: privileged containers, the host's process namespace. |
restricted | The strictest one: no root, no privilege escalation, read-only filesystem. The target for applications. |
It applies in three modes: enforce (reject), audit (log it) and warn (warn the user). The usual adoption path is to start with warn, see what would break, and then switch to enforce. For more complex rules, people use OPA Gatekeeper or Kyverno.
13.3 NetworkPolicy
By default, in Kubernetes any pod can talk to any pod, in any namespace. It's the most surprising default in the whole system, and a good answer to “what would you improve in a legacy cluster”. NetworkPolicy changes that: you describe what traffic is allowed, and everything else is blocked. One catch: it only has an effect if the network plugin supports it (Calico, Cilium; some simple setups silently ignore it).
13.4 A namespace is not a security boundary
A namespace is a logical partition: it groups objects, supports resource quotas, and gives you a point to apply RBAC and NetworkPolicy. But pods from different namespaces run on the same nodes and share the same kernel. Real isolation between different tenants means separate nodes, or separate clusters. Saying this in an interview immediately sets you apart from people who've only read tutorials.
13.5 The rest of the hygiene checklist
- Disable automatic mounting of the ServiceAccount token in pods that don't talk to the API (
automountServiceAccountToken: false). - Encrypt etcd at rest, or keep secrets in an external vault.
- Don't use
:latest. You can't know what's actually running, and you can't reproduce an issue. - Scan images in the pipeline and block on critical vulnerabilities.
- Keep the API audit log turned on. It's the only source that tells you who changed what.
14 Cluster troubleshooting — the most-asked part
In 2026 Kubernetes interviews, “what is a pod”-type questions are just the warm-up. The part that matters is the scenario: “you’re told the service is down, what do you do?”. They’re testing the method, not the memorization.
14.1 The order you investigate in
# 1. What state are the pods in? (this gets you 60% of the answer) kubectl get pods -n prod -o wide # 2. Why is it in that state? Look at EVENTS, at the end of the output. kubectl describe pod api-7d9f-x2k -n prod # 3. What is the app saying? kubectl logs api-7d9f-x2k -n prod kubectl logs api-7d9f-x2k -n prod --previous # the DEAD container's logs — essential for CrashLoop # 4. What happened recently in the whole namespace? kubectl get events -n prod --sort-by=.lastTimestamp | tail -30 # 5. Does the service have endpoints? (empty list = wrong selector or pods not ready) kubectl get endpoints api -n prod # 6. From inside a pod: DNS and connectivity kubectl exec -it api-7d9f-x2k -n prod -- sh nslookup api.prod.svc.cluster.local wget -qO- http://api:8080/healthz # 7. Actual consumption kubectl top pods -n prod kubectl top nodes # 8. Are the nodes healthy? kubectl get nodes kubectl describe node nod-3 | grep -A5 Conditions
14.2 The diagnostic table
| What you see | What it means | Common causes | Where you look |
|---|---|---|---|
Pending | The pod wasn’t scheduled on any node. | Requests too high; no node that tolerates the taint; unbound PVC; wrong zone. | kubectl describe pod → the Events section tells you exactly why the scheduler refused. |
ImagePullBackOffErrImagePull | Can’t pull the image. | Wrong name or tag; private registry without imagePullSecret; pull rate limit hit. | Events; manually check the image name. |
CrashLoopBackOff | The container starts and dies repeatedly; Kubernetes waits longer and longer between attempts. | Config error; missing environment variable; can’t connect to the database; wrong command; liveness probe too aggressive. | kubectl logs --previous. Without --previous you only see the new container, which has barely started. |
OOMKilled (137) | It exceeded the memory limit and was killed by the kernel. | Limit too low; memory leak; a batch of data too large loaded into memory. | kubectl describe pod → Last State: Terminated, Reason: OOMKilled. Compare with kubectl top. |
Running but 0/1 READY | It’s running, but the readinessProbe isn’t passing. It isn’t receiving traffic. | Wrong probe path or port; the app genuinely isn’t ready yet; a dependency is missing. | kubectl describe pod → the probe message; then manually hit the path from inside the pod. |
Terminating stuck | Can’t be deleted. | A finalizer waiting on something; a volume that can’t unmount; the process is ignoring SIGTERM. | kubectl get pod -o yaml → the finalizers field. |
Evicted | The node ran out of resources and kicked the pod out. | Memory or disk pressure on the node; the pod was BestEffort. | kubectl describe node → Conditions. |
Node NotReady | The node has stopped reporting. | kubelet stopped; disk full; network; the machine is gone. | Cloud console; kubelet logs. |
“First I confirm what ‘down’ means — errors, slowness, or just some users — and since when. Then I look at the pods: how many are ready out of how many should be. If the pods are fine, the problem is higher up: service, ingress, DNS, certificate. If the pods aren’t fine, I describe the pod and read the events, because the state already tells me the class of the problem.
In parallel I ask what changed in the last few hours, because in most cases there is a change. If it’s a recent deploy, my first move is to roll back to the previous version and only then understand why — I restore the service first, diagnose after. I communicate on the incident channel at the start and at every important step, so nobody has to interrupt me to ask.”
The method above is exactly what you did 24/7 at Metro and at Volvo, just with different tools. Say it as a method you have, not as a list you memorized: “My order has stayed the same for years — confirm the symptom, restore the service, then look for the cause. On Kubernetes the commands change, not the method.”
14.3 A network troubleshooting path, start to finish
“Service A can’t reach service B.” The order that finds the problem without guessing:
kubectl get pods -l app=b— do the pods exist and are they READY? A pod that isn’t ready isn’t in service.kubectl get endpoints b— is the list empty? Then the service’s selector doesn’t match the pods’ labels. That’s the most common cause.- From pod A:
nslookup b.namespace.svc.cluster.local— if DNS doesn’t answer, the problem is CoreDNS, not your service. - From pod A:
wget -qO- http://b:8080/— if DNS works but the connection is refused, check the port: theportin the Service must match thetargetPortand the app’s actual port. kubectl get networkpolicy -n namespace— is there a policy blocking it? If it’s the first policy in the namespace, it flipped the default from “everything allowed” to “only what’s written.”
15 Terraform and infrastructure as code
15.1 The idea
Infrastructure isn’t created by hand from the cloud console. It’s described in files, the files live in Git, and the tool applies the difference between what’s written and what exists. The consequences: environments are reproducible, changes go through review like any code, and there’s a history of the reasons behind them.
15.2 The cycle
15.3 State — the subject that separates people
Terraform keeps a state file that ties what you wrote to what exists in reality. Without it, it doesn’t know whether to create or modify.
- The state doesn’t live locally. It lives in a shared store (S3 with DynamoDB locking, an Azure Storage container, GCS), otherwise two colleagues applying at the same time step on each other’s state.
- Locking is mandatory. Two
applyruns in parallel on the same state can corrupt it. - The state contains secrets in plain text. The password of a database created by Terraform shows up in the state file. So: encryption, restricted access, and never in Git.
- Drift happens when someone changes something manually in the console.
terraform planshows it. The technical fix isimportor realignment; the real fix is organizational — you remove manual write access to production.
terraform { required_version = "~> 1.9" backend "s3" { # the state, remote and lockable bucket = "tf-state-company" key = "prod/network.tfstate" region = "eu-central-1" dynamodb_table = "tf-locks" encrypt = true } required_providers { aws = { source = "hashicorp/aws", version = "~> 5.60" } } } variable "environment" { type = string default = "prod" } module "network" { # module = reuse across environments source = "./modules/network" environment = var.environment cidr = "10.20.0.0/16" } resource "aws_db_instance" "main" { identifier = "db-${var.environment}" engine = "postgres" instance_class = "db.t4g.medium" allocated_storage = 50 lifecycle { prevent_destroy = true # safety net: can't be deleted by mistake } } output "db_endpoint" { value = aws_db_instance.main.endpoint }
How do you separate environments? Separate directories with separate state files and shared modules. Workspaces are useful for small variations, but dev and prod in the same state file is a recipe for disaster.
What do you do with a resource created manually? terraform import brings it under control, then I write the definition to match it.
How do you avoid a destructive apply? The plan is read by a human, prevent_destroy on the resources that hold data, and apply only from the pipeline, with an identity that has no delete rights on anything critical.
At RapidBid you wrote Terraform for AWS and Azure. So you also have an extra topic that many don’t: what it means to support two providers — separate modules per provider, because the resources aren’t equivalent, with a common interface on top. If they ask you about multi-cloud, that’s real experience, not theory.
16 Cloud: what you need to know from each one
You don’t need a certification to answer at an interview. You need the map of the core services and the equivalences, because job posts ask for “AWS or Azure or GCP,” and whoever knows one can learn the other.
| What you need | AWS | Azure | Google Cloud |
|---|---|---|---|
| Virtual machines | EC2 | Virtual Machines | Compute Engine |
| Managed Kubernetes | EKS | AKS | GKE |
| Image registry | ECR | ACR | Artifact Registry |
| Object storage | S3 | Blob Storage | Cloud Storage |
| Relational databases | RDS / Aurora | Azure SQL / DB for PostgreSQL | Cloud SQL |
| Private network | VPC | VNet | VPC |
| Identity and access | IAM | Entra ID + RBAC | IAM |
| Secrets | Secrets Manager | Key Vault | Secret Manager |
| Monitoring | CloudWatch | Azure Monitor | Cloud Monitoring |
| Serverless functions | Lambda | Functions | Cloud Functions / Run |
| Native CI/CD | CodePipeline | Azure DevOps | Cloud Build |
The bolded entries in the table are from your résumé: GKE and Cloud SQL from DISH, Azure DevOps from RapidBid, plus Terraform on AWS. That means you’ve touched all three clouds. The honest and strong way to put it: “I’ve worked on Google Cloud in production — GKE and Cloud SQL — and on AWS and Azure through Terraform and Azure DevOps. The core services map to each other; what differs is the identity and access model, which is the part you need to be careful about when you move from one to another.”
Cost, the topic that comes up at senior level
A senior DevOps engineer gets asked about cost, because it’s a real lever. The main sources of waste and what you do about them:
- Resources reserved far above actual consumption. Fixed by measuring actual usage and adjusting
requests. This is usually where the biggest savings are. - Test environments running at night and on weekends. Shut them down on a schedule.
- Orphaned disks and IP addresses left behind after the resources that used them were deleted.
- Cross-availability-zone traffic, which gets billed and which nobody notices until they look at the invoice.
- Tagging by team and environment, otherwise the bill can’t be attributed to anyone and nobody optimizes it.
17 Monitoring, Logs and SLOs
17.1 The three pillars
📊 Metrics
Numbers over time: requests per second, latency, errors, memory. Cheap to keep for a long time. Answer “what isn’t going well and since when”.
📜 Logs
Events with text. Expensive to keep. Answer “why”. Written structured (JSON) so they can be searched.
🔗 Traces (tracing)
The path of a request through all services, with the time spent in each one. Answer “where time is lost”. Indispensable for microservices.
17.2 What actually gets monitored
Two sets of rules that get quoted a lot:
The four golden signals
Latency (how long it takes), traffic (how much is requested), errors (how much fails), saturation (how full the most limited resource is).
The RED method, for services
Rate — requests per second. Errors — how many fail. Duration — the distribution of response times.
Don’t look at average latency, look at the percentiles: p50, p95, p99. With an average of 200 ms you can still have 1% of users waiting 8 seconds — and those are the ones who call. Interview phrasing: “I track p95 and p99, not the average, because the average hides exactly the users who are affected.”
17.3 SLI, SLO, SLA and the error budget
| Term | What it is | Example |
|---|---|---|
| SLI | The measured indicator. | The percentage of requests that respond in under 300 ms. |
| SLO | The internal target on that indicator. | 99.9% of requests under 300 ms, over 30 days. |
| SLA | The contractual promise to the customer, with penalties. | 99.5% availability per month. Always looser than the SLO. |
| Error budget | How much you’re allowed to get wrong: 100% minus the SLO. | An SLO of 99.9% per month gives you about 43 minutes of non-compliance. |
“The error budget turns a discussion of opinions into one with numbers. If we still have budget left, the team can ship aggressively. If we’ve burned through it, we stop new features and work on stability until it recovers. Nobody has to convince anybody else that ‘it’s too risky’ — the measurement makes the call.”
17.4 Alerts
The most common problem on real teams isn’t too few alerts, it’s too many. When you get forty alerts a day and thirty-nine of them don’t need anything, you stop reading them — and you miss the fortieth one. The rules that fix it:
- Alert on symptom, not cause. “Error rate went above the threshold” is useful. “CPU at 90%” isn’t — it might be perfectly normal.
- Every alert that fires at night has to require immediate human action. If it can wait until morning, it’s not an alert, it’s a ticket.
- Every alert has a runbook — a short document that says what to check and what to do.
- Alert on the rate at which the error budget is being consumed, not on every spike.
17.5 Tools
Prometheus collects metrics by pulling them periodically from applications and is queried with PromQL. Grafana draws them. Alertmanager routes alerts and groups them. Loki or ELK for logs. Jaeger or Tempo for traces. OpenTelemetry is the instrumentation standard that feeds all of them. The commercial alternative that does all of it in one product: Datadog, New Relic, Dynatrace.
You’ve worked with Nagios and the Tivoli suite. They’re the previous generation, but the idea is identical: collect signals, define thresholds, route the alert to whoever’s on call. If they ask you about Prometheus, don’t say “I haven’t worked with it.” Say what you did and where the difference is: “I did monitoring with Nagios and Tivoli. The model is the same — signal, threshold, alert, who responds. What’s different with Prometheus is that it pulls metrics from applications, and querying is its own language, PromQL, instead of checks written one by one.”
18 Incidents and On-Call
18.1 Roles in a major incident
- Incident commander — coordinates, doesn’t troubleshoot. Makes the decisions and keeps order.
- Communications lead — updates stakeholders and customers at fixed intervals, so no one interrupts the technical team to ask “how’s it going.”
- The people fixing it — actually look at the system.
- The scribe — records the timeline with exact times. Without it, the postmortem is a discussion from memory.
On small teams, one person can wear several hats. What you never do: have the commander also get hands-on in the system, because then no one has the overall picture.
18.2 The blameless postmortem
The principle: people aren’t the cause; the system that allowed the mistake is. If a wrong command could delete production, the problem isn’t the person who typed it, it’s that the system accepted it without confirmation, without a blast-radius limit, and without a way back. If a name shows up in the postmortem instead of a systemic cause, people will hide incidents next time, and you lose exactly the information you needed.
Document structure: summary, impact (who, how long, what broke), timeline with times, contributing causes (plural), what went well, what went wrong, actions with an owner and a deadline. Without that last section, it’s just literature.
24/7 production control at Metro, on-call, escalation, customer communication. That’s not something you learn from a course, and it’s exactly what a team with a live service is looking for. When you get to the incidents topic, slow down and give a concrete example — this is your home turf.
19 Security in the Pipeline (DevSecOps)
The idea in one sentence: security checks move earlier into the development process, where they’re cheap, instead of being an audit at the end.
| Where | What you check | With what |
|---|---|---|
| Before commit | Secrets accidentally written into code. | gitleaks, git-secrets, as a local hook. |
| On every PR | Vulnerabilities in your own code (SAST). | CodeQL, Semgrep, SonarQube. |
| On every PR | Dependencies with known vulnerabilities (SCA). | Dependabot, Snyk, npm audit. |
| After build | Vulnerabilities in the container image. | Trivy, Grype. |
| After build | Full component list (SBOM) and the image’s signature. | Syft for SBOM, Cosign for signing. |
| Before apply | Insecure configurations in Terraform and in the Kubernetes manifests. | Checkov, tfsec, kube-score. |
| In production | What’s actually running and whether it’s behaving abnormally. | Falco, provider security agents. |
If a key got pushed to the repo, it stays in the history and in every existing clone. The right order is: first invalidate it with the provider and generate a new one, only then clean up the history. Do it the other way around and cleaning the history is pointless — the key’s already been copied. It’s a question that comes up, and it separates the learned answer from the lived one.
Software supply chain
A topic that’s come up more and more after the last few years’ incidents. The main points: pin actions and images to a hash, not a tag; sign artifacts and verify the signature on delivery; generate an SBOM so you know within five minutes whether you’re affected by a new vulnerability; give the pipeline minimal permissions, obtained through OIDC instead of permanent keys.
20 Bash and Python
Almost every job posting asks for “scripting.” You’re not being asked to be a programmer: you’re being asked to automate something repetitive and not cause damage when the script fails halfway through.
20.1 The header that saves you
#!/usr/bin/env bash set -euo pipefail # -e : stop at the first command that fails # -u : error when using an undefined variable (catches typos) # -o pipefail : a pipeline fails if ANY element fails, not just the last one IFS=$'\n\t' # no longer split on spaces — protects paths with spaces # guaranteed cleanup, whatever happens TMP=$(mktemp -d) trap 'rm -rf "$TMP"' EXIT # required argument, with a clear message if it's missing ENVIRONMENT=${1:?usage: $0 <environment>} log() { printf '%s [%s] %s\n' "$(date +%FT%T)" "$1" "$2" >&2; } # retry with increasing backoff — mandatory for any network call retry() { local n=0 max=5 delay=2 until "$@"; do n=$((n+1)) [ "$n" -ge "$max" ] && { log ERROR "failed after $max attempts: $*"; return 1; } log WARN "attempt $n failed, retrying in ${delay}s" sleep "$delay"; delay=$((delay*2)) done } retry curl -fsS "https://api.internal/$ENVIRONMENT/status" > "$TMP/status.json" log INFO "done"
curlwithout-freturns success even on a 500 response, because it successfully downloaded the error page. Almost everyone gets this wrong."$@"keeps the arguments separate;"$*"glues them into a single string. With paths that have spaces, the difference is between “it works” and “it deletes something else.”- Quotes around variables aren’t a style choice, they’re correctness.
- Idempotency: running the script twice gives the same result. And
flockon a file so two instances don’t run at once from cron. shellcheckin the pipeline catches most of these mistakes automatically.
20.2 When you switch to Python
The practical rule: Bash up to a hundred lines or until you start parsing serious JSON; past that, Python. Signs you’ve outgrown it: you need data structures, error handling for different cases, HTTP calls with headers and retries, tests.
import os, sys, json, logging import requests from tenacity import retry, stop_after_attempt, wait_exponential logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s") TOKEN = os.environ["API_TOKEN"] # from env, NOT from code @retry(stop=stop_after_attempt(5), wait=wait_exponential(min=1, max=30)) def get_status(environment: str) -> dict: r = requests.get(f"https://api.internal/{environment}/status", headers={"Authorization": f"Bearer {TOKEN}"}, timeout=10) # timeout ALWAYS — otherwise it hangs forever r.raise_for_status() return r.json() if __name__ == "__main__": try: print(json.dumps(get_status(sys.argv[1]), indent=2)) except Exception as e: logging.error("could not read status: %s", e) sys.exit(1) # correct exit code, so the pipeline knows
20.3 Linux — the order you investigate a server in
| Question | Command |
|---|---|
| Is the process running? | systemctl status service · ps aux | grep name |
| Is it listening on a port? | ss -ltnp · lsof -i :443 |
| Can you reach it? | curl -v · telnet host port · traceroute |
| Does the name resolve? | dig name · nslookup |
| Has it run out of something? | df -h (disk) · free -h (memory) · top · iostat |
| What does the log say? | journalctl -u service -n 100 --no-pager · dmesg -T | tail |
| Who’s using the disk? | du -sh /* 2>/dev/null | sort -h |
| What changed? | last · shell history · the deploy log |
“The disk.” It’s mundane and it’s by far the most common cause: a full disk stops almost every service, and often in ways that look unrelated — the database stops writing, the logs stop, processes hang. It’s the first command I run, because it takes a second and rules out the most likely cause.
21 Your experience, translated into the language of the role
This is the section you need most. You’re not short on experience — you’re short on the habit of naming it in the words the person across from you is listening for. Someone who kept production up at Metro and migrated message queues at Volvo has exactly the background the DevOps role is built on; if he tells it with 2012’s terms, the listener hears “system administrator,” not “reliability engineer.”
| What you did | What it’s called today | The phrase you say |
|---|---|---|
| 24/7 production control, on-call, escalation (Metro, IBM) | Reliability engineering, incident response, on-call | “I worked for years in round-the-clock production control. I picked up incidents, escalated by severity, and communicated with the client for as long as the issue was open.” |
| Nagios, the Tivoli suite | Monitoring and alerting | “I did monitoring with Nagios and Tivoli — signals, thresholds, alert routing. The model is the same as Prometheus, what differs is the collection method and the query language.” |
| WebSphere MQ v8 → v9 migration, queue managers across multiple instances (Volvo Cars) | Platform modernization, high availability, zero-downtime migration | “I migrated the messaging platform from one major version to another, on a multi-instance setup — exactly the case where you’re not allowed to stop the flow.” |
| TADDM (Etihad) | Automated discovery, configuration inventory, dependency mapping | “I worked on automated infrastructure discovery and mapping dependencies between applications — what talks to what, which matters enormously when you’re planning a change.” |
| AIX, Linux, shell | System administration, scripting | “I administered UNIX systems in production. System-level troubleshooting — disk, memory, processes, network — I do without looking at the docs.” |
| GKE, 17 pods, 44 Cloud SQL databases (DISH Digital Solutions) | Managed Kubernetes, operating on Google Cloud | “I operated on Google Kubernetes Engine, with 17 pods and 44 Cloud SQL databases behind them, and I shipped releases with no service interruption.” |
| Terraform on AWS and Azure (RapidBid) | Infrastructure as code, multi-provider | “I wrote infrastructure as code in Terraform, across two providers — AWS and Azure. What differs between them is the providers and the authentication method; the plan–apply flow and the state discipline are identical.” |
| Azure DevOps, Jenkins with blue-green delivery (RapidBid) | CI/CD, delivery strategies | “I built pipelines in Azure DevOps and in Jenkins, including a blue-green rollout, switching traffic after the new environment passed its checks.” |
| Own servers, n8n, integration with language models (recent years) | Automation, systems integration, end-to-end operation | “In recent years I’ve built and kept running systems of my own, from the server up through automations and integrations. I’ve done every role involved, including the ones I’m after here.” |
Every row above is verifiable against your CV. If in the interview you catch yourself adding a detail that sounds good — a percentage, a mechanism, a tool — and you’re not sure it happened exactly like that, stop. A technical interviewer digs exactly where the detail sounds impressive, and if the second question catches you without an answer, you lose what was true before it too. “I don’t remember exactly, I’d need to check” is an answer that has never cost anyone a job.
21.1 Where you’re on solid ground and where you’ll be pushed
🟢 Solid ground
Incidents and on-call. System-level troubleshooting. Migrations on systems that aren’t allowed to go down. Client communication under pressure. Tenure at a multinational, with real procedures. Operating on managed Kubernetes.
🟠 Where they’ll dig
How much Kubernetes you’ve written, not just operated. How many pipelines you’ve built from scratch. GitHub Actions. How much code you actually write. What you did during the gap between multinationals.
“That’s an area I’ve worked less on than the rest. What I’ve actually done is [the exact thing]. What I haven’t done is [the exact thing]. I know the model, and the gap to what you need I can close in a few weeks of hands-on work — I’ve made that same jump before, when I moved onto Google Cloud.”
The structure has three parts and each one does a job: you acknowledge it (so they don’t catch you out), you draw the line precisely (showing you know what you know, which is rarer than it sounds), you show the path (giving the reason their risk is small). What you never do: just say “I haven’t worked with that” and go quiet.
22 Learning plan
You don’t need to know everything. You need what you say to be backed by something you’ve touched with your own hands. The plan below is built on that: every block ends with something that exists on your machine and that you can talk about in the past tense.
| Block | What you do | What’s left behind |
|---|---|---|
| 1 | Install a local cluster (kind or minikube). Write by hand a Deployment, a Service, and an Ingress for a plain app. | You can say “I wrote the manifests” without exaggerating. |
| 2 | Break things on purpose: a nonexistent image, a readiness probe pointed at the wrong path, a memory limit set too low. Fix each case by reading describe and logs. | You’ve seen ImagePullBackOff, CrashLoopBackOff, and OOMKilled with your own eyes. These are safe interview questions. |
| 3 | A GitHub repo with a workflow that builds an image, pushes it to ghcr, and runs tests. You add needs, a matrix, and caching. | The honest answer to “have you worked with GitHub Actions?” gets stronger. |
| 4 | Connect the workflow to the local cluster or a free one and do a rolling update. Then trigger a failure and run rollout undo. | You’ve taken a full chain from commit to production. |
| 5 | Terraform: one simple resource on one provider, with remote state and locking. Make a manual change in the console and watch the drift show up in plan. | The story about drift becomes lived, not read. |
| 6 | Prometheus and Grafana on top of the local cluster. A dashboard with the four golden signals and an alert that actually fires. | You can tie Nagios to Prometheus in one credible sentence. |
| 7 | Rehearsal out loud: chapters 9–14 of this page, explained as if you were teaching. | The sentences come out fluent. That’s the whole point of this document. |
One thing broken and fixed with your own hands is worth ten chapters read. Technical interviewers spot instantly the difference between someone who has seen a pod stuck in Pending and someone who has read what it means. The first says “I looked at the events and there was no node that could take it.” The second recites the definition.
Certifications, if you want something on paper
CKA (Certified Kubernetes Administrator) is the best regarded of all of them, because it’s a hands-on exam — you sit at a terminal and solve things. KCNA is the entry-level version, questions only. Terraform Associate is cheap and quick. Vendor certifications (AWS, Azure, Google Cloud) matter mostly when the company runs on that provider. No certification replaces a cluster you broke and fixed, but CKA gets past the automated screening filters, which is a real and separate problem.
23 Glossary
Terms that come up in conversation and are assumed to be known. If one of them stops you mid-interview, you’ve lost the thread of the sentence — that’s why they’re here.
Infrastructure
Idempotent — run it twice, get the same result. Declarative — you describe the desired state, not the steps. Imperative — you give the steps. Drift — reality has moved away from what the code says. Provisioning — creating resources. Ephemeral — disappears without a trace.
Delivery
Artifact — the result of the build (image, archive). Registry — the image repository. Immutable tag — a tag that never gets overwritten again. Rollback — reverting to the previous version. Blast radius — how much breaks if the change is wrong.
Reliability
Availability — the percentage of time the service responds. Redundancy — spare copies. Single point of failure — the piece that, if it goes down, stops everything. Graceful degradation — works partially instead of falling over. Backpressure — the system refuses requests so it doesn’t collapse. Thundering herd — all the clients retry at once and take down the service that was just coming back.
Kubernetes
Manifest — the YAML file that describes an object. Reconciliation — the loop that brings the actual state to the desired one. Label — the key-value pair that selectors use. Annotation — metadata with no selection role. Sidecar — a helper container in the same pod. Eviction — removing a pod from a node under pressure. Cordoning — the node stops taking new pods.
Networking
Reverse proxy — sits in front and distributes to services. TLS termination — the point where encryption ends. Virtual host — several domains on one IP. CIDR — the notation for address ranges. Egress / Ingress — traffic going out / coming in.
Process
Blameless — analysis of the system, not the person. Runbook — the instructions for a known situation. Change freeze — the period when nothing goes to production (holidays, traffic peaks). Feature flag — you turn a feature on or off without shipping code.
“Could you repeat the question? I want to make sure I’m answering what you actually asked.”
That sentence buys you five seconds, sounds professional, and costs you nothing. Nobody, ever, has been turned down for asking for a clarification. People get turned down for starting to talk without knowing where they’re headed. If you take away one single thing from this whole page, take away this.