⚙️ DevOps manual

Part two · updated 25 September 2026

The interview. What they ask, what you answer, and what you do when you freeze.

The first page teaches you the job. This page prepares you for the forty minutes in which you have to show it. It contains the hiring process step by step, the concrete techniques against freezing up, the stories from your CV already written out in full sentences, the technical questions with written answers, and the short list of things you do not say.

The goal is not to know everything. The goal is to carry every sentence through to the end, calmly, and to say “I don’t know” when you genuinely don’t, without it wrecking the rest of the interview.

processanti-freezeSTAR stories technical questionsbehavioralthe market, verified
← back to the technical part
1 · What you know it counts, but less than you think 2 · How you say it full sentences, order, calm 3 · What you do when you don’t know this is where most senior interviews are decided

1 What the hiring process looks like

Almost every company hiring DevOps engineers in Romania uses the same sequence. If you know which stage you're at, you know what's being measured there — and that alone cuts the tension in half, because you stop trying to prove everything, all the time.

StageWhoHow longWhat's actually being measured
Screening callRecruiter20–30 minYou're real, you're available, you fit the budget, you speak English. Almost nothing technical.
Technical interview 1Engineer from the team45–60 minFundamentals: Linux, containers, CI/CD, Kubernetes. Open questions, not a quiz.
Practical exerciseTake-home or live1–3 hoursYou write a pipeline, a Dockerfile, a manifest, or debug something broken on purpose.
System design interviewPrincipal engineer / architect60 min“Design the delivery for a service with X users.” What's measured is the reasoning, not the answer.
Behavioral interviewManager45 minHow you work with people, how you react under pressure, how you handled a conflict.
FinalDirector / client30 minMotivation, fit, sometimes an alignment discussion with the end client.
💡 What this changes for you

At the first call you're not being evaluated technically. Many people burn themselves out emotionally there, trying to prove competence to a recruiter who's actually just checking whether you're available and whether you can hold a conversation in English. Treat the first stage as an administrative conversation. Save your energy for the second stage.

1.1 Outsourcing firm versus product company

At outsourcing firms — Luxoft, GlobalLogic, Nagarro, Endava — there's often one more round with the end client, and the first internal round is gentler. At product companies, the technical interview is usually deeper and comes earlier. At both, the question that decides things most often is “tell me about an incident”, and there you have more material than a candidate with half your seniority.

2 Anti-freeze: what to do when you get stuck

The freeze doesn't come from a lack of knowledge. It comes from something very concrete: you start talking before you know where the sentence is going, you realize halfway through that you don't have an ending, and then your mind starts looking for the exit instead of continuing. That's where the sentences left hanging, the filler words, and the feeling of making a fool of yourself come from.

The solution isn't to know more. It's to not start the sentence until you have its ending. Below are the techniques, in the order you use them.

2.1 The two-second rule

🗣️ What you say before any answer

“Good question. Let me think about that for a second.”

It sounds trivial and it works for three reasons. It gives you time to find the end of the sentence. It signals that you're taking the question seriously, which reads as professional maturity, not hesitation. And it breaks your reflex to start talking out of panic. Two seconds of silence in an interview feel long to you and go unnoticed by the other person. No one has ever been rejected for a two-second pause.

2.2 Announce the structure before you get into it

Instead of jumping straight into the content, you first say how many parts the answer has. “There are three things here” or “first I'll tell you what it is, then where I used it.” The effect is twofold: the listener follows you more easily, and you now have a skeleton. When you have a skeleton, you can't stay hanging in the air — you know part two is coming even if part one came out limping.

2.3 When you get lost halfway through

🗣️ Three phrases that fix any botched sentence

“Wait, let me rephrase that, so it's clearer.”

“The main point is that…” (you go straight back to the conclusion and drop the path there)

“Let me back up a little — I started too far into the detail.”

All three are things senior engineers do in real meetings, every day. They're not signs of weakness, they're signs that you're listening to your own answer. What reads badly is something else: continuing a sentence that has obviously gone astray, or stopping abruptly and going silent.

2.4 When you don't know the answer

This is the situation you fear the most, and it's actually the easiest one — because it has a scripted answer you can learn by heart. It has three parts, in this order:

🗣️ The full formula

“I haven't worked directly with that. What I know is [the part you do have a handle on]. How I'd find out is I'd check [the documentation / such-and-such log / I'd ask someone on the team]. What I've done that's similar is [the thing from your experience that resembles it].”

Notice that the answer doesn't end with “I don't know.” It ends with something you did. A good interviewer notices exactly that: the candidate acknowledges the limit, has a method for getting past it, and has a real point of support. A candidate who makes something up, on the other hand, gets caught at the second question — and from there, everything they said before gets called into question.

🛑 The trap that would cost you the most

You already have a concrete example: the optimization at DISH, with 40–50% better resource efficiency, whose exact mechanism you no longer remember. Don't reconstruct it in the interview. Any plausible-sounding explanation you make up leads straight to the next question — “how did you measure that?” — where you have nowhere left to go. The prepared answer is in chapter 5, under the Kubernetes questions, and it's short, honest, and entirely acceptable.

2.5 The physical technique, for the first few minutes

  • 4–6 breathing. Inhale for four seconds, exhale for six. Three times, right before you go in. A longer exhale than inhale lowers your pulse — it's not a suggestion, it's a physiological reflex.
  • Speak more slowly than feels natural. Under tension, everyone speeds up. If you feel like you're speaking too slowly, you're probably at a normal pace.
  • Drink water. Keep a glass next to you. A sip is a legitimate three-second pause, any time you need one.
  • Write down the question on paper as it's being asked, if it's online. It anchors your attention and leaves you, in writing, a visible reminder of what you need to answer — this is where half of all sentence-wandering comes from.
  • Stand up, if it's a video call and you have room. Your voice sounds more confident and you breathe better.
💡 The reframe that changes the most

You're not a candidate who's asking for something. You're an engineer with over twenty years in production who's checking whether this job is worth his time. You're both checking something. The difference isn't rhetorical: it changes your posture, the pace of your voice, and the kind of questions you ask. And if it goes badly, you've lost forty minutes — that's it. The rest is imagination.

3 The structure of any answer

3.1 For technical questions: three layers

1 · The definition, brief One or two sentences. Don't recite, explain. “Readiness means I can receive traffic.” 2 · Why it matters What breaks without it. This is where it shows you understand. “Without it, traffic gets routed into pods that aren't up.” 3 · Where you've seen it An example from your work. This is what separates you from theory. If you don't have one, say where you've seen it done wrong. Three layers ≈ 45 seconds. You stop and leave room for the next question.
The structure still works when you don't have the third layer: “I haven't run into this case in production, but I've seen what happens when it's missing” is still a complete answer.
⚠️ Don't go past a minute without stopping

The three-minute monologue is the most common mistake experienced candidates make. It reads as insecurity, not mastery — it looks like you're trying to cover the question until you hit whatever they wanted. Answer in forty to sixty seconds, stop, and let the person dig wherever they want. If they wanted more, they'll ask. The silence after a short answer is theirs to fill, not yours.

3.2 For experience questions: STAR

S — Situation

Where, when, what system, who was affected. Two sentences. The context, not your life story.

T — Task

What needed to be achieved and what was your responsibility, specifically.

A — Action

The longest part. What you did, in the first person singular. “I did,” not “it was done” and not “the team decided.”

R — Result

How it ended and what you learned. With numbers only if you have them verified.

⚠️ “We” erases your contribution

The most costly habit in interviews at large companies: the candidate tells the story in the first person plural, out of politeness, and the interviewer is left not knowing what you actually did. Say “I migrated,” “I noticed,” “I decided.” If it was teamwork, say so once, at the end: “there were three of us, my part was X.” Not the other way around.

4 Your stories, already written

A senior technical interview is won with four or five stories you know so well you can tell them without thinking about the structure. Below are your stories, built only on what's documented in your CV. Each one has, at the end, what you should add from memory to make it come alive — and if you don't remember, the story still stands without those details. What you never do: you never fill the gap with something plausible.

4.1 The WebSphere MQ migration at Volvo Cars — the baseline story

🗣️ Full version

S. “At Volvo Cars, through IBM, I worked on the WebSphere MQ messaging platform. It was the layer that carried messages between applications, set up across multiple instances, specifically so there'd be no single point of failure.”

T. “It had to be upgraded from version 8 to version 9. In a messaging system you can't afford to lose messages and you can't afford to stop the flow, because business processes behind it grind to a halt.”

A. “I prepared the migration first in the test environment, checked the configuration compatibility, and worked out the exact order for moving the instances, with a rollback path for every step. I did the migration instance by instance, not all at once, so the service stayed available on the other one. I watched it through monitoring, with Nagios, so I'd see right away if anything looked off.”

R. “The platform moved to the new version without losing any messages. What stayed with me from that is the discipline of never starting a migration without a rollback path that's prepared and tested — not just described in a document.”

⚠️ What's worth adding, if you remember

How long the work window was, how many instances there were, whether you hit anything unexpected and how you resolved it. If you don't remember, don't make it up — the story above is complete and holds up under questions, because it doesn't contain anything you can't back up.

4.2 24/7 production at Metro — the incident story

🗣️ Full version

S. “At Metro I worked in production control, round the clock, on shifts and on-call. We watched over systems that had to run continuously, and when something went down, we were the first to know.”

T. “My job was to catch the problem early, assess how serious it was, restore the service, and keep the client informed for as long as it was open.”

A. “The order I followed was always the same: restore first, find the cause after. I take ownership of the incident explicitly, so two people don't end up working the same thing in parallel without knowing it. I check what changed recently, because in most cases that's where the cause is. I escalate based on severity, not on how loud it sounds. And I give updates on my own initiative, at intervals, so nobody has to ask how it's going.”

R. “Those years taught me something you don't learn from documentation: under pressure, what matters most isn't speed, it's order. If you have an order to follow, you don't panic, and if you don't panic, you don't make a second mistake on top of the first.”

💡 Why this is your most valuable story

A thirty-year-old candidate who's read about incident management can't say that closing line. You can say it because you lived it. When you get here, slow your pace down — this is the point in the interview where you have the most advantage and the least to prove.

4.3 DISH Digital Solutions — the Kubernetes and cloud story

🗣️ Full version

S. “At DISH Digital Solutions I worked on Google Cloud, on managed Kubernetes — GKE — with 17 pods and 44 Cloud SQL databases behind it.”

T. “My part was operations: making sure deployments went out without taking the service down, the applications stayed healthy, and resources matched what the workload actually needed.”

A. “I did the deployments progressively, bringing up new pods before the old ones came down, with checking the rollout status as a mandatory step — if the rollout doesn't complete, the deployment has to fail, not go through silently. On the resource side, I did a configuration optimization on the allocation.”

R. “The deployments ran without service interruption, and the optimization gave between 40 and 50 percent better resource efficiency, with no impact on performance.”

🛑 If they ask exactly what you changed

The answer, exactly like this:

“It was a configuration optimization on resource allocation, and I'd rather not reconstruct the exact parameters from memory — it was a while ago and I don't have the records in front of me. What I can tell you is the shape of it: the configuration didn't match the workload's real profile, and fixing it gave between 40 and 50 percent better efficiency, with no impact on performance. If it's useful, I can tell you how I'd size it today.”

The last sentence matters: it moves the discussion from memory to competence, which is exactly where you're strong. From there you move into requests, limits, throttling and OOMKilled — chapter 12 on the technical page — and you answer with solid material.

4.4 RapidBid — the infrastructure-as-code story

🗣️ Full version

S. “At RapidBid I worked on infrastructure written as code, in Terraform, across two cloud providers in parallel — AWS and Azure.”

T. “The infrastructure had to be reproducible and the deployments predictable, and both environments had to look the same no matter who stood them up.”

A. “I wrote the resources in Terraform and kept the delivery pipelines in Azure DevOps and Jenkins. One of the deployments was blue-green: I'd stand up the new environment completely, verify it, and only then switch the traffic over to it. The advantage is that rolling back just means switching the traffic back, not a new deployment under pressure.”

R. “What stayed with me from that is that the difference between two cloud providers is much smaller than it looks from the outside. The providers change, and the way you authenticate; the plan–apply flow, the state discipline and the way you think about modules stay identical.”

💡 The question that comes after blue-green

They'll almost certainly ask: “and what do you do about the database?” The answer: “The database doesn't get duplicated — it stays shared, and schema changes have to be backward and forward compatible, so both the old and new version can work on it at the same time. In practice, you add columns, you don't rename or drop them in the same deployment. Dropping comes in a later deployment, once nothing reads from them anymore.” That's a sentence that immediately separates someone who's done it from someone who's read about it.

4.5 Etihad — the dependencies and inventory story

🗣️ Short version

“At Etihad, through IBM, I worked with TADDM — automatic infrastructure discovery and dependency mapping between applications. Basically, you built the picture of what talks to what. Sounds boring until the day you need to shut down a server and nobody knows what depends on it. The idea is still the same today, it's just called something else now: configuration inventory and service maps.”

4.6 The last few years, on your own — the story you're afraid of

They'll ask you what you did during the period when you weren't at a corporation. It's not a trick question, it's a question to fill in the picture. The answer needs to be short, flat, in the past tense, with no justifications and no forced enthusiasm.

🗣️ Full version

“I worked on my own. I built and kept my own systems running — servers, automations, integrations with language models. I did every role, from infrastructure to what the customer sees. It was useful because I saw the whole chain, but I missed what I now intend to get back: a team, a bigger scale, and serious procedures. That's why I want back into a multinational.”

Four sentences. You stop. You don't explain why you left back then, you don't apologize, you don't add any vision statements. If they want more, they'll ask.

5 Technical Questions, With Answers

These are written as spoken answers, not textbook definitions. Read them out loud; if a sentence sounds unnatural in your mouth, rewrite it in your own words — a badly memorized answer sounds worse than one said simply. Click the question to open it.

5.1 Fundamentals

What is DevOps, in your own words?

“It's the way of working where the people who write the application and the people who keep it running work toward the same goal, instead of handing off responsibility. In practice it means delivery, infrastructure and monitoring are automated and versioned, so anyone on the team can deliver safely, and when something breaks you know right away what changed.”

What you don't say: “it's a culture.” It's true and it says nothing. Give an operational definition.

What does a normal day in DevOps look like for you?

“In the morning I check what happened overnight — alerts, deployments, failed jobs. During the day there are three kinds of work: building automation, helping developers when something gets stuck in the pipeline, and dealing with what's broken. On top of that comes planned work — version upgrades, cleanup, cost reduction. When there's an incident, everything else stops.”

What is “shift left”?

“It means the checks move as early as possible in the process. Tests, security analysis, configuration checks happen on every change, not in an audit at the end. The reason is simple: a problem caught in the editor costs minutes, the same problem caught in production costs days and, sometimes, customers.”

How do you explain to a manager why investing in delivery automation is worth it?

“Through four numbers you can measure: how often we deploy, how long it takes from commit to production, how often a deployment causes a problem, and how long it takes us to recover. Automation improves all four at once, and the last one is the one that matters to the customer. Counterintuitively, teams that deploy often are also the most stable ones — because they deploy small changes that are easy to understand and easy to revert.”

What is SRE and how is it different from DevOps?

“DevOps is the way of working. SRE is a concrete implementation of it, focused on measured reliability: you define targets, you measure, and you have an error budget that decides whether the team ships more aggressively or stops to fix things. In practice, at most companies in Romania, the title is DevOps and the job covers both.”

What's the most important quality in a DevOps engineer?

“Not assuming. Most long incidents drag on not because the problem was hard, but because someone started from an assumption they never checked. The discipline of verifying before acting is what separates a ten-minute incident from a three-hour one.”

5.2 CI/CD and GitHub Actions

The difference between continuous delivery and continuous deployment.

“With continuous delivery, any change that passed the tests is ready to ship, but someone presses the button. With continuous deployment there's no button — if everything passed, it reaches production on its own. The difference is one person, and the choice depends on how well you've covered the risk with automated tests and the ability to roll back quickly.”

What steps does a delivery pipeline have?

“I pull the code, install the dependencies, check style and run static analysis, run the unit tests, build the artifact exactly once, scan it for vulnerabilities, push it to a registry, deploy it to the test environment, run the integration tests, wait for approval if needed, deploy to production and verify the deployment's status. The rule I stick to is that the artifact is built once and the same one goes through every environment — the only thing that differs is the configuration.”

Why “build once, deploy everywhere”?

“Because if you rebuild for every environment, what you test isn't what you ship. A dependency can resolve differently, a cache can be different, and you end up in the classic situation where it works in test and doesn't work in production. I build once, push to the registry, and deploy from there everywhere.”

The structure of a GitHub Actions workflow.

“A workflow is a YAML file in .github/workflows, triggered by an event. It contains jobs; each job runs on a runner, on its own clean machine, and has steps. A step is either a command or a reusable action. Jobs run in parallel unless you say otherwise; with needs you put them in order.”

How do you pass files from one job to another?

“Through artifacts — upload-artifact in one job, download-artifact in the other. It's a common trap for people coming from Jenkins, where stages share the same workspace. In Actions, each job starts on a fresh machine, so nothing is passed implicitly. Small values are passed through outputs.”

How do you handle secrets?

“On three levels: organization-level for what's shared, repository-level for what's specific, and environment-level for production — there I can also require human approval before the job starts. But if I'm going toward a cloud provider, I'd rather not have secrets at all and use OIDC.”

What is OIDC and why is it better than a key?

“Instead of keeping a permanent key in the repository's secrets, the workflow receives an identity token signed by GitHub, and the cloud provider accepts it based on a trust relationship configured once. The token lives for a few minutes and is tied to the repository and branch that requested it. There's nothing left to leak, nothing left to rotate. The detail that matters: the trust condition is written against identifiers that can't change, not against the repository name — the name can be renamed and taken over by someone else.”

How do you secure third-party actions?

“I pin them to the full commit hash, not to a tag, because a tag can be moved. Then I let Dependabot propose the updates, so I don't stay on an old version. And I give the workflow minimal permissions — read-only by default, and I add write access only where it's actually needed.”

Why is pull_request_target dangerous?

“Because it runs with the repository's secrets, but in the context of a pull request that can come from anyone. If you check out the PR's code inside it and execute it, you've just given a stranger access to your secrets. It's only used for things that don't touch the PR's code — for example, adding labels or comments.”

Have you worked with GitHub Actions?

“I built my pipelines in Azure DevOps and in Jenkins, including a blue-green deployment. The model is the same — trigger, stages, artifact, approval, deploy — and it translates directly: what's agent in Jenkins is runs-on in Actions; what's stages become jobs, with the important difference that in Actions each job starts clean, so artifacts have to be passed explicitly. I've worked with Actions at the level of being able to read and modify a workflow; what I haven't done is roll them out across a whole organization, and that's something I'd learn along the way.”

Honest phrasing. Doesn't claim more than you can back up on the follow-up question, but doesn't throw away real experience either.

What do you do when a pipeline fails intermittently?

“First I check whether the failure is in the test or in the infrastructure — a network call with no timeout and no retry is the most common cause. If it's a flaky test, I isolate it and flag it, because a test nobody trusts anymore is worse than no test at all. I never let the intermittent failure just “pass,” because that trains the team to ignore red.”

5.3 Kubernetes

What is Kubernetes and what problem does it solve?

“It's a system that keeps containers running across a group of machines. You describe the desired state — I want three instances of this application, with this configuration — and it constantly compares reality to what you asked for and corrects the difference. If a container dies, it restarts it. If a node dies, it moves the workload to another one. That's the whole thing; everything else is detail on top of that idea.”

What is a pod?

“It's the smallest unit that Kubernetes schedules. Usually one container, but it can have several that share the same network address and the same volumes. Important: a pod is disposable. It doesn't get repaired, it gets replaced. It has an IP address that disappears with it, which is why traffic never goes directly to a pod, only to a Service.”

Deployment, ReplicaSet, Pod — what's the relationship?

“The Deployment is what I write. It creates a ReplicaSet, which makes sure the requested number of pods exists. On every image change, the Deployment creates a new ReplicaSet and scales it up gradually while scaling the old one down — that's where the rolling update comes from. The old ReplicaSet stays around, empty, so I can roll back to it.”

The types of Service.

“ClusterIP is the default and is only visible inside the cluster. NodePort opens the same port on every node. LoadBalancer requests a load balancer from the cloud provider. ExternalName is just a DNS alias. In practice I use ClusterIP everywhere with a single Ingress in front, because a LoadBalancer per service costs money and adds nothing.”

The Service isn't sending traffic to the pods. What do you check?

“The first command is kubectl get endpoints. If the list is empty, the Service isn't finding any pod — and nine times out of ten it's a mismatch between the pod's label and the service's selector, one extra letter somewhere. If the list has pods but it's still not working, then it's the port: targetPort has to be the port the application is actually listening on, not the one the service exposes.”

Ingress vs. Service.

“The Service works at the network level and routes traffic to a set of pods. The Ingress works at the HTTP level: it routes by domain and by path, terminates TLS, and all the services come in through a single door. Ingress on its own does nothing — you need a controller installed that reads the rules and applies them.”

ConfigMap vs. Secret.

“Both bring configuration into the pod, as environment variables or as mounted files. The declared difference is that Secret is for sensitive data. The real difference is much smaller than it sounds: by default, a Secret is only encoded in base64, not encrypted. Anyone with read access to it reads it in plain text. That's why, for anything serious, you either turn on encryption in etcd or keep the secret outside — Vault, External Secrets Operator, Sealed Secrets.”

The line about base64 is one of the few that, said spontaneously, immediately shows you've actually worked with it.

Deployment vs. StatefulSet.

“Deployment is for stateless applications: the pods are interchangeable, they have random names, they start and stop in any order. StatefulSet is for anything with an identity: stable, numbered names, its own disk that stays its own, ordered startup and shutdown. Databases, queues, anything with local data. The practical difference people often get tripped up on: if a node becomes unreachable, Deployment pods get recreated somewhere else right away, whereas a StatefulSet pod doesn't get recreated automatically — because the system can't be sure the old one actually died, and two instances with the same identity would corrupt the data.”

Readiness, liveness, startup.

“Readiness means “I can receive traffic” — if it fails, the pod is taken out of service, but it doesn't restart. Liveness means “I'm alive” — if it fails, the container restarts. Startup is for applications that start up slowly, and it holds off the other two until startup finishes. The mistake I've seen is a liveness probe pointed at an endpoint that checks the database: the database has a ten-second hiccup, all the pods restart at once, and a slowdown turns into a total outage.”

Requests and limits.

“Request is how much you reserve, and it's what decides which node the pod lands on. Limit is the ceiling at runtime. The asymmetry matters: on CPU, if you go over the limit, you get throttled — the application keeps running, just badly, and it's hard to diagnose. On memory there's no throttling. You go over and you get killed — OOMKilled, exit code 137.”

A pod is stuck in Pending.

“The scheduler has nowhere to put it. Most often there aren't enough free resources for what it's requesting, but it could also be a taint on the nodes with no matching toleration on the pod, a node selector that doesn't match, or a volume that can't attach in that zone. kubectl describe pod writes the exact reason in the events, at the bottom.”

CrashLoopBackOff.

“The container starts, dies, and Kubernetes restarts it with longer and longer pauses. The first command is kubectl logs --previous, to see why the previous instance died, not the one that just started. Then the exit code: 137 means it was force-killed, usually memory; 1 or 2 is usually an application error or a missing config or secret.”

The pod is Running but READY shows 0/1.

“The container is running, but the readiness probe isn't passing. Either the application is started and not ready yet, or the probe is configured wrong — wrong path, wrong port, or too short a timeout for how long startup takes. I check what the probe says in describe and try the endpoint from inside the pod.”

How do you do a zero-downtime update?

“Rolling update with maxUnavailable: 0 and maxSurge: 1 — meaning I bring up a new pod before taking down an old one, so capacity never drops. I need a correct readiness probe, otherwise traffic goes into pods that aren't started yet, and graceful shutdown in the application, so in-flight requests finish. And in the pipeline I always put a kubectl rollout status with a timeout, because otherwise the update command returns success immediately and the deployment looks successful even if no new pod ever started.”

The update is stuck. What's happening and what do you do?

“Nothing bad is happening, and that's exactly the danger. If the new pods don't pass the readiness probe, the update simply stops: the old ones keep serving traffic, the new ones sit there not ready. There's no automatic rollback. I check with kubectl rollout status, look at the events and the new pod's logs, and if there's nothing to fix on the spot, I run kubectl rollout undo.”

How do you scale?

“Horizontally, with Horizontal Pod Autoscaler, on CPU or on a custom metric, with node autoscaling underneath so there's somewhere for it to fit. It only works if the application is stateless. The trap is scaling on CPU for a service that isn't CPU-bound — if the bottleneck is the number of database connections or the length of a queue, scaling on CPU either doesn't trigger, or it triggers and makes things worse. There, you scale on the real metric.”

Is a namespace a security boundary?

“No. It's a logical division — names, resource quotas, permissions. But network traffic flows freely between namespaces unless you set a NetworkPolicy, and some things are cluster-level anyway. This confusion is common and leads to clusters where the test environment can talk to production.”

5.4 Docker and Containers

Container versus virtual machine.

“A virtual machine has its own kernel and its own operating system, so it boots in minutes and takes up gigabytes. A container is an ordinary process on the host machine, isolated through kernel mechanisms, sharing the kernel with everything else. It starts in milliseconds and takes up as much as the application. The practical consequence: isolation is weaker than with a virtual machine, which is why you don't put workloads with very different trust levels into the same cluster.”

Image versus container.

“The image is the template — read-only, immutable layers. The container is the image started up, with a writable layer on top. From one image you can start a hundred containers. What you write inside the container disappears when you stop it, which is why data goes on a volume and logs go to standard output.”

How do you reduce image build time?

“I order the instructions from the ones that change rarely to the ones that change often. I copy the dependency file first and install them, and only then copy the code — that way the dependency layer stays cached and doesn't get reinstalled on every code change. The opposite, meaning COPY . . before installing, invalidates everything on every commit, and it's the most common mistake.”

What is a multi-stage build?

“In a first stage I use a large image, with a compiler and tools, and build the application. In the second stage I start from a minimal image and copy over only the output from the first one. It's a double win: the final image is tens of times smaller, and I don't ship compilers and tools into production, which are, in effect, attack surface.”

How do you secure an image?

“A minimal base image, ideally distroless, pinned to a version, not to latest. Running as a non-privileged user, not root. Read-only filesystem where possible. No secrets at build time — if one made it into a layer, it stays recoverable even if you delete it in a later layer. And automated scanning in the pipeline, with Trivy or equivalent.”

Why don't you use the latest tag?

“Because it's not a version, it's a pointer that moves. Two identical deploys can end up running different images, and when something breaks you can no longer tell what was running an hour ago. I use the tag with the commit hash: it's unique, immutable, and tells me exactly what code is inside.”

Where do you write logs from inside a container?

“To standard output, never to a file inside the container. The file disappears along with the container, fills up the node's disk, and nobody sees it. On standard output, the platform picks it up and it ends up where it needs to.”

5.5 Linux and Scripting

A server isn't responding. What do you check first?

“Disk space. df -h. It's trivial and it's the most common cause — a full disk stops almost everything, and often in ways that seem unrelated: the database stops writing, logs stop, processes hang. It takes a second and rules out the most likely culprit. After that I check whether the process is running, whether it's listening on the port, what the log says, and what changed recently.”

How do you see what's listening on a port?

“ss -ltnp — shows the listening ports, with the process. lsof -i :443 does the same thing a different way. If nothing is listening, the problem is in the application or the service, not the network, and I go into journalctl -u.”

The disk is full. How do you find what's taking up the space?

“du -sh /* | sort -h, and I work down directory by directory. But first I check something that catches a lot of people out: a deleted file that a process still has open doesn't free up the space. lsof | grep deleted shows it, and it's fixed by restarting the process, not by deleting something else.”

What does set -euo pipefail do?

“-e stops the script at the first failed command. -u errors out on an undefined variable, which catches typos before they do damage. -o pipefail makes a pipeline fail if any element in it fails, not just the last one. Without them, a script that fails halfway through keeps going and can leave the system in a worse state than if you hadn't run it at all.”

Why do you add -f to curl in scripts?

“Because without it, curl returns a success code even on a 500 response — it successfully downloaded the error page. With -f it fails on HTTP error codes, so set -e actually stops the script. I also add --max-time, otherwise a call can hang indefinitely inside a cron job.”

What does it mean for a script to be idempotent?

“That running it twice gives the same result as running it once. It matters because in operations you never have a guarantee that something ran exactly once — a retry, a job started twice, someone clicking again. A script that appends a line to a config file every time it runs is a ticking bomb.”

How do you make sure two instances don't run at the same time?

“flock on a file. It's one line and it completely solves the problem of cron jobs overlapping when one run takes longer than the interval.”

5.6 Terraform and Cloud

What is infrastructure as code?

“It means infrastructure is described in files kept in Git, not created by hand in the console. You gain three things: you can rebuild the environment identically, you can see in the history who changed what and why, and you can review a change before it happens. The fundamental difference from a script is that Terraform is declarative — you describe what you want to exist, not the steps.”

What is state in Terraform and why does it matter?

“It's the record of what Terraform created and how it maps to the real resources. Without it, Terraform can't know what to change and what to delete. Three rules: it lives remotely, not on a laptop, so the whole team works off the same one; it has locking, otherwise two simultaneous runs corrupt it; and it contains sensitive values in plain text, so it's treated as a secret.”

What is drift?

“When someone changes something directly from the console, reality no longer matches the code. terraform plan shows it as a diff. What I do in practice is remove the cause: write access only from the pipeline, people get read access. If drift is already there, I either bring it into the code or let Terraform correct it — but that decision gets made consciously, not through a rushed apply.”

Have you ever run terraform apply and broken something?

“The rule I hold myself to is that I don't run apply without reading the plan all the way through, and I look first at anything marked “destroy” or “replace”. A seemingly harmless rename can mean, in the plan, destroying and recreating a resource. On resources with data I set prevent_destroy, so the system refuses even if the person makes a mistake.”

Have you worked with multiple cloud providers?

“Yes — Terraform on AWS and Azure, in parallel, at RapidBid, plus Google Cloud at DISH, on GKE and Cloud SQL. What I took away is that the difference is smaller than it looks from outside: the providers change, the service names and the authentication method change, but the plan–apply flow, the state discipline and the way you split things into modules stay the same. The services map almost one to one.”

How do you reduce cost in a cloud?

“The biggest win, almost always, comes from resources requested far above what's actually used — you pay for the reservation, not the consumption. After that: test environments shut down at night and on weekends, disks and IP addresses left orphaned after deletions, traffic between availability zones that costs money and can often be avoided. And tagging resources, without which you can't even tell who's spending what.”

5.7 Monitoring and Incidents

What do you monitor for a service?

“The four golden signals: latency, traffic, errors, saturation. For latency I look at p95 and p99, not the average — the average hides exactly the affected users. You can have an average of 200 milliseconds and a percentage of users waiting eight seconds, and those are the ones who call in.”

SLI, SLO, SLA.

“SLI is what you measure — for example, the percentage of requests under 300 milliseconds. SLO is the internal target on that indicator. SLA is the contractual promise to the customer, always more relaxed than the SLO, so there's margin. The SLO gives you the error budget: how much you're allowed to get wrong. If there's still budget left, the team ships aggressively; if it's used up, they stop and fix things. The measurement makes the decision, not whoever talks loudest in the meeting.”

How do you avoid alert fatigue?

“I alert on symptom, not on cause — the error rate, not CPU at 90%, which can be perfectly normal. Every alert that goes off at night has to require immediate human action; if it can wait until morning, it's a ticket, not an alert. And every alert has a runbook. When you get forty alerts a day and thirty-nine of them need nothing, you miss the fortieth.”

The service is down. What do you do in the first few minutes?

“First I restore, then I look for the cause. I confirm that I'm taking ownership of the incident, so two people aren't working the same thing. I check what changed recently — a deploy, a config change, a certificate — because that's where the cause is in most cases. If it's a recent deploy, I roll back to the previous version without digging further; I do the diagnosis after people have the service back. And I communicate proactively, at intervals, so nobody has to interrupt the team to ask how it's going.”

This is your home turf. Tie it immediately to Metro: “this is exactly the order I used for years in production control.”

What is a blameless postmortem and why does it matter?

“It's the post-incident analysis where you look for the cause in the system, not in the person. If a wrong command was able to wipe out production, the problem isn't the person who typed it, it's that the system accepted it without confirmation and without a way back. It matters for a very practical reason: if names show up in the analysis, next time people hide incidents, and you lose exactly the information you need.”

MTTR, and why it matters more than MTBF.

“MTTR is the average time to restore, MTBF is the average time between failures. In distributed systems, something is always failing — you can't push MTBF to infinity. What you can control is how fast you recover. That's why you invest in fast recovery, in small deploys, and in the ability to roll back a change in a few minutes, not in the promise that nothing will ever fail again.”

💡 If you only have time for one thing

Learn ten answers well, not fifty superficially. The ten: what DevOps is, what Kubernetes is, pod and Deployment, Service and Ingress, ConfigMap and Secret, readiness and liveness, requests and limits, CrashLoopBackOff, rolling update with rollback, and “the service is down, what do you do”. Say those ten fluently and you'll pass most first-round technical interviews.

6 Behavioral questions

These are underestimated, and they decide more often than the technical ones, especially at large companies. They all share the same hidden structure: the interviewer wants to see how you behave when things go wrong. Answer in STAR, first person singular, in under two minutes.

Tell me about an incident you handled.

Use the Metro story from chapter 4.2. Don't improvise a different one — this one is rehearsed and holds up under questioning, because it's yours.

Tell me about a mistake you made in production.

This question isn't looking for the mistake, it's looking for whether you're able to admit it. A candidate who says “I don't remember ever making a mistake” disqualifies themselves. Structure: what I did, what broke, what I did right away, what I changed so it couldn't happen again. The emphasis falls on the last part — that's where maturity shows.

⚠️ This one you have to prepare yourself

This is the only answer in the whole document that I can't write for you, because I don't know what happened to you. Pick a real case, preferably an old one that's resolved, where you changed something afterward. Write it in four sentences and rehearse it out loud. If you walk into the interview without it prepared, you'll improvise — and this is the one place where improvising shows.

Have you ever had a disagreement with a developer or a colleague?

What's being tested is whether you can disagree without turning it into conflict. The structure that works: what the disagreement was, what my argument was, what their argument was, how it got decided, and what I did after it was decided differently than I wanted. The last part is what's being evaluated. “I said what I thought, it was decided differently, I supported the decision, and I put a safeguard in place for the risk I was worried about” is a very good answer.

How do you prioritize when you have several urgent things at the same time?

“By impact on the customer, not by who asked loudest. If something affects production, it takes absolute priority and everything else waits — and I tell the others it's waiting, I don't let them think I'm working on it. My years in production control taught me that communicating what you're not doing is just as important as doing.”

How do you stay current?

Answer concretely, not with “I read a lot.” What exactly, how often, and the last thing you learned and actually used. One concrete example is worth ten general statements.

How do you work in a distributed team?

“I write down what I do, so nobody depends on me being present on a chat channel. I document what I broke and what I fixed. And I flag early when something is taking longer than I estimated — a delay said in time is information, said late it's a problem.”

Where do you see yourself in three years?

Don't invent a spectacular career plan. “Still on infrastructure and reliability, with more depth in Kubernetes and automation, on a team where I can also pass on what I know. I don't have management ambitions.” It's short, credible, and doesn't expose you.

7 Motivation and uncomfortable questions

This is where people lose interviews they'd already won technically. The rule for all the answers below: flat, in the past tense, no long justifications, no vision statements. A long explanation sounds like an excuse. Answer in three to four sentences and stop talking.

Why do you want to go back to a large company?

“Because I miss what I had there: a team, a larger scale, and serious procedures. On my own I've done every role and learned a lot, but I work alone. I want to go back to an environment where the systems are large and I can specialize instead of covering everything.”

Why did you leave IBM?

Answer briefly and factually, with no criticism of them at all. Don't blame anyone, don't tell the internal backstory, don't explain it three times. One sentence about what happened, one about what you did next. Anything you add beyond that works against you.

Why do you have a gap from the corporate world?

“I worked independently during that time. I built and kept my own systems running — infrastructure, automations, integrations. It wasn't a break from work, it was a different kind of work. What I missed was the team and the scale, and that's why I'm coming back.”

You have over twenty years of experience. Aren't you overqualified?

“I don't think I am. The operations, incidents, and large-systems side I have solid from the earlier years. The modern-tooling side — Kubernetes, infrastructure as code, pipelines — I have, and I want to go deeper into it. I'm interested in this role, not a management one, and technical work doesn't bore me.”

The hidden question is: “will you leave quickly?” or “will you be hard to manage?” The answer needs to put exactly those two at ease.

What are your salary expectations?

Ask first: “Do you have a range allocated for the role? I'd rather we start from there, so neither of us wastes time.” If they insist you give a number first, give a range, not a figure, and add “depending on the full package.” Don't apologize for the number and don't explain why you need it — personal reasons have no place in a negotiation.

⚠️ What we've verified about money

Of the five open listings verified on September 23, 2026, none show a salary. The only verified figure comes from a listing that's already closed, Ketryx Vienna, with a minimum of €90,000 gross per year for someone with over five years of experience — but that's Austria and it's not a benchmark for Romania. In other words: we don't have a verified benchmark for the local market, so don't go into the discussion with a figure pulled from somewhere. Ask them first.

What don't you know that you should know for this role?

It's a good question and rarely asked, and an honest answer impresses. “I've operated Kubernetes, but I've written fewer manifests than I'd like. And I haven't taken a whole organization onto GitHub Actions. Both are a few weeks of actual work away, and I've made that kind of transition before, when I moved onto Google Cloud.” Name exactly two things, not ten, and don't pick something essential to the job.

Do you have other hiring processes going on?

Tell the truth, briefly, no names. “Yes, I have a few others in discussion, but none finalized.” It's not a threat and it's not a weakness — it's normal information, and it positions you as an in-demand professional.

8 What you ask them

At the end you'll be asked “do you have any questions for us?” A “no, I think you've told me everything” is the cheapest way to lose points in the whole interview: it reads as disinterest. Have two or three questions prepared, written down on paper, and don't be embarrassed to look at it.

About the work, concretely

  • What does a typical week look like for the person taking this role?
  • What happens from commit to production — how many steps, and how long does it take?
  • Who's on call, and how is on-call organized?
  • How many environments do you have, and who can change them?

About the system

  • Is the cluster managed by a vendor or self-managed?
  • What hurts most right now in your infrastructure?
  • How much of the infrastructure is in code, and how much is still done manually?
  • What does a post-incident review look like on your team?

About the team

  • How many people are on the team, and how is the work divided?
  • Is the role new, or is it replacing someone?
  • How does a technical decision get made — who has the final say?

About the process

  • What's the next step, and in how long?
  • What would someone need to demonstrate in the first three months to be considered a success?
💡 The question that changes the tone of the conversation

“What hurts most right now in your infrastructure?” Almost nobody asks it, and almost everybody answers honestly. From the answer you find out whether the job is what the listing says, and you can immediately connect it: “I had something similar at…” From there it stops being an interview and becomes a conversation between two engineers — which is exactly the state where you stop freezing up.

⚠️ What you don't ask in the first round

Vacation days, flexible schedule, when you can work from home, benefits, how fast you can get promoted. They're all legitimate questions — but at the offer round, not the first technical discussion. In the first round you ask about the work.

9 The red list: what you don't say

🛑 The eight
  1. Don't invent numbers or mechanisms. If you don't remember exactly, say you don't remember. A made-up detail leads straight into the next question, where you have nowhere left to go — and from there, even what was true gets doubted.
  2. Don't badmouth a former employer, client, or colleague. No matter how justified you'd be. The person across from you has no way to check who was right, but they'll remember that you're the kind of person who talks like that.
  3. Don't say “we” when you're talking about what you did. It erases your contribution at the exact moment it was being measured.
  4. Don't just say “I don't know” and then go silent. Use the formula from chapter 2.4: what I do know, how I'd find out, what I've done that's similar.
  5. Don't talk for three minutes on a one-minute question. It reads as insecurity. Answer, stop, let them dig further.
  6. Don't apologize for your age, for the gap, or for what you don't know. State it, don't justify it. “I haven't worked with that” is a complete sentence; whatever you add after it, as an excuse, weakens you.
  7. Don't promise what you can't back up. “I learn fast” is an empty phrase unless it comes with an example where you actually learned something concrete, fast.
  8. Don't ask about money and benefits in the first round, unless they bring it up.
🗣️ What you say instead, when you feel out of your depth

“That part I haven't done. What I did do is [the concrete thing]. How would you solve it now?”

You admit it, anchor yourself in what you do have, and turn it back. It's rare for someone not to answer that last question — and while they're answering, you regroup and find out how they think.

10 What the market asks for, verified

⚠️ Read this first, so you know how much weight the numbers carry

The data below comes from a search done on September 23, 2026. Many matching ads were already closed by the time they were actually opened — Nagarro Bucharest, Allianz Technology Bucharest, Bitdefender twice, GlobalLogic Bucharest, Greenbone. That left five ads verified as open. Five is a signal, not a statistic, and the percentages below should be read as such.

The five: Luxoft (Senior DevOps Engineer, Bucharest), CommIT (DevOps AWS/GitHub Actions, remote from Latin America), GlobalLogic (Senior DevOps with Azure, remote), Deutsche Postbank Group (DevOps, remote from Romania), Bet On Talent (Senior DevOps, remote from Europe).

RequirementIn how many ads out of 5Do you have it?
Cloud platform (any)5✅ AWS, Azure, Google Cloud
CI/CD with GitHub Actions named explicitly4🟠 partial — see below
Terraform or other infrastructure as code4✅ Terraform on AWS and Azure
Docker3✅
Python3🟠 to confirm
Monitoring (Prometheus, Grafana, Datadog, New Relic)3🟠 Nagios, Tivoli — previous generation
Kubernetes explicit2✅ GKE
Bash explicit2✅
Linux administration explicit2✅ Linux and AIX
High availability, as a phrase1✅ queues across multiple instances, 24/7 production
ArgoCD or GitOps0—

What I take from this, concretely

  • GitHub Actions shows up almost everywhere, but almost nowhere alone. It's listed alongside Jenkins, GitLab CI and Azure DevOps. That confirms the answer “I built pipelines in Jenkins and Azure DevOps, the concepts are the same” is accepted on the market, not an excuse.
  • Terraform is asked for just as often as GitHub Actions, and you have it. Put it forward on your own initiative, even if the ad doesn't emphasize it. It's the best ratio between how much it's asked for and how much you have.
  • Python shows up in 3 out of 5, more often than Bash. If you have it at an operational scripting level, say so. If not, it's the next investment after Bash, not Kubernetes.
  • GitOps and ArgoCD don't show up at all in the open ads. Don't spend prep time on them.
  • Monitoring shows up in 3 out of 5 and is the place where you need one well-built sentence, the one from chapter 21 of the technical page: the model is the same, the collection differs.
💡 The open thread you haven't closed

At the time of the search, Luxoft had an open position for Senior DevOps Engineer in Bucharest — career.luxoft.com/jobs/senior-devops-engineer-26131 — requiring over eight years, AWS, Kubernetes and Terraform, and for CI/CD it accepted GitLab, Jenkins or GitHub Actions. It matches your current CV better than the ad you were preparing for. You have an open thread there that you haven't closed.

11 Rehearsal plan

This document doesn't help you if you read it. It helps you if you say it out loud. The difference between knowing an answer and being able to say it fluently under pressure is exactly the difference between a good interview and one where you freeze up — and the second one is trained, not understood.

SessionWhat you doHow long
1Read chapters 9–14 of the technical page. Don't memorize, just understand.90 min
2Say the ten essential answers from 5.7 out loud, without looking. Record yourself on your phone.30 min
3Listen to the recording. Note where you hesitated and where you left sentences unfinished. It's the unpleasant part and the most useful one.20 min
4Repeat the four stories from chapter 4, timed. Each under two minutes.40 min
5Write the answer about your mistake in production (6.2) and repeat it until it comes out naturally.30 min
6Open the technical questions from chapter 5 at random and answer on the spot, without reading the answer first.45 min
7Full simulation with someone, or alone with a timer: 45 minutes without a break, without looking at documents.45 min
8Reread only chapter 2 (anti-freeze) and chapter 9 (red list), the day before the interview.15 min
💡 Recording on your phone

It's the most effective tool in the whole plan and the one almost everyone skips. When you listen to yourself, you hear things you don't notice while you're speaking: where you speed up, where you repeat yourself, where the sentence breaks. Two listening sessions change more than ten hours of reading.

⚠️ What you don't do in the last few days

Don't start learning new technologies two days before. Don't read about ArgoCD, service mesh or other things that don't appear in the ad. Last-minute panic shows up as a sudden appetite for new material, and the effect is the opposite: it fragments what you already knew. In the last few days you consolidate, you don't add.

12 On the day of the interview

One hour before

  • Check the camera, microphone and connection. If it's on a program you haven't used before, open it now, not at zero hour.
  • Close everything that could ring or pop up on screen.
  • Put on paper, next to you: the three questions for them, the four story titles, and the formula from 2.4.
  • A glass of water.
  • Reread only chapter 2. Nothing else.

In the first two minutes

  • 4–6 breathing, three times, before you go in.
  • Speak slower than feels natural.
  • On the first question, use the two-second pause even if you know the answer — it sets your pace for everything else.
  • If the first question goes badly, don't keep analyzing it while you're answering the second. No one is rejected based on the first answer.

Along the way

  • Write down the question while it's being asked.
  • Announce the structure: “there are three things here”.
  • Answer in forty to sixty seconds and stop.
  • When you don't know, the formula from 2.4. Never improvise a mechanism.
  • “Can you repeat the question?” is allowed any time, as often as you need.

At the end

  • Ask the two to three prepared questions.
  • Ask what the next step is and how long it will take.
  • Thank them briefly. No closing speech.
  • Write down right after, while your memory is fresh, what they asked and where you stumbled. You prepare the next interview from that list.
🗣️ The last thing, and the only one to remember if you forget the rest

You don't have to be perfect. You have to be coherent.

No one hires someone who knows everything — that person doesn't exist. They hire someone they can work with: who says what they know, says what they don't know, finishes their sentences and doesn't make things up. You already have the raw material — over twenty years of real systems that had to work. All you have left to do is say it calmly.