DevOps interview questions that test incident behaviour, not tool lists

Tool lists are the least useful thing on a DevOps CV — the tools change every two years and anyone competent picks up a new one in weeks. What doesn't change is how someone behaves during an incident: whether they communicate while firefighting, whether they reach for a rollback before a clever fix, and whether their post-mortem looks for a system failure or a person to blame. Those three predict your incident duration far better than any tool.

What the job actually involves

A DevOps or SRE engineer builds and runs the systems other engineers deploy onto: pipelines, infrastructure, monitoring, alerting and the on-call rotation that catches what the monitoring missed. Much of the work is invisible when it is going well, which makes it easy to underfund and easy to undervalue at interview.

The high-stakes portion is small and concentrated. Most of the year is steady improvement. A handful of hours are incidents where the cost of a bad decision is measured in customer minutes. Screening for how someone behaves in those hours is most of what a DevOps interview should do.

The questions, and what a good answer sounds like

Ask these as written. Each one is paired with what a strong answer sounds like, so two people rating the same candidate reach the same conclusion for the same reason.

  1. 1. Production is down, you don't know why, and you have a hypothesis that would take twenty minutes to test.

    A good answer: Restores service first (rollback, failover, feature flag) and investigates after. Strong answers say plainly that understanding isn't the priority while customers are down. Debugging forward on a live outage is the most expensive instinct in this role.

  2. 2. Who do you tell, and when, during an incident?

    A good answer: Early and on a cadence, with impact in customer terms rather than technical detail. Naming a comms owner separate from the person fixing it is a strong answer. Silence while heroically fixing is the failure mode.

  3. 3. Walk me through a post-mortem you wrote.

    A good answer: Timeline, contributing factors, systemic actions with owners. And no individual named as the cause. If a person appears as the root cause, that's the answer, and it tells you what their previous culture was like.

  4. 4. A developer wants to deploy something on Friday afternoon.

    A good answer: Depends on the maturity of the rollback and observability instead of on the day. Strong answers reframe it as a question about confidence in reverting. Blanket rules in either direction are weaker answers.

  5. 5. Your alerting is noisy and people have started ignoring it.

    A good answer: Treats it as an emergency, because it is — an ignored alert is worse than no alert. Deletes or tunes aggressively instead of adding documentation about which alerts matter.

  6. 6. You're asked to build something you think is over-engineered for the stage.

    A good answer: Says so with the trade-off, and builds the simpler thing if overruled. Over-engineering is the characteristic failure of the discipline and self-awareness about it is a good sign.

  7. 7. How do you keep on-call sustainable?

    A good answer: Concrete measures — alert budgets, follow-the-sun or fair rotation, time back after a bad night, fixing the cause of repeat pages. Candidates with no view here burn out and take the team with them.

  8. 8. Something you automated caused an outage.

    A good answer: Owned, with what guardrail was missing. Automation failures are the discipline's own category of mistake and how someone talks about theirs is informative.

A scorecard you can rate against

Rate every candidate 1 to 4 on each row, and write the rating down before you watch the next one. Ratings drift badly when they're relative to whoever you just saw.

CriterionWhat a 4 looks like
Restore before diagnoseRolls back or fails over first. A 2 debugs forward while customers are down.
Incident communicationCommunicates early on a cadence in customer terms. A 2 goes quiet while fixing.
Blameless analysisPost-mortem names systemic contributing factors. A 2 names a person.
Alert hygieneTreats noisy alerting as urgent and deletes; a 2 documents around it.
On-call sustainabilityNames concrete measures. A 2 treats burnout as inevitable.

How to run the screen

  1. Use the recorded round for the incident and post-mortem questions. Incident communication is literally a spoken skill and this is one of the few technical contexts where a recording is close to a work sample.
  2. Keep a practical or systems-design round for the technical assessment. Nothing here tests whether someone can actually build the pipeline.
  3. Weight the restore-versus-diagnose question hardest. It is the clearest single predictor of incident duration.
  4. Listen for a person named as root cause. It is a fast read on the culture the candidate came from and the one they will bring.
  5. Don't screen on tool lists. Ask what they would choose now and why, which tests judgement instead of exposure.

Running this on a large applicant pool is where it gets expensive. VoxScreen sends these questions as a single link, transcribes and scores every answer against criteria you set, and hands back a ranked list. Free for 50 candidates a month, no card.

Try it free: 50 candidates a month, no card

How one-way video interviews work →

Common hiring mistakes for this role

  • Interviewing on tools. They churn every couple of years and competent engineers switch in weeks.
  • Not testing incident behaviour. It is where the cost is concentrated and it is the least likely thing to appear on a CV.
  • Hiring someone who blames people in post-mortems. It suppresses reporting, and suppressed reporting is how small incidents become large ones.
  • Ignoring on-call sustainability. Attrition in this discipline is driven by pager load more than by anything else.
  • Using a recorded round as the technical screen. It shows judgement and communication. The systems round does the rest.

Common questions

What should I ask a DevOps engineer in an interview?

Ask what they do when production is down and their hypothesis would take twenty minutes to test, who they tell and when during an incident, and to walk you through a post-mortem they wrote. Those three cover restore-first instinct, communication and whether their analysis is blameless — which together predict incident duration better than any tool question.

Should I test DevOps candidates on specific tools?

Not as a filter. Tooling churns every couple of years and competent engineers transfer in weeks, so a tool list mostly measures where someone last worked. Ask instead what they would choose for your situation and why, which tests judgement and is much harder to prepare for.

What does a good post-mortem look like?

A timeline, the contributing factors, and systemic actions with named owners — and no individual identified as the root cause. If a candidate's post-mortem story ends with a person who made a mistake, that tells you about the culture they came from, and blame suppresses the incident reporting you need most.

How do I screen for on-call sustainability?

Ask directly how they keep on-call sustainable and listen for concrete mechanisms — alert budgets, fair rotation, time back after a bad night, fixing repeat pages rather than documenting them. Attrition in this discipline is driven by pager load more than by pay, and a candidate with no view on it won't fix yours.

Interview questions for other roles