Interview guide for DevOps roles: a loop with a live incident and a pipeline design
On this page
- The loop at a glance
- Stage 1: the screen (30 minutes)
- Stage 2: the incident simulation (60 minutes)
- Stage 3: the pipeline design session (60 minutes)
- Stage 4: infrastructure code review (30 minutes)
- Stage 5: manager and developer conversation (45 minutes)
- Scorecard competencies and weights
- Legal points for DevOps hiring
- The decision rule
- Adjusting the loop
- Common mistakes in DevOps loops
- Questions people ask
DevOps, site reliability and platform engineers are hired to keep software shipping safely and to restore service fast when it breaks. Their worst day is visible to every customer. A loop that asks which tools they know, or has them reverse a linked list, misses what matters: how they debug a live system with incomplete information, how they design a delivery path that is fast without being fragile, and how they communicate when the pressure is highest. This guide gives engineering managers a loop for a mid-level DevOps engineer: five stages, owners, a live incident simulation, a pipeline design session, an infrastructure code review, example weights and a decision rule. Adjustments for SRE, platform and junior roles are at the end.
The first call, including how to tell who built a pipeline from who pushed code through one, is in DevOps engineer screening questions. For cloud-focused roles, see cloud engineer screening questions.
The loop at a glance
| Stage | Interviewer | Owns | Length | Pass rule |
|---|---|---|---|---|
| 1. Screen | Recruiter or manager | Environment, scale, on-call pattern, logistics | 30 min | Has operated production systems or a credible route in |
| 2. Incident simulation | Senior SRE or DevOps engineer | Debugging method; incident communication | 60 min | At least 3 on debugging method |
| 3. Pipeline design | Staff engineer or platform lead | Delivery design; security; cost awareness | 60 min | At least 3 on delivery design |
| 4. Infrastructure code review | DevOps engineer | Infrastructure as code; security judgment | 30 min | No unaddressed critical security issue |
| 5. Manager and developer conversation | Engineering manager plus an application developer | Collaboration; ownership; blameless learning | 45 min | No competency below 2 |
Stages 2 to 5 fit in one day or two half-days. The general structure follows the interview guide for software engineers, with operations replacing the coding stage.
Stage 1: the screen (30 minutes)
- "Describe the system you support: what runs where, how it is deployed, and how many deploys a day?" Listen for: concrete numbers and their own part in it.
- "Tell me about your on-call: rotation, how often you were paged, and the last page you took." Listen for: specifics, not a summary.
- "This role is on call one week in [n] with [compensation or time off]. How does that fit?" State it plainly; on-call surprises are a common reason engineers leave early.
Stage 2: the incident simulation (60 minutes)
The interviewer plays the monitoring system, the logs and the rest of the team. The candidate says what they would check or run; the interviewer returns prepared output. Write the scenario and outputs once and use them for every candidate.
The scenario
At 14:05, error rates on the checkout API jump from under 1% to 18%. Latency is up. A deploy of the checkout service went out at 13:50. The database team mentions they rotated credentials "sometime this afternoon." Customer support is asking what to tell customers.
What the outputs reveal, step by step
- Errors are only on pods from the new deploy that restarted after 14:00.
- Logs on those pods show authentication failures to the database.
- The new deploy reads the database secret at startup; older pods still hold the old credential in memory.
What to listen for
- Mitigate first: considering a rollback or stopping the rollout early, before the root cause is certain, and saying why.
- Method: forming a hypothesis, checking it with one query, and not changing three things at once. Noticing that two changes happened close together.
- Communication: a short status update to support at a stated interval, an incident channel and a named lead.
- Afterward: "What goes in the postmortem?" Listen for: contributing factors across both teams, not blame, and concrete follow-ups such as secret rotation that does not break running services.
Stage 3: the pipeline design session (60 minutes)
Brief: "Our team of 30 developers deploys a monolith once a week with a two-hour manual checklist. We want to deploy several services daily. Design the path from merged code to production." Give constraints: one cloud provider, a regulated customer that needs an audit trail, and a budget owner who watches costs.
- Listen for: build once and promote the same artifact; automated tests at the right stages; progressive delivery such as canary or blue-green with automatic rollback on health checks; secrets management; who approves what; and an audit trail.
- Probe: "A developer says the pipeline is too slow. What do you change first?" and "What do you not automate yet, and why?"
- A 4: sequences the migration in steps the team can absorb, rather than drawing an ideal end state only.
Stage 4: infrastructure code review (30 minutes)
Share a short invented infrastructure-as-code change, in whatever language your team uses, that opens a storage bucket to public read, gives a service role wildcard permissions, hard-codes a password, and has no tags for cost tracking. The candidate reviews it as a pull request and writes comments.
Score: finding the security issues, explaining them in a tone a colleague would accept, and suggesting specific fixes. Missing the public bucket and the hard-coded secret is a gate failure for this role.
Stage 5: manager and developer conversation (45 minutes)
- "Tell me about an outage you caused." Listen for: plain ownership, the fix and the systemic change, with no hiding.
- "Tell me about a time developers resisted a change you made to their workflow." Listen for: listening, adjusting and measuring whether it helped.
- From the developer: "What would you need from my team to support our service well?" Listen for: shared ownership rather than a ticket queue.
- "What did you automate away from your own job recently?" Listen for: reducing repetitive work, with a measurable result.
Scorecard competencies and weights
Example weights for a mid-level DevOps engineer. Lock yours before the first candidate.
| Competency | Owned by | Example weight |
|---|---|---|
| Debugging method under pressure | Incident simulation | 25% |
| Delivery and system design | Pipeline design | 20% |
| Incident communication | Incident simulation | 15% |
| Infrastructure as code | Code review | 15% |
| Collaboration with developers | Manager and developer conversation | 15% |
| Ownership and blameless learning | Manager and developer conversation | 10% |
| Security judgment | Code review; design | Pass/fail gate |
The scorecard builder checks that the weights add up to 100 and prints a sheet per interviewer.
Legal points for DevOps hiring
Checked against the linked primary sources as of October 2026. Not legal advice.
- Sponsorship. The Justice Department's Immigrant and Employee Rights Section, in a September 2022 technical assistance letter, restated two questions it considers appropriate: whether the applicant is legally authorized to work in the United States, and whether they will now or in the future require sponsorship for employment visa status. It cautioned against going further.
- Background checks for privileged access. If a background check company provides the report, 15 U.S.C. 1681b(b) requires a clear disclosure in a document consisting solely of the disclosure and the candidate's written authorization before the report is obtained, and a copy of the report and a summary of rights before adverse action. The steps are in the FCRA background check process.
- Exercises as selection procedures. The Uniform Guidelines at 29 CFR 1607.3 treat a selection procedure with adverse impact as discriminatory unless it is validated. Keep the simulation tied to the real job and the same for everyone.
The decision rule
- Scorecards first, including the written review comments, before the debrief.
- Gate: no unaddressed critical security issue in the code review.
- Floor: debugging method at 3 or above.
- Weighted total: in this example, 2.8 or higher on a 1–4 scale is an offer.
- Split panel: if the incident and design interviewers differ by two points, compare notes on what the candidate actually said before anything else.
Adjusting the loop
| Role | What changes |
|---|---|
| Site reliability engineer | Add a service level objective exercise: set objectives and an alerting policy for the checkout service from a page of invented traffic data; weight reliability design higher. |
| Platform engineer | Make the design stage an internal developer platform; the developer interviewer scores whether they would want to use it. |
| Junior or career changer from IT support | Simpler incident with more prompts; accept home labs; weight learning. The IT support interview guide covers the adjacent loop. |
| Systems administrator | Replace pipeline design with a patching and backup design. |
Common mistakes in DevOps loops
- Tool checklists. Tools change; method and judgment transfer. Score the incident.
- Leetcode-style coding rounds. Script-level coding can be checked inside the code review.
- No communication score. In a real incident, silence is a failure even when the fix is right.
- Hiding on-call. State the rotation in the screen.
After the hire, adapt the new hire 30-60-90 day plan template with a shadow on-call rotation in the first 30 days.
Questions people ask
How do you test a DevOps engineer in an interview?
Put them in a realistic incident and a realistic design problem. A 60-minute incident simulation, where an interviewer plays the monitoring system and the team, shows debugging method, communication and judgment under pressure. A pipeline design session shows how they balance speed, safety and cost. Both beat questions about tools.
Should a DevOps interview include a hands-on lab?
A lab in a disposable cloud account or container environment is useful if you can make it reliable for every candidate. If you cannot, an interviewer-run simulation, where the candidate says what they would run and the interviewer returns realistic output, gives most of the same evidence without setup failures.
How much should on-call experience count?
A lot for SRE and production-facing DevOps roles. Ask for a specific incident, what the candidate did in the first fifteen minutes, and what changed afterward. Candidates who have only deployed onto platforms that others ran will have thinner answers, which is fine for junior roles if the rest of the loop is strong.
Can we run background checks for engineers with production access?
Yes, if you follow the Fair Credit Reporting Act when you use a background check company: a stand-alone written disclosure and written authorization before the report, and a copy of the report and summary of rights before any adverse action. State and local laws may add limits. As of October 2026; not legal advice.