DevOps Engineer interview questions and sample answers

A devops engineer interview usually mixes two kinds of question: ones that test whether you know the job, and behavioral ones that ask for proof from your past. Below are 10 of the first with a short model answer, and 5 of the second with what the interviewer is really checking. Rewrite every answer in your own words and with your own numbers.

Job-specific questions

1. How do you approach setting up a CI/CD pipeline for a new service?

I start with automated build and test stages that run on every commit, then add a staging deploy gated on test pass, and a manual or automated production promotion depending on the team's risk tolerance. I make sure rollback is as simple as the deploy itself, usually by keeping the previous artifact ready to redeploy.

2. Walk me through how you'd debug a production incident where a service is returning 500 errors.

I check the service's recent deploy history first, since a recent change is the most common cause, then pull logs and metrics like error rate and latency from our monitoring stack to isolate the failing component. If it's a recent deploy I roll back immediately to restore service before doing root cause analysis.

3. How do you manage secrets and credentials across your infrastructure?

I use a secrets manager like Vault or AWS Secrets Manager rather than hardcoding or storing them in environment files in the repo, and I rotate credentials on a schedule, especially for anything with broad access. Secrets in plaintext config files are one of the most common security findings I look for in a review.

4. Describe your approach to infrastructure as code and why you use it.

I define infrastructure in Terraform or similar so changes go through code review and version control rather than manual console changes, which prevents drift and makes the environment reproducible. I always run a plan before apply and review the diff carefully, since a wrong apply can take down production.

5. How do you approach setting up monitoring and alerting for a new service?

I instrument the four golden signals, latency, traffic, errors, and saturation, and set alert thresholds based on what actually indicates user impact rather than arbitrary numbers. I tune alerts over the first few weeks to reduce noise, since an ignored alert from too many false positives is worse than no alert.

6. Walk me through how you'd handle a database migration with zero downtime requirements.

I use a backward-compatible schema change approach, like adding new columns before removing old ones, and deploy application code that works with both old and new schema during the transition. I test the migration on a staging environment with production-like data volume before running it live.

7. How do you decide between containerizing a service versus running it on a VM?

I consider the service's scaling needs, dependency isolation requirements, and team familiarity, generally leaning toward containers for stateless services that need to scale horizontally. For something stateful with unusual resource requirements, a VM or managed service might still be simpler and more reliable.

8. Describe how you approach capacity planning for an application expecting a traffic spike, like a product launch.

I load test against expected peak traffic beforehand, usually at 1.5 to 2 times the projected number for margin, and set up autoscaling policies with a fast enough scale-up trigger to handle a sudden spike. I also make sure the database, not just the app servers, can handle the increased load since that's often the actual bottleneck.

9. How do you handle a situation where a teammate wants to skip code review to ship a hotfix faster?

For a true production emergency I'll do an expedited review, focusing just on the specific change rather than blocking on style nits, but I don't skip it entirely since a rushed unreviewed fix has caused outages before. A five-minute review is rarely the actual bottleneck in an incident.

10. What's your approach to reducing cloud infrastructure costs without hurting reliability?

I start by identifying underutilized resources through actual usage metrics, like rightsizing instances running at 10% CPU, before cutting anything that affects redundancy or headroom. I also check for orphaned resources like unattached volumes or old snapshots, which are often pure waste with zero tradeoff.

Behavioral questions

Answer these with STAR: the Situation in one sentence, the Task, the Action you took (most of the answer), the Result with a number.

11. Tell me about a result you are proud of.

What they are checking: A specific outcome with a number, and what you personally did to get it.

A result to build the answer around: Moved 60 services to Kubernetes with Terraform, taking deploys from weekly to 40 a day.

12. Describe a time you improved how something was done.

What they are checking: That you notice waste and fix it without being told, then measure the difference.

A result to build the answer around: Cut AWS spend by 34% ($410k a year) with rightsizing, spot instances and storage lifecycle rules.

13. Tell me about a time you had to deliver under pressure.

What they are checking: How you prioritise, communicate early and still finish to standard.

A result to build the answer around: Raised uptime from 99.5% to 99.95% by adding health checks, autoscaling and runbooks.

14. Give an example of working with a difficult colleague or customer.

What they are checking: Calm, the other person's view stated fairly, and a result that helped both sides.

A result to build the answer around: Built a CI pipeline that runs 3,000 tests in 7 minutes, down from 28.

15. What is something you learned from a mistake?

What they are checking: Ownership without excuses, and the habit you changed so it did not happen again.

A result to build the answer around: Led on-call rotation for 12 engineers and halved mean time to recovery.

Before the interview

Interviewers read your resume just before they walk in, and most questions come from it. Every bullet on it should be one you can expand into a two-minute STAR story. Check that it matches the posting with the free ATS Match Score, and look at the devops engineer resume keywords the ad is likely to use.

Turn your notes into a devops engineer resume, free

No signup, no card. See the full devops engineer resume example and the cover letter example.

Interview questions for related jobs

All interview question lists · CV Forge — free resume generator