Data Scientist interview questions and sample answers

A data scientist interview usually mixes two kinds of question: ones that test whether you know the job, and behavioral ones that ask for proof from your past. Below are 10 of the first with a short model answer, and 5 of the second with what the interviewer is really checking. Rewrite every answer in your own words and with your own numbers.

Job-specific questions

1. How do you decide between a simple model like logistic regression and a more complex model like gradient boosting?

I start with the simple model as a baseline since it's interpretable and fast to iterate on, then check if the complex model's accuracy gain justifies the added latency and maintenance cost. For something like fraud detection where marginal accuracy matters a lot, I'll push toward XGBoost; for a regulated credit decision, interpretability often wins even at some accuracy cost.

2. Walk me through how you'd handle a dataset with significant class imbalance.

I avoid relying on accuracy and use metrics like precision, recall, and PR-AUC instead, then address the imbalance with class weighting or techniques like SMOTE rather than just downsampling blindly. I also check that my train and test splits preserve the imbalance ratio so I'm evaluating on a realistic distribution.

3. How do you validate that a model isn't overfitting before deployment?

I use k-fold cross-validation, typically 5-fold, and check that performance is stable across folds rather than trusting a single train-test split. I also hold out a true final test set that I only touch once, at the very end, to get an honest read on generalization.

4. Describe how you would explain a model's prediction to a non-technical stakeholder.

I use SHAP values or feature importance to identify the top three or four drivers of a specific prediction, then translate them into plain business language, like 'this customer's churn risk is high mainly due to a drop in usage over the last 30 days.' I avoid showing raw coefficients or technical jargon in that conversation.

5. How do you approach feature engineering for a time series forecasting problem?

I create lag features and rolling averages at intervals that match the business cycle, like weekly and monthly, and I'm careful to only use information that would have actually been available at prediction time to avoid leakage. I also test for seasonality and trend before choosing between something like ARIMA or a tree-based model with time features.

6. What's your process when a production model's performance starts degrading over time?

I check for data drift first, comparing the distribution of current input features against the training distribution using something like population stability index. If drift is confirmed, I retrain on more recent data rather than assuming the original model architecture is the problem.

7. How do you handle missing data in a dataset before modeling?

I first check whether the missingness is random or systematic, since that changes the right approach, then use imputation like median or model-based methods for random gaps, or a missing indicator flag if the absence itself is informative. I never just drop rows without understanding why the data is missing first.

8. Describe how you would set up an A/B test to evaluate a new recommendation model.

I define the primary metric upfront, like click-through rate, calculate the required sample size for statistical power before starting, and randomize at the user level to avoid contamination between groups. I run it for a full business cycle, usually at least one to two weeks, to account for day-of-week effects.

9. How do you decide what to log and monitor for a model in production?

I track prediction distribution, input feature distributions, and the actual business metric like conversion rate, not just model accuracy, since ground truth labels are often delayed. I set alert thresholds so I know within a day if something looks off rather than finding out weeks later from a stakeholder.

10. Walk me through how you'd approach a request from a stakeholder who wants '99% accuracy' on a hard prediction problem.

I'd ask what decision the number actually drives, then show them the realistic performance ceiling given the data and explain tradeoffs like precision versus recall for their use case. Setting an arbitrary accuracy target without understanding the base rate or business cost of errors usually leads to a model optimized for the wrong thing.

Behavioral questions

Answer these with STAR: the Situation in one sentence, the Task, the Action you took (most of the answer), the Result with a number.

11. Tell me about a result you are proud of.

What they are checking: A specific outcome with a number, and what you personally did to get it.

A result to build the answer around: Built a churn model (AUC 0.86) that targeted retention offers and saved $1.2M a year.

12. Describe a time you improved how something was done.

What they are checking: That you notice waste and fix it without being told, then measure the difference.

A result to build the answer around: Designed an A/B testing framework adopted by 6 product teams, cutting test setup from 2 weeks to 2 days.

13. Tell me about a time you had to deliver under pressure.

What they are checking: How you prioritise, communicate early and still finish to standard.

A result to build the answer around: Improved demand forecast error (MAPE) from 18% to 11%, reducing overstock by $400k.

14. Give an example of working with a difficult colleague or customer.

What they are checking: Calm, the other person's view stated fairly, and a result that helped both sides.

A result to build the answer around: Moved 14 models from notebooks to scheduled Airflow pipelines with monitoring.

15. What is something you learned from a mistake?

What they are checking: Ownership without excuses, and the habit you changed so it did not happen again.

A result to build the answer around: Presented findings monthly to leadership, turning 3 analyses into funded roadmap items.

Before the interview

Interviewers read your resume just before they walk in, and most questions come from it. Every bullet on it should be one you can expand into a two-minute STAR story. Check that it matches the posting with the free ATS Match Score, and look at the data scientist resume keywords the ad is likely to use.

Turn your notes into a data scientist resume, free

No signup, no card. See the full data scientist resume example and the cover letter example.

Interview questions for related jobs

All interview question lists · CV Forge — free resume generator