For hiring managers

Interview guide for data scientists: a loop that tests judgment, not just modeling

On this page
  1. The loop at a glance
  2. Stage 2: the experiment readout review (45 minutes)
  3. Stage 3: the modeling working session (90 minutes)
  4. Stage 4: the stakeholder session (40 minutes)
  5. Stage 5: hiring manager interview (45 minutes)
  6. Scorecard competencies and weights
  7. Legal points for data science loops
  8. The decision rule
  9. Adjusting the loop
  10. Common mistakes in data science loops
  11. Questions people ask

Data scientists rarely fail because they cannot fit a model. They fail because they answer the wrong question, trust an experiment that was broken, ship a feature that leaks the label, or present a result nobody can act on. Most loops test the first thing and miss the rest. This guide gives data science managers a loop for a mid-level product data scientist working on experimentation and modeling: five stages, who owns which competency, an experiment readout review, a modeling case with planted problems, a stakeholder session, example weights and a decision rule. Adjustments for analytics-leaning, machine-learning-leaning and senior roles are at the end.

The first call, including how to tell analytics, applied machine learning and research-leaning seats apart, is in data scientist phone screen questions. If the role is mostly SQL and reporting, use the interview guide for data analysts instead.

The loop at a glance

StageInterviewerOwnsLengthPass rule
1. ScreenRecruiter or managerSeat type, past projects end to end, logistics30 minOne project explained end to end with their own part clear
2. Experiment readout reviewSenior data scientistStatistics and experiment design; skepticism45 minAt least 3 on experiment design
3. Modeling working sessionSenior data scientist or ML engineerProblem framing; modeling; data hygiene; code90 minAt least 3 on problem framing; leakage found or reasoned about
4. Stakeholder sessionProduct manager or business partnerCommunication; influence; business sense40 minAt least 3 on communication
5. Hiring manager interviewData science managerOwnership; collaboration; learning45 minNo competency below 2

Stages 2 to 5 fit into one day or two half-days. Keeping them close together matters because strong candidates are usually in more than one process.

Stage 2: the experiment readout review (45 minutes)

Send nothing in advance. In the session, share a two-page invented A/B test readout: a new checkout design, a headline claim of a 4% lift in conversion, a p-value, a chart and a recommendation to ship. Plant three problems:

  • The test was stopped early, on the day the result first crossed significance.
  • The treatment group has noticeably more mobile users than control, suggesting a randomization or assignment issue.
  • Conversion rose, but average order value fell, and the readout does not mention revenue.

Ask: "Your product manager wants to ship tomorrow. What do you tell them?" Then probe: "How would you have designed this test?" and "What would make you comfortable shipping?"

  • A 4: finds all three, explains why each matters in plain words, proposes a sample size and duration decided in advance, checks the assignment, and suggests a guardrail metric.
  • A 2: recalculates the statistics carefully but accepts the setup.

Stage 3: the modeling working session (90 minutes)

A prepared notebook with an invented customer dataset and the brief: "Predict which customers will cancel their subscription in the next 30 days so the retention team can call them." The candidate may use any library and search documentation. The interviewer answers questions about the data as a colleague would.

Planted problems

  • A feature called last_contact_reason that is filled in by the retention team after a customer calls to cancel: it leaks the label.
  • Cancellations are about 5% of rows, so accuracy is a misleading metric.
  • The retention team can call only 200 customers a week, which should shape the evaluation.

Running it

  1. Framing (15 minutes): "Before you write code, what questions do you have?" Listen for: how the prediction will be used, when it is made, and what the team can act on.
  2. Working (55 minutes): exploration, a baseline, a model. The interviewer stays quiet unless asked.
  3. Review (20 minutes): "How good is this model, and would you put it in front of the retention team?" Listen for: a metric tied to 200 calls a week, such as precision in the top 200, and a plan for checking it after launch.

Candidates who miss the leakage but notice their score is suspiciously high and investigate still show the habit you need. Candidates who report a near-perfect score with pride do not. If you prefer a take-home, keep the same brief and rubric and read how to evaluate a take-home assignment.

Stage 4: the stakeholder session (40 minutes)

A product manager runs it in role. Part one: the candidate explains the churn model from stage 3 to the head of customer success, who is not technical, in ten minutes. Part two: the product manager pushes for a conclusion the data does not support ("So the new pricing page caused the churn, right?").

  • Listen for: leading with the decision and what the model can and cannot do, avoiding jargon, and saying no to the causal claim while offering a way to test it.
  • A low score: a methods tour with no recommendation, or agreeing to the claim to keep the peace.

Stage 5: hiring manager interview (45 minutes)

  1. "Tell me about a model or analysis of yours that was wrong after it shipped. How did you find out?" Listen for: monitoring, owning it and the change they made.
  2. "Tell me about a project you stopped or talked a stakeholder out of." Listen for: a business reason, not just a technical preference.
  3. "Tell me about working with engineers to get a model into production." Listen for: what they handed over, what they owned and how they shared responsibility for monitoring.
  4. "What is a method you learned in the last year, and where did you use it?" Listen for: a real application, not a course list.

Scorecard competencies and weights

Example weights for a mid-level product data scientist. Lock yours before the first candidate.

CompetencyOwned byExample weight
Problem framingModeling session20%
Statistics and experiment designReadout review20%
Modeling and data hygieneModeling session20%
Communication and influenceStakeholder session20%
Code qualityModeling session10%
Ownership and collaborationHiring manager interview10%

The scorecard builder prints a sheet per interviewer and checks the weights add up to 100.

Checked against the linked primary sources as of October 2026. Not legal advice; ask your counsel how they apply to your hiring.

  • Sponsorship questions. In a September 2022 technical assistance letter, the Justice Department's Immigrant and Employee Rights Section repeated the two questions it has identified as appropriate to ask applicants: whether they are legally authorized to work in the United States, and whether they will now or in the future require sponsorship for employment visa status (for example, H-1B). It cautioned employers against pre-screening questions beyond those two. More in work authorization questions in interviews.
  • Tests are selection procedures. Under the Uniform Guidelines at 29 CFR 1607.3, a selection procedure with adverse impact on any race, sex or ethnic group is considered discriminatory unless validated under the guidelines. Keep the exercises tied to the actual work, give everyone the same materials, and watch pass rates for each group.
  • Confidential work. Do not ask candidates to show code, models or data from a current or former employer. Ask them to describe the approach without proprietary details; the exercises give you the hands-on evidence.

The decision rule

  1. Scorecards first, including the notebook, before the debrief.
  2. Floors: problem framing at 3 or above and communication at 3 or above. A strong modeler who cannot frame or explain will build the wrong thing well.
  3. Weighted total: in this example, 2.9 or higher on a 1–4 scale is an offer.
  4. Split panel: if the two technical interviewers differ by two points, they review the notebook together before discussing anything else.

Adjusting the loop

SeatWhat changes
Analytics-leaningReplace the modeling session with a metrics design case and a SQL working session; raise experiment design to 30%.
Machine-learning-leaningAdd a production design stage on serving, monitoring and retraining; see machine learning engineer screening questions.
Senior or staffAdd a past-project deep dive led by the candidate and a stage on setting direction for a team's roadmap.
New graduateShorter notebook with more prompts; accept research and coursework projects as evidence.

Common mistakes in data science loops

  • Algorithm trivia. Recall of formulas is easy to look up. Score judgment on messy data.
  • Unbounded take-homes. They filter for free time, not skill.
  • No business partner in the loop. The person who will use the work should judge whether they understand it.
  • Celebrating high scores. A model that looks too good in an interview is a test of skepticism.

After the hire, the 30-60-90 day plan for data scientists covers the first quarter.

Questions people ask

Should data scientist candidates do a take-home assignment?

Only if it is short and respects the candidate's time: a few hours at most, with a clear brief, the same data for everyone and a scoring rubric written in advance. Many teams get the same evidence from a 90-minute live working session on a prepared notebook, which also removes the question of who did the work.

What is the most common gap in data scientist interviews?

Testing modeling technique and not judgment. Strong loops check whether the candidate notices a broken experiment, leakage in a feature, or a metric that does not match the business question, and whether they can explain what the result means to someone who will act on it.

Can I ask a data scientist candidate whether they need visa sponsorship?

Yes. The Justice Department's Immigrant and Employee Rights Section has identified two questions as appropriate: whether the applicant is legally authorized to work in the United States, and whether they will now or in the future require sponsorship for employment visa status. It has cautioned against going beyond those two. As of October 2026; not legal advice.

How senior should the interviewers be?

At least one interviewer per technical stage should be at or above the level being hired and should have shipped work of the same type, whether experimentation, forecasting or production machine learning. A product or business partner should run the stakeholder stage.