← All articles
13 min read

Data Scientist Interview Questions and Prep Plan for 2026

Data science interviews split into two very different loops: product analytics and ML-heavy. This guide shows which one you face, the questions each asks, and a worked A/B test case you can reuse.

The most common data scientist interview questions fall into five buckets: SQL, statistics and probability, A/B testing, product sense (metrics), and machine learning concepts. Which buckets dominate depends on the role. Product analytics loops lean heavily on SQL, experimentation, and metrics, while ML-heavy data science loops replace much of the product sense work with modeling, evaluation, and Python coding.

This guide shows you how to tell the two loops apart, what each round asks, how to work a full A/B test case out loud, and a six-week plan to get ready.

Key Takeaways

  • Read the job description before you study. "Product analytics", "decision science", or "experimentation" means a metrics-heavy loop. "Machine learning", "modeling", or "applied scientist" means an ML-heavy loop.
  • SQL is the one round almost every data science interview shares. Window functions, self-joins, and conditional aggregation show up far more than exotic syntax.
  • A/B testing questions reward a fixed structure: hypothesis, metrics, randomization unit, sample size, duration, analysis, and threats to validity.
  • Product sense questions are graded on structure, not on landing the "right" metric. Use a goal, journey, metric, guardrail sequence every time.
  • Statistics questions test interpretation more than formulas. Be able to explain a p-value and a confidence interval in one plain sentence each.
  • Six weeks at around ten hours per week is a realistic prep window if you already use SQL and Python at work.

What Is the Difference Between Product and ML Data Scientist Interviews?

A product data scientist (often titled product analytics, decision scientist, or data scientist, analytics) uses data to guide product decisions. An ML data scientist builds and ships predictive models. The titles overlap, so the loop is the better signal of what the job really is.

This table compares the two loops as they typically run at larger tech companies. Exact rounds vary by company and team, so confirm the format with your recruiter.

DimensionProduct / analytics DS loopML-heavy DS loop
Core question the role answersShould we ship this, and why did this metric move?Can we predict or rank this better?
Screening roundSQL plus a short stats or metrics questionPython or SQL coding plus ML basics
SQL depthHigh: window functions, funnels, retentionMedium: joins and aggregations for feature prep
Statistics depthHigh: hypothesis tests, power, varianceMedium to high: distributions, estimation, likelihood
A/B testingDedicated round, often the deciding oneUsually one question, about model launch tests
Product senseDedicated round on metrics and diagnosisLight, framed around model impact
ML depthConcepts: regression, trees, evaluationDeep: model choice, features, training, failure modes
Codingpandas or SQL manipulationPython data structures, sometimes LeetCode-style easy/medium
Take-homeCommon: analysis of a dataset with a write-upCommon: build and evaluate a model
Who does wellAnalysts with strong business judgmentEngineers and researchers with modeling depth

If the role sits closer to the right column and includes model deployment, the machine learning engineer interview guide covers the production side in more depth. If it leans toward pipelines and warehouses, compare it with the data engineer interview guide.

Statistics and Probability Interview Questions

Statistics rounds test whether you can reason about uncertainty without hiding behind formulas. Interviewers ask short conceptual questions, then push with a follow-up that checks real understanding.

The questions that come up again and again:

  • What is a p-value? It is the probability of seeing a result at least as extreme as yours if the null hypothesis were true. It is not the probability that the null is true.
  • What does a 95% confidence interval mean? If you repeated the experiment many times, about 95% of the intervals built this way would contain the true value.
  • What is the difference between Type I and Type II errors? A Type I error is a false positive (rejecting a true null). A Type II error is a false negative. Power equals one minus the Type II error rate.
  • When would you use a t-test instead of a z-test? Use a t-test when the population variance is unknown and samples are small. With large samples the two converge.
  • Explain the Central Limit Theorem and why it matters for A/B testing. The sampling distribution of the mean approaches normal as sample size grows, which is why you can run normal-based tests on conversion rates.
  • What is the difference between correlation and causation, and how would you establish causation without an experiment? Mention difference-in-differences, regression discontinuity, instrumental variables, or matching.

A Probability Question Worked Through

Probability questions are often a Bayes' rule problem in disguise. A classic version: a disease affects 1% of people. A test catches 95% of true cases and has a 10% false positive rate. If someone tests positive, what is the chance they have the disease?

Work it with counts, which is easier to say out loud than the formula. In 10,000 people, 100 have the disease and 95 of them test positive. Of the 9,900 without it, 990 test positive. So 95 out of 1,085 positives are real, about 8.8%. The interviewer wants to see you notice that a low base rate swamps a decent test.

Saying each step aloud matters as much as the answer. The habits in how to think out loud in a coding interview apply directly to stats rounds.

A/B Testing Interview Questions: A Worked Case

A/B testing questions are the deciding round in most product data science loops. The interviewer gives you a product change and asks how you would test it. They are grading your structure, your trade-offs, and whether you spot the traps.

Here is a full worked case you can adapt.

Prompt: "We want to change the checkout button on our e-commerce app from 'Continue' to 'Buy now'. How would you test it?"

Step 1: Hypothesis and metrics

State the hypothesis plainly: a clearer call to action will raise checkout conversion without increasing returns or cancellations.

  • Primary metric: checkout conversion rate (orders divided by users who reached the cart).
  • Secondary metric: revenue per cart visitor.
  • Guardrails: order cancellation rate, refund rate, page latency, and support contacts.

Step 2: Randomization unit

Randomize by user, not by session. Session-level randomization lets one person see both buttons, which contaminates the result. Mention that you would use a stable user ID and hash it into buckets.

Step 3: Sample size

Assume a 10% baseline conversion and a minimum detectable effect of 1 percentage point (10% to 11%). With a two-sided alpha of 0.05 and 80% power, the standard formula gives:

n per group = (z_alpha/2 + z_beta)^2 * (p1(1-p1) + p2(1-p2)) / (p2 - p1)^2
            = (1.96 + 0.84)^2 * (0.09 + 0.0979) / 0.01^2
            ≈ 14,700 users per group

Interviewers rarely need the exact number. They want to see that you know sample size grows with variance and shrinks with the square of the effect size, so halving the MDE roughly quadruples the users you need.

Step 4: Duration

Divide the needed sample by daily eligible traffic, then round up to full weeks. Weekday and weekend shoppers behave differently, so a test should cover at least one full weekly cycle and often two. Commit to the end date before launch.

Step 5: Reading the result

Suppose you run a smaller test by mistake: 10,000 users per arm, with 1,000 control conversions (10.0%) and 1,080 treatment conversions (10.8%). A two-proportion z-test looks like this:

from math import sqrt
from scipy.stats import norm

n_c, x_c = 10_000, 1_000
n_t, x_t = 10_000, 1_080

p_c, p_t = x_c / n_c, x_t / n_t
p_pool = (x_c + x_t) / (n_c + n_t)
se = sqrt(p_pool * (1 - p_pool) * (1 / n_c + 1 / n_t))
z = (p_t - p_c) / se
p_value = 2 * (1 - norm.cdf(abs(z)))

print(round(z, 2), round(p_value, 3))

This prints a z of about 1.85 and a p-value of about 0.064. The right answer is not "the button failed." The lift is positive, but the test was underpowered for an effect this size. You would either extend the test to the planned sample or report the result as inconclusive with its confidence interval.

Step 6: Threats to validity

Name these before the interviewer asks:

  • Sample ratio mismatch: if a 50/50 split arrives as 50.8/49.2 at scale, something in assignment or logging is broken. Check with a chi-square test before reading any metric.
  • Peeking: checking daily and stopping at the first significant result inflates false positives. Use a fixed horizon or a sequential testing method.
  • Novelty and primacy effects: users may click a new button out of curiosity. Look at the effect over time.
  • Network effects: in marketplaces or social products, treated users affect control users. Consider cluster or geo randomization.
  • Multiple comparisons: testing many metrics or segments raises the odds of a false win. Pre-register the primary metric.

Variance reduction is a strong bonus point. Mentioning CUPED (using pre-experiment data as a covariate to shrink variance) shows you have run real experiments.

A/B testing rounds move fast, and blanking on a sample size formula or a guardrail metric costs you the round. TechScreen is an invisible AI interview assistant that gives you real-time prompts during live data science interviews on Zoom, Google Meet, and Teams. Start with 3 free tokens, no credit card required.

Get started free →

SQL and Python Rounds

SQL is the most shared round across both data science loops. Expect 30 to 45 minutes on a shared editor such as CoderPad or a company tool, often without the ability to run the query. Clean logic and clear explanation matter more than syntax trivia.

Common SQL tasks in data science interviews:

  • Daily or weekly active users, and the ratio between them.
  • Funnel conversion from one event to the next.
  • Day-1 and day-7 retention by signup cohort.
  • Top N items per group (a window function with ROW_NUMBER or DENSE_RANK).
  • Month-over-month growth with LAG.
  • Users who did X but never did Y (anti-join).

Here is a retention query that shows several of those skills at once:

WITH signups AS (
  SELECT user_id, DATE(created_at) AS signup_date
  FROM users
),
activity AS (
  SELECT DISTINCT user_id, DATE(event_time) AS active_date
  FROM events
)
SELECT
  s.signup_date,
  COUNT(DISTINCT s.user_id) AS cohort_size,
  COUNT(DISTINCT CASE
    WHEN a.active_date = s.signup_date + INTERVAL '7 days' THEN s.user_id
  END) * 1.0 / COUNT(DISTINCT s.user_id) AS d7_retention
FROM signups s
LEFT JOIN activity a ON a.user_id = s.user_id
GROUP BY s.signup_date
ORDER BY s.signup_date;

Talk through the LEFT JOIN (so users with no activity still count in the denominator) and the DISTINCT inside the conditional count. Those two details are where most candidates lose points. For a wider question bank, use the SQL interview questions guide and the deeper SQL window functions interview guide.

Python rounds differ by loop. Product loops usually ask for pandas work: grouping, merging, reshaping, and handling missing values. ML loops may ask you to implement something from scratch, such as k-means, a train/test split, or a simple gradient descent, or to solve an easy or medium array and hash map problem. The Python interview questions guide covers the language fundamentals that come up in both.

Product Sense and Metrics: A Framework That Works

Product sense rounds ask open questions like "How would you measure the success of Instagram Stories?" or "Daily active users dropped 8% yesterday. What happened?" There is no single right answer. You are graded on structure, judgment, and whether your metrics connect to a real user goal.

The GAMES framework for defining metrics

Use this five-step sequence for any "how would you measure" question:

  1. Goal: What is the product for, and what does the business want from it? Say it in one sentence.
  2. Actors: Who are the users? Name both sides in a marketplace (riders and drivers, buyers and sellers).
  3. Moments: Map the user journey into three to five stages, such as discover, try, adopt, and return.
  4. Evaluation metrics: Pick one north-star metric tied to the goal, then one input metric per journey stage.
  5. Safeguards: Add guardrails that catch harm: latency, crashes, spam reports, churn, or cannibalization of another product.

Applied to Instagram Stories, that might be: goal of more frequent sharing among friends; actors are posters and viewers; north-star of daily users who post or view at least one story; input metrics like story creation rate, completion rate, and replies per story; guardrails like feed engagement (to catch cannibalization) and time to load.

Diagnosing a metric drop

For "metric X dropped" questions, work outside in:

  1. Check the data first. Was there a logging change, a pipeline delay, or a metric definition change?
  2. Check the calendar. Holidays, seasonality, and day-of-week effects explain many drops.
  3. Segment. Break the drop down by platform, app version, country, new versus existing users, and acquisition channel.
  4. Look for internal causes: a release, an experiment ramp, an outage.
  5. Look for external causes: a competitor launch, a policy change, an app store issue.

The segment that holds most of the drop usually points to the cause. If a drop is isolated to one Android version released yesterday, you have your lead.

Machine Learning Interview Questions for Data Scientists

ML questions in data science loops focus on judgment: which model, why, and how you would know it works. Product DS candidates need clear explanations of core concepts. ML-heavy candidates need the same plus depth on features, training, and failure modes.

Core questions to be ready for:

  • Explain the bias-variance trade-off and how regularization (L1 vs L2) affects it. L1 can push coefficients to zero and acts as feature selection; L2 shrinks them evenly.
  • When would you use logistic regression instead of a gradient boosted tree? Logistic regression wins on interpretability, speed, and small data. Boosted trees usually win on tabular accuracy with nonlinear interactions.
  • How do you handle class imbalance? Resampling, class weights, threshold tuning, and choosing a metric like precision-recall AUC over accuracy.
  • Precision vs recall: which matters more for fraud detection, and which for a spam filter that might hide real email?
  • What is data leakage, and how have you caught it? A feature that would not exist at prediction time, such as a refund flag used to predict refunds.
  • How would you evaluate a model offline and then online? Offline metrics on a time-based holdout, then an A/B test on the business metric.

ML-heavy loops may also ask about large language models, such as evaluating LLM outputs or choosing between fine-tuning and retrieval. If your target role touches that, review the AI engineer and LLM interview questions.

Take-Home Assignments and Behavioral Rounds

Many data science loops include a take-home: a dataset, a business question, and a few days to return a notebook or short deck. Graders look for a clear answer up front, honest handling of data quality, and a recommendation a non-technical reader can act on. A beautiful model with no recommendation scores poorly. The take-home coding assignment guide covers how to scope and present this kind of work.

Behavioral rounds for data scientists focus on influence. Expect "Tell me about a time your analysis changed a decision" and "Tell me about a time you disagreed with a product manager about what the data showed." Prepare four to six stories in STAR format using the top behavioral interview questions as a checklist.

A Six-Week Data Scientist Interview Prep Plan

This plan assumes about ten hours per week and that you already use SQL and Python at work. Shift time toward your weak areas and toward the column of the comparison table your target role sits in.

WeekFocusConcrete output
1SQL fundamentals and window functionsSolve 25 to 30 SQL problems, timed, without running them first
2Statistics and probabilityExplain 15 core concepts aloud in one sentence each; solve 20 probability problems
3A/B testingWork 5 full cases end to end using the six-step structure above
4Product sense and metricsDo 6 metric-definition and 4 metric-drop cases with GAMES
5ML concepts (or ML coding for ML-heavy roles)Write one-paragraph answers to 20 ML questions; implement 2 algorithms in Python
6Mocks and behavioral3 full mock loops with a peer; finalize 5 STAR stories

Two habits make this plan work. First, practice out loud, because every round in a data science loop is graded on how you explain your reasoning. Second, keep a mistakes log: every time you get a stats interpretation or SQL edge case wrong, write down the correct version and review the log in week six.

If you are still deciding between data roles, the backend engineer interview guide is a useful contrast for how much heavier the coding bar is in pure engineering loops.

Data scientist interviews jump between SQL, statistics, experiments, and product judgment in a single hour. TechScreen runs invisibly during screen shares and gives you real-time help with queries, test design, and metric frameworks when you need it. Try it with 3 free tokens before your next loop.

Get started free →

Frequently Asked Questions

What questions are asked in a data scientist interview?

Most data scientist interviews cover five areas: SQL (joins, aggregations, window functions), statistics and probability (hypothesis tests, p-values, Bayes' rule, distributions), experimentation (designing and reading A/B tests), product sense (defining metrics and diagnosing metric changes), and behavioral questions. ML-focused roles add model selection, evaluation metrics, feature engineering, and sometimes Python coding on data structures or pandas. The mix depends on whether the role is product analytics or machine learning.

Is a data science interview harder than a software engineering interview?

It is broader rather than harder. A software engineering loop goes deep on algorithms and system design, while a data science loop spreads across SQL, statistics, experimentation, product judgment, and sometimes ML. The product sense round is often the least predictable because there is no single correct answer. The coding bar is usually lower than a SWE loop, but sloppy SQL or a wrong statistical interpretation is weighted heavily.

How do I answer an A/B testing interview question?

Use a fixed sequence. State the hypothesis and the primary metric, name two or three guardrail metrics, choose the randomization unit, estimate the sample size from baseline rate and minimum detectable effect, set the duration to cover full weekly cycles, then explain how you would read the result. Finish by naming threats such as novelty effects, network interference, sample ratio mismatch, and peeking at results before the planned end date.

How much machine learning do product data scientists need to know?

Product analytics data scientists need working fluency, not research depth. Expect questions on the bias-variance trade-off, regularization, overfitting, precision versus recall, logistic regression interpretation, and tree-based models. You should be able to explain when you would pick a simple model over a complex one and how you would evaluate it. Deriving backpropagation or tuning deep networks is rarely asked outside ML-heavy data science roles.

How long should I prepare for a data scientist interview?

Six weeks of focused preparation at roughly ten hours per week is enough for most candidates who already work with data. Spend the first two weeks on SQL and statistics fundamentals, weeks three and four on experimentation and product sense cases, and the final two weeks on ML concepts, mock interviews, and behavioral stories. Candidates switching from analyst roles should add extra time on statistics and experiment design.

Should I use Python or R in a data science interview?

Python is the safer default in 2026 because most interview platforms, take-homes, and ML rounds commonly assume it, and pandas plus NumPy cover nearly every manipulation task you will be asked to do. R is still accepted by some analytics and biostatistics teams. Ask the recruiter which languages the coding round supports, and practice in the exact environment, since some rounds disable autocomplete and package installs.

What is a good framework for product sense questions in data science interviews?

Start by clarifying the product goal and the user, then map the user journey into stages. Pick one north-star metric tied to the goal, add input metrics for each stage, and add guardrails that catch harm such as latency or churn. For a metric drop question, check data and definition issues first, then segment by time, platform, geography, and user cohort before proposing causes.

Ready to use AI assistance in your next interview?

TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.

Ace your next interview →