Data Analyst Interview Questions: 25 Questions & Answers
By Rehearsa Team · 2026-10-01
Data Analyst Interview Questions: 25 Questions & Answers
Data analyst interviews test much more than whether you can write a SQL query. A strong candidate must turn messy data into decisions, question the numbers before presenting them, explain findings to people who do not work with data, and show sound judgement about accuracy versus speed.
This guide covers 25 questions that appear across recruiter screens, hiring-manager interviews, technical rounds, case-study discussions, and behavioural interviews. Each includes what the interviewer is assessing, a model answer, and a practical preparation note. Use the answers as structures, not scripts: replace the details with your own data, projects, tools, and results.
For a complete preparation sequence alongside these questions, read our data analyst interview preparation guide. If your process includes a timed technical test, the coding exercise guide covers how to practise under pressure.
How Data Analyst Interviews Are Usually Assessed
| Area | What strong evidence sounds like | What weakens an answer |
|---|---|---|
| SQL and technical skill | You write correct logic and explain why you chose it | You recite syntax but cannot reason through a scenario |
| Analytical judgement | You question data quality, definitions, and assumptions before concluding | You accept the first number you see |
| Business impact | You connect analysis to decisions, costs, or revenue | You describe charts and queries without an outcome |
| Communication | You explain findings simply and state your confidence level | You hide behind jargon or overwhelm with detail |
| Stakeholder skills | You manage conflicting requests and agree scope | You treat every request as equally urgent |
| Integrity | You flag limitations and correct your own errors | You present uncertain numbers as certain |
General and Motivation Questions
1. "Tell me about yourself."
What they are assessing: Whether you can present a relevant, coherent career story rather than recite your CV.
Model answer: "I am a data analyst with three years of experience turning product and marketing data into decisions. In my current role I own the reporting layer for a subscription business: I maintain our SQL models, build dashboards in Looker, and deliver a weekly retention analysis that leadership uses to prioritise feature work. Last year my churn analysis identified onboarding friction that, once fixed, improved 30-day retention by six percentage points. I am now looking for a role with larger datasets and closer partnership with decision-makers, which is why this position stood out."
Prepare: Build a 60–90 second answer around present role, one quantified proof point, and why this opportunity is the logical next step. Our tell me about yourself guide gives a fuller structure.
2. "Why do you want to be a data analyst?"
What they are assessing: Whether your motivation fits the day-to-day reality of the job, not just the title.
Model answer: "I enjoy the moment when a vague question becomes a measurable one. In my first role, support kept saying customers were 'unhappy' — I framed it as ticket themes and time-to-resolution, and we could finally see which fixes mattered. What keeps me in this work is that small analytical decisions, like how you define an active user, change what a whole company believes about itself. I like being the person who makes those definitions careful and transparent."
Prepare: Name a specific moment when analysis changed a decision, and show you understand the job includes stakeholder work and data cleaning, not just interesting puzzles.
3. "Walk me through a data project you are proud of."
What they are assessing: Your end-to-end process: how you scoped the question, handled the data, and delivered impact.
Model answer: "Marketing wanted to know which channels drove customers who stayed. I started by agreeing the definition of 'stayed' with them — active at 90 days, not just a completed signup. I then joined attribution data with usage logs in SQL, handled duplicate tracking IDs by keeping the earliest verified touch, and built a cohort comparison in Python. The finding was counter-intuitive: our biggest-volume channel had the worst 90-day retention, while a small referral programme had three times the retention. We shifted 20% of budget, and 90-day retention rose by nine points over two quarters. I documented the definitions so the analysis could be rerun monthly."
Prepare: Use the structure question → data → method → finding → decision → result. Have one project where you pushed back on a flawed assumption.
4. "How do you approach a new dataset you have never seen before?"
What they are assessing: Whether you explore systematically before drawing conclusions.
Model answer: "First I ask where the data came from and what each row represents — most mistakes happen at that level. Then I profile it: row counts over time to spot gaps, null rates per column, distinct values for anything categorical, and ranges for numeric fields. I sanity-check a few totals against a trusted source, like a finance report, because if the basics do not reconcile, every conclusion on top is unsafe. Only after that do I start the actual analysis, and I keep a written log of every assumption."
Prepare: Have a personal checklist ready. Interviewers want evidence of habit, not theory.
5. "Which tools do you use, and how do you decide which one for a task?"
What they are assessing: Practical competence and judgement about fit, not a long tool list.
Model answer: "My core stack is SQL for extraction and transformation, Python with pandas for analysis that repeats or needs statistics, and either Looker or Excel for communicating, depending on the audience. The choice follows the task: if it is a one-off question, I stay in SQL and a notebook; if it will be rerun every week, I build it as a documented, tested model rather than a query someone's laptop depends on. I recently added dbt to make transformations version-controlled, which cut our reconciliation time in half."
Prepare: Mention one tool you learned recently and why. It signals you can keep up as the stack changes.
SQL and Technical Questions
6. "Explain the difference between INNER JOIN, LEFT JOIN, and FULL OUTER JOIN."
What they are assessing: Whether you understand what each join does to your row counts — a common source of silent errors.
Model answer: "An INNER JOIN keeps only rows that match in both tables. A LEFT JOIN keeps every row from the left table and fills the right side with NULLs where there is no match — so you see customers with no orders instead of losing them. A FULL OUTER JOIN keeps unmatched rows from both sides, useful for reconciliation, for example comparing two systems' customer lists. The practical risk is fan-out: if the right table has multiple matches, a join multiplies rows, so I always check row counts before and after any join."
Prepare: Be ready for a follow-up like "when would a LEFT JOIN silently change your averages?" — aggregates on joined data need care.
7. "How would you find the second-highest value in a table?"
What they are assessing: Comfort with window functions versus quick hacks.
Model answer: "I would use a window function: DENSE_RANK() OVER (ORDER BY amount DESC) and filter for rank 2. I prefer DENSE_RANK to RANK here because if two rows tie for the highest amount, RANK skips to 3 while DENSE_RANK gives the next distinct value as 2. In an interview I would also ask a clarifying question: do you want the second row by order, or the second distinct value? Those give different answers with duplicates."
Prepare: Know ROW_NUMBER, RANK, and DENSE_RANK well enough to explain when they differ — interviewers probe exactly that.
8. "Write a query to get each customer's most recent order."
What they are assessing: Whether you can solve a classic scenario cleanly.
Model answer: "The cleanest general form is a window function: ROW_NUMBER() OVER (PARTITION BY customer_id ORDER BY order_date DESC) and filter to 1, which returns exactly one row per customer even with ties on the timestamp. GROUP BY customer_id with MAX(order_date) is shorter but only gives the date, not the whole order row. If the table is large, I would check whether filtering on an indexed date window first performs better. I would also ask whether 'most recent' should exclude cancelled orders — a definition question hidden inside a technical one."
Prepare: Practise this one out loud, including the clarifying question. Seeing you question definitions matters as much as the SQL.
9. "How do you handle duplicate records in your analysis?"
What they are assessing: Whether you find duplicates before they distort results.
Model answer: "I first define what a duplicate means in context — same order ID is obvious, but same customer with two emails is a fuzzy case needing rules. For exact duplicates I check with GROUP BY all key columns HAVING COUNT(*) > 1, then decide whether to deduplicate with ROW_NUMBER() keeping the newest record, or to keep both and count them once through a distinct ID. I never just delete — I keep a before count, document the rule, and rerun a headline metric to show the impact of cleaning. If duplicates come from the source system, I raise it, because cleaning in a notebook hides the problem rather than fixing it."
Prepare: Have one real story of duplicates changing a number. It makes the answer memorable.
10. "A query you run weekly has become very slow. How do you fix it?"
What they are assessing: Structured troubleshooting, not memorised tuning tricks.
Model answer: "I start by measuring: run the query with the execution plan on and find where the time actually goes. Common culprits are functions wrapped around indexed columns, which stop index use, a join fan-out multiplying rows early, or pulling columns the analysis never uses. Fixes in order of preference: filter and aggregate as early as possible, join only the needed keys, check whether an index matches the filter and sort pattern, and consider pre-aggregating if the grain is much finer than the question. I would confirm the result set is identical before and after — a faster wrong answer is worse than a slow right one."
Prepare: Learn to read one execution plan format well. Naming your measurement step first is what separates strong candidates.
11. "What is the difference between WHERE and HAVING?"
What they are assessing: Core SQL reasoning that interviewers use as a filter question.
Model answer: "WHERE filters rows before grouping; HAVING filters groups after aggregation. If I want orders over £100 from customers with more than five orders, the £100 condition goes in WHERE — it applies to each row — and the count condition goes in HAVING because it exists only after grouping. I also watch for the aggregate-in-WHERE mistake, which fails, and for conditions that could live in either place on non-aggregated columns, where WHERE is clearer and usually cheaper because it shrinks the data before grouping."
Prepare: Give one example of each in your answer rather than a definition alone.
Statistics and Analytical Judgement Questions
12. "What is the difference between correlation and causation, and how do you handle it in practice?"
What they are assessing: Whether this is a memorised phrase or a working habit.
Model answer: "Correlation means two variables move together; causation means changing one changes the other. In practice I handle it three ways. I look for confounders — a variable that drives both — and control for them. I prefer experiments when the decision justifies the cost: an A/B test turns correlation into evidence. And when neither is possible, I say what the analysis cannot claim. In my last role, users who contacted support churned more — not because support caused churn, but because broken onboarding caused both. Framing it that way moved the fix to onboarding, where it belonged."
Prepare: A real example where the naive correlation was wrong is far stronger than the textbook definition.
13. "How would you know if a metric change is real and not random noise?"
What they are assessing: Statistical literacy applied to everyday analytics work.
Model answer: "I start with volume and variance — a 2% move on ten conversions is noise, the same move on 100,000 is a signal. For a proper check I would compare against the same period historically, account for weekly and seasonal patterns, and where it matters, run an A/B test with a pre-calculated sample size and check that the result reaches significance. I also watch for multiple comparisons: if you check twenty segments, one will look significant by chance. My rule is to state the confidence I actually have — 'directional, needs another week of data' — rather than borrowing certainty from a small sample."
Prepare: Be ready to explain a p-value or confidence interval in one plain sentence to a non-technical stakeholder.
14. "How do you deal with outliers?"
What they are assessing: Whether you investigate anomalies instead of mechanically deleting them.
Model answer: "First I find out why the outlier exists — that matters more than what to do with it. Data entry errors, like an order of 10,000 units meant to be 10, get corrected at the source. Genuine extremes, like one wholesale customer, are real and should be analysed, sometimes separately. My method: detect with a distribution check or IQR rule, investigate the cause, then choose the treatment that matches the question — medians instead of means for skewed revenue data, or a segment split for reporting. I document every exclusion, because unexplained filtering is how numbers lose trust."
Prepare: Mention one case where an outlier was the story, not the noise — fraud and failures usually are.
15. "How do you decide which metrics matter for a product or campaign?"
What they are assessing: Whether you anchor metrics to decisions rather than collecting numbers.
Model answer: "I work backwards from the decision. For a campaign, the question is usually not clicks but what happens after: qualified signups and their downstream retention or revenue. I define one primary metric the team would actually act on, guardrail metrics that catch harm — like unsubscribe rate or support tickets — and diagnostic metrics to explain movement. Then I check the definition is measurable and honest: 'engagement' that counts a login means little, so I push for definitions tied to real value. Writing the definition down, with its exclusions, prevents the metric meaning different things to different teams."
Prepare: Reference a North Star or guardrail concept by name — it shows familiarity with how product teams work.
16. "How would you forecast next quarter's revenue?"
What they are assessing: Structured thinking under an ambiguous, classic case-study prompt.
Model answer: "First I would clarify what the forecast drives — hiring plans need more rigour than a blog post. Then I would choose the method to fit the data: with two years of history and seasonality, a baseline like a seasonal decomposition or a simple exponential-smoothing model; with a subscription business, I would build it bottom-up from churn rate and pipeline rather than extrapolating a line. I would present a range with assumptions, not a single number, and track accuracy against actuals so the next forecast improves. The most common forecast failure I watch for is assuming last quarter's growth rate continues unchanged."
Prepare: You are not expected to be a data scientist. Showing that you match method to purpose and communicate uncertainty is the win.
17. "Two dashboards show different numbers for the same metric. What do you do?"
What they are assessing: Whether you debug definitions and pipelines methodically.
Model answer: "I treat it as a definition or pipeline problem, not a mystery. First I compare the definitions — time zones, date filters, excluded test accounts, and whether 'revenue' means booked or recognised. That resolves most cases. If definitions match, I trace the data lineage back to the source and run both queries against the raw tables to find where they diverge. Once found, I fix it at the source of truth and document the canonical definition, because if two teams can quietly build competing versions, the conflict will return. Reconciliation drills like this are why I keep metric definitions in one place."
Prepare: Have a short real story of a discrepancy you traced. This question rewards specific war stories.
Dashboards, Data Quality and Stakeholder Questions
18. "How do you design a dashboard people actually use?"
What they are assessing: Whether you design around decisions and audiences, not around available charts.
Model answer: "I start with the question the viewer brings: a leadership dashboard should answer 'are we on track, and what changed?' in the first screen. That means few metrics, clear comparisons to target and last period, and a consistent layout so change is easy to spot. Detail lives one click down, not on the front page. I also set expectations about freshness — stamped update times prevent arguments about numbers. My test is whether the weekly meeting uses the dashboard live; if people bring their own spreadsheets instead, the design has failed and I ask what it is missing."
Prepare: Name a design trade-off you made — for example, dropping a chart a stakeholder liked because it did not support any decision.
19. "How do you check data quality before sharing an analysis?"
What they are assessing: Whether quality checks are a habit, not an afterthought.
Model answer: "I run checks at three levels. Completeness: expected row counts, no unexpected gaps in dates, null rates in key fields. Validity: ranges and categories make sense, no negative prices, no customers created after their orders. Consistency: totals reconcile with a known source like finance or the production system. For anything recurring I automate the checks — a model that fails loudly is better than a silent wrong number. And before sending, I read my own conclusion and ask what would make it wrong; if a plausible failure mode exists that I have not checked, I check it or say so in the write-up."
Prepare: Give one example of a check that caught a real problem. Prevention stories build trust.
20. "Three stakeholders need analyses and you only have time for one. How do you prioritise?"
What they are assessing: Whether you can manage demand rather than silently overloading.
Model answer: "I prioritise by decision impact and deadline, not by who asks loudest. I ask each requester what decision the analysis feeds and what happens if it arrives late — that alone often reveals that one is urgent and two are 'nice to have next week'. I bring the trade-off back to my manager or the stakeholders themselves with a proposed order, and I look for cheaper versions: often a rough cut today serves the decision better than a polished version next week. Over a year of doing this, we also built a self-serve dashboard for recurring questions, which cut ad-hoc requests by roughly a third."
Prepare: Show that you escalate transparently rather than absorb everything — that is the behaviour managers are screening for.
21. "How do you explain a technical finding to a non-technical stakeholder?"
What they are assessing: Translation skill, which separates analysts who influence from analysts who produce reports.
Model answer: "I lead with the decision, not the method: 'we should pause channel X because it loses money on acquisition' comes before any mention of models. I use one simple visual, avoid statistical jargon — 'we would see this pattern by chance less than one time in twenty' instead of naming a p-value — and I state what I am confident about and what is still uncertain. Then I check understanding by asking what they would do with the result, which surfaces confusion early. Method details go in an appendix for anyone who wants them."
Prepare: Practise explaining one of your real analyses in 30 seconds to a friend outside your field. If they can repeat the takeaway, your framing works.
Behavioural and Collaboration Questions
22. "Tell me about a time you made a mistake with data."
What they are assessing: Honesty, ownership, and whether you build defences against repeats.
Model answer: "Early in my current role, I reported signups using a table that logged events on processing date rather than event date, so the last two days of every month looked artificially low. A director questioned the dip, and I initially defended the number before checking. Once I found the cause, I corrected the report, told everyone who had received the wrong version, and quantified the impact of the error. Then I fixed the system: the event-vs-processing distinction went into our metric definitions, and I added a week-over-week sanity alert. The mistake taught me to distrust convenient explanations for anomalies and to verify before defending."
Prepare: Choose a real mistake with a genuine fix. Never blame a tool or a colleague in this answer.
23. "Tell me about a time you disagreed with a stakeholder about the analysis."
What they are assessing: Whether you can hold a position on evidence while staying collaborative.
Model answer: "A product manager wanted to launch based on a positive trend in weekly usage. When I dug in, the growth was almost entirely one enterprise account testing the product — excluding it, usage was flat. I did not want to just contradict him, so I showed him the same chart split by account type and walked through it together. He agreed the evidence did not support the launch, and we ran a small pilot with segmentation instead, which later justified a wider rollout with better data. I learned that disagreement lands better when you bring the person into the analysis rather than delivering a verdict."
Prepare: Structure with the STAR method: situation, task, action, result — with the result showing a preserved relationship and a better decision.
24. "How do you handle pressure to deliver a number that supports a decision already made?"
What they are assessing: Integrity — many analysts face this, and interviewers want to hear a real strategy.
Model answer: "I hold the line on facts while staying helpful on process. If someone wants a number to justify a decision, I offer the honest version plus the strongest legitimate case for their position — better segmentation, the right comparison window — because often the data genuinely supports their goal and the framing was the problem. If the data truly contradicts the decision, I say so clearly to them, and if pressure continues, I escalate to my manager. I would rather have one uncomfortable conversation than attach my name to an analysis built backwards from a conclusion — that reputation is worth more than any single deliverable."
Prepare: Keep the tone constructive: you are showing how you protect the company from bad decisions, not lecturing about ethics.
25. "What questions do you have for us?"
What they are assessing: Curiosity about the role and whether you think like an owner.
Model answer: "Three questions. First, how is data quality handled upstream — is there a data engineering function, or does the analyst own the pipelines too? Second, what decision will this role's analysis influence in the first ninety days, and what does success look like? Third, how do teams here agree on metric definitions — is there a shared semantic layer or documentation, or has that been a pain point?"
Prepare: These questions do double duty: they show maturity, and the answers tell you whether data work here is trusted or constantly disputed.
Questions to Ask the Interviewer
- "What does the path from raw data to a shipped analysis look like here — who owns collection, cleaning, and definitions?"
- "What is the most important decision an analyst influenced in the last six months?"
- "How does the team handle conflicting requests from different departments?"
- "What happened the last time an analysis was wrong, and how was it caught?"
- "What would make you say, one year in, that hiring for this role was a great decision?"
A Seven-Day Preparation Plan
Day 1 — Story and motivation. Write your tell me about yourself answer and your reason for this role. Record yourself once.
Day 2 — SQL drills. Practise joins, window functions, and "top N per group" problems. Say your reasoning out loud while writing, exactly as you will in the interview.
Day 3 — Statistics and judgement. Prepare plain-language answers for correlation, significance, and outliers, each with one real example from your work.
Day 4 — Project walkthroughs. Pick two projects and rehearse the full arc: question, data, method, finding, decision, result — with quantified outcomes.
Day 5 — Stakeholder scenarios. Practise the prioritisation, disagreement, and integrity answers. These decide hiring-manager rounds more than SQL does.
Day 6 — Company research. Read their product and any public metrics. Decide which metrics you would track in the role and why.
Day 7 — Full rehearsal. Run a complete mock interview under real conditions. A free Rehearsa practice session scores your answers, pacing, and delivery, so you walk in knowing your weak spots rather than guessing.
Common Mistakes to Avoid
- Reciting SQL definitions without scenarios. Interviewers ask "how would you…" on purpose. Answer with a method, not a textbook line.
- Presenting findings without caveats. Saying what the analysis cannot claim builds more trust than false certainty.
- Ignoring the clarifying question. Asking "what does 'active' mean here?" is a strength, not a delay.
- Listing every tool you have touched. One tool used well with a judgement story beats a ten-item list.
- Skipping the business outcome. Always close with what changed: the decision, the number, the result.
- Defending a number under pressure. When someone questions your analysis, verify first — then answer.
Final Checklist Before Your Interview
- 60–90 second introduction with one quantified result
- Two project walkthroughs with question → method → decision → impact
- SQL: joins, window functions, deduplication, slow-query reasoning
- Plain-language explanations of correlation, significance, and outliers
- One integrity or disagreement story with a positive outcome
- Questions for the interviewer that probe data maturity
- One full mock interview completed with feedback reviewed
Strong data analysts are hired for judgement as much as technique — and judgement is exactly what practice reveals. Start your free Rehearsa trial to rehearse these answers aloud with an AI interviewer that scores your content, clarity, and delivery, then see what to fix in your Interview Dashboard.