Data engineering interviews are unusually easy to prepare badly for. Candidates revise SQL syntax and Spark internals, walk in, and get asked how they would handle a pipeline that silently produced wrong numbers for three weeks.
The technical questions are real, but they are the entry fee. The decision usually comes down to something else: do you think about data as a product other people depend on, or as a job that runs?
Here is what actually gets asked, and what each question is looking for underneath.

The SQL round
Expect window functions. Almost every data engineer interview has at least one.
Be comfortable with ROW_NUMBER, RANK and DENSE_RANK, and — more importantly — be able to say when you would use each. "Deduplicate by keeping the latest row per key" using ROW_NUMBER() OVER (PARTITION BY id ORDER BY updated_at DESC) is close to a standard exercise.
The two other reliable areas:
Joins that produce more rows than you expected. Be ready to explain why a join fanned out and how you would find the duplicate keys.
Slowly changing dimensions. Know type 1 and type 2, and be able to say plainly why type 2 exists: because "what did this customer's plan look like in March" is a question the business will eventually ask.
The pipeline design round
Usually phrased as "design a pipeline that…" — daily sales, event tracking, something ingesting from an API.
They are watching for whether you ask questions before designing. How much data? How fresh does it need to be? What happens if the source is late or sends the same batch twice? Who consumes it, and what breaks for them if it is wrong?
A candidate who asks about volume and lateness before drawing boxes is already ahead.
Then cover, in your answer:
- Idempotency. If this runs twice, do we get double rows? The single most common real-world failure.
- Late and out-of-order data. What is your watermark, and what happens to the record that arrives four hours late?
- Backfill. Can you re-run last month without hand-editing anything?
- Schema change. A source adds a column, or changes a type. Does the pipeline fail loudly or corrupt quietly? Loudly is the right answer.
The question that decides it
"A dashboard has been showing wrong numbers for three weeks. Nobody noticed. What now?"
This is the one. The answer they want is not just the fix.
"First I would stop the bleeding — flag the dashboard as unreliable so nobody makes another decision on it. Then find when it diverged, using a known-good period to compare against. Then find the cause, and only then backfill. Afterwards, the real work: why did three weeks pass without anyone noticing? That is a monitoring gap, not a pipeline gap. I would add a freshness check and a row-count or distribution check so the next one surfaces in a day."
Communicating to the affected teams belongs in that answer too. Data engineers who mention it stand out immediately.

Batch, streaming and the honest answer
You will be asked whether you would use batch or streaming. The trap is choosing streaming to sound modern.
The right answer starts with the requirement. If the business looks at the number once a morning, streaming adds cost, complexity and a new class of failure for no benefit. If someone is making a decision within seconds, batch cannot serve it. Say which one you would pick and what would change your mind.
Have one story about something breaking
Every data engineer has one — the job that failed silently, the timezone that shifted a day's worth of rows, the upstream team that changed a field without telling anyone.
Tell it in two minutes: what broke, how you found it, what you fixed, and what you put in place so it could not happen the same way twice. That last part is the whole point.
Then say it out loud
Pipeline answers are long. That is what makes them hard to deliver — you know the material and still spend four minutes wandering because you have never had to compress it into speech with someone watching.
Take the wrong-numbers scenario and the pipeline design question, and answer each aloud in under three minutes. Once each. You will hear exactly where you lose the thread.
If you want the follow-ups too, RehearseAI asks them out loud for the role you are actually applying for, then tells you where the answer sagged. Every account gets one free 5-minute interview a month, no credit card required.
Quick answers
How much SQL do data engineer interviews test? A lot, and usually live. Window functions and de-duplication are the most common.
Do I need Spark specifically? Only if the role uses it. Understanding partitioning, shuffles and skew transfers across most engines.
Is dbt worth mentioning? If you have used it, yes — especially for tests and lineage. Do not claim it otherwise.
How do I answer if I have never built a streaming pipeline? Say so, then explain the trade-off you would weigh. Honest plus reasoning beats a rehearsed answer you cannot defend.
Sources
The technical ground here is the cloud vendors' own material; these are the pages I checked it against:
- What is a data engineer? — Google Cloud
- Data architecture guide — Azure Architecture Center, Microsoft Learn
- Microsoft Certified: Azure Data Engineer Associate — Microsoft Learn
If the interview opens with "tell me about a time…", use the STAR method. More on why I built this.



