Amazon Data Engineer Interview Experience | 2026
Summary
I interviewed for a Data Engineer role at Amazon and completed a technical assessment followed by a technical interview covering SQL, ETL, and data‑warehouse concepts.
Full Experience
Hi everyone,
I interviewed with Amazon a couple of weeks ago for a Data Engineer role requiring 1+ years of experience. I wanted to share my interview experience in case it helps others prepare.
Round 0: Technical Assessment This was a 90‑minute online assessment consisting of:
- 10 MCQs
- 2 medium‑level SQL questions
The SQL questions were focused on query‑writing skills and required a good understanding of SQL fundamentals.
Round 1: Technical Interview (Elimination Round)
Topics Covered: ETL workflow design Database concepts and architecture
The interview started with a brief introduction and a discussion about my resume. Since I had mentioned Apache Spark, the interviewer immediately asked questions related to Spark joins and data processing concepts. Questions Asked:
- Difference between: Physical Joins, Hash Join, Broadcast Nested Loop Join
- Difference between GROUP BY and Window Functions
- Explain Slowly Changing Dimensions (SCD Types)
- Difference between Star Schema and Snowflake Schema
- How do you optimize SQL queries?
- Difference between: Data Warehouse, Data Lake, Lakehouse
- Difference between traditional batch architecture and event‑driven incremental ingestion
- Questions related to Idempotency
- Questions related to Change Data Capture (CDC)
SQL Question At the end, I was given a live SQL coding problem on consecutive user logins.
Read and understand every concept thoroughly. Make sure your fundamentals are crystal clear, especially around SQL, ETL, data warehousing concepts.
All the best !
Interview Questions (10)
Physical Joins vs Hash Join vs Broadcast Nested Loop Join
Explain the differences between Physical Joins, Hash Join, and Broadcast Nested Loop Join, including when each join type is used and their performance characteristics.
GROUP BY vs Window Functions
Describe the difference between using GROUP BY and Window Functions in SQL, and give examples of scenarios where a window function is preferred.
Slowly Changing Dimensions (SCD Types)
Explain the concept of Slowly Changing Dimensions and detail the different SCD types (Type 0, Type 1, Type 2, Type 3, etc.) and their use cases.
Star Schema vs Snowflake Schema
Compare Star Schema and Snowflake Schema in data warehousing, highlighting the structural differences and trade‑offs in query performance and normalization.
SQL Query Optimization Strategies
Discuss various techniques to optimize SQL queries, such as indexing, query rewriting, using appropriate joins, and analyzing execution plans.
Data Warehouse vs Data Lake vs Lakehouse
Explain the differences between a Data Warehouse, a Data Lake, and a Lakehouse architecture, covering storage, schema enforcement, and typical use cases.
Batch Architecture vs Event‑Driven Incremental Ingestion
Contrast traditional batch processing architectures with event‑driven incremental ingestion pipelines, focusing on latency, complexity, and scalability.
Idempotency in Data Processing
Define idempotency in the context of data pipelines and discuss how to achieve idempotent operations in ETL workflows.
Change Data Capture (CDC)
Explain what Change Data Capture is, how it works, and common techniques/tools used to implement CDC in data pipelines.
Consecutive User Logins SQL Query
Write a SQL query to find users who have logged in on consecutive days (e.g., day N and day N+1). The query should return the user id and the dates of the consecutive logins.