Senior Data Engineer (software Engineer II) - JP Morgan Chase - HYD
Summary
I cleared three technical rounds and an HR round for a Senior Data Engineer role at JPMorgan Chase and received an offer of 30 LPA fixed plus 12% variable.
Full Experience
Role: Senior Data Engineer (Software Engineer II)
Experience: 5+ years Skills: PySpark, SQL, Python, AWS, Databricks Current CTC: 16 LPA Offer: 30 LPA Fixed + 12% Variable
Pre-Screening – Mobile Call
The initial pre-screening was conducted over a mobile call.
Round 1 – Technical Interview
Interviewed by: 2 interviewers
- Self-introduction and project discussion, followed by questions related to the projects.
- Python: Longest Substring problem using different approaches — brute force, better, and optimal/greedy approaches.
- Spark: Spark architecture and related concepts.
- SQL: Queries involving window functions, aggregations, and other concepts. Difficulty level: Easy to Medium.
- PySpark: Questions on joins, optimization techniques, and related concepts.
- Cluster Designing (30 minutes):
- Given 100 GB of data, design a Spark cluster.
- Questions covered the number of nodes, cluster size, number of tasks, stages, partitions, and detailed execution flow.
Round 2 – In-Person Technical Interview
Interviewed by: 3 interviewers
- Self-introduction and detailed follow-up questions about the projects.
- AWS Lambda Automation + ServiceNow Integration:
- Discussed an end-to-end architecture where a team onboarding request is submitted through a ServiceNow form.
- The request automatically triggers an AWS Lambda function.
- Lambda then creates the required resources in AWS and Databricks.
- Discussed the complete architecture, workflow, integrations, and implementation details.
- Follow-up questions covered coding, connection details, Terraform, resource provisioning, and automation.
- SQL and PySpark:
- Employee salary greater than the average salary.
- Highest-earning employee at the department level.
- Other SQL and PySpark-based problems.
- ETL Architecture:
- Designed an end-to-end ETL architecture.
- Discussed design decisions, data flow, scalability, optimization, and follow-up scenarios.
- Discussion about the company, role, and offer.
- Discussion about the JPMorgan project and related experience.
Final Round – HR
- Discussion about the offer and compensation.
- Discussion about the organization, role, and overall opportunity.
Interview Questions (4)
Longest Substring
Given a string, find the length of the longest substring without repeating characters. Implement three approaches: brute force, an improved method, and an optimal greedy solution.
Design Spark Cluster for 100 GB Data
Given 100 GB of data, design a Spark cluster. Discuss the number of nodes, cluster size, number of tasks, stages, partitions, and the detailed execution flow required to process the data efficiently.
Employees with Salary Greater Than Average
Write a SQL query to list all employees whose salary is greater than the average salary of all employees in the company.
Highest Earning Employee per Department
Write a SQL query to find the employee with the highest salary in each department.