Amazon | Applied Scientist Intern (6 Months) | Interview Experience
Summary
I cleared two interview rounds consisting of DSA and ML breadth & depth questions and received an offer.
Full Experience
Role: Applied Scientist Intern (6 Months) Company: Amazon Rounds: 2 Rounds (1 DSA + 1 ML Breadth & Depth)
Round 1: Data Structures & Algorithms (2-DSA questions)
1. Given an array of integers, maximize the score of the array by reversing at most one contiguous subarray. The score of an array is defined as the sum of absolute differences between consecutive elements.
2. Given a grid where 0 represents a wall and 1 represents free space, find the minimum number of moves to reach from the top-left corner (0,0) to the bottom-right corner (R-1, C-1). You can move in 4 directions and are allowed to break at most one wall.
Round 2: ML Breadth & Depth
Brief introduction and deep-dive into project architecture (combining CNN, GRU, and Attention mechanisms). What are residual connections, and why did you use them in your architecture? What is the vanishing gradient problem? What are the common solutions to mitigate it? Multiple follow-up in-depth questions based on the specific project implementation details, design decisions, and trade-offs.
Why do we avoid using the Sigmoid activation function in the hidden layers of neural networks? Compare different activation functions (Sigmoid, Tanh, ReLU, Leaky ReLU) along with their pros and cons. In Leaky ReLU, why do we keep small alpha value? What happens if we set a large alpha value?
What happens to the number of channels across dense blocks in DenseNet? How do you compute the output dimension when applying a specific filter size to an image of given dimensions (gave a numerical question, don't remember exactly). Can small filters (1 x 1 convolutions) learn meaningful features in an image?
What are different types of loss functions, and when/why are they used? Explain Categorical Cross-Entropy (CCE). How does it mathematically penalize wrong predictions severely?
What are the different Transformer architectures you know (Encoder-only, Decoder-only, Encoder-Decoder)? How does the attention mechanism differ across these architectures? In masked attention, what is the preferred sequence of operations: applying masking before attention, or applying attention before masking? Why? Where exactly does learning happen (where are the trainable parameters) in a Transformer?
What is the concept of kernels in Support Vector Machines (SVM)? What are the different types of kernels, and when would you use each? Additional technical depth questions extending from project discussions into core Machine Learning concepts.
Interview Questions (11)
Maximize Array Score by Reversing a Subarray
Given an array of integers, maximize the score of the array by reversing at most one contiguous subarray. The score of an array is defined as the sum of absolute differences between consecutive elements.
Minimum Moves in Grid with One Wall Break
Given a grid where 0 represents a wall and 1 represents free space, find the minimum number of moves to reach from the top-left corner (0,0) to the bottom-right corner (R-1, C-1). You can move in 4 directions and are allowed to break at most one wall.
Residual Connections in Neural Architecture
What are residual connections, and why did you use them in your architecture?
Vanishing Gradient Problem and Solutions
What is the vanishing gradient problem? What are the common solutions to mitigate it?
Avoiding Sigmoid Activation in Hidden Layers
Why do we avoid using the Sigmoid activation function in the hidden layers of neural networks? Compare different activation functions (Sigmoid, Tanh, ReLU, Leaky ReLU) along with their pros and cons. In Leaky ReLU, why do we keep small alpha value? What happens if we set a large alpha value?
Channel Growth in DenseNet Dense Blocks
What happens to the number of channels across dense blocks in DenseNet?
Computing Output Dimensions After Convolution
How do you compute the output dimension when applying a specific filter size to an image of given dimensions?
Effectiveness of 1x1 Convolutions
Can small filters (1 x 1 convolutions) learn meaningful features in an image?
Loss Functions and Categorical Cross-Entropy
What are different types of loss functions, and when/why are they used? Explain Categorical Cross-Entropy (CCE). How does it mathematically penalize wrong predictions severely?
Transformer Architectures and Attention Details
What are the different Transformer architectures you know (Encoder-only, Decoder-only, Encoder-Decoder)? How does the attention mechanism differ across these architectures? In masked attention, what is the preferred sequence of operations: applying masking before attention, or applying attention before masking? Why? Where exactly does learning happen (where are the trainable parameters) in a Transformer?
Kernels in Support Vector Machines
What is the concept of kernels in Support Vector Machines (SVM)? What are the different types of kernels, and when would you use each?