Amazon | Applied Scientist Intern (6 Months) | Interview Experience

interview experience logo
interview experience
· Applied Scientist Intern (6 Months)
September 19, 2026 · 2 reads

Summary

I cleared two interview rounds consisting of DSA and ML breadth & depth questions and received an offer.

Full Experience

Role: Applied Scientist Intern (6 Months) Company: Amazon Rounds: 2 Rounds (1 DSA + 1 ML Breadth & Depth)

Round 1: Data Structures & Algorithms (2-DSA questions)

1. Given an array of integers, maximize the score of the array by reversing at most one contiguous subarray. The score of an array is defined as the sum of absolute differences between consecutive elements.

2. Given a grid where 0 represents a wall and 1 represents free space, find the minimum number of moves to reach from the top-left corner (0,0) to the bottom-right corner (R-1, C-1). You can move in 4 directions and are allowed to break at most one wall.

Round 2: ML Breadth & Depth

Brief introduction and deep-dive into project architecture (combining CNN, GRU, and Attention mechanisms). What are residual connections, and why did you use them in your architecture? What is the vanishing gradient problem? What are the common solutions to mitigate it? Multiple follow-up in-depth questions based on the specific project implementation details, design decisions, and trade-offs.

Why do we avoid using the Sigmoid activation function in the hidden layers of neural networks? Compare different activation functions (Sigmoid, Tanh, ReLU, Leaky ReLU) along with their pros and cons. In Leaky ReLU, why do we keep small alpha value? What happens if we set a large alpha value?

What happens to the number of channels across dense blocks in DenseNet? How do you compute the output dimension when applying a specific filter size to an image of given dimensions (gave a numerical question, don't remember exactly). Can small filters (1 x 1 convolutions) learn meaningful features in an image?

What are different types of loss functions, and when/why are they used? Explain Categorical Cross-Entropy (CCE). How does it mathematically penalize wrong predictions severely?

What are the different Transformer architectures you know (Encoder-only, Decoder-only, Encoder-Decoder)? How does the attention mechanism differ across these architectures? In masked attention, what is the preferred sequence of operations: applying masking before attention, or applying attention before masking? Why? Where exactly does learning happen (where are the trainable parameters) in a Transformer?

What is the concept of kernels in Support Vector Machines (SVM)? What are the different types of kernels, and when would you use each? Additional technical depth questions extending from project discussions into core Machine Learning concepts.

Interview Questions (11)

1.

Maximize Array Score by Reversing a Subarray

Data Structures & Algorithms

Given an array of integers, maximize the score of the array by reversing at most one contiguous subarray. The score of an array is defined as the sum of absolute differences between consecutive elements.

2.

Minimum Moves in Grid with One Wall Break

Data Structures & Algorithms

Given a grid where 0 represents a wall and 1 represents free space, find the minimum number of moves to reach from the top-left corner (0,0) to the bottom-right corner (R-1, C-1). You can move in 4 directions and are allowed to break at most one wall.

3.

Residual Connections in Neural Architecture

Other

What are residual connections, and why did you use them in your architecture?

4.

Vanishing Gradient Problem and Solutions

Other

What is the vanishing gradient problem? What are the common solutions to mitigate it?

5.

Avoiding Sigmoid Activation in Hidden Layers

Other

Why do we avoid using the Sigmoid activation function in the hidden layers of neural networks? Compare different activation functions (Sigmoid, Tanh, ReLU, Leaky ReLU) along with their pros and cons. In Leaky ReLU, why do we keep small alpha value? What happens if we set a large alpha value?

6.

Channel Growth in DenseNet Dense Blocks

Other

What happens to the number of channels across dense blocks in DenseNet?

7.

Computing Output Dimensions After Convolution

Other

How do you compute the output dimension when applying a specific filter size to an image of given dimensions?

8.

Effectiveness of 1x1 Convolutions

Other

Can small filters (1 x 1 convolutions) learn meaningful features in an image?

9.

Loss Functions and Categorical Cross-Entropy

Other

What are different types of loss functions, and when/why are they used? Explain Categorical Cross-Entropy (CCE). How does it mathematically penalize wrong predictions severely?

10.

Transformer Architectures and Attention Details

Other

What are the different Transformer architectures you know (Encoder-only, Decoder-only, Encoder-Decoder)? How does the attention mechanism differ across these architectures? In masked attention, what is the preferred sequence of operations: applying masking before attention, or applying attention before masking? Why? Where exactly does learning happen (where are the trainable parameters) in a Transformer?

11.

Kernels in Support Vector Machines

Other

What is the concept of kernels in Support Vector Machines (SVM)? What are the different types of kernels, and when would you use each?

📣 Found this helpful? Please share it with friends who are preparing for interviews!

Discussion (0)

Share your thoughts and ask questions

Join the Discussion

Sign in with Google to share your thoughts and ask questions

No comments yet

Be the first to share your thoughts and start the discussion!