Virtusa | Data Scientist | On-Campus | Rejected at Final L2 Round (Deep Technical ML Intuition)
Summary
I interviewed for a Data Scientist L1 role at Virtusa, completed two technical rounds but was rejected after the final L2 interview.
Full Experience
Status: Final year B.Tech CSE student at IIIT Bhagalpur Position: Data Scientist (L1 Role) Result: Rejected after L2 (Only 1 out of 4 final candidates pushed to HR)
Online Assessments (OAs)
- Round 1: General Technical/Coding Assessment covering foundational concepts, Aptitude, and CS fundamentals. Out of around 100 students, only 7 were shortlisted.
- Round 2 (SpeechX): An automated behavioral and communications test evaluating spoken English, parsing abilities, clarity, and articulation.
Technical Round 1 (L1 Interview) - ~27 Mins
The interviewer was highly supportive and friendly. The round felt like a rapid-fire technical assessment covering my projects and core Machine Learning theory.
- Project Deep Dive: Asked me to explain my unsupervised machine learning project (Earthquake Clustering System). We discussed the codebase, the exact data pipeline, and why I selected specific clustering metrics.
- Core ML Cross-Questions:
- DBSCAN vs. K-Means (When to use density-based vs. centroid-based approaches?).
- Write out/explain the inner mathematical algorithm of K-Means.
- Supervised vs. Unsupervised Learning (What fundamentally constitutes labeled vs. unlabeled data?).
- Linear Regression vs. Logistic Regression (Equations, cost functions, and use cases).
- Decision Trees vs. Random Forests (Entropy, Gini Index, and handling variance).
- GenAI / RAG Discussion: I had mentioned learning Generative AI on my resume. He asked for the end-to-end architecture flow of a Retrieval-Augmented Generation (RAG) system. I walked him through: User Query -> Embedding Models -> Semantic Similarity Search via Vector Databases -> Top-K context retrieval -> LLM generation.
- Live Coding Question: A basic Python data structures / manipulation problem (finding duplicates/manipulating list data) that I managed to solve seamlessly.
- Closing: I mentioned my current learning path regarding advanced AI agents (LangGraph, Model Context Protocol - MCP servers). The interviewer was highly impressed by the initiative and communication.
Technical Round 2 (L2 Interview) - ~1 Hour
This round was conducted by an industry veteran with 18+ years of experience. There were no standard coding, DSA, SQL, or OOP definition questions. The entire interview was heavily practical, business-centric, and focused strictly on systemic ML engineering.
- Project Discussion: I initially walked through my RAG-based PDF answering platform pipeline.
- The Trap: He immediately asked precisely which embedding model I used. My mind completely went blank under the pressure and I couldn't recall the exact name.
- Dimensionality Reduction Question:
- Scenario: Suppose you have a highly dimensional data representation, and I show you a plot where the X-axis is the Population of India and the Y-axis is Gender. What technique would you use to translate complex high-dimensional semantic spaces into a clean visual?
- Answer: PCA (Principal Component Analysis). He probed into why and how the variance mapping works.
- Aptitude Question:
- If x = A% of B and y = B% of A, what is the exact mathematical relationship between x and y?
- Answer: x = y. He calmly asked me three separate times in different phrasings to check if I was guessing or confident in the math.
- The Core Case Study: Real Estate Flat Price Prediction (30 Mins):
- Scenario: He explained he is a broker selling flats in Bangalore and needs an ML system to predict the price of a flat today versus its trajectory over 10 years.
- Question 1: Regression or Classification? (Answer: Regression).
- Question 2: List exactly 4 baseline features to collect. (Answer: Size in sq ft, Property Age, BHK, Locality/Location).
- Question 3: Write the predictive model equation. I began writing out the generic mathematical equation (y = b0 + b1x1...). He stopped me immediately: "I don't want a generic math textbook definition. You are solving a real business problem. Write it out in terms of actual product features." (Correct expected approach: Predicted_Price = w0 + w1(Size) + w2(BHK)...).
- Question 4: Identify a feature that contributes negatively to the price. (Answer: Property Age).
- Question 5: How does the machine automatically figure out if a dynamic categorical variable like Location is positive or negative without hardcoding? I struggled here, providing abstract answers about the model calculating weights automatically during backpropagation/error minimization. The precise framework he was fishing for was Feature Engineering (specifically structural implementations like Target Encoding, Mean Encoding, or designing external numeric domain scores like proximity to transit hubs, crime indices, or regional infrastructure metrics).
My Takeaway & Key Learnings
- Forget rote definitions for Senior interviews: Senior interviewers do not care if you can define what Linear Regression is. They want to know if you can convert a chaotic business problem into a clean, operational ML pipeline.
- Know every single line of your tech stack: Blanking on the specific model name in my GenAI project cost me crucial confidence points early in the L2 round.
- Double down on Feature Engineering: Knowing algorithms isn't enough; you must know how to restructure raw real-world data format types so a model can actually learn from them effectively.
Interview Questions (7)
Visualize high-dimensional data using PCA
Scenario: You have a highly dimensional data representation and need to translate complex high-dimensional semantic spaces into a clean visual (e.g., a plot with X-axis as Population of India and Y-axis as Gender). What technique would you use?
Relationship between x = A% of B and y = B% of A
Aptitude question: If x = A% of B and y = B% of A, what is the exact mathematical relationship between x and y?
Select model type for flat price prediction
Case study scenario: Predict the price of a flat today versus its trajectory over 10 years. Should you use a regression or classification model?
Baseline features for flat price prediction
Case study: List exactly four baseline features to collect for predicting flat prices.
Write predictive model equation using actual product features
Case study: Write the predictive model equation for flat price prediction using real product features instead of a generic formula.
Feature negatively impacting flat price
Case study: Identify a feature that contributes negatively to the flat price prediction.
Automatic handling of categorical variable impact
Case study: How does the machine automatically determine if a dynamic categorical variable like Location positively or negatively affects price without hardcoding?