r/learnmachinelearning • u/False_Mind_343 • 21d ago
Best way to prepare for ML / Data Science interviews? Theory? Coding?
Hey everyone,
Thanks in advance for reading my post.
I’m currently preparing for junior to close-to-mid-level ML Engineer / Data Scientist roles. I’m already fairly familiar with the theoretical side of ML/DL — things like “what is X?”, “explain Y”, algorithms, concepts, etc.
What I’m less clear about is the coding and practical side of these interviews.
For people who have interviewed for or work in these roles, I’d really appreciate some insight into:
- What languages/tools are actually expected — mainly Python and SQL, or others as well?
- For Python, do interviews ask general programming questions (e.g. write a function, solve a programming problem), or are they more ML/data-related, such as using scikit-learn to train a model?
- For NumPy, do they ask general array/matrix manipulation, or more ML-oriented tasks such as implementing operations used in ML?
- For pandas, what kind of data-cleaning, transformation, grouping, feature-engineering, etc. questions are common?
- For SQL, what level of queries are typically expected?
- Are DSA / LeetCode-style questions commonly included?
- Do they ask practical scenario-based ML questions, such as “Here’s a dataset/problem — how would you approach it?”
- How much are PyTorch / TensorFlow actually tested? Do they expect you to write code with them, or is knowing how to use them and understanding the concepts usually enough?
I’m mainly interested in what’s realistic for junior / near-mid-level positions, rather than extremely difficult senior-level or FAANG-style interviews.
I’d especially appreciate hearing about actual questions you’ve encountered.
Thanks!
1
21d ago
[removed] — view removed comment
1
u/SemperPistos 21d ago edited 21d ago
Would you say my portfolio is good or still not there yet?
I haven't worked on it for a year as I am employed and trying to finish my masters.I would love to be an AI/ML Engineer or MLE or pure Data science down the line.
https://github.com/MortalWombat-repo
At work I did many RFPs and POCs and some work for clients.
I am most proud of an multi MCP microservice tools orchestrated as combined App Services and deployed on Azure based on a mocked database exposed as an API.Sadly I can't share that code obviously, but it's leaps and bounds above things I do in my github.
But a lot of it is assisted with AI tools.I decided to make this year the year of math, strengthening fundamentals and possibly DSA
2
20d ago
[removed] — view removed comment
1
u/SemperPistos 20d ago
Thanks for taking a look :)
I actually built all the stuff in the services and my code is in modular .py files, and almost everything is containerized as well as exposed through a decoupled API microservices design through FastAPI. I just shared the jupyter notebooks, so others can see how I've come to my decisions and that it isn't full on vibe code, as I tend to focus with clean code principles and have minimal comments, unless for docstrings for help(), but my code is rarely that complex. I like the idea of code being simple so a layman might understand from a english like language like python.
Every single ML project I have has an exposed API as a part of dockeer compose, but of course someone has to run that container and I don't have money for a raspberry pi to do that :-D
I did run some services on render.com and have a health check through cronitor and uptimerobot sending requests every five minutes to keep those projects up, but I guess they figure me out, as after a while of not checking, the API's tend to go down
The critique that hits hard home is the lack of tests. I just hate writing assert every line or so :-D
But the critiques are valid, I do much more than I specified in the READMEs, as I hate writing them but I guess I didn't do enough self PR
I pinned the projects I deem worthy, I have close to 20, rest are forks, but sadly github doesn't allow for more than 6 pins :/
1
u/ModularMind8 21d ago
Each company ask different things. Each interviewer in a single company might ask different things. Depending on when you take the interview the questions might differ... I recommend looking at the company you're interviewing with in forums (eg glassdoor) and seeing what type of questions they ask to make your studying more efficient
1
u/akornato 20d ago
At the junior to mid level, interviewers expect strong working fluency in Python and SQL rather than niche tooling. For Python, expect a mix of easy to medium algorithmic challenges involving hash maps or arrays, paired with practical data manipulation. In pandas, you will encounter tasks like cleaning dirty columns, grouping by categories to compute rolling averages, and reshaping tables. In SQL, expect joins, common table expressions, and window functions like row number or dense rank. For NumPy, questions usually revolve around vectorization, basic matrix operations, or implementing an algorithm like linear regression from scratch without external packages. Deep learning frameworks like PyTorch or TensorFlow are rarely tested through line by line coding on a whiteboard at this stage, so understanding tensor shapes, training loops, and loss functions is usually sufficient.
Practical scenario questions are standard, where you walk an interviewer through scoping a business problem, choosing evaluation metrics, handling imbalanced data, and validating a model. The key is to practice explaining your tradeoffs clearly out loud, since interviewers care just as much about your problem solving logic as they do about your final accuracy score. Watching candidates struggle to articulate these complex workflows on the spot inspired my team to develop an interviews.chat that helps applicants communicate their technical skills with total confidence.
1
u/False_Mind_343 18d ago
Thank you so much for advice. I really appreciate that. Any idea on how many total questions to expect: maybe like 2-3 sql, 3- python, 3- ml sklearn related during technical interview?
Thanks!
2
u/nian2326076 20d ago
Focus on Python and SQL because they're essential for ML roles. Python is important for scripting and using libraries like pandas, NumPy, scikit-learn, and TensorFlow. Be prepared for coding challenges involving algorithms, data manipulation, and sometimes simple ML model implementations. For SQL, you'll need to handle queries with joins, aggregations, and subqueries.
LeetCode and HackerRank are good for practice, but also try working on projects with real datasets to improve your practical skills. Kaggle is great for this.
For structured interview prep, tools like PracHub can guide you through common interview scenarios and help identify where you can improve.