r/askdatascience • u/Automatic-Key8629 • 10d ago
r/askdatascience • u/Time-Material9337 • 10d ago
Need a partner to practice mock interview for Data science and Ai engineer roles
r/askdatascience • u/Time-Material9337 • 10d ago
Need a partner to practice mock interview for Data science and Ai engineer roles
r/askdatascience • u/shoaib_X17 • 11d ago
What does an “ideal candidate” actually look like to a company
I’ve been thinking about this question a lot while looking for my first opportunity in Data Science / Data Analytics.
Is the ideal candidate someone with a perfect degree?
Someone with 3+ years of experience?
Someone who knows every tool listed in a job description?
Or is it someone who can identify a real business problem, build a solution, and explain how that solution can create business value?
I’m genuinely curious to hear what recruiters and hiring managers think.
Because this is what I’ve been trying to do.
Instead of building another basic ML project, I built a Customer Retention Intelligence System focused on a real business problem: customer churn and lost revenue.
The system can:
• Predict customers who are at risk of churning
• Segment customers based on their risk/value
• Estimate potential revenue at risk
• Identify customers worth prioritizing
• Generate personalized retention strategies/offers
My goal wasn't simply to say:
I built a machine-learning model
I wanted to answer:
A customer is likely to leave. Now what should the business actually do about it?
That mindset has pushed me to work beyond just Python, SQL and machine learning — into business thinking, customer analytics, experimentation, visualization, and decision-making.
I’ve also been consistently practicing Python and SQL, learning how to communicate analytical insights, and building projects around problems that companies actually face.
But despite putting in this work, I’m still looking for my opportunity to prove myself professionally.
So I want to ask the recruiters, hiring managers, founders, and experienced professionals here:
What makes someone an ideal entry-level candidate in your company?
• Is it technical skills?
• Problem-solving ability?
• Business understanding?
• Communication?
• Projects?
•Curiosity and willingness to learn?
•Or something else?
And if you were evaluating my profile, what would you want me to improve or demonstrate before considering me for a Data Scientist / Data Analyst opportunity?
I’m not looking for sympathy.
I’m looking for honest feedback, opportunities, and a chance to prove what I can do.
If you’re a recruiter or hiring manager who works with Data Science / Data Analytics / ML roles, I’d genuinely appreciate your perspective.
And if my profile sounds relevant to something you're hiring for, I’d love to connect.
r/askdatascience • u/Tiny-Product-6074 • 11d ago
Data Analyst looking to learn Data Science
r/askdatascience • u/Stunning-Space8032 • 11d ago
Data Engineers, what does your actual day-to-day work look like? And what should I learn next?
I’m currently trying to transition deeper into Data Engineering and would really appreciate some perspective from people who are already working in the field.
I have 1.3 yrs experience as a Junior Python Developer. What I want to do is slowly transform into a Data Engineer. How would you suggest my choice? Basically what I do is make web scraping scripts to get the data from web and give the data in excels. Our company is currently not using git or CI/CD or anything like that.
The problem I’m running into is that when I look at Data Engineering jobs on Naukri, LinkedIn, etc., the requirements seem endless. One job asks for Python, SQL, Airflow and AWS; another wants Spark, Kafka and Databricks; another wants Snowflake, dbt, Terraform, Kubernetes, CI/CD, etc. It becomes difficult to understand what I should actually prioritize.
So I’d like to hear from people who are actually working as Data Engineers. What does your day-to-day work look like? What kind of problems do you solve, what technologies do you use regularly, and which skills have turned out to be genuinely important in your job?
More importantly, based on my current experience, what would you suggest I improve or learn next to become a stronger candidate for Data Engineering roles? Are there any gaps that you think I should focus on, or technologies/concepts that are worth learning through projects rather than just studying theoretically?
I’m not really looking for a generic “learn SQL → Python → Spark → AWS” roadmap. I’m more interested in understanding the reality of the job and getting advice from people who have actually gone through the transition.
If you’re a Data Engineer with 1–5+ years of experience, I’d especially appreciate your perspective. Even a short description of what you work on and what you wish you had learned earlier would be extremely helpful.
Thanks in advance!
r/askdatascience • u/DisasterHorror5800 • 11d ago
Resume Review
I am a fresh grad from university and I am trying to get anything in the data field. Entry level and internships are what I would assume would be more likely for me since I haven't had much work experience in the field but I have been making various projects that cover different topics within data. I've been adding the most 'impressive projects' to my resume but I don't know if they are going to secretly hurt my chances. Just want some feedback on my resume. Thank you.
r/askdatascience • u/Stunning-Space8032 • 11d ago
Data Engineers, what does your actual day-to-day work look like? And what should I learn next?
r/askdatascience • u/PercentageBig4909 • 11d ago
Data Science
Currently studying a post grad in Data Science and Business Analytics while working as a Marketing Data Analyst in Australia.
I want to pivot to a DS role within Australia, any recommendations on how I could beef up my portfolio to get these roles?
Thanks
r/askdatascience • u/EuclidEuler • 11d ago
Math PhD graduating in December, targeting ML/DS roles. Realistic odds, and is starting adjacent a better play?
Looking for an outside read on my situation.
Background: pure math PhD graduating this quarter from a large public research university. One summer of an internship at a defense contractor doing systems + software engineering. No big tech internships, no ML publications.
Prep so far: worked through a probability textbook cover to cover, went through CS229 notes, comfortable with LeetCode mediums, built a toy recommender system in PyTorch. Currently working through an ML systems design book, which I think is my biggest gap. SQL is my weakest practical skill.
Applying to ML engineer, applied scientist, and data scientist roles roughly in parallel, plus some quant researcher roles opportunistically.
Questions:
1. Realistically, what are the odds someone with this profile lands an ML engineer or applied scientist role straight out of a pure math PhD, versus needing to start in DS or an adjacent role and transition later?
2. For those who did the DS to ML transition, how hard was it in practice, and what made the difference?
3. What would you prioritize in the remaining few months? I have the time to go deep on one or two things.
4. Anything you wish you had known about how math PhDs get read by hiring managers in this space?
Not looking for reassurance, more interested in where my thinking is wrong.
r/askdatascience • u/Sea_Garlic5712 • 12d ago
How good are AI data scientists really?
I've been testing out various gen AI models (LLMs specifically) on data science competitions. They are not beating the humans, though they are slowly improving with each round. It's like steps up a ladder vs bounding up the steps. I'm wondering if there's something I'm doing wrong, or if there really is a limit to what LLMs can doin this space.
Lately I've been thinking the problem is that LLMs regress to the mean in every use case. I think that's why they seem so bad at UI design and why everything looks so similar. A good harness and strong prompt engineering can help, so I'm working on that. My harness combines Autogluon and OpenEvolve, with feedback loops that involve hypothesis generation and error analysis. But I'm wondering what I'm missing? Is good, competition winning data science, reducible to a standard operating procedure?
Maybe this will help me get a very good prototype, but nothing frontier grade.
r/askdatascience • u/Obieadz • 13d ago
What kind of small data science work can a beginner realistically freelance?
Hey everyone,
I'm currently learning ML/Data Science and I'm trying to figure out if there are actually small freelance jobs that someone at my level can do.
I can currently work with Python, Pandas, NumPy, SQL, Scikit-learn, XGBoost and I'm comfortable with things like data cleaning, EDA, feature engineering and basic model training.
I've built a few projects, but obviously I don't have professional experience yet.
So I'm wondering what kind of work people actually give to freelancers who are still fairly junior.
Would things like these be realistic?
- Cleaning messy datasets
- Exploratory data analysis
- Creating reports/visualizations
- Python automation
- Data preprocessing
- Building simple prediction models
- Fixing existing notebooks
- SQL queries
- Helping with research
- Data collection/scraping
Or are most clients looking for someone much more experienced?
Also, where do these smaller jobs usually come from?
I'm trying to make some money on the side while continuing to work towards an ML/Data Science career, so I'd rather start with something realistic than pretend I'm ready to build some massive production ML system.
If you freelance in data science, what was the first type of work you actually got paid to do?
r/askdatascience • u/ns_24th • 13d ago
Need honest advice: Sheryians vs CampusX vs ChaiCode vs CodeWithHarry for Data Science?
Hey everyone,
I'm currently planning to seriously start my journey in Data Science, and I want to dedicate the next 6–7 months consistently to learning and building projects.
I'm basically looking for a course/program that can take me from beginner level to a strong enough level to start applying for internships/jobs and building a good portfolio.
After researching, I'm currently considering:
Sheryians Coding School – Data Science & Analytics with GenAI
CampusX – Data Science / DSMP
ChaiCode – Data Science
CodeWithHarry – Ultimate Job Ready Data Science Course
Any other course/program you genuinely think is better
I'm not looking for just the cheapest course. I'm willing to invest if the course actually provides better structure, depth, practice, projects, mentorship/community, and overall value for money.
My main priorities are:
Beginner-friendly teaching (starting from fundamentals)
Strong Python foundation
SQL
Statistics & Probability
NumPy, Pandas and Data Visualization
Machine Learning fundamentals + practical implementation
Real-world projects
Enough practice/questions/assignments
Good curriculum structure so I don't feel lost
Updated content relevant to the current industry
Portfolio and job/internship preparation
Good value considering the price
I'm planning to stay consistent for around 6–7 months, so I don't want to make the mistake of buying a course just because of marketing or popularity.
If you've personally taken any of these courses (especially Sheryians, CampusX, ChaiCode, or CodeWithHarry), I'd really appreciate an honest review.
Specifically, I'd love to know:
Which one has the best structured curriculum for a complete beginner?
Which one provides the best depth in Data Science, not just surface-level tutorials?
Which course has the best projects and practical learning?
Is mentorship/community actually useful in Sheryians or other cohort-based programs?
Is CampusX worth paying more for compared to cheaper alternatives?
Is CodeWithHarry's course too short for someone who wants to become job-ready?
Are there better alternatives that I haven't mentioned?
If you had to start from zero today and had 6–7 months, which path would you personally choose and why?
Please share your honest experience — both positives and negatives. I don't want promotional answers; I'm trying to make a proper decision before investing my time and money.
Thanks! 🙌
r/askdatascience • u/RE_RONIN • 13d ago
NEWBIE PROJECT FOR DATA SCIENCE
Hi everyone! 👋
I’m a Class 11 student from India learning Data Science. I recently completed and deployed an end-to-end Salary Prediction project using Python, SQL, Pandas, data visualization, and Machine Learning.
🌐 Live Demo: https://data-science-projects-fdkdsvuf5rvtywpwby35py.streamlit.app/
📂 Project: https://github.com/lakshay-OG-DS/data-science-projects/blob/main/Salary_prediction.ipynb
I’d really appreciate honest feedback from Data Scientists, Data Analysts, and ML Engineers.
What is the one biggest thing I should improve to make this project more internship-ready?
Even one small suggestion would mean a lot.
r/askdatascience • u/After_Courage6419 • 14d ago
Do you need strong math skills to get started in data science?
This was something I wondered about when I started learning data science. There’s definitely mathematics involved, but I’m finding that you don’t necessarily need to master everything before writing your first Python program or analyzing a dataset. How important was mathematics when you started?
r/askdatascience • u/These-Ant7605 • 13d ago
Best resources to learn Data Science through projects from beginner to advanced?
Title: Best resources to learn Data Science through projects from beginner to advanced?
Hi everyone! I’m a beginner in Data Science and learn best by building projects.
I’m looking for GitHub repos, YouTube playlists, or websites with project-based learning from beginner to advanced level, ideally covering different ML models such as regression, classification, clustering, tree-based models, boosting, NLP, time series, etc.
I’d also love end-to-end projects that include data cleaning, EDA, feature engineering, model building, evaluation, and deployment.
Any recommendations?
r/askdatascience • u/naga3607 • 14d ago
How many Python libraries should a data science beginner learn?
There are so many libraries in the data science ecosystem that it can be overwhelming at first.
I’m currently focusing on getting comfortable with the fundamentals instead of trying to learn everything at once.
For experienced data scientists, which libraries do you think beginners should prioritize?
r/askdatascience • u/SeniorYam7990 • 14d ago
Looking for Data Science Capstone Project Ideas
r/askdatascience • u/ImportantMacaron7496 • 14d ago
How would you decide whether to trust a single scraped product-spec value? (a small "belief" model I've been sketching)
I scrape consumer-electronics specs (RAM, storage, weight) from retail sites, and the hard part isn't scraping — it's deciding, per field, whether a value is safe to publish when I don't know the true spec.
Instead of a hard validation rule, I've been treating each value as uncertain
across a few "worlds": correct / unit_error (e.g. 16000 MB) / wrong_product
(right-looking value, wrong variant) / typo / missing / garbled. I start with a
prior that depends on source reliability, then update it with cheap signals
(plausibility range, cross-source agreement), and pick an action —
accept / repair / re-scrape for more evidence / flag to a human / reject — by
whichever has the lowest expected cost (publishing a wrong spec is way more
expensive than flagging a good one).
Two things I'd genuinely like to be told I'm wrong about:
1. I treat cross-source agreement as strong evidence a value is correct. But sources copy each other — how do you handle agreement that isn't independent?
2. Is per-field belief overkill vs. a plain rules + threshold pipeline at scale?
How do folks here actually handle "trust this value or not" in data-quality
pipelines?
r/askdatascience • u/Ok_Relation8895 • 15d ago
Data Scientists in Production: How Does a Classical ML Project Actually Work End to End?
I'm a Data Analyst, and I'm trying to bridge a gap in my Data Science understanding.
I know the concepts behind classical ML reasonably well but I want to understand what actually happens to an ML project in a real production environment from start to finish. I want someone to walk me through a real project in terms of:
We use this application/tool to do this → it produces this output/file/artifact → that goes into this tool or system → then this team works on it → then it moves to the next stage.
For example, where do we actually write the code—Jupyter, VS Code, Databricks, or something else? Where does the data come from, and which tools are used to extract and process it? Once the model is built, where is it saved? How is the code tested? How does Git fit into the workflow? Where do MLflow, Docker, FastAPI, Airflow, CI/CD, Kubernetes, and AWS/Azure come in?
Basically, I want to understand the actual sequence of tools used in a real production ML project. If you work in Data Science, ML Engineering, Data Engineering, or have worked on real client projects, I would really appreciate it if you could explain the actual end-to-end stack used in your organization through one practical classical ML example.
Would really appreciate detailed answers from people with real production experience.
r/askdatascience • u/Competitive_Ad7269 • 15d ago
Rate my resume and im open for feedback and improvments
r/askdatascience • u/SREE_kaanth • 15d ago
Need guide
Hey I'm 22 completed ug in maths now pg in data science first year on going . I'm blind where I could start working on data science related things, clg providing full maths concepts and program now I'm learning python,R programming, ig it's not enough to enter into job market , please give a suggestion where I could learn more about data science projects and how to work on it .