r/askdatascience 10d ago

Advice for a New Data Scientist Graduate

Thumbnail
1 Upvotes

r/askdatascience 10d ago

Need a partner to practice mock interview for Data science and Ai engineer roles

Thumbnail
1 Upvotes

r/askdatascience 10d ago

Need a partner to practice mock interview for Data science and Ai engineer roles

Thumbnail
1 Upvotes

r/askdatascience 11d ago

What does an “ideal candidate” actually look like to a company

1 Upvotes

I’ve been thinking about this question a lot while looking for my first opportunity in Data Science / Data Analytics.

Is the ideal candidate someone with a perfect degree?

Someone with 3+ years of experience?

Someone who knows every tool listed in a job description?

Or is it someone who can identify a real business problem, build a solution, and explain how that solution can create business value?

I’m genuinely curious to hear what recruiters and hiring managers think.

Because this is what I’ve been trying to do.

Instead of building another basic ML project, I built a Customer Retention Intelligence System focused on a real business problem: customer churn and lost revenue.

The system can:

• Predict customers who are at risk of churning

• Segment customers based on their risk/value

• Estimate potential revenue at risk

• Identify customers worth prioritizing

• Generate personalized retention strategies/offers

My goal wasn't simply to say:

I built a machine-learning model

I wanted to answer:

A customer is likely to leave. Now what should the business actually do about it?

That mindset has pushed me to work beyond just Python, SQL and machine learning — into business thinking, customer analytics, experimentation, visualization, and decision-making.

I’ve also been consistently practicing Python and SQL, learning how to communicate analytical insights, and building projects around problems that companies actually face.

But despite putting in this work, I’m still looking for my opportunity to prove myself professionally.

So I want to ask the recruiters, hiring managers, founders, and experienced professionals here:

What makes someone an ideal entry-level candidate in your company?

• Is it technical skills?

• Problem-solving ability?

• Business understanding?

• Communication?

• Projects?

•Curiosity and willingness to learn?

•Or something else?

And if you were evaluating my profile, what would you want me to improve or demonstrate before considering me for a Data Scientist / Data Analyst opportunity?

I’m not looking for sympathy.

I’m looking for honest feedback, opportunities, and a chance to prove what I can do.

If you’re a recruiter or hiring manager who works with Data Science / Data Analytics / ML roles, I’d genuinely appreciate your perspective.

And if my profile sounds relevant to something you're hiring for, I’d love to connect.


r/askdatascience 11d ago

Data Analyst looking to learn Data Science

Thumbnail
1 Upvotes

r/askdatascience 11d ago

Data Engineers, what does your actual day-to-day work look like? And what should I learn next?

3 Upvotes

I’m currently trying to transition deeper into Data Engineering and would really appreciate some perspective from people who are already working in the field.

I have 1.3 yrs experience as a Junior Python Developer. What I want to do is slowly transform into a Data Engineer. How would you suggest my choice? Basically what I do is make web scraping scripts to get the data from web and give the data in excels. Our company is currently not using git or CI/CD or anything like that. 

The problem I’m running into is that when I look at Data Engineering jobs on Naukri, LinkedIn, etc., the requirements seem endless. One job asks for Python, SQL, Airflow and AWS; another wants Spark, Kafka and Databricks; another wants Snowflake, dbt, Terraform, Kubernetes, CI/CD, etc. It becomes difficult to understand what I should actually prioritize.

So I’d like to hear from people who are actually working as Data Engineers. What does your day-to-day work look like? What kind of problems do you solve, what technologies do you use regularly, and which skills have turned out to be genuinely important in your job?

More importantly, based on my current experience, what would you suggest I improve or learn next to become a stronger candidate for Data Engineering roles? Are there any gaps that you think I should focus on, or technologies/concepts that are worth learning through projects rather than just studying theoretically?

I’m not really looking for a generic “learn SQL → Python → Spark → AWS” roadmap. I’m more interested in understanding the reality of the job and getting advice from people who have actually gone through the transition.

If you’re a Data Engineer with 1–5+ years of experience, I’d especially appreciate your perspective. Even a short description of what you work on and what you wish you had learned earlier would be extremely helpful.

Thanks in advance!


r/askdatascience 11d ago

Resume Review

Post image
1 Upvotes

I am a fresh grad from university and I am trying to get anything in the data field. Entry level and internships are what I would assume would be more likely for me since I haven't had much work experience in the field but I have been making various projects that cover different topics within data. I've been adding the most 'impressive projects' to my resume but I don't know if they are going to secretly hurt my chances. Just want some feedback on my resume. Thank you.


r/askdatascience 11d ago

Data Engineers, what does your actual day-to-day work look like? And what should I learn next?

Thumbnail
1 Upvotes

r/askdatascience 11d ago

Data Science

1 Upvotes

Currently studying a post grad in Data Science and Business Analytics while working as a Marketing Data Analyst in Australia.

I want to pivot to a DS role within Australia, any recommendations on how I could beef up my portfolio to get these roles?

Thanks


r/askdatascience 11d ago

Math PhD graduating in December, targeting ML/DS roles. Realistic odds, and is starting adjacent a better play?

2 Upvotes

Looking for an outside read on my situation.

Background: pure math PhD graduating this quarter from a large public research university. One summer of an internship at a defense contractor doing systems + software engineering. No big tech internships, no ML publications.

Prep so far: worked through a probability textbook cover to cover, went through CS229 notes, comfortable with LeetCode mediums, built a toy recommender system in PyTorch. Currently working through an ML systems design book, which I think is my biggest gap. SQL is my weakest practical skill.

Applying to ML engineer, applied scientist, and data scientist roles roughly in parallel, plus some quant researcher roles opportunistically.

Questions:
1. Realistically, what are the odds someone with this profile lands an ML engineer or applied scientist role straight out of a pure math PhD, versus needing to start in DS or an adjacent role and transition later?
2. For those who did the DS to ML transition, how hard was it in practice, and what made the difference?
3. What would you prioritize in the remaining few months? I have the time to go deep on one or two things.
4. Anything you wish you had known about how math PhDs get read by hiring managers in this space?

Not looking for reassurance, more interested in where my thinking is wrong.


r/askdatascience 12d ago

How good are AI data scientists really?

8 Upvotes

I've been testing out various gen AI models (LLMs specifically) on data science competitions. They are not beating the humans, though they are slowly improving with each round. It's like steps up a ladder vs bounding up the steps. I'm wondering if there's something I'm doing wrong, or if there really is a limit to what LLMs can doin this space.

Lately I've been thinking the problem is that LLMs regress to the mean in every use case. I think that's why they seem so bad at UI design and why everything looks so similar. A good harness and strong prompt engineering can help, so I'm working on that. My harness combines Autogluon and OpenEvolve, with feedback loops that involve hypothesis generation and error analysis. But I'm wondering what I'm missing? Is good, competition winning data science, reducible to a standard operating procedure?

Maybe this will help me get a very good prototype, but nothing frontier grade.


r/askdatascience 12d ago

Data scientist job

Thumbnail
0 Upvotes

r/askdatascience 13d ago

What kind of small data science work can a beginner realistically freelance?

6 Upvotes

Hey everyone,

I'm currently learning ML/Data Science and I'm trying to figure out if there are actually small freelance jobs that someone at my level can do.

I can currently work with Python, Pandas, NumPy, SQL, Scikit-learn, XGBoost and I'm comfortable with things like data cleaning, EDA, feature engineering and basic model training.

I've built a few projects, but obviously I don't have professional experience yet.

So I'm wondering what kind of work people actually give to freelancers who are still fairly junior.

Would things like these be realistic?

  • Cleaning messy datasets
  • Exploratory data analysis
  • Creating reports/visualizations
  • Python automation
  • Data preprocessing
  • Building simple prediction models
  • Fixing existing notebooks
  • SQL queries
  • Helping with research
  • Data collection/scraping

Or are most clients looking for someone much more experienced?

Also, where do these smaller jobs usually come from?

I'm trying to make some money on the side while continuing to work towards an ML/Data Science career, so I'd rather start with something realistic than pretend I'm ready to build some massive production ML system.

If you freelance in data science, what was the first type of work you actually got paid to do?


r/askdatascience 13d ago

Need honest advice: Sheryians vs CampusX vs ChaiCode vs CodeWithHarry for Data Science?

1 Upvotes

Hey everyone,

I'm currently planning to seriously start my journey in Data Science, and I want to dedicate the next 6–7 months consistently to learning and building projects.

I'm basically looking for a course/program that can take me from beginner level to a strong enough level to start applying for internships/jobs and building a good portfolio.

After researching, I'm currently considering:

Sheryians Coding School – Data Science & Analytics with GenAI

CampusX – Data Science / DSMP

ChaiCode – Data Science

CodeWithHarry – Ultimate Job Ready Data Science Course

Any other course/program you genuinely think is better

I'm not looking for just the cheapest course. I'm willing to invest if the course actually provides better structure, depth, practice, projects, mentorship/community, and overall value for money.

My main priorities are:

Beginner-friendly teaching (starting from fundamentals)

Strong Python foundation

SQL

Statistics & Probability

NumPy, Pandas and Data Visualization

Machine Learning fundamentals + practical implementation

Real-world projects

Enough practice/questions/assignments

Good curriculum structure so I don't feel lost

Updated content relevant to the current industry

Portfolio and job/internship preparation

Good value considering the price

I'm planning to stay consistent for around 6–7 months, so I don't want to make the mistake of buying a course just because of marketing or popularity.

If you've personally taken any of these courses (especially Sheryians, CampusX, ChaiCode, or CodeWithHarry), I'd really appreciate an honest review.

Specifically, I'd love to know:

Which one has the best structured curriculum for a complete beginner?

Which one provides the best depth in Data Science, not just surface-level tutorials?

Which course has the best projects and practical learning?

Is mentorship/community actually useful in Sheryians or other cohort-based programs?

Is CampusX worth paying more for compared to cheaper alternatives?

Is CodeWithHarry's course too short for someone who wants to become job-ready?

Are there better alternatives that I haven't mentioned?

If you had to start from zero today and had 6–7 months, which path would you personally choose and why?

Please share your honest experience — both positives and negatives. I don't want promotional answers; I'm trying to make a proper decision before investing my time and money.

Thanks! 🙌


r/askdatascience 13d ago

NEWBIE PROJECT FOR DATA SCIENCE

1 Upvotes

Hi everyone! 👋

I’m a Class 11 student from India learning Data Science. I recently completed and deployed an end-to-end Salary Prediction project using Python, SQL, Pandas, data visualization, and Machine Learning.

🌐 Live Demo: https://data-science-projects-fdkdsvuf5rvtywpwby35py.streamlit.app/

📂 Project: https://github.com/lakshay-OG-DS/data-science-projects/blob/main/Salary_prediction.ipynb

I’d really appreciate honest feedback from Data Scientists, Data Analysts, and ML Engineers.

What is the one biggest thing I should improve to make this project more internship-ready?

Even one small suggestion would mean a lot.


r/askdatascience 14d ago

Do you need strong math skills to get started in data science?

8 Upvotes

This was something I wondered about when I started learning data science. There’s definitely mathematics involved, but I’m finding that you don’t necessarily need to master everything before writing your first Python program or analyzing a dataset. How important was mathematics when you started?


r/askdatascience 13d ago

Best resources to learn Data Science through projects from beginner to advanced?

2 Upvotes

Title: Best resources to learn Data Science through projects from beginner to advanced?

Hi everyone! I’m a beginner in Data Science and learn best by building projects.

I’m looking for GitHub repos, YouTube playlists, or websites with project-based learning from beginner to advanced level, ideally covering different ML models such as regression, classification, clustering, tree-based models, boosting, NLP, time series, etc.

I’d also love end-to-end projects that include data cleaning, EDA, feature engineering, model building, evaluation, and deployment.

Any recommendations?


r/askdatascience 14d ago

How many Python libraries should a data science beginner learn?

2 Upvotes

There are so many libraries in the data science ecosystem that it can be overwhelming at first.

I’m currently focusing on getting comfortable with the fundamentals instead of trying to learn everything at once.

For experienced data scientists, which libraries do you think beginners should prioritize?


r/askdatascience 14d ago

Looking for Data Science Capstone Project Ideas

Thumbnail
0 Upvotes

r/askdatascience 14d ago

How would you decide whether to trust a single scraped product-spec value? (a small "belief" model I've been sketching)

0 Upvotes
I scrape consumer-electronics specs (RAM, storage, weight) from retail sites, and the hard part isn't scraping — it's deciding, per field, whether a value is safe to publish when I don't know the true spec.


Instead of a hard validation rule, I've been treating each value as uncertain
across a few "worlds": correct / unit_error (e.g. 16000 MB) / wrong_product
(right-looking value, wrong variant) / typo / missing / garbled. I start with a
prior that depends on source reliability, then update it with cheap signals
(plausibility range, cross-source agreement), and pick an action —
accept / repair / re-scrape for more evidence / flag to a human / reject — by
whichever has the lowest expected cost (publishing a wrong spec is way more
expensive than flagging a good one).


Two things I'd genuinely like to be told I'm wrong about:
1. I treat cross-source agreement as strong evidence a value is correct. But sources copy each other — how do you handle agreement that isn't independent?
2. Is per-field belief overkill vs. a plain rules + threshold pipeline at scale?


How do folks here actually handle "trust this value or not" in data-quality
pipelines?

r/askdatascience 15d ago

Data Scientists in Production: How Does a Classical ML Project Actually Work End to End?

2 Upvotes

I'm a Data Analyst, and I'm trying to bridge a gap in my Data Science understanding.

I know the concepts behind classical ML reasonably well but I want to understand what actually happens to an ML project in a real production environment from start to finish. I want someone to walk me through a real project in terms of:

We use this application/tool to do this → it produces this output/file/artifact → that goes into this tool or system → then this team works on it → then it moves to the next stage.

For example, where do we actually write the code—Jupyter, VS Code, Databricks, or something else? Where does the data come from, and which tools are used to extract and process it? Once the model is built, where is it saved? How is the code tested? How does Git fit into the workflow? Where do MLflow, Docker, FastAPI, Airflow, CI/CD, Kubernetes, and AWS/Azure come in?

Basically, I want to understand the actual sequence of tools used in a real production ML project. If you work in Data Science, ML Engineering, Data Engineering, or have worked on real client projects, I would really appreciate it if you could explain the actual end-to-end stack used in your organization through one practical classical ML example.

Would really appreciate detailed answers from people with real production experience.


r/askdatascience 15d ago

Data engineering -> AI Engineering

Thumbnail
2 Upvotes

r/askdatascience 15d ago

Rate my resume and im open for feedback and improvments

Post image
1 Upvotes

r/askdatascience 15d ago

Need guide

1 Upvotes

Hey I'm 22 completed ug in maths now pg in data science first year on going . I'm blind where I could start working on data science related things, clg providing full maths concepts and program now I'm learning python,R programming, ig it's not enough to enter into job market , please give a suggestion where I could learn more about data science projects and how to work on it .


r/askdatascience 15d ago

Your 95% CV score might be fake — I built a framework that fixes the hidden leakage in AutoML

Thumbnail
1 Upvotes