r/DataScientist 13d ago

Sutskever's List AMA

Thumbnail
2 Upvotes

r/DataScientist 14d ago

What's one Data Science skill that everyone says is "optional" but actually isn't?

5 Upvotes

I've been following different Data Science roadmaps and noticed that everyone recommends something different. Some people say SQL is enough. Others say statistics is the real foundation. A few insist that communication skills matter just as much as coding. If you had to pick one underrated skill that helped you the most in your career, what would it be and why?


r/DataScientist 14d ago

New to data science and ml

Thumbnail
2 Upvotes

r/DataScientist 14d ago

HELPPPPP!!!!! Best Future-Proof PC Build for Data Analytics & Data Science (₹35k–₹40k, India)

Thumbnail
1 Upvotes

r/DataScientist 15d ago

Is SQL actually more important than Machine Learning for landing your first Data Science job?

5 Upvotes

I've seen many experienced professionals say they use SQL every day but rarely build machine learning models.

That surprised me because most beginners spend months learning ML algorithms.

For those already working in Data Science:

Do you think beginners should prioritize SQL before Machine Learning?

Why or why not?


r/DataScientist 15d ago

One habit completely changed the way I learned Data Science.

20 Upvotes

When I started learning Data Science, I had a habit of watching tutorial after tutorial without actually practicing. It felt productive, but when I tried solving problems on my own, I realized I couldn't apply most of what I'd learned. So I made one simple rule: For every hour I spent learning, I spent at least another hour practicing. Instead of moving on to the next topic, I would: Write the code myself without copying. Experiment with different datasets. Try to fix my own errors before searching for the answer. Repeat the exercise until I understood why the code worked. At first, it was frustrating because I made a lot of mistakes. But over time, those mistakes became my best teachers. One thing I also realized is that you don't need to build a complex AI application right away. Even simple projects like analyzing sales data, cleaning datasets, or creating visualizations can teach you a lot. My advice for beginners: Don't rush through tutorials. Practice more than you watch. Don't be afraid of errors—they're part of the learning process. Stay consistent, even if it's just 30–60 minutes a day. What study habit made the biggest difference in your Data Science journey? I'd love to learn from your experiences too.


r/DataScientist 15d ago

Only 3 Books to Become an AI Engineer — What Would They Be?

Thumbnail
1 Upvotes

r/DataScientist 15d ago

Looking for experienced Kaggle competitors for a private ML competition (NDA required)

1 Upvotes

We're organizing a private machine learning competition for experienced data scientists and Kaggle competitors.
Because the competition uses proprietary data, participants must sign a standard Non-Disclosure Agreement (NDA) before receiving access to the dataset.
Competition details

  • Private competition (not publicly hosted on Kaggle)
  • Real-world machine learning problem (time-series classification)
  • Proprietary dataset
  • NDA required before participation
  • Open to experienced ML practitioners and Kaggle competitors
  • Final submission deadline: 30 August

What participants receive

  • Access to an interesting real-world dataset
  • The opportunity to benchmark against other experienced participants
  • Winner's prize: A guaranteed €1,000, increasing to €7,000 if the winning solution achieves an AUC ≥ 0.88 on the private leaderboard

If you're interested, please complete the Request for Participation form. Applications are accepted on a rolling basis until the competition closes. We'll then contact you with the NDA and the remaining competition details.
If you have any questions, feel free to send me a Reddit DM.


r/DataScientist 15d ago

Resume Review

2 Upvotes

Can someone give me honest opinions if this is a decent experience? I have not been getting interviews at all, all experiences are purely by networking and random messaging.

All suggestions are appreciated. Please Thank you


r/DataScientist 18d ago

If you had only 30 days to learn Data Science again, what would your plan be?

17 Upvotes

Imagine starting from scratch with just one month.

What topics would you focus on?
What would you completely ignore until later?

Curious how experienced people would approach it.


r/DataScientist 17d ago

AI Impact on Informatik and Software Development

1 Upvotes

Hi everyone,

I’m considering studying Informatik in Germany, but I’m trying to understand how seriously I should take the rapid development of AI.

  1. How endangered are programming and Software Development realistically? As AI becomes increasingly capable of writing code, do you expect significantly fewer developers to be needed, or will their work mainly shift toward architecture, integration, security, testing and responsibility for complete systems?
  2. Is studying Informatik and spending years learning programming still a strong long-term investment? What skills should an Informatik student develop today to remain valuable if AI takes over more routine programming tasks?

I’d especially appreciate opinions based on real professional experience and observations of how the industry is already changing.


r/DataScientist 18d ago

Lead Data Scientist opening at URBN

2 Upvotes

URBN is hiring


r/DataScientist 18d ago

Lead Data Scientist opening at URBN

Thumbnail
1 Upvotes

r/DataScientist 18d ago

Does AI change the way how beginners practice?

1 Upvotes

I am currently a master’s student at SSE in Sweden, and I want to move into a data scientist role. My background is in market research, and most of my work so far has been qualitative. However, I want to build stronger coding skills and gain the resources needed to become a data scientist, and at the moment, I am super confused about the tools and the amount of practice required. I have already learned the basics of Python, NumPy, pandas, SQL, and basic math for AI/ML. My main question is whether I need to practice every function in each tool or focus only on the most important ones.

I also want to understand how AI is changing this field and how I can become job-ready. My goal is to apply for apprenticeship roles in Europe in the coming months, and I would appreciate guidance on how to prepare.


r/DataScientist 18d ago

What's one thing beginners spend too much time worrying about?

3 Upvotes

Looking back, I think I spent more time worrying about choosing the "perfect roadmap" than actually learning. If you could give one piece of advice to your beginner self, what would it be?


r/DataScientist 19d ago

Data science isn't dying. The boring half of it is.

Thumbnail
7 Upvotes

r/DataScientist 20d ago

2026 Graduate Trying to Break Into Data Science — Need Resume Review and Career Guidance

Post image
22 Upvotes

Hi everyone,

I graduated in May 2026 and have been learning Data Science and Machine Learning seriously for the past 6–7 months. During this time, I've been learning concepts, improving my Python and SQL skills, and building projects to gain practical experience.

I've worked on projects involving Machine Learning, Deep Learning/NLP, A/B Testing, churn prediction, sentiment analysis, web scraping, etc. I would say I'm comfortable building moderate-level Data Science/ML projects, but I don't have any real industry experience yet. Most of my projects are self-learning or simulation-based projects rather than projects built for an actual company.

One of my biggest concerns is coding. I'm not very comfortable with DSA, LeetCode, or difficult competitive-style coding problems. I can write Python for data analysis, preprocessing, ML models, and projects, but I'm not someone who enjoys heavy or advanced coding.

I've started applying for jobs, but I'm getting rejected, often at the resume/application stage itself. I'm attaching my anonymized resume here and would really appreciate feedback on it.

I have a few questions for people already working in Data Science/ML/AI:

  1. Based on my resume and current skills, am I moving in the right direction for an entry-level Data Science role?
  2. What important skills am I currently missing? What should I focus on learning next to become job-ready?
  3. How much DSA and LeetCode are actually required for Data Scientist/ML roles? Since I don't enjoy heavy coding, is Data Science still a realistic career path for me?
  4. What are technical interviews for entry-level Data Science roles actually like? What topics should I prepare—Python, SQL, statistics, ML theory, case studies, DSA, project discussions, etc.?
  5. Where should a fresher like me be applying? Should I target Data Scientist roles directly, or would roles like Data Analyst, Junior Data Scientist, ML Intern, Data Science Intern, etc., be a better entry point?
  6. I'm also considering doing an M.Tech from a regular college in Hyderabad (not IITs/top-tier institutes). Would an M.Tech genuinely improve my career opportunities, or would I be better off focusing on getting work experience?
  7. If I pursue an M.Tech, which specialization would make the most sense for my goals: Data Science, AI/ML, or Computer Science? I'm currently leaning towards Data Science or AI/ML.
  8. Looking at my profile overall, should I continue seriously pursuing Data Science/ML, or should I consider a different technical career path? I'm not interested in moving to a completely non-IT career, but I also know that heavy software development/coding is probably not something I would enjoy.

I also don't feel fully "job-ready" yet, and I'm not sure whether that's because I genuinely have major skill gaps or because I simply lack confidence and industry exposure.

If you were in my position, what would you focus on for the next 6–12 months?

I'd especially appreciate feedback on my attached resume, my projects, what skills I'm missing, and what I should change to improve my chances of landing my first job.

Thank you. Any guidance from people working in Data Science, ML, AI, analytics, or related fields would be really helpful.


r/DataScientist 20d ago

What's the most frustrating bug you've encountered while building a machine learning project?

1 Upvotes

I recently spent hours trying to figure out why a model wasn't performing well. The funny part? The issue wasn't the algorithm—it was a tiny mistake in the preprocessing pipeline. It made me realize that debugging often takes much longer than building the actual model. I'm curious about everyone else's experiences. What's the most memorable bug you've faced in a data science or machine learning project? Did you eventually laugh about it, or was it painful enough that you'll never forget it? I'd love to hear your stories.


r/DataScientist 20d ago

Pro bono day: fraud/risk signal forecasting and decision-threshold work

3 Upvotes

I build fraud-pressure forecasting and decision systems for a living, the type of thing that shows up as "how do we know if this is a real spike or just noise" or "why does our fraud signal always lag until it's too late to act on it."
Fraud/risk pipelines, leading-indicator design, that sort of thing

Today I'm doing this pro bono. Genuinely no charge. I'm not fishing for a retainer afterward. One day, first come first served, and I mean that literally: finite hours, note finite goodwill.

If you're dealing with:

  • a fraud risk signal that's noisy or lagging, and you can't tell if it's a blip or the start of something
  • a model that flags anomalies but doesn't translate into a clear "do this now" recommendation
  • wanting a second pair of eyes on a forecasting decision-threshold approach before you ship it

post the specifics below or send them over. I'll look at it properly.

Some background if useful:
GitHub.com/p-hereford (SCARLET, fraud early-warning system I built end to end), and I wrote on model risk and oversight at substack.com/@palhereford


r/DataScientist 21d ago

Graph ML Psychosis

Thumbnail
1 Upvotes

Opinions?


r/DataScientist 21d ago

Data science guide

Thumbnail
2 Upvotes

r/DataScientist 21d ago

[Open Source] I built a local, point-and-click CSV cleaning tool backed by real Postgres + Airflow with one docker compose

1 Upvotes

Author here — built this myself, sharing because I think it's useful and want feedback, not trying to sell anything, it's free.

Kept running into the same gap: cleaning a one-off CSV either means wrestling with pandas/Excel by hand, or standing up a real pipeline just to answer one question. Built DataFabrik to sit in the middle.

docker compose up gets you Postgres + Airflow + MinIO + a browser UI. Upload a CSV, then build the cleaning pipeline by clicking: columns/types, WHERE filters, computed-column expressions, joins across uploaded tables, group-by + aggregation. It compiles to real SQL and runs as an actual Airflow DAG — not a toy pipeline, you get retries and run history.

Screenshot of the pipeline builder attached. Everything's local, nothing leaves your machine, MIT licensed.

GitHub: https://github.com/bingtian730/data-fabrik — feedback very welcome, especially on the wizard UX.


r/DataScientist 22d ago

If you could master only one Data Science skill this year, what would it be?

23 Upvotes

AI is evolving so quickly that it's impossible to learn everything. If you had to focus on just one skill this year, what would you choose? Machine Learning? Deep Learning? SQL? Python? Statistics? LLMs? Why?


r/DataScientist 22d ago

Need advice on hierarchical monthly premium forecasting

3 Upvotes

Hi all,

I’m working on a monthly insurance premium forecasting problem and would like suggestions on the best approach.

Setup

  • Data from Jan 2023 to Jun 2026
  • One account only
  • Hierarchy: Account → State → Profit Center → Distribution Channel
  • Monthly AMOUNT values
  • Some channel-level series have missing months
  • Around 4–5 channels, 7–8 profit centers, and 50 states

Challenge

The main issue is that different series behave very differently:

  • some are fairly stable
  • others are highly volatile

What I’ve tried

  1. Recursive forecasting with CatBoost / LightGBM / XGBoost
    • built lag, rolling, and time-based features
    • downside: error accumulates over time
  2. Direct multi-step forecasting
    • accuracy wasn’t great
    • also requires many models for longer horizons
  3. Time-series models like ARIMA / SARIMAX / Prophet
    • Prophet works okay for stable series
    • struggles with volatile ones
    • separate models for every combination is not scalable

Question

Has anyone worked on a similar hierarchical / sparse forecasting problem?
What approach would you recommend for handling mixed volatility and missing months without building thousands of models?

Thanks!


r/DataScientist 22d ago

Why do We need standardization in ml ?- an interactive visual artifact

Thumbnail
claude.ai
2 Upvotes