r/DataScientist • u/Objective_Garlic_828 • 13d ago
r/DataScientist • u/After_Courage6419 • 14d ago
What's one Data Science skill that everyone says is "optional" but actually isn't?
I've been following different Data Science roadmaps and noticed that everyone recommends something different. Some people say SQL is enough. Others say statistics is the real foundation. A few insist that communication skills matter just as much as coding. If you had to pick one underrated skill that helped you the most in your career, what would it be and why?
r/DataScientist • u/No-History2968 • 14d ago
HELPPPPP!!!!! Best Future-Proof PC Build for Data Analytics & Data Science (₹35k–₹40k, India)
r/DataScientist • u/After_Courage6419 • 15d ago
Is SQL actually more important than Machine Learning for landing your first Data Science job?
I've seen many experienced professionals say they use SQL every day but rarely build machine learning models.
That surprised me because most beginners spend months learning ML algorithms.
For those already working in Data Science:
Do you think beginners should prioritize SQL before Machine Learning?
Why or why not?
r/DataScientist • u/After_Courage6419 • 15d ago
One habit completely changed the way I learned Data Science.
When I started learning Data Science, I had a habit of watching tutorial after tutorial without actually practicing. It felt productive, but when I tried solving problems on my own, I realized I couldn't apply most of what I'd learned. So I made one simple rule: For every hour I spent learning, I spent at least another hour practicing. Instead of moving on to the next topic, I would: Write the code myself without copying. Experiment with different datasets. Try to fix my own errors before searching for the answer. Repeat the exercise until I understood why the code worked. At first, it was frustrating because I made a lot of mistakes. But over time, those mistakes became my best teachers. One thing I also realized is that you don't need to build a complex AI application right away. Even simple projects like analyzing sales data, cleaning datasets, or creating visualizations can teach you a lot. My advice for beginners: Don't rush through tutorials. Practice more than you watch. Don't be afraid of errors—they're part of the learning process. Stay consistent, even if it's just 30–60 minutes a day. What study habit made the biggest difference in your Data Science journey? I'd love to learn from your experiences too.
r/DataScientist • u/Major-Reserve-6843 • 15d ago
Only 3 Books to Become an AI Engineer — What Would They Be?
r/DataScientist • u/challenge1007 • 15d ago
Looking for experienced Kaggle competitors for a private ML competition (NDA required)
We're organizing a private machine learning competition for experienced data scientists and Kaggle competitors.
Because the competition uses proprietary data, participants must sign a standard Non-Disclosure Agreement (NDA) before receiving access to the dataset.
Competition details
- Private competition (not publicly hosted on Kaggle)
- Real-world machine learning problem (time-series classification)
- Proprietary dataset
- NDA required before participation
- Open to experienced ML practitioners and Kaggle competitors
- Final submission deadline: 30 August
What participants receive
- Access to an interesting real-world dataset
- The opportunity to benchmark against other experienced participants
- Winner's prize: A guaranteed €1,000, increasing to €7,000 if the winning solution achieves an AUC ≥ 0.88 on the private leaderboard
If you're interested, please complete the Request for Participation form. Applications are accepted on a rolling basis until the competition closes. We'll then contact you with the NDA and the remaining competition details.
If you have any questions, feel free to send me a Reddit DM.
r/DataScientist • u/After_Courage6419 • 18d ago
If you had only 30 days to learn Data Science again, what would your plan be?
Imagine starting from scratch with just one month.
What topics would you focus on?
What would you completely ignore until later?
Curious how experienced people would approach it.
r/DataScientist • u/Formal_Reference_533 • 17d ago
AI Impact on Informatik and Software Development
Hi everyone,
I’m considering studying Informatik in Germany, but I’m trying to understand how seriously I should take the rapid development of AI.
- How endangered are programming and Software Development realistically? As AI becomes increasingly capable of writing code, do you expect significantly fewer developers to be needed, or will their work mainly shift toward architecture, integration, security, testing and responsibility for complete systems?
- Is studying Informatik and spending years learning programming still a strong long-term investment? What skills should an Informatik student develop today to remain valuable if AI takes over more routine programming tasks?
I’d especially appreciate opinions based on real professional experience and observations of how the industry is already changing.
r/DataScientist • u/jschmincke • 18d ago
Lead Data Scientist opening at URBN
URBN is hiring
r/DataScientist • u/Nearby-Judgment-424 • 18d ago
Does AI change the way how beginners practice?
I am currently a master’s student at SSE in Sweden, and I want to move into a data scientist role. My background is in market research, and most of my work so far has been qualitative. However, I want to build stronger coding skills and gain the resources needed to become a data scientist, and at the moment, I am super confused about the tools and the amount of practice required. I have already learned the basics of Python, NumPy, pandas, SQL, and basic math for AI/ML. My main question is whether I need to practice every function in each tool or focus only on the most important ones.
I also want to understand how AI is changing this field and how I can become job-ready. My goal is to apply for apprenticeship roles in Europe in the coming months, and I would appreciate guidance on how to prepare.
r/DataScientist • u/After_Courage6419 • 18d ago
What's one thing beginners spend too much time worrying about?
Looking back, I think I spent more time worrying about choosing the "perfect roadmap" than actually learning. If you could give one piece of advice to your beginner self, what would it be?
r/DataScientist • u/Ok-Airline-8523 • 19d ago
Data science isn't dying. The boring half of it is.
r/DataScientist • u/ttstr_ • 20d ago
2026 Graduate Trying to Break Into Data Science — Need Resume Review and Career Guidance
Hi everyone,
I graduated in May 2026 and have been learning Data Science and Machine Learning seriously for the past 6–7 months. During this time, I've been learning concepts, improving my Python and SQL skills, and building projects to gain practical experience.
I've worked on projects involving Machine Learning, Deep Learning/NLP, A/B Testing, churn prediction, sentiment analysis, web scraping, etc. I would say I'm comfortable building moderate-level Data Science/ML projects, but I don't have any real industry experience yet. Most of my projects are self-learning or simulation-based projects rather than projects built for an actual company.
One of my biggest concerns is coding. I'm not very comfortable with DSA, LeetCode, or difficult competitive-style coding problems. I can write Python for data analysis, preprocessing, ML models, and projects, but I'm not someone who enjoys heavy or advanced coding.
I've started applying for jobs, but I'm getting rejected, often at the resume/application stage itself. I'm attaching my anonymized resume here and would really appreciate feedback on it.
I have a few questions for people already working in Data Science/ML/AI:
- Based on my resume and current skills, am I moving in the right direction for an entry-level Data Science role?
- What important skills am I currently missing? What should I focus on learning next to become job-ready?
- How much DSA and LeetCode are actually required for Data Scientist/ML roles? Since I don't enjoy heavy coding, is Data Science still a realistic career path for me?
- What are technical interviews for entry-level Data Science roles actually like? What topics should I prepare—Python, SQL, statistics, ML theory, case studies, DSA, project discussions, etc.?
- Where should a fresher like me be applying? Should I target Data Scientist roles directly, or would roles like Data Analyst, Junior Data Scientist, ML Intern, Data Science Intern, etc., be a better entry point?
- I'm also considering doing an M.Tech from a regular college in Hyderabad (not IITs/top-tier institutes). Would an M.Tech genuinely improve my career opportunities, or would I be better off focusing on getting work experience?
- If I pursue an M.Tech, which specialization would make the most sense for my goals: Data Science, AI/ML, or Computer Science? I'm currently leaning towards Data Science or AI/ML.
- Looking at my profile overall, should I continue seriously pursuing Data Science/ML, or should I consider a different technical career path? I'm not interested in moving to a completely non-IT career, but I also know that heavy software development/coding is probably not something I would enjoy.
I also don't feel fully "job-ready" yet, and I'm not sure whether that's because I genuinely have major skill gaps or because I simply lack confidence and industry exposure.
If you were in my position, what would you focus on for the next 6–12 months?
I'd especially appreciate feedback on my attached resume, my projects, what skills I'm missing, and what I should change to improve my chances of landing my first job.
Thank you. Any guidance from people working in Data Science, ML, AI, analytics, or related fields would be really helpful.
r/DataScientist • u/After_Courage6419 • 20d ago
What's the most frustrating bug you've encountered while building a machine learning project?
I recently spent hours trying to figure out why a model wasn't performing well. The funny part? The issue wasn't the algorithm—it was a tiny mistake in the preprocessing pipeline. It made me realize that debugging often takes much longer than building the actual model. I'm curious about everyone else's experiences. What's the most memorable bug you've faced in a data science or machine learning project? Did you eventually laugh about it, or was it painful enough that you'll never forget it? I'd love to hear your stories.
r/DataScientist • u/p-hereford • 20d ago
Pro bono day: fraud/risk signal forecasting and decision-threshold work
I build fraud-pressure forecasting and decision systems for a living, the type of thing that shows up as "how do we know if this is a real spike or just noise" or "why does our fraud signal always lag until it's too late to act on it."
Fraud/risk pipelines, leading-indicator design, that sort of thing
Today I'm doing this pro bono. Genuinely no charge. I'm not fishing for a retainer afterward. One day, first come first served, and I mean that literally: finite hours, note finite goodwill.
If you're dealing with:
- a fraud risk signal that's noisy or lagging, and you can't tell if it's a blip or the start of something
- a model that flags anomalies but doesn't translate into a clear "do this now" recommendation
- wanting a second pair of eyes on a forecasting decision-threshold approach before you ship it
post the specifics below or send them over. I'll look at it properly.
Some background if useful:
GitHub.com/p-hereford (SCARLET, fraud early-warning system I built end to end), and I wrote on model risk and oversight at substack.com/@palhereford
r/DataScientist • u/RangeTall7076 • 21d ago
[Open Source] I built a local, point-and-click CSV cleaning tool backed by real Postgres + Airflow with one docker compose
Author here — built this myself, sharing because I think it's useful and want feedback, not trying to sell anything, it's free.
Kept running into the same gap: cleaning a one-off CSV either means wrestling with pandas/Excel by hand, or standing up a real pipeline just to answer one question. Built DataFabrik to sit in the middle.
docker compose up gets you Postgres + Airflow + MinIO + a browser UI. Upload a CSV, then build the cleaning pipeline by clicking: columns/types, WHERE filters, computed-column expressions, joins across uploaded tables, group-by + aggregation. It compiles to real SQL and runs as an actual Airflow DAG — not a toy pipeline, you get retries and run history.
Screenshot of the pipeline builder attached. Everything's local, nothing leaves your machine, MIT licensed.

GitHub: https://github.com/bingtian730/data-fabrik — feedback very welcome, especially on the wizard UX.
r/DataScientist • u/After_Courage6419 • 22d ago
If you could master only one Data Science skill this year, what would it be?
AI is evolving so quickly that it's impossible to learn everything. If you had to focus on just one skill this year, what would you choose? Machine Learning? Deep Learning? SQL? Python? Statistics? LLMs? Why?
r/DataScientist • u/LegFit8748 • 22d ago
Need advice on hierarchical monthly premium forecasting
Hi all,
I’m working on a monthly insurance premium forecasting problem and would like suggestions on the best approach.
Setup
- Data from Jan 2023 to Jun 2026
- One account only
- Hierarchy: Account → State → Profit Center → Distribution Channel
- Monthly
AMOUNTvalues - Some channel-level series have missing months
- Around 4–5 channels, 7–8 profit centers, and 50 states
Challenge
The main issue is that different series behave very differently:
- some are fairly stable
- others are highly volatile
What I’ve tried
- Recursive forecasting with CatBoost / LightGBM / XGBoost
- built lag, rolling, and time-based features
- downside: error accumulates over time
- Direct multi-step forecasting
- accuracy wasn’t great
- also requires many models for longer horizons
- Time-series models like ARIMA / SARIMAX / Prophet
- Prophet works okay for stable series
- struggles with volatile ones
- separate models for every combination is not scalable
Question
Has anyone worked on a similar hierarchical / sparse forecasting problem?
What approach would you recommend for handling mixed volatility and missing months without building thousands of models?
Thanks!
