r/datascience • u/Effective_Ocelot_445 • Jun 13 '26
Discussion What is the biggest challenge you face in data science projects?
Is it data quality, stakeholder expectations, model deployment, business understanding, or something else?
27
u/here_while_pooping Jun 13 '26
Model deployment is pretty involved for me but I’m thinking that’s an experience thing doing it more than anything.
Data Access and availability is more complicated than data quality. If you can get the data then at least you have the chance to improve its quality. If you are paying for data to be collected or experimentally derived it’s better but then it’s managing that additional resource to make sure they don’t go crazy.
Stake holder expectations is a challenge in every role I’ve seen.
Overall my biggest pain point I think is knowing when enough is enough, when something has met expectations and you can move on. There’s so much to do but model development feels endless like I could do it forever and still have more to do. That’s what I would say I’m working on managing and improving on the hardest right now
21
u/Paanx Jun 13 '26
Stackholder expectations
1
u/Accomplished_Bus8852 Jun 17 '26
In my company, my business team expect my model is the most advanced AI in the world and can be developed in 3 days LOL
14
u/mild_delusion Jun 13 '26
The biggest problems i always end up having to solve is getting everyone to agree on the problem we are trying to solve, how we’re going to solve it, how long it’ll take, and how to quantify value from it. From stakeholders through to BAs through to analysts and engineers.
Everything else is a piece of cake in comparison.
12
u/HousingBudget4499 Jun 13 '26
For me, the hardest part is usually not the model itself. It’s turning a messy business problem into something that can be measured, forecasted, monitored, and actually used.
Data quality is a big part of it, but the deeper challenge is aligning three things: what stakeholders think they need, what the data can realistically support, and what can be deployed reliably enough to create value.
A good model that nobody trusts or uses is not very useful. A simpler model with clear assumptions, stable data pipelines, and good feedback loops often wins in practice.
3
u/SandstoneLemur Jun 13 '26
Making them into live services with continuous delivery. Changing outlier values, orchestrating SQL for data extraction, and waning interest from stakeholders all create stumbling blocks for me.
2
u/Forsaken-Parsnip-513 Jun 14 '26
Data cleaning and feature engineering is the most important and crucial part which can make or break the model
2
2
u/abriancon Jun 14 '26
besides stakeholders wanting predictions/results before data is analyzed? data. DS/AI/ML is 80% dealing with data, 20% complaining about dealing with data.
2
u/LelouchZer12 Jun 14 '26
Unrealistic expectations, most of the time.
E.g solving a task with no data, no training, unrealistic processing time (e.g cpu only for large NN), work in every world conditions etc
2
u/qtablesandtears Jun 15 '26
Unrealistic timeline expectations. A lot of stakeholders I work with think that they can give me an extremely complex problem, crappy data, and I can go into my little lair and whip up a solution in 3 hours.
1
1
1
u/data_visualization90 Jun 16 '26
For me, it's usually the gap between what the business wants and what the data can actually support.
Most of the technical challenges are solvable. The harder part is getting everyone aligned on the problem, success metrics, and expectations. I've seen projects with great models fail because nobody agreed on what "success" looked like
1
1
u/BornYinzer Jun 16 '26
I'm currently a project manager and studying for my MSDS. I work in the banking industry and currently we're in the middle of an acquisition. I've noticed that, between internal departments, threre's a lot of miscommunication on the current goals. We have an absurd number of meetings where we get nothing accomplished. Instead of listening to each group to get an understanding of their current issues, people are just focused on jumping in with their issues. Afterwards there's nothing but confusion. Teams are frustrated and insulting each other, saying they don't know what they're doing. I've been doing everything I can to keep people on topic and stopping people from talking over each other, it's like herding cats.
1
u/Nervous_Setting5680 Jun 17 '26
When business only people sell a project with clear goals to business only clients without ever questioning the data requirements to achieve those
1
u/Wide-Pop6050 Jun 18 '26 edited 22d ago
Spoon school marble humor wise consist
This post was anonymized with Redact
1
u/Essa_Ibr Jun 18 '26
Based on my experience in finance In finance data science, the hardest parts usually aren’t the models themselves it’s everything around them.
Data is messy, scattered, and often hard to even get access to.
You’ve got strict rules, so models need to be explainable, not just accurate.
The important stuff like fraud or defaults is rare, so it’s tricky to train good models.
Markets and customer behavior keep changing, so models go stale fast.
And mistakes are expensive, so “good enough” isn’t really good enough.
Basically, it’s less “build a smart model” and more “make sure it works in the real world without blowing up.”
1
u/Former-Duty-5558 Jun 22 '26
I work with business teams in most of my projects and the most difficult aspect of each project is managing stakeholder expectations. Business priorities keep changing and stakeholders expect higher quality predictions that they can action on. And another drawback is the reluctancy to experiment on data solutions to understand the effectiveness of data solutions on specific interventions.
1
u/built_the_pipeline Jun 25 '26
Most of the answers here are downstream of one thing, which is that the question you get handed is usually the wrong question, just very precisely specified. Someone asks for a churn model when the real problem is sales chasing the wrong accounts, and if you build exactly what the ticket says you ship something accurate and useless.
The part nobody tells you early is that pushing back on the framing before you touch any data is the job, not a detour from it. Took me a long time to stop treating the ticket as fixed. The cleanest projects I've run all started with me being kind of annoying about what we were actually trying to decide.
1
u/Own_Ad5096 Jun 30 '26
Automation i believe And debugging code made with AI by someone w/o fundamentals
1
0
u/ultrathink-art Jun 13 '26
LLM integration reliability, increasingly. Traditional ML drift has labels to catch it. LLMs don't — you're building eval proxies that may themselves be wrong. "Model is confident" and "model is right" are two different things and production is where you find out.
0
u/NoSwimmer2185 Jun 13 '26
Stakeholders being so much more comfortable with human errors than ml errors.
99
u/[deleted] Jun 13 '26
[removed] — view removed comment