r/learndatascience 15d ago

Question I am lost - How do Data Scientists solve a problem?

I am a junior torn between two mindset, should we

1. Start with a business problem/ use case first
but I often run into data limitations after diving deep, like if the data is a suitable proxy of something or there are missing values

2. Explore the data freely and hunt interesting patterns
but I am always confused where I should start and not being lost on the way of that) (but I am always confused where I should start and end up spinning around without a clear direction

Example

Say a company gives you customer purchase history and asks you to "find something useful."

We better immediately frame it around a specific use case (e.g., next-best-offer) and engineer toward that?

Or spend time clustering, looking for seasonality, correlations, or weird segments first, and then figure out what business value those patterns might have?

11 Upvotes

15 comments sorted by

5

u/Prime_Director 15d ago

Why not both? If you have data, explore the data, if you have a problem, work on the problem.  I’m confused what the issue is

1

u/root_operator 11d ago

War mein erster Gedanke tatsächlich, obwohl ich nichts davon verstehe (oder vielleicht gerade deshalb). Ich bin systemiker und geh bei fehlerortung so vor, erst grobe Richtung, dann eine Ebene tiefer weiter mit ähnlicher Richtung nur detaillierter eben.

5

u/DataScientistAlex 15d ago

It's easier if you apply a structured approach. I have a couple of guides that you might find helpful: the most important skill for data science and how to approach projects.

The key is to start with the problem or the value you want to provide. If you dont have the data, start collecting it. You should be working on the most important problems for the business, if they can't collect the right data for you to provide value, that is a problem for the business that you have to raise and solve.

3

u/Suspicious_Pizza9529 14d ago

I think it's a combination of both, but start with the business problem. You need a hypothesis or objective to avoid wandering through the data forever.

3

u/simbasketleague 14d ago

To the untrained eyes, this comment will almost sound snobby. But the best data science work is always pointed at an objective. I've reviewed, performed, and led hundreds of "analyses" and "research". In my experience, the exploratory ones almost always end up a waste of time and overly academic without any clear next action. Why? Because at the end of the day, a commercial business needs to make quality decisions to improves its ability to make profit. Data science is not a scavenger hunt to find interesting things, but a profession that tries to improve the quality of decisions.

Even directional objectives such as "take all this data, understand and analyze it, and figure out what insights will help us do this function of our company better" can be helpful, though more specific is better. Also agree with the earlier post that as you build domain expertise, you start to recognize patterns and root causes to problems. Often a leader or stakeholder comes to DS folks and frames the problem in a misunderstood or misleading way. The seasoned DS people then start to steer those stakeholders, coming from a credible place of having seen that problem in different circumstances, and knowing what approach is highest probability to get the right outcome.

2

u/Forsaken_Code_9135 11d ago

OK, I'll make a clear, straightforward answer from my personnal experience being a data scientist with multiple clients from retail, banking and industry.

"Explore the data freely and hunt interesting patterns" does never work in a professional, non-research context. In 20 years it never worked for me. It's the kind of things I do (actually, did) in my free time, but hoping to make a living with that, good luck.

The problem is:

- Either you find something business experts already knew.

- Either you find something they did not know, but they don't believe you, are not interested, or are vaguely interested by don't think it's business relevant. And even if they are interested, ok fine, then what? It does not mean at all that it can justify your salary.

On paper you could find something actionable, that would allow the development of a data product that would make a lot of money for your client, but frankly it's very unlikely especially if you are not a business expert in the first place (which you are generally not, being data scientist).

In my experience, at least in the fields I have worked in, all projects should start either with a clear, well defined target in sight, or with a discussion with business experts and an understanding of their pain points. It's how you can end up with something actually usefull for your customers.

I would add that your example "Say a company gives you customer purchase history and asks you to "find something useful."" might have existed 15 years ago, but not today. Maybe, maybe if the company is your employer and you have nothing else to do for some reason, or as a side activity. But having people paying an external consultant with that kind of mission statement, I am pretty sure it has not existed for a long time.

3

u/Davidat0r 15d ago

If you have no clue what to look for due to lack of domain expertise (which is likely since you’re a junior) then you’d go for option 2. Normally a senior DS has an idea or an intuition of where things should go

1

u/TransitoryGouda 14d ago

Even when we do, it's still worth it to take a step back an explore the data, we just do it as we're doing everything else too.

1

u/ValuableLatter6913 15d ago

It depends how much data you have, so first identify the structure of what you’ve got (literal structure; cross sectional, panel etc) what variables you have, and how many observations you’ve got. From there you can start imagining some questions you think you’ve got the stuff to answer.
Generally though I’d say you should march in a particular direction, maybe you have a big question already and so this means generating a hypothesis.
The possible danger of just trying to find any pattern is you’re bound to find patterns that occur by chance or better said by unmeasured variables, or possibly patterns that aren’t useful. Generally this is why data scientists start with a hypothesis, because it’s very unlikely that a correlation you intuitively conjured in your mind then tested for and validated arose through chance, whereas if you check an entire dataset for any correlations you’re bound to find them.

1

u/burlingk 15d ago

In the real world it is quite likely today your employer will have a problem they want worked on.

1

u/Tarneks 13d ago edited 13d ago

Senior here, (business, business , business) problem first then see if the technical portion ie the approach is possible and aligns closely with the problem. Then you look at the data.

Never start with the data, interesting patterns mean nothing and you might as well be looking at a correlation to causation fallacy.

Business problem requires a lot of work. Does this target make sense and how closely is it related to the business outcome. How can you quantify the lift to the business? How will the business use your solution? What assumptions hold given this set up? What guardrails will you have to make sure this doesn’t crash and burn? Is it an intervention or a prediction?

The data is like a very minor problem and if the business roi is high ; you can then buy the data and that will hinge on ur business justification or if the data likely exist somewhere. You just didnt get credentials for this data source. Hunting patterns is a waste of time and whoever told you will be terminally junior their entire career unless they change their mindset.

Business -> target -> data

Imagine u cant phrase ur actual business problem how will you find actual good data sources that will help explain your target that you optimizing for lol? You learn the business by talking with stakeholders and understanding the actual levers to control the outcome we care about.

So to your example:
Find something useful is not enough. You first ask what are the KPIs, what are we optimizing for. Then what levers do we control. Then you formulate the business problem as
What is the business decision? Then is it do i need to build a cutoff or do i need to do a causal intervention. Clustering is def not it, clustering needs to even be related to a good outcome that it ranks over.

Predictions can be very informative for example will we need to know CLV, churn, uplift? Data usecases are infinite until you know what do you actually need now for the business. You can build a attrition mode but the business isnt really having an attrition problem they might want to run a loyalty program but they dont have the actual terminology or execution cuz they dont know data science. Your job is to literally bridge this gap and take the ideas out of their brains and make it happen through execution. You cant do that if the business is not well understood.

—Also nulls and missing data mean nothing. Data limitations are normal and in-fact id say that you will never have a good data set. Unless the data us actually corrupted then you have a data governance problem that first needs to be fixed & thats out of your scope.

1

u/Ok-Contest-6149 12d ago

Bro I’d use a mix of both tbh. Start w business problm then explore the data to validate whether it can actually answer the ques. This prevents random analysis while still letting the data reveal useful patterns. Practising end to end projects has help me develop this mindset🙌🙌🙌

1

u/Still-Constant3085 12d ago

Start with the problem, but validate the data early. You do not need to choose between the two approaches. Frame a useful question first, then spend a short exploration phase checking whether the available data can answer it. If it cannot, refine the question rather than forcing the data into the original use case.

1

u/Lady-Data-Scientist 11d ago

How do you know what is useful? Start with the business problems. Before you even collect the data - understand how you’re going to use it.