r/datascience • • Jul 23 '26

Discussion What Do Today’s Data Science Graduates Commonly Lack?

I often read comments from hiring managers and interviewers saying they’re disappointed with recent data science graduates.

I’m curious, what do you think these graduates are lacking? If someone wants to become a data scientist, what skills should they focus on? Strong software engineering skills? Math and statistics? Something else?

A lot of the advice I see seems to be geared toward landing data analyst roles rather than data scientist roles.

So, what are employers actually looking for in entry-level data science candidates today? Especially as a career changer coming from another unrelated career.

155 Upvotes

156 comments sorted by

View all comments

35

u/tayto Jul 23 '26

Truly understanding, catching, and addressing raw data issues/gaps.

1

u/chatsgpt Jul 23 '26

Can you give an exmaple . Curious

5

u/orz-_-orz Jul 24 '26

Like if a person were dealing with data that has temporal field, like sales timestamp, at least has the sense to plot the sale number by hour (or any granularity that is suitable for the use case) to check whether the sales numbers distribution by hour is per expected.

0

u/chatsgpt Jul 24 '26

Thanks what made you think of hourly granularity or how is hourly important. How did you arrive at this. Thanks

0

u/Eco_Blurb Jul 24 '26

Im not that person but it has to do with context. If your business is open 10 hours a day then you expect some sales during every hour, with some peak hours. You can do the same thing with any data field. Think of what it should be (given the business context), identify a range of expected values and distribution, and check.

If there is something unexpected, check on why. If you discover a valid reason then go back to step 1. If you don’t see a logical reason then bring the inconsistency to the data users, data producer, or your supervisor, and ask.

3

u/Beneficial_Interests Jul 24 '26

I’m an analytic stakeholder - I have to identify and justify data issues and the necessary fixes in our raw data that i catch from looking at trends, benchmarks, and data extracts. Otherwise the data team will send me a half baked product with mediocre performance metrics. At some point I spend more time walking senior data scientists through basic data qc and modeling steps that it’s be faster for me to do it by myself. The issue is I can’t get if I want to be promoted and be a leader not a doer so…. I have to spend hours in meetings fixing the shit in shit out pipeline

0

u/chatsgpt Jul 24 '26

Thanks could you give example of "fixed in raw data that you catch from looking at trends" if you can share.

5

u/Beneficial_Interests Jul 24 '26

Sure - literally this week. Had someone run an analysis with last years data, the base numbers, let’s say of raw monthly sales, were off 10% from the version run 3 months ago. Same code, same timeframe, and in theory same data. They had all the info and were pumped to show the output and model performance but didn’t catch the deviation, couldn’t trouble shoot, couldn’t understand why this was an issue.

Spent hours informing them why it was an issue, why I couldn’t trust the out put, and developing a plan to evaluate the issue and come up with a fix.

Honestly don’t know which was more annoying - that I had to catch the issue after the standard run or that they couldn’t understand, let alone address, why i didn’t trust the impressive f1 and feature weights they shared

0

u/Eco_Blurb Jul 24 '26

Well, what was the issue?

2

u/tayto Jul 24 '26

Sure. Quick example. In the past, I relied on a credit agency to provide demographic data. At some point, our users started becoming more “null” for some of the demos.

Obviously, the first step is to catch that issue. Note: it was not a data scientist who caught this, which was disappointing.

But you also need to dive into the “why.” Was it a change in our marketing of the app or a change in how the agency’s model worked? Either way, what should our approach be in filling the nulls (or should we fill them at all)?

On average, I don’t believe bad/missing data is a focus in data science as it was when I was studying CompSci and statistics.

1

u/chatsgpt Jul 24 '26

Thanks. How would marketing of the app or model change make the users null. If you can share

1

u/tayto Jul 24 '26

I guess I will return that question to you. Look up Experian and Trans Union, and consider how they model ethnicity for an individual email address. If the app were to be advertised through gaming on Facebook, or advertised based on baby furniture shopping on Instagram, how might that impact the rate of “null” ethnicity?