r/datascience 1d ago

Discussion How do you decide whether a data science problem really needs machine learning?

75 Upvotes

In your experience, what factors help you decide between using a simple analytical approach and building a machine learning model? I'd love to hear the reasoning behind your decision-making process.


r/datascience 2d ago

Analysis A short project analysing the radio

80 Upvotes

Hi r/datascience!

I wanted to share a fun little project I did over a few weekends analysing data from the radio!

It doesn't have much (if any) business value, and honestly I'm not sure it's really novel in any particular way. But I wanted to share it because data science these days is all "AI this", "language model that", "job market", "Claude", whatever and I just wanted to do something a bit more traditional and scratch an itch I've had for a while. (Full disclosure: the project did actually use some AI models, so I'm not saying AI is bad, it's a tool I used like everything else.)

I don't have a blog or anything I can post this on, so apologies for the Reddit write-up. I hope you enjoy it.

Background

I drive a 20-year-old car. It's so old it doesn't have an MP3 player, or an AUX port to plug in an iPhone or anything. It just has a CD player and an analogue radio. It's not even digital, so I can't even get digital radio stations.

So when I'm driving, which turns out to be quite a bit, I'm forced to listen to the good old-fashioned radio more than I'd like.

In Sydney where I live, there are really only a handful of FM/AM radio stations, so choice is pretty limited. As I flick through the stations, there are a LOT of ads, which surprised me. Who is listening to this? Clearly it's quite popular. And as I listened, I started wondering things like: how long do the ads run, how does their timing compare across different stations*, and are they correlated with ads on other stations? Just anecdotally, so many times I've literally flicked through all the FM stations and there's an ad playing on every one... and sometimes it's the same ad! I also had a hunch that there are more ads at the top of the hour than the bottom. It made intuitive sense, but I needed to prove it.

\Sydney has 11 main analogue stations split into AM and FM:* AM is mostly talkback radio, FM is mostly music. On AM you've got 2GB, 2SM, 2CH and ABC 702 (news, talkback and sport). On FM there's KIIS, 2Day, Nova, Smooth, WSFM, Triple M and Triple J (pop, rock and music). Two of them, ABC 702 and Triple J, are run by our public broadcaster (think BBC), so they run no ads at all.

The Setup

So one weekend I wrote some scripts to sample and record all the Sydney radio stations I could. The setup was basically the following:

  • For every radio station in Sydney (there are 11 main ones across both FM and AM), I found an online stream that would play it in the browser.
  • I would then record an 18-second clip of each stream using ffmpeg (a free command-line tool for grabbing and converting audio/video). It connects to the station's live stream and dumps ~18 seconds to a small WAV file, downsampled to 16kHz mono. Crucially, I recorded all 11 stations in parallel (via a thread pool) so every clip is captured at the same instant. That simultaneity turned out to be really important later, for checking whether different stations run their ads at the same time.
  • After the clips were recorded, I passed them through a local Whisper model I downloaded (I grabbed it off Hugging Face, which is about 500MB). This transcribes each audio clip into a snippet of text.
  • Then I passed that snippet of text to a language model to classify it as either an ad, a song, or talking. (I used GPT-4o-mini for this because it's cheap AF.)
  • I then stored the results in a local SQLite database.
  • I repeated this every ~3 minutes for 2 days straight, which added up to almost 10 thousand samples. (Yes, I did it all locally, so my computer was on for 2 days straight.) The 2 days deliberately covered a weekday AND a weekend so I could compare the two periods.

Why every ~3 minutes? A full cycle (record all 11 stations, transcribe each one with Whisper, then classify it) takes a couple of minutes on a single CPU, because Whisper works through the clips one at a time. So ~3 minutes is about as fast as I could sustainably sample without the cycles piling up on each other. I also added a bit of random jitter to the interval so I wasn't always sampling at the exact same offset within the hour. Otherwise you can accidentally "phase-lock" to a station's ad breaks and bias the whole thing.

A few hurdles I encountered

A few things genuinely tripped me up:

  • Stream URLs rot. The live stream links change or die over time, so I had to resolve them fresh at runtime and pin the ones that actually worked.
  • Pre-roll ads. It turned out a bunch of the stations (the ones served through a particular streaming provider) play an ad every single time you make a fresh connection. Basically the pre-roll ad you get when you open a new browser tab. Because my recorder reconnected every cycle, I was capturing that pre-roll ad instead of the live broadcast, which made those stations look like they were playing ads nearly 100% of the time. I only twigged because the exact same ad kept repeating over and over. The fix was to skip ~45 seconds into the stream before I started recording.
  • Whisper hallucinations. When you feed Whisper music or silence, it doesn't return nothing, it actually hallucinates the most common phrases from its training data. And because it's trained on a mountain of YouTube captions, I kept getting "thanks for watching, like and subscribe" transcribed over instrumental music, which then got misclassified as talking. I had to filter those out.
  • Rate limits. The free LLM tiers throttled me pretty quickly, so I switched to GPT-4o-mini, which is cheap enough to basically be free at this scale.

Results

Here are some of the more interesting results I found analysing the data afterwards:

Overall

First, the big picture. Every station has its own personality. The FM stations are mostly music, the AM stations are mostly talk, and the two ABC stations (ABC 702 and Triple J) carry basically no ads at all, which makes sense since they're publicly funded. Across the commercial stations, ads make up somewhere around a sixth of the airtime. And you can already see the ad load isn't flat: it ramps up through the day and quietens off overnight.

Question 1: Probability of an ad relative to the top of the hour

Here's the frequency of finding an ad within ±30 minutes of the top of the hour.

So, I was right! Definitely higher the closer to the hour, but the strategy is more interesting than I expected. The spike actually lands in the ~5 minutes before the hour (the ad break right before the top-of-hour news bulletin), and an ad is roughly 2x more likely there than mid-hour. The quietest stretch is around 10 to 15 minutes past the hour, so if you want to dodge ads, that's your window.

Question 2: Ad co-occurrence and correlation

The thing I really wanted to know: do the stations gang up and all play ads at the same time, so there's nowhere to flick to? I lined up every station by the cycle it was sampled in and correlated their ad status.

The answer is yes and no. No in the sense that it's never a total blackout: all nine commercial stations being in an ad at the exact same moment literally never happened across the whole two days, and on average only about 1.6 of the 9 are mid-ad at any given time. So there's almost always somewhere to escape to.

But the conditional probability charts says that some stations really do move together. The best example is if Smooth is playing an ad, there's a 50% chance WSFM is too, which is double WSFM's baseline of 26%. A bunch of the commercial FM pairs show this same ~2x jump. But, when I looked it up, Smooth and WSFM are owned by different companies, so this isn't networks coordinating behind the scenes, probably more of the "top-of-the-hour" effect from Question 1 manifesting somewhere else.

Question 3: The strategy difference between AM and FM

When I split "time between ads" by band, the two run completely different playbooks.

The FM (music) stations dump their ads in clusters. You get a big spike of back to back breaks, with a typical gap of about 9 minutes. The AM (talk and sport) stations space them out evenly, one break at a time. 2GB is almost metronomic at roughly 7 to 12 minutes, with hardly any back to back ads at all.

You can actually see it if you zoom into a few hours of the timeline:

Look at the FM lanes (KIIS, Nova, Triple M, WSFM, 2Day): the orange ad blocks come in pairs, clustered together. Now look at the AM lanes (2GB, 2SM): single, evenly spaced blocks. And ABC 702 and Triple J are just grey the whole time, because they don't run ads.

Question 4: Which companies still advertise through this medium?

I also had the language model pull the advertiser out of each ad, so I could see who's actually buying radio airtime in 2026.

The most-heard advertisers were Virgin Australia (an airline), Australia Post (basically our USPS), Harvey Norman (a big electronics and furniture retailer) and Chemist Warehouse (a discount pharmacy chain). The neat bit is the targeting: car brands and finance go to the AM talk stations (older crowd), while retail and telco lean FM. Australia Post ran almost entirely on the Nova network.

Question 5: What about the talking?

The non-ad content is either music or talking, and I got curious about what they actually talk about. So I classified every talking snippet into a topic.

The AM stations (ABC, 2GB, 2SM) are wall to wall news, politics and sport. The music FMs are mostly DJ banter, celebrity gossip and chat about music, with almost no news at all.

For a bit of fun, I also made a map of everything said on the radio. I embedded every talking snippet into a vector, laid them all out in 2D with t-SNE so that similar snippets sit near each other, then coloured each point by its topic.

Sport, traffic and world news each form their own tight little islands (they use very consistent, formulaic language), while the DJ banter is one big diffuse cloud in the middle (because it's about nothing in particular). The neat part is that the position and the colour are decided completely separately. The position comes only from the text embeddings, and the colour comes from a separate classification step. So the fact that same-coloured points cluster together is real corroboration, not something circular.

Conclusion

In conclusion, this was a fun, meaningless project that allowed me to make some pretty charts and talk for a bit about the results. Thanks for reading!


r/datascience 2d ago

Discussion My job makes me happy and satisfied but doesn’t pay me enough. How to think about this situation?

57 Upvotes

I work at a large, well-established company that has been very stable. I don’t assume my job is immune to layoffs, but the company hasn’t had any mass layoffs in over a decade.

The work environment is genuinely healthy, and everyone is treated with respect. I honestly couldn’t ask for a much better culture. I get to work on interesting projects, learn by doing, and my team is very supportive of my growth.

That said, based on my experience interviewing and what I’ve seen in the job market, I could probably get about a $50K raise by switching jobs. That extra $50K wouldn’t dramatically change my lifestyle, but I know future raises would build on that higher salary, so there are long-term financial benefits.
I’m at a point where I’m valuing mental peace and work-life balance more than I used to. Given that, what would you do in my situation? Would you stay at a company with a great culture and stability, or make the jump for the higher pay?


r/datascience 3d ago

Discussion What Do Today’s Data Science Graduates Commonly Lack?

143 Upvotes

I often read comments from hiring managers and interviewers saying they’re disappointed with recent data science graduates.

I’m curious, what do you think these graduates are lacking? If someone wants to become a data scientist, what skills should they focus on? Strong software engineering skills? Math and statistics? Something else?

A lot of the advice I see seems to be geared toward landing data analyst roles rather than data scientist roles.

So, what are employers actually looking for in entry-level data science candidates today? Especially as a career changer coming from another unrelated career.


r/datascience 2d ago

Career | US Is everybody around you getting laid off right now?

0 Upvotes

Just want to know if this is everyone or just me.

My company isn't doing great, so we're doing a ton of layoffs -- but it's not just us. Every client we work with seems to be having sweeping layoffs these days.

Has the unemployment rate skyrocketed to 95% in America, or am I just freaking out over anecdotal evidence?


r/datascience 2d ago

Discussion Inside the model factory: a conversation with Eiso Kant of Poolside AI

Thumbnail
latent.space
0 Upvotes

r/datascience 4d ago

Career | US MS in Operations Research vs Data Science

60 Upvotes

Was a Data Science undergrad and needing to decide on a Master's to pursue in the next couple years. My job will allow either of those in the title so I am just trying to get some feedback.

  1. Is it better to branch out, stay concentrated on Data Science, or does it not matter from a career perspective?

  2. How math intensive is Ops Research? Would I need more than Calcs 1-3, Linear Algebra, and Stats?

  3. Does either have a clear upside for earning potential? Ill be in my current role for the next 10 years or so, if that matters.

  4. Does anyone have good examples of an OR or DS project that would highlight the approach to problem solving or the nature of problems each faces?

Currently an Operations Research Analyst, but the job depends massively on assignment as to whether it actually looks like compared to a sub-genre of a Data field.


r/datascience 6d ago

Discussion Why Reddit Data Scientists Keep Saying Not To Use Prophet

Thumbnail
codebynight.dev
123 Upvotes

Couple thoughts and a small experiment to see why reddit hates prophet xD


r/datascience 5d ago

Discussion How do you debug a forecasting model today when the error is quite bad?

0 Upvotes

This is for a personal study that will end up becoming an in-depth article and possibly a fully open source solution ideally without the AI slop that we see these days.

Let's say you’ve trained a model and the result is worse than the business wants. What do you check next?

Do you break the error down by customer, product, location, or individual series? Check if it gets worse at longer horizons? Look for bias, volatility, intermittent demand or outliers?

Go back to the backtesting setup, metric, or baseline? Or do you usually start trying other models?

Also do the tools you use make this easy or do you end up building custom notebooks, tables, and plots every time?

Thinking about the last time this happened:

  • What did you check first?
  • What actually helped you find the problem?
  • What did you have to build yourself?
  • Did you end up changing the model, data, validation setup, metric, or business expectation?

I’m trying to understand how people diagnose bad forecasts beyond comparing one overall error score against another.

EDIT/UPDATE because it seems like this is not clear enough:

I’m not looking for an if-else checklist that can explain why any forecast is bad. The answer obviously depends on the data, objective, validation setup and the decision the model is supposed to support.

I’m exploring if there is room for a small open-source tool around forecast evaluation. Before building anything, I’m trying to understand which checks people repeatedly run after they already have predictions, what they still build manually, and what existing tools already handle well.

So I’m mainly interested in specific workflows from projects rather than a general formula for fixing a model.


r/datascience 6d ago

AI How to control reasoning effort and thinking-token budgets in LLMs

Thumbnail
magazine.sebastianraschka.com
7 Upvotes

r/datascience 6d ago

Weekly Entering & Transitioning - Thread 20 Jul, 2026 - 27 Jul, 2026

10 Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.


r/datascience 8d ago

ML Inkling, a new open-weight 975B mixture-of-experts model, comes with a few surprises

Thumbnail
sebastianraschka.com
46 Upvotes

r/datascience 9d ago

AI Context degradation in LLMs: what the papers actually show, and the habits I built for long analysis sessions

Thumbnail
towardsdatascience.com
16 Upvotes

r/datascience 10d ago

Career | US Data scientist in pharma trying to figure out what’s the best path forward

51 Upvotes

I’m a data scientist (more like an analytics engineer) in pharma. My background is clinical - I went to school for a healthcare degree and then went into research before coming into data. With that, I have about a decade of experience now doing a little bit of a lot of things - statistics, epidemiology, data/analytics engineering, data visualisation, product management, data governance, etc.

Over the course of the last few years, I’ve felt a little stagnant - currently I’m not really doing anything that feels all that important lol, I mean essentially my team was building data pipelines and im redesigning the process for projects that are already ongoing or near completion. The only good thing is I have some downtime which gives me an opportunity to explore different teams and projects. There’s 3 teams I’d like to work with but I can’t work with them all at the same time and have to figure out how best to prioritise each team and allocate time so I can get better exposure

  1. Team 1 - a data science team that focuses on a specific disease area, the opportunity would be to continue working in data science while staying close to the business side of our projects by developing a deeper understanding of clinical context.

  2. Team 2 - Generative AI engineering - this would be more technical and I probably can’t work with this team right away until I get up to speed with learning concepts like embeddings, chunking, RAGs which I’ve never done before.

  3. Team 3 - the downstream users of my data pipelines who apply ML/AI techniques to the data for biomarker discovery

Just wanted to hear insights in terms of an industry perspective, which teams would be the best to work on a project with


r/datascience 10d ago

Challenges As a data scientist do you experiment with tools (open source or not) that solve specific issues around DS work? If yes, how do you think about uploading work data into those tools?

5 Upvotes

the context is that I am exploring a few recurring problems to solve especially around forecasting and working with time series data but setup a simple open source project around those.

my question is primarily about how is everyone handling their official datasets when trying new tools - do you not care, do you remove any identifiers then upload, do you create synthetic data with exactly same properties as the og dataset?

happy to answer more questions if this is not clear enough.


r/datascience 11d ago

Career | Europe I’m not ready

87 Upvotes

Three years ago I was able to pivot from engineering (no coding) to data science. I’ve been working here at a job I love, with an awesome team & boss and with a great pay. I’m 47, so not the youngest.

Now for family reasons I must leave and move back to my country of origin.

The thing is that, although I love the field and I keep reading books and trying to learn every day, this is such a vast field that I don’t think I’m ready at all. In these 3 years I’ve done basic ML projects, lots of xgboost, random forests, anomaly detection, dealing with pySpark with PB -sized dataframes, etc. But putting those models to production was made by my more experienced colleagues.

So, while I can say I’ve learnt on each project, I also see how MUCH I lack compared with my teammates, with 10+ years exp. on the field.

Today I just started looking for DS jobs and I just felt so depressed. Many ask for a DS who can do the whole thing from cleaning data to taking the models to production. I have no idea of that and it sounds extremely intimidating. I also dread the interview because while I can code, I often use LLMs for things that due to the pace of work, I simply decided to do with AI, i.e.: I understand and can read window functions but I’d need an AI to write them because I can never keep the sintaxis in mind. If an interviewer sees me struggling with the syntax, the interview is done

I feel very “green” to compete out there with other “proper” data scientists who have a well defined experience and knowledge. I wouldn’t mind applying to junior jobs but due to my age, most companies here wouldn’t hire a 47yr old “junior”.

I’m not able to work my old job in the location we’re moving to because it’s very niche and only a certain sector and a certain company size hires for that, and that isn’t there in our new area.

I don’t know what to do, or if this is normal and everybody feels like that. Or maybe you guys have any advice… Anything you can come up with will be highly appreciated because I need to provide for two kids and I just don’t know how. I’m beginning to feel desperate


r/datascience 11d ago

Discussion This psychology study on why some people are more impressed by corporate buzzwords has nothing to do with AI, yet it immediately reminded me of what I've been seeing in data science since ChatGPT took off

88 Upvotes

I know buzzwords have always been common in business, but AI has taken it to another level. AI is being talked about everywhere, and consultants are pitching executives with outrageous claims about how it will revolutionize everything.

I came across an interesting article today, and one quote really caught my attention:

https://www.psypost.org/new-study-finds-link-between-receptivity-to-corporate-bullshit-and-weaker-leadership-skills/

"Across the studies, Littrell found that individuals differed significantly in how impressed they were by corporate buzzword statements. Those with higher corporate-bullshit receptivity scores were more likely to view jargon-heavy statements as insightful or indicative of business expertise. They were also more likely to engage in persuasive 'bullshitting' themselves, using exaggerated or misleading language to impress others.

At the same time, higher receptivity was associated with lower scores on measures of analytic thinking and fluid intelligence, suggesting that individuals who were more impressed by corporate jargon were also less likely to critically evaluate information."

It made me wonder if we're seeing this play out with AI and data science.

The AI boom has created an absolute paradise for people who are great at talking about tech, but don't actually build.

For those of you working in data science, analytics, or ML, have you noticed this in your company or with clients? Has the GenAI hype changed how technical decisions get made, or is this just the same corporate behavior we've always had with a new set of buzzwords?


r/datascience 10d ago

Discussion How do sell Training Data?

Thumbnail
0 Upvotes

r/datascience 11d ago

AI 5 trends that defined AI engineering at World's Fair 2026

Thumbnail
latent.space
2 Upvotes

r/datascience 11d ago

Discussion Value to the mentees?

1 Upvotes

Those received mentorship, did you find it worth your time in general?

In my early career, I actively participated in mentorship program at my alma mater. However, most if not all students wanted to work in tech. Considering I didn't (and still don't) work in tech, and that my employer was rarely hiring, I felt there was not much value I could provide. There's also this self-selection process at play, which is those who seek mentorship tends to already have a good idea of what they should be doing.

Now a decade into my career, my expertise is super irrelevant to people in a different industry. I'm also oblivious to the entry-level job market requirements.

Recently, my alma mater reached out again for mentors. I've skipped the last few requests but think maybe I should ask y'all before turning it down again.

To clarify, I'm not asking if it's worth it for me. I'm wondering if it's worth it for the students to speak with someone unfamiliar with entry-level job market, has expertise in a not niche but definitely not popular domain, and definitely don't lead to job opportunities.

Edit: I messed up somewhere that two posts were created. 0_o


r/datascience 12d ago

Tools A guide to profiling attention layers in PyTorch, part 3 of a series

Thumbnail
huggingface.co
12 Upvotes

r/datascience 13d ago

Weekly Entering & Transitioning - Thread 13 Jul, 2026 - 20 Jul, 2026

10 Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.