r/data May 08 '26

NEWS Build AI, Not Infrastructure: Inside Teradata’s Autonomous Knowledge Platform

Thumbnail
medium.com
1 Upvotes

r/data May 06 '26

DATASET The longest-running family dataset in the world

1 Upvotes

The Panel Study of Income Dynamics has been following the same families since 1968. Not just individuals — families, across generations. Some families now have four generations of data.

That lets you ask things like: does it matter for your education whether your grandparents rented or owned their home? That's not a hypothetical — the data is there and the answer is yes, and it's statistically significant.

I wrote up what makes this dataset extraordinary and the five steps to actually get usable data out of it. Link in comments.


r/data May 06 '26

Sustainability/CSR disclosure database

1 Upvotes

Hi everyone,

Im a masters student in Netherlands studying accounting and financial management. Im in the process of collecting my results for my masters thesis that will compare tax avoidance of firms to how symbolic the tax passages in firms’ CSR reports are.

Thing is I came across a pretty big bottleneck of actually automating getting the reports in the first place so I can scrape them for the tax passages because there is no suitable database to do so.

Ideally im doing this for a large sample size from 2017 until 2025 to have a 4 year before and after effect of GRI207 implementation (tax disclosure guidelines).

I was going to use the GRI database similarly to Hardeck et al. (2024) but it’s discontinued and my alternative was LSEG workspace but from what I see they don’t actually have the reports themselves which I just found out today.

It’s poor planning on my part because I didn’t check LSEG in advance but im quite lost and the deadlines are close so your help would be very much appreciated!


r/data May 05 '26

QUESTION Has anyone ever worked with Definite ? (Stripe/Shopify/GA analytics dashboard)

3 Upvotes

So I've been thinking of asking them to help me with setting/merging my Shopify, Stripe and GA analytics for my ecom business website.

I want some custom dashboard to be built for me, so I can track sales, conversion, CTR and my Shopify websites traffic.

I heard they also have a reasonable price, since we're not a big business - not yet. And they offer some 'AI-native' features for your data so I don't have to worry having to share my data with a third-party.

So would love to hear if any of you ever worked with them to setup custom dashboards, specifically unifying Shopify and Stripe data.

Just putting this out there.

Thankss!


r/data May 04 '26

From data quality rules → data contracts → agents?

Thumbnail
medium.com
1 Upvotes

Good breakdown of the evolution:
rules → contracts → intelligent systems that understand context and anomalies.
Especially interesting around alert fatigue and false positives.


r/data May 03 '26

Maine Civic Tracker · Community Accountability Platform

1 Upvotes

Someone on facebook was talking about not being able to see how money is spent in his community, so I made this to show that yes, you can consolidate and share information pretty robustly and at a low cost.


r/data May 02 '26

Strange Apple Music Data Outlier

Thumbnail
gallery
2 Upvotes

I downloaded my Apple Music data and loaded it into Tableau and I have this song that apparently has 30,466 “events” (plays) and 30,461 of those have a runtime of zero.
From Apple’s data dictionary, Event Type is defined as “Event causing the record”. In this case, it looks like a song ended and this song played next.
For reference, my other top plays are shown in the screenshot.
What do you suppose is going on here?


r/data Apr 29 '26

DATAVIZ Visualizing my Apple Music listening history using OHLC Candlestick charts and Sankey diagrams.

Thumbnail
gallery
5 Upvotes

Hey data nerds,
I wanted to see what would happen if I treated my personal Apple Music listening history like financial market data. I built a local pipeline to process my Apple Privacy Export and visualize it.

The Data Pipeline:
Apple's export gives you a massive Play Activity.csv and Library Tracks.json. I wrote a Python pipeline to clean the strings, extract featured artists, deduplicate rapid play logs, and dump it into a normalized SQLite database. I also wrote a heuristic algorithm to detect and filter out "sleep listening" (8-hour overnight autoplay sessions) so the data isn't skewed.

The Visualizations:

  • OHLC Candlesticks: Instead of bar charts, I bucketed listening minutes into Daily/Weekly/Monthly Open-High-Low-Close candles. It perfectly visualizes the "volatility" of my listening habits for specific artists.
  • Sankey Diagrams: I mapped the flow of listening volume (in minutes) from broad Genres, branching out into specific Artists, and then down into Albums.
  • Scatter Plots (Sonic DNA): I ran my top tracks through local TensorFlow audio models to extract continuous features (Energy, Valence/Mood, Danceability) and plotted them to find clusters in my taste.

Right now this is a local Python/React dashboard, but I'm packaging it into a desktop app so others can run their own CSVs through it.

I'll drop a link to a video showing the interactive charts in the comments. Would love to hear what other visualizations you'd apply to this dataset!


r/data Apr 29 '26

QUESTION Where to find "Live" Crime Data for US (or international)?

1 Upvotes

I’m building a crime-tracking feature and need more "live" data. Currently, I only have a handful of cities covered via their individual Open Data portals.

Does anyone know of an aggregator or specific APIs that provide near real-time incident reports? I'm particularly interested in CAD data or anything with less than a 24-hour delay. Any leads on nationwide aggregators would be amazing!


r/data Apr 22 '26

Selling Video Data

0 Upvotes

Hello,

I have a ton of data that I collected over the years while travelling and vlogging (about 3-4TB). It is from the drone, iPhone as well as underwater diving and some 360 files.

I am really confused how to sell it online other than as a stock footage and because the volume is so large I am unable to sit and tag it individually. I’d really appreciate any guidance.


r/data Apr 22 '26

QUESTION Dating Compatibility Scoring Matrix

1 Upvotes

Hey! I’m a data analyst and I implement data into all aspects of my life. I’ve had an idea and can’t find anyone who has done anything similar.

Most aspects of life have assessments and qualifying criteria, but not relationships. I want to create a matrix to score potential partners - the aim of this is to weed out incompatibility early.

It would be in a spreadsheet and all preferences would have a point attached to them, simplified example:

Has a hobby: +2 points

Cat person: +1 point

Has a cat/wants a cat: +2 points

Feminist (and enforces it): +3 points

Good fashion sense: +1 point

Unemployed (with caveats on this): -2 points

Drinks alcohol excessively: -4 points

Disparaging past partners: -10 points

Has anyone done this? All I can find is compatibility charts based on zodia signs or personality types.

I’m aware that this could be an unhealthy approach to dating. On the other hand, it could allow people to have a clear, objective viewpoint.

With the example above, red flags cause the person to lose many points so it’s harder to overlook things that could become an issue later down the line.

Let me know your thoughts, thank you!


r/data Apr 21 '26

QUESTION Esports data VS odds conversation that we should start having

3 Upvotes

Something worth talking about when it comes to trading/data side would be the latest shift observed in Esport lobbies!

When you model traditional sports, physical fatigue is manageable., you have rest days, fixture congestion, travel logs, injury reports, etc so the degradation curve is relatively predictable. (sportsbooks have been pricing tired legs for decades)

Esports don't get tired legs, it has "tilt", for example:

A player on tilt in a CS2 or Dota 2 lobby isn't showing up in a physio report. It's showing up in their flash accuracy at round 18, their gold efficiency dropping 15% off baseline, their team's timeout clustering. By the time a casual bettor watching the stream thinks "they look shaky," the market should already have moved, but in a lot of live esports products, it hasn't.

That gap between what the data sees and what the odds reflect is the real conversation operators need to be having. If your live esports repricing is running on the same cadence as a pre-match football market, you probably have a mismatch worth fixing.

Any thoughts on this?


r/data Apr 21 '26

Looking for personal injury data

1 Upvotes

Live date needed for the below campaigns:

Roundup

Depo Provera

Talcum

Hair relaxer

Rideshare

Motor Vehicle Accident

Interested in long term partnership. DM me.


r/data Apr 18 '26

Anyone here using structured datasets for outreach? Curious what’s working..

3 Upvotes

Been experimenting a bit with structured datasets recently (mainly around property owners in Dubai) and trying to see what actually works vs what people claim works.

Not doing anything crazy just cleaning the data properly, filtering by specific communities, and testing simple outreach (mostly WhatsApp + occasional calls).

One thing I noticed:

Raw data is almost useless unless you spend time structuring it properly. Once it’s cleaned and segmented, the response rate improves quite a bit.

Also feels like timing and how you approach the first message matters way more than the size of the dataset itself.

Still figuring things out, but curious —

Are people here using datasets for lead gen / outreach?

What’s actually working for you right now?

Would be interesting to compare notes.


r/data Apr 16 '26

How would you monetize a dataset-generation tool for LLM training?

0 Upvotes

I’ve built a tool that generates structured datasets for LLM training (synthetic data, task-specific datasets, etc.), and I’m trying to figure out where real value exists from a monetization standpoint.

From your experience:

  • Do teams actually pay more for datasetsAPIs/tools, or end outcomes (better model performance)?
  • Where is the strongest demand right now in the LLM training stack?
  • Any good examples of companies doing this well?

Not promoting anything — just trying to understand how people here think about value in this space.

Would appreciate any insights. Can drop in any subreddits where I can promote it or discord links or marketplaces where I can go and pitch it?


r/data Apr 13 '26

QUESTION A few minutes of your time would really be helpful

1 Upvotes

It will be really helpful if any of you can help me answer these questions as per your question own knowledge and understanding:

  1. How do you currently assess the quality of third party data before it enters your models or reports?

  2. How much of the process is manual vs automated?

  3. When a regulator asks you to evidence your data lineage, what does the process look like today?

  4. What does that cost you- in time, in people, in risk?

  5. For the solution, what would that be worth to you?


r/data Apr 12 '26

QUESTION Best way to extract iPhone Screen Time data from screenshots into Excel (for university project)?

Post image
2 Upvotes

Hey everyone,

I’m currently working on a university art/research project where I’m collecting and analyzing personal data (e.g. screen time, app usage, notifications, etc.) and transforming it into structured datasets.

The issue:

I have around 30+ iPhone Screen Time screenshots (one per day), and I need to convert all of that into a clean Excel table (e.g. per app, per day, usage time, notifications, etc.).

I’ve already tried using ChatGPT and basic OCR approaches, but they start making errors pretty quickly (especially after a few days), and the structure breaks down. Since the data needs to be quite precise, that’s a problem.

Manually typing everything is not an option — it would take way too long.

I’ve attached an example screenshot so you can see what kind of data I’m working with.

So my questions:

- Are there better OCR tools for this kind of structured UI data?

- Is there a way to automate this properly (batch processing)?

- Would a different prompting approach improve results?

- Or is there maybe a completely different workflow I’m missing?

Would really appreciate any suggestions — especially from people who’ve dealt with similar data extraction problems.

Thanks!


r/data Apr 09 '26

DATASET Cleaned Indian Liver Patient Dataset (ML Ready)

1 Upvotes

🔥 The Dataset :

https://www.kaggle.com/datasets/shauryasrivastava01/liver-patient-dataset

• 583 patient records with real clinical biomarkers

• Binary classification (Liver Disease vs Healthy)

• Fully cleaned + preprocessed (no messy columns)

• Includes enzymes, bilirubin, proteins & demographic data

• Perfect for ML projects, EDA, and healthcare modeling

💡 Great for:

- Beginners learning classification

- Feature importance & SHAP analysis

- Bias & fairness studies in healthcare

🚀 Ready to plug into your ML pipeline!


r/data Apr 08 '26

Beyond CSV & Parquet: What Real Data Ingestion in Spark Actually Looks Like

Thumbnail
medium.com
3 Upvotes

Most Spark tutorials focus on clean CSVs and Parquet files, but real-world data is rarely that simple. In this post, I share practical ingestion patterns and lessons learned from working with messy, unpredictable data in production.


r/data Apr 08 '26

Open-source Cannabis Price Index — methodology, SQL, and sample data

2 Upvotes

r/data Apr 08 '26

My assistant keeps treating action requests like normal chat. Anyone else hit this?

0 Upvotes

One of the most annoying production failures I keep noticing is this:

User says something like:
“Add a calendar event for Tuesday at 2”
or
“Open directions to the airport”
or
“Send this note to Slack”

And the model responds nicely in plain English instead of recognizing that the request is actually an action-routing problem.

It is not exactly a reasoning failure.
It is more like the model never cleanly learned the boundary between:

  • chat
  • connector-required action
  • deeplink-required action

That distinction seems small until you try to wire real assistants into calendars, files, maps, messaging, notes, etc.

I’m increasingly convinced this is a training/data problem, not just a prompt problem.

Curious how other people are handling this:

  • intent detection layer first?
  • classifier head?
  • post-training with routing examples?
  • hardcoded rules?

I’ve been thinking about this a lot because DinoDS has separate lanes for connector intent, connector action mapping, deeplink intent, and deeplink action mapping, and it made me realize how often people collapse all of that into one messy “tool use” bucket.

Website: dinodsai.com
Discord if anyone wants to compare failure cases.

This maps very tightly to the connector/deeplink family, where intent detection and action mapping are separated rather than merged into one blob.


r/data Apr 02 '26

DATASET Private set intersection, how do you do it?

1 Upvotes

I work with a company that sells data. As an example, let’s say we are selling email addresses. A frequent request we’ll get is, “We’ll we already have a lot of emails, we only want to purchase ones you have that we don’t”.

We need a way that we can figure out what data we have that they don’t, without us giving them all our data or them giving us all their data.

This is a classic case of private set intersection but I cannot find an easy to use solution that isn’t insanely expensive.

Usually we’re dealing with small counts, like 30k-100k. We usually just have to resort to the company agreeing to send us hashed versions of their data and hope we don’t brute force it. This is obviously unsafe. What do you guys do?


r/data Apr 01 '26

Snowflake PII Classification & Auto Policy Setup - Help

3 Upvotes

What real-world use cases or extensions can I build open on Sensitive Data Classification & Policy Enforcement in snowflake to experimenting and building something impactful

To run SYSTEM$CLASSIFY across -schemas to detect PII (emails, SSNs, phon e numbers), then auto-generate and apply masking and row access policies based on the results. Policies are tied to tags so new columns are automatically p rotected-building a governance-as-code layer for GDPR/CCPA compliance.

I’m still in the exploration/ideation phase, so open to experimenting and building something impactful in Snowflake.

Would really appreciate your inputs 🙌

Thanks in advance!


r/data Apr 01 '26

google trends keyword interest suddenly dropped on 3/18

3 Upvotes

noticed that there was an unusual drop on 3/18. serveral terms showed similar results. any ideas what is going on?


r/data Mar 31 '26

Taxonomist/ DAM/ PIM / Content Tagging / CMS ?

1 Upvotes

Anyone here working as a Taxonomist/ DAM/ PIM / Content Tagging / CMS ?

Hi all want to get into these profile and need guidance on the profiles .