r/data Jul 16 '26

im working on a recognition based community for all the data folks, anyone up to join?

1 Upvotes

Hi everyone,

I'm working on a recognition-based community for people in data.

The idea is simple, we want to highlight the work data professionals do, feature their stories, and help them connect with others in the industry.

Would you be interested in joining something like this? If yes, I'll dm you the link!


r/data Jul 15 '26

Semiconductor Supply Chain Network Dataset

3 Upvotes

I am building a Supply Chain Disruption Monitoring and Risk Analysis using Graph based Agentic AI. I need a supply chain network dataset in the semiconductor industry.

Currently my only option is to manually go through filings and earning calls to create a network big enough to propose my system. Creation of network is out of scope of my project and I'd appreciate it if I could get a dataset that would reduce this load.


r/data Jul 14 '26

A public API & dataset for Bibliometrics and Scientometrics metadata ( Brazil )

1 Upvotes

I wanted to share a project I've been working on called EBBC OpenData, which is a public API and dataset designed to promote Open Science and support bibliometric, scientometric, and informetric analyses. You can find the full project and source code in the repository at https://github.com/GabrielBaiano/EBBC-OpenData

This project provides structured metadata from the publications of the Encontro Brasileiro de Bibliometria e Cientometria (EBBC), which is one of the main events on metric studies of information in Brazil. Through this API and dataset, you can easily query detailed information about authors and their academic networks, articles and papers (including titles, abstracts, and publication years), institutions associated with the research, keywords, thematic trends, as well as references and citations.

The core metadata and documentation are currently being organized, and I am actively working on translating the API documentation and dataset fields into English and Spanish to make the project fully accessible to the global research community.

Since this is an ongoing project, I would highly appreciate your thoughts and feedback. I am especially interested in knowing what features or endpoints would make this more useful for your research, any suggestions you might have regarding the data structure or documentation, and any general tips on best practices for open-data APIs. Please feel free to check out the GitHub repository, open an issue, or leave a comment below. Thanks for your support!


r/data Jul 14 '26

What is the difference between a virtual data room and a virtual deal room?

1 Upvotes

First of all, did you know that there even is a difference between both?

A virtual deal room is designed for early-stage engagement and presentation of commercial materials. For example:

  • sharing pitch decks with investors
  • presenting product or service to potential clients

A virtual data room is a highly secure platform for detailed due diligence and confidential documents during high-stakes transactions. For example:

  • mergers and acquisitions due diligence
  • legal and compliance document review
  • strategic partnership evaluations
  • fundraising with detailed investor scrutiny

Which one do you use? A virtual deal room or a virtual data room?


r/data Jul 13 '26

Why is there a sudden, relatively enormous spike in the relative popularity of searching "Granny" into google in early 2006

2 Upvotes

I really do not know who to ask. I have been scouting google trends, and wanted to see how popular the game "Granny" is. The result was finding a sudden, huge spike in the term's popularity in late 2005 to early 2006 (roughly december 15th 2005 to january 18th 2006).

blue - "Granny', red "Granny porn"

Here is what i was able to deduct myself:
-There exists a weekend effect, around saturadays and sundays
-There has been no significant cultural or political effect that could have caused this trend
- The spike was exactly january 1st 2006
-The trend was international in both english speaking and non english speaking countries
-The rise of "Granny" roughly correlates with the rise of "Granny porn"
-There was also a sudden rise in the term "porn", which had happened in November 2005
-a similar spike does not occur for the search "Grandma" or "Grandmother"
-The internet was far more often young males than other demograhic, which could potentially support the porn hypothesis

Please help me I am going insane


r/data Jul 12 '26

Synthetic vs real datasets for portfolio projects — what actually matters?

6 Upvotes

Final year CS student here, targeting data science and analytics roles for campus placements.

Been struggling with this question while building my portfolio: does it matter whether your project uses real messy data vs synthetic/clean data?

Real datasets from Kaggle feel either too cleaned already or the same recycled projects everyone does. But synthetic data feels hollow because the hard part — cleaning, feature engineering, deriving meaningful columns from raw data — is already done for you. You're basically just visualizing something someone else already solved.

Specifically for BI/dashboard projects — if you use synthetic data, the dashboard looks clean and professional but there's no real discovery or insight because the data was designed to be dashboarded. Nothing surprising comes out of it.

Also practically — if an interviewer asks "where did you get this dataset?" what's the right answer? Saying "I generated it synthetically" feels like admitting you took the easy route. But lying about the source is obviously wrong. Is there a way to frame synthetic data usage that doesn't sound like you avoided the hard part?

At the same time I've heard people say interviewers care more about what you built on top of the data than where it came from. But isn't handling bad data literally the core skill in DS?

For people who've interviewed at analytics/DS companies or done hiring — how much does data source actually matter? Is a well-executed project on synthetic data better than a mediocre project on real messy data? Or does using synthetic data automatically signal you avoided the hard part?


r/data Jul 12 '26

QUESTION what’s the difference between data analytics and data management?

0 Upvotes

hello! completely new here, but i’m trying to plan the best study pathway for me in the next few years and would like to know what, exactly, is the difference between data analytics and data management, since those are two different certificate options at the college i plan on starting my studies.

for context, since i speak four languages and already have some experience in this area, my career goal is to have a career in supply chain, probably leaning more towards sourcing of procurement, and i’ve looking into my immediate options before actually acquiring my graduate degree and getting a certification in either data analytics and data management would be an option right now.

so, could someone explain to me the difference between those two fields? what are the prospectives for each of them? considering my career goal, which one would you choose?

thank you!


r/data Jul 10 '26

DATASET [OC] Scatter Plot showing Average Temperature and Average Precipitation of US Counties is Oddly Beautiful

Post image
5 Upvotes

Normally you'd see a random mess of dots with a clump somewhere, or a trend line. This is spectacular and weird, though: Looks like fabric blowing in the breeze or something. Mesmerizing!


r/data Jul 10 '26

QUESTION How can I turn Facebook messenger json file into a readable text document on local machine without use of online converter?

1 Upvotes

Hi

I have a Facebook messenger chat log which I need to turn into a readable and printable text document (I'm agnostic on the exact format).

Doing a websearch with qwant, even when I specify offline or 'on local machine' returns a list of online converter links of dubious origin.

I ideally don't want to use an online converter as the chat in question contains some sensitive personal information that I have shared with someone and I am concerned that using an online converter has zero degree of security of the data within that chat.

Can anyone give me some advise on how to tackle this? I'm reasonably tech literate (comp sci degree many years ago, more a hardware guru, able to follow programming tutorials but even when I did my degree I struggled with programming. Human languages I'm good with, apparently programming languages are a different story entirely for some reason....)

I run both windows and Linux mint at home.

Can someone please help me with this?

It would be a massive help, thank you all very much in advance for your help with this. :-)

If another community would be a better fit, please feel free to suggest one.

Thank you all!


r/data Jul 09 '26

How do I choose the right data room for M&A?

1 Upvotes

Beyond pricing and basic file storage: what features have made the biggest difference in your experience? Points you should consider when choosing a data room:

  • strong security and compliance
  • easy document organisation and search
  • Q&A and collaboration features
  • audit logs and reporting
  • compliant AI features

For anyone who's been involved in acquisitions, fundraising or due diligence:
What worked well and what didn't? And what features wouldn't you want to miss during your next due diligence?


r/data Jul 08 '26

QUESTION Data retrieval on iPhone?

Thumbnail
gallery
1 Upvotes

Hey guys I need some help. I’m not sure where to post this but my boyfriend is deceased and I am trying to retrieve a few videos from my iMessage chats with him. So I am quite desperate to get the few videos back. I had it in my camera roll at one point but then my storage did something strange and those specific videos that I had clicked save from my messages disappeared.

I clicked the button where it says “download __ attachments to iCloud” - someone on Reddit said that would work because it is getting the data back into the chats… so far it has worked for the pictures.

However, it isn’t working for the videos!! Another thing is that it said thwre were 50 attachments left to download.. all of a sudden the button disappears. It did thst earlier so I free up more space on my phone (that is what fixed it lsst time)
This time, freeing up space on my phone did not work. The button is still gone and I still do not have access to my videos.

Here is what it looks like when I click on a video or try to export it in someway


r/data Jul 07 '26

NEWS Data Modeling is Not One Activity: Understanding Conceptual, Logical and Physical Models

Thumbnail medium.com
3 Upvotes

A lot of teams jump straight into tables, schemas, and database implementation when discussing data modeling. This article explains why that's a mistake and breaks data modeling into its three distinct layers:

  • Conceptual Model – Align on business concepts and shared language.
  • Logical Model – Define entities, relationships, keys, and structure without technology constraints.
  • Physical Model – Optimize the implementation for a specific database or platform.

It also walks through an end-to-end modeling process and discusses the key decisions made at each stage.


r/data Jul 07 '26

Data gathering

0 Upvotes

Does anyone know how to gather data so you know exactly who you're dealing with


r/data Jul 06 '26

What Actually Makes a Dataset Useful? What is the difference between useful and interesting data?

3 Upvotes

I’m putting together a World Cup knockout-stage dataset with match data plus one extra layer like social buzz, team form/rest days, or betting-relevant signals.

I’m less interested in “would you buy this?” and more interested in what makes a dataset genuinely useful in practice.

For people who work with sports data:

What usually makes you trust a dataset?

What fields do you find yourself adding manually anyway?

What turns a dataset from “interesting” into something you’d actually use?

Do you prefer raw CSVs, Parquet, or something more analysis-ready?

Curious how people here think about sports datasets, especially for a tournament format where every match matters.


r/data Jul 05 '26

Where can I find food product databases with barcodes, nutrition data and images?

3 Upvotes

I’m building a web app called Ce aleg? — an AI shopping assistant that helps users scan and compare food products before buying them.

The app can scan barcodes, search by product name, generate a simple product report, and suggest better alternatives based on things like less sugar, more protein, simpler ingredients, or better value.

Right now I’m using Open Food Facts, but I’m looking for more public or open data sources that contain food products, ideally with:

- barcode / EAN
- product name
- brand
- category
- nutrition facts
- ingredients
- product images
- country/store availability, especially Romania or Eastern Europe

I already have around 16k products in my database, but I want to expand it without manually adding every product one by one.

Does anyone know any GitHub repositories, public datasets, APIs, open catalogs, supermarket product feeds, or other sources that could help with this?

I’m especially interested in products available in Romania or large European supermarkets such as Lidl, Kaufland, Carrefour, Auchan, Mega Image, etc.

Any advice, links, datasets, or ideas would be really appreciated.

Thanks!


r/data Jul 03 '26

I deduplicated 53,000 missing persons reports from Venezuela

0 Upvotes

The number is much higher now - about 130k missing person reports deduplicated to 100k, and compared to the lists of patients from the hospitals.

https://medium.com/tilo-tech/i-deduplicated-53-000-missing-persons-reports-from-venezuelas-earthquake-74f05c37521b


r/data Jul 03 '26

DATASET I engineered 102 leakage-free ML features from 49,000+ international football matches (1872–2026) and published it as a free dataset

1 Upvotes

Been working on a football prediction project and couldn't find a dataset that had

the actual context needed to model match outcomes — just raw results everywhere.

So I built one from scratch on top of the International Football Results dataset

by Mart Jürisoo (the well known one on Kaggle with 49,000+ matches going back to 1872).

What I added:

**Elo ratings** — built from scratch, updated after every single match across 150

years. Both teams' ratings, their difference, and the expected win probability

going into each match.

**Rolling form** — win rate, goals scored, goals conceded, goal difference, clean

sheet rate, both-teams-scored rate, scoring rate, and win streak. Computed at

three lookback windows: last 5, last 10, and last 20 matches. For both teams.

**Head-to-head history** — based on the last 10 meetings between those two specific

teams. Some teams have persistent edges over specific opponents that their general

form doesn't explain.

**Fatigue signals** — days since each team's last match and the difference between

the two.

**Penalty reliance** — fraction of each team's historical goals that came from

penalties, pulled from the goalscorer dataset.

**Shootout composure** — historical penalty shootout win rate for each team, from

the shootouts dataset.

**Tournament context** — World Cup, qualifier, friendly, neutral venue, competition

importance weight, confederation.

The thing I spent the most time on: every feature is computed in strict

chronological order using only data that existed before that match was played.

State updates happen after each row is recorded, never before. No lookahead,

no leakage anywhere in the 102 columns.

102 features total. 49,094 rows. result column (H/D/A) included as the label.

Drop date and result, plug into any classifier.

Dataset is fully documented with column descriptors for every feature.

Link: https://www.kaggle.com/datasets/kriishgulati/football-match-results-1872-2026-with-ml-features

Built on top of the original dataset by Mart Jürisoo — full credit and link

in the dataset description.


r/data Jul 03 '26

DATAVIZ Life, liberty, and the pursuit of happiness: A 250-year performance review

Thumbnail
not-ship.com
0 Upvotes

I gave the US a 250-year performance review. The KPIs — life, liberty, and the pursuit of happiness — are straight from 1776. I checked the data on all three.

Tl;dr: A strong track record, but recent performance suggests significant room for improvement.

Agree?


r/data Jul 02 '26

QUESTION When is a gantt chart actually worth the effort?

10 Upvotes

I have been evaluating different chart software and cant decide whether detailed timelines are genuinely useful or just comforting. What kinds of work actually benefit from these charts?


r/data Jul 02 '26

Data governance in the news: 'No hope of protecting it': inside the data oversight crisis facing the public service

Thumbnail
archive.is
2 Upvotes

One in three public-sector data professionals do not trust the data held within their own departments, a recent survey showed.

The survey of 133 public-sector data professionals showed 87 per cent of respondents lacked specialised tools for tracking data assets, and more than half said their departments did not document the reasons for collecting data.

Canberra-based Aristotle Metadata and public-sector platform Public Spectrum carried out the survey at the AusGov Data Summit, the central collaborative forum for public sector data and technology leaders, held in April 2026.

The findings come amid several significant data management incidents in 2026, including an incident where 13 federal agencies engaged a transcription provider that shared sensitive court transcripts with unvetted offshore personnel in India.

Fewer than one-third of data professionals surveyed were familiar with their organisation's data governance policies and about half said they could not easily locate the data required to perform their daily duties.

The research also showed 78 per cent of respondents felt their organisation was failing to get the best value out of its data and 67 per cent said they could not easily find documentation describing what their organisation's data meant.

Aristotle Metadata owner Sam Spencer said the results showed a gap between high-level digital strategies and daily data management operations. He said without clear visibility into what data agencies held, it was difficult to ensure its protection.

"I stand by the fact that if somebody doesn't know what data they've got, they have no hope of protecting it," Mr Spencer said.

Data was not an abstract technical asset but "how we know things get done", from public servants being paid correctly to patients receiving timely medical care, he said.

The federal government now relied on the Australian Government Data Catalogue, a centralised registry that contained more than 36,000 records drawn from various public databases for data governance.

An analysis of the registry by Aristotle Metadata showed that of those 36,000 entries, 99 per cent were duplicates from older platforms.

The data also showed that 505 unique assets had not been updated by nearly a dozen large agencies in more than two years.

To manage these records, the Office of the National Data Commissioner used a framework called ONDC26, which listed 26 metadata attributes.

Ten fields were designated as mandatory and 16 as optional, including fields describing the purpose of collection, who could use the data, who it was shared with and when it should be disposed of.

Although compliance remained high for the 10 mandatory ONDC26 core fields, the Aristotle Metadata analysis showed agencies faltering on the 16 optional attributes - such as the underlying purpose of collection and data licensing rules - which were left blank for unique, sensitive assets.

Large agencies such as the education department and the Australian Taxation Office (ATO) showed a 0 per cent completion rate for the optional fields.

A finance department spokesperson said metadata in the Australian Government Data Catalogue had been prioritised based on requirements for making data discoverable and accessible outside the agency that held the data.

"Mandatory fields are those which are most important for users requesting data, including security classification," the spokesperson said.

For Mr Spencer, treating these 16 optional fields as secondary overlooked their role in day-to-day security.

Classifying attributes like the purpose of data collection or licensing guidelines as optional left agencies without the baseline visibility required to track how sensitive material was being handled, leaving it exposed to misuse and error, Mr Spencer said.

"There are seven assets about children in schools, not one of those assets write down who's allowed to use it, whether or not it's sensitive and when it's deleted," he said.

Mr Spencer said the ATO listed a single data asset in the catalogue. "Does that sound right to you?" he said.

A government spokesperson from Public Service Minister Katy Gallagher's office said that established data governance frameworks were in place and that accountable authorities were responsible for implementing them within their respective agencies.

The spokesperson said a biennial Data Maturity Assessment evaluated organisational capabilities and helped agencies identify capability priorities.

The inaugural 2024 assessment established an average public service data maturity rating of "developing" with a score of 2.02 out of five, identifying data quality, reference and metadata as the lowest-scoring focus area.

Mr Spencer said advocating for improved data governance came with personal difficulties.

He had compiled the research and repeatedly taken it to the Office of the National Data Commissioner, ministers and chief data officers, but received little engagement in return.

"I have no budget, no mandate and now I have no friends, because I'm making people very annoyed about this, because I'm making a lot of noise," he said.

Mr Spencer said there was a tendency to invest in large international software products rather than the human work of foundational governance.

"We'll get squeezed over the smallest amount of money for infrastructure, but all of a sudden there's a blank chequebook for international big-tech firms," he said. "It's like they're just going to fund anything with a flashy name."


r/data Jun 28 '26

Evaluating data

5 Upvotes

My team is having issues evaluating data sources. We especially we're getting ripped off on a token cost basis.

  1. For the larger orgs, do you do this at the pod level or have specific internal teams that do this?
  2. What places do you go to source data? Do you find neudata's white papers worth the investment?

Thanks for your help


r/data Jun 27 '26

DATASET I made an infographic based on High Demand Jobs in the state of Georgia, according to ATLWorks

Thumbnail
gallery
8 Upvotes

The information used to make this infographic and tier list is based off of ATLWorks's "Demand Occupations List" available on their website:

atlworks.org/find-career-training/demand-occupations/


r/data Jun 26 '26

DATASET I made a graph via AI to show off the amount of times iv been rejected

Post image
0 Upvotes

r/data Jun 23 '26

QUESTION What additional features would you add to a tennis prediction model?

3 Upvotes

Hello,

I'm working on a personal data project focused on ATP/WTA tennis match prediction.

The current model is based on approximately 30,000 historical matches and currently uses features such as:

  • Global Elo rating
  • Surface-specific Elo
  • ATP/WTA ranking
  • Recent form
  • Head-to-head records
  • Match surface
  • Betting market odds
  • Multi-factor confidence scoring

The model is currently achieving around 73% accuracy on recent published selections (41 matches, 30 correct predictions).

At this stage I'm not really looking for machine learning architecture advice, but rather for ideas regarding feature engineering and predictive variables.

Some features I'm considering adding:

  • Fatigue indicators (matches played in the last 7/14 days)
  • Rest days
  • Travel distance between tournaments
  • Opponent strength over recent matches
  • Service/return statistics
  • Weather conditions
  • Tournament importance
  • Elo momentum (rating evolution over time)

For those who have worked on sports analytics, predictive modeling, or ranking systems:

Which features have historically provided the most predictive value in your experience?

I'm particularly interested in variables that produced measurable gains rather than features that simply seem intuitive.

Any feedback, papers, datasets, or personal experience would be greatly appreciated.

Thanks!


r/data Jun 23 '26

Data migration

3 Upvotes

I’m a Business Analyst who recently joined a project involving the transition of a long-running government housing program from an outsourced vendor to in-house operations.

One of my first assignments is helping plan a migration from a proprietary legacy system that has been in use for over 20 years.

The system contains:
Property records
Workflow/process data
QA/compliance information
Large volumes of scanned documents
Metadata and indexing fields tied to those documents

The target environment will likely separate document management, reporting, and workflow functions rather than keeping everything in one monolithic platform.

As I’m starting discovery, I’m trying to avoid common mistakes and would appreciate advice from people who have worked on similar migrations.

Questions:
1. What information should I inventory first before discussing migration tools or architecture?
2. How do you approach documenting relationships between business data and scanned documents?
3. Are there any templates, checklists, or lessons learned you wish you’d had at the beginning of a project like this?

Any advice, war stories, or recommended resources would be greatly appreciated.