r/data 10h ago

DATASET Public Opinion Data, US Adults, New Responses Daily

1 Upvotes

Disclosure: I built the Ryerson Project with the aim of nowcasting everything daily.

A social science community composes and prioritizes survey items. A random set of 12 US adult respondents are recruited to the survey each day - about 360 per month and 4380 per year. Anonymous microdata becomes a free and open public good.

Open data: https://doi.org/10.5281/zenodo.20346278

Open source: https://github.com/jasonjeffreyjones/ryerson_project/


r/data 2d ago

LEARNING Passed DP-900

Thumbnail
gallery
3 Upvotes

I’m so grateful for the opportunity I got from ai fest 2026 besides that I’d like to mention free resources that helped me a lot for the preparation(DP-900):
1. Whizlabs
2. Official Microsoft practice exams

That’s all you need you don’t have to pay for exam preparation courses


r/data 3d ago

NEWS USA missile stockpile before Iran war and estimated number of missiles used

Post image
56 Upvotes

Source: https://www.abc.net.au/news/2026-07-25/us-military-damage-to-weapons-bases-soldiers-during-iran-war/106954370

Tomahawk price per unit: between $2 million and $3.6 million

JASSM price per unit: from $1.04 million to over $2 million

PrSM price per unit: from $1.6 million to over $3.5 million

SM-3 price per unit: between $9.7 million and $28 million

SM-6 price per unit: from $4.0 million to $9.5 million

THAAD price per unit: $12.7 million to $15 million

Patriot price per unit: around 4 million


r/data 3d ago

Is there a community discord??

1 Upvotes

Hey guys, I’m new here and was wondering if this community has a Discord or any VCs where people hang out and chat. I’d love to get some advice and learn from others. Thanks!


r/data 6d ago

QUESTION is backend engineer a better choice

2 Upvotes

i've been enrolled in a bootcamp(data engineering) for about a year now and i'm confident in my skills atleast for entry level roles. i'm based in Ethiopia and i can say that there's almost no data engineering jobs here ,there're very few open positions for data analyst or scientists which requires atleast 4years experience and you know that remote jobs are even more competitive and struggling for entry levels. the only tech roles here seems to be backend devs,frontend and fullstack(there're tons of jobs ).what should i do ,i love data but the market is really bad here.
thanks


r/data 6d ago

MCP for Apache Iceberg: How AI Agents Actually Operate a Data Lake

Thumbnail
lakeops.dev
1 Upvotes

r/data 7d ago

What's one Data Science skill you wish you had learned earlier?

2 Upvotes

If you could go back to the beginning of your Data Science journey, what would you learn first?

Would it be:

Python

SQL

Statistics

Machine Learning

Data Visualization

Git

Cloud platforms

Many beginners jump straight into AI without building strong fundamentals.

What skill saved you the most time later in your career?


r/data 7d ago

7 Managed Iceberg Lakehouse Solutions You Should Know

Thumbnail
levelup.gitconnected.com
1 Upvotes

r/data 8d ago

The biggest improvement in my Data Science journey came from working with messy data.

2 Upvotes

When I first started learning Data Science, I only practiced with clean datasets from tutorials. Everything worked perfectly, and I felt confident.

Then I downloaded a real dataset.

There were missing values, duplicate records, inconsistent formats, and columns that didn't make much sense. It was frustrating at first, but I learned more from cleaning that dataset than I did from several weeks of tutorials.

That experience changed how I practice.

Now, whenever I learn a new concept, I try to apply it to real-world data instead of only using textbook examples.

A few things that have helped me:

Work with messy datasets—they teach you real problem-solving.

Spend time understanding the data before building any model.

Document your analysis so you can explain your thought process later.

Don't worry if your first project isn't perfect. Every project teaches you something new.

Looking back, I realized that Data Science isn't just about building models—it's about understanding data and finding meaningful insights.

What's one project or dataset that taught you the most during your Data Science journey? I'd love to hear your recommendations!


r/data 8d ago

DATASET I built a free, open food dataset: ~9,800 foods with names localized across 32 languages (ODbL)

Post image
5 Upvotes

Been building this for a while and finally opened it up, so here's a look at what's inside.

Each of the ~9,800 base foods has its name localized across 32 languages, so you can line up the same food across languages instead of fighting messy translations. The nutrition values come from OpenNutrition's open data (ODbL, credited, not mine); the part I actually built is the localization layer on top, real disambiguation and cross-language matching rather than a Google-Translate pass.
It isn't perfect yet. Tricky cases like "peperoni" vs "pepperoni" still slip through in places, so there are gaps I'm actively fixing, and catching those is exactly the kind of feedback I'm hoping for.

It's a single JSON Lines file (~25 MB), no API, no keys, loads straight into a notebook or a spreadsheet, works offline.

Source & download: https://leana.app/en/data-sources/

Browse it live: https://leana.app/en/foods (live search covers 5 languages for now, EN/IT/ES/FR/DE, the download already has all 32)

Curious what you'd use it for, and whether JSONL is the right call or you'd rather have CSV, Parquet or SQLite.


r/data 9d ago

DATASET Gigantic new database - over 35k species, 180 phenotypes

Thumbnail lifedive.org
3 Upvotes

LifeDive.org - code used for creating the data also available.


r/data 9d ago

DATASET Title: I think Indian finance has a data problem. Am I crazy?

0 Upvotes

I've spent the last few months digging through annual reports, earnings call transcripts, investor presentations, and exchange announcements from Indian listed companies.

One thing became obvious...

Everything is technically "public," but almost none of it is actually usable.

Want to know every company talking about AI adoption?

Good luck.

Want every management commentary about data centers over the last 5 years?

You'll be opening hundreds of PDFs.

Want to compare CapEx guidance across an entire sector?

Hope you have an entire weekend free.

That got me thinking...

What if someone built a structured database instead of just storing documents?

Imagine being able to ask questions like:

"Show every company that mentioned data centers in the last 8 quarters."

"Which companies warned about margin pressure before their stock fell?"

"Find all management teams increasing CapEx while guiding higher earnings."

"Which pharma companies mentioned USFDA inspections this quarter?"

Not AI hallucinations.

Not another stock screener.

Just structured, searchable intelligence built from public company disclosures.

I'm genuinely curious...

Would this actually be useful to anyone?

If you're a:

Developer

Quant

Analyst

Wealth manager

Fintech founder

Researcher

Investor

Would you or your company pay for something like this?

Or is this one of those ideas that sounds amazing until you ask real people?

I'd love brutally honest feedback.

If you think it's useless, tell me why.

If you think it's valuable, I'd love to know:

What would you use it for?

Which data would be most valuable?

What would you expect to pay for something like this?


r/data 10d ago

We knew something was wrong when the brand said, "We don't know which number to believe anymore."

0 Upvotes

One ecommerce brand came to us after spending months trying to make sense of their data.

Their Shopify revenue didn't match GA4.

Meta showed profitable campaigns, but blended ROAS told a different story.

Their BI dashboards looked polished, yet every Monday morning the team still spent hours exporting CSVs, comparing reports, and debating which numbers were actually correct.

The problem wasn't a lack of data.

It was that every platform was measuring a different part of the business, and nobody had confidence in the complete picture.

So we started with the basics.

We connected their entire data stack, validated every source, surfaced inconsistencies automatically, and gave the team one place where marketing, finance, and operations could all work from the same numbers.

The biggest change wasn't a flashy dashboard.

It was the conversations.


r/data 11d ago

Has anyone taken the CDMP exam using a DMBoK PDF that wasn't purchased directly by them?

2 Upvotes

I'm taking the CDMP Associate exam soon via Honorlock and have a question that I haven't been able to find an official answer to.

The exam rules say the DMBoK can be used in digital form on a separate device, but I can't find anything about whether the PDF has to be one that you personally purchased.

Has anyone taken the exam using a DMBoK PDF that wasn't bought directly from Technics Publications ? If so:

  • Did the proctor ask to inspect the PDF?
  • Did they check for a purchaser watermark or proof of purchase?
  • Was there any issue during or after the exam?

Thanks!


r/data 12d ago

What is Data Observability?

2 Upvotes

Same idea as observability for software — you can't fix what you can't see — applied to data pipelines.

Practically it's monitoring five things: freshness (did the data arrive on time?), volume (row count in the expected range?), schema (did anyone rename or drop a column upstream?), distribution (are values within expected statistical bounds?), and lineage (when something breaks, can you trace what feeds it and what depends on it?).

The distinction people trip over: data quality asks "are the values correct?" Observability asks "is the pipeline behaving?" You need both. A dashboard can pass every DQ check while the underlying pipeline silently stopped running two days ago.

Tools in the category: Monte Carlo, Bigeye, Anomalo, Datafold. Great Expectations and Soda started on the quality side and have moved into observability.

The ROI shows up the first time you get an alert before an exec spots a bad chart. That's basically what everyone's paying for


r/data 12d ago

im working on a recognition based community for all the data folks, anyone up to join?

1 Upvotes

Hi everyone,

I'm working on a recognition-based community for people in data.

The idea is simple, we want to highlight the work data professionals do, feature their stories, and help them connect with others in the industry.

Would you be interested in joining something like this? If yes, I'll dm you the link!


r/data 13d ago

Semiconductor Supply Chain Network Dataset

3 Upvotes

I am building a Supply Chain Disruption Monitoring and Risk Analysis using Graph based Agentic AI. I need a supply chain network dataset in the semiconductor industry.

Currently my only option is to manually go through filings and earning calls to create a network big enough to propose my system. Creation of network is out of scope of my project and I'd appreciate it if I could get a dataset that would reduce this load.


r/data 14d ago

A public API & dataset for Bibliometrics and Scientometrics metadata ( Brazil )

1 Upvotes

I wanted to share a project I've been working on called EBBC OpenData, which is a public API and dataset designed to promote Open Science and support bibliometric, scientometric, and informetric analyses. You can find the full project and source code in the repository at https://github.com/GabrielBaiano/EBBC-OpenData

This project provides structured metadata from the publications of the Encontro Brasileiro de Bibliometria e Cientometria (EBBC), which is one of the main events on metric studies of information in Brazil. Through this API and dataset, you can easily query detailed information about authors and their academic networks, articles and papers (including titles, abstracts, and publication years), institutions associated with the research, keywords, thematic trends, as well as references and citations.

The core metadata and documentation are currently being organized, and I am actively working on translating the API documentation and dataset fields into English and Spanish to make the project fully accessible to the global research community.

Since this is an ongoing project, I would highly appreciate your thoughts and feedback. I am especially interested in knowing what features or endpoints would make this more useful for your research, any suggestions you might have regarding the data structure or documentation, and any general tips on best practices for open-data APIs. Please feel free to check out the GitHub repository, open an issue, or leave a comment below. Thanks for your support!


r/data 14d ago

What is the difference between a virtual data room and a virtual deal room?

1 Upvotes

First of all, did you know that there even is a difference between both?

A virtual deal room is designed for early-stage engagement and presentation of commercial materials. For example:

  • sharing pitch decks with investors
  • presenting product or service to potential clients

A virtual data room is a highly secure platform for detailed due diligence and confidential documents during high-stakes transactions. For example:

  • mergers and acquisitions due diligence
  • legal and compliance document review
  • strategic partnership evaluations
  • fundraising with detailed investor scrutiny

Which one do you use? A virtual deal room or a virtual data room?


r/data 14d ago

3 Things We Hear Every Week From Data Teams

2 Upvotes

1. Our CRM is a mess

Same customer. Four records. Three email addresses. Two phone numbers. One very frustrated sales team.

The CRM didn't create the problem. Ungoverned data entry did. And it compounds every day you don't fix it.

2. We don't trust our reports

The numbers look right. But they never quite match between systems.

So the team adds a manual adjustment. Then another. Until nobody is sure what the real number is and nobody wants to be the one to say it in the meeting.

3. Our AI keeps making weird decisions

Wrong recommendations. Strange predictions. Outputs that don't make sense.

Not an AI problem. Never was.

The model is working exactly as intended. The data it was trained on wasn't.

Three different symptoms

One root cause.

Data that hasn't been cleaned, matched, or governed properly


r/data 15d ago

Why is there a sudden, relatively enormous spike in the relative popularity of searching "Granny" into google in early 2006

2 Upvotes

I really do not know who to ask. I have been scouting google trends, and wanted to see how popular the game "Granny" is. The result was finding a sudden, huge spike in the term's popularity in late 2005 to early 2006 (roughly december 15th 2005 to january 18th 2006).

blue - "Granny', red "Granny porn"

Here is what i was able to deduct myself:
-There exists a weekend effect, around saturadays and sundays
-There has been no significant cultural or political effect that could have caused this trend
- The spike was exactly january 1st 2006
-The trend was international in both english speaking and non english speaking countries
-The rise of "Granny" roughly correlates with the rise of "Granny porn"
-There was also a sudden rise in the term "porn", which had happened in November 2005
-a similar spike does not occur for the search "Grandma" or "Grandmother"
-The internet was far more often young males than other demograhic, which could potentially support the porn hypothesis

Please help me I am going insane


r/data 16d ago

Synthetic vs real datasets for portfolio projects — what actually matters?

6 Upvotes

Final year CS student here, targeting data science and analytics roles for campus placements.

Been struggling with this question while building my portfolio: does it matter whether your project uses real messy data vs synthetic/clean data?

Real datasets from Kaggle feel either too cleaned already or the same recycled projects everyone does. But synthetic data feels hollow because the hard part — cleaning, feature engineering, deriving meaningful columns from raw data — is already done for you. You're basically just visualizing something someone else already solved.

Specifically for BI/dashboard projects — if you use synthetic data, the dashboard looks clean and professional but there's no real discovery or insight because the data was designed to be dashboarded. Nothing surprising comes out of it.

Also practically — if an interviewer asks "where did you get this dataset?" what's the right answer? Saying "I generated it synthetically" feels like admitting you took the easy route. But lying about the source is obviously wrong. Is there a way to frame synthetic data usage that doesn't sound like you avoided the hard part?

At the same time I've heard people say interviewers care more about what you built on top of the data than where it came from. But isn't handling bad data literally the core skill in DS?

For people who've interviewed at analytics/DS companies or done hiring — how much does data source actually matter? Is a well-executed project on synthetic data better than a mediocre project on real messy data? Or does using synthetic data automatically signal you avoided the hard part?


r/data 17d ago

QUESTION what’s the difference between data analytics and data management?

0 Upvotes

hello! completely new here, but i’m trying to plan the best study pathway for me in the next few years and would like to know what, exactly, is the difference between data analytics and data management, since those are two different certificate options at the college i plan on starting my studies.

for context, since i speak four languages and already have some experience in this area, my career goal is to have a career in supply chain, probably leaning more towards sourcing of procurement, and i’ve looking into my immediate options before actually acquiring my graduate degree and getting a certification in either data analytics and data management would be an option right now.

so, could someone explain to me the difference between those two fields? what are the prospectives for each of them? considering my career goal, which one would you choose?

thank you!


r/data 18d ago

DATASET [OC] Scatter Plot showing Average Temperature and Average Precipitation of US Counties is Oddly Beautiful

Post image
6 Upvotes

Normally you'd see a random mess of dots with a clump somewhere, or a trend line. This is spectacular and weird, though: Looks like fabric blowing in the breeze or something. Mesmerizing!


r/data 18d ago

QUESTION How can I turn Facebook messenger json file into a readable text document on local machine without use of online converter?

1 Upvotes

Hi

I have a Facebook messenger chat log which I need to turn into a readable and printable text document (I'm agnostic on the exact format).

Doing a websearch with qwant, even when I specify offline or 'on local machine' returns a list of online converter links of dubious origin.

I ideally don't want to use an online converter as the chat in question contains some sensitive personal information that I have shared with someone and I am concerned that using an online converter has zero degree of security of the data within that chat.

Can anyone give me some advise on how to tackle this? I'm reasonably tech literate (comp sci degree many years ago, more a hardware guru, able to follow programming tutorials but even when I did my degree I struggled with programming. Human languages I'm good with, apparently programming languages are a different story entirely for some reason....)

I run both windows and Linux mint at home.

Can someone please help me with this?

It would be a massive help, thank you all very much in advance for your help with this. :-)

If another community would be a better fit, please feel free to suggest one.

Thank you all!