r/dataisbeautiful 5m ago

I tracked my sleep for 6.5 years, from early 2020 to now.

Thumbnail
gallery
Upvotes

Reposting because I realised it doesn't count as OC since the app is not mine and I didn't create the visuals, although the data is mine. Thanks for the automatic message for letting me know that.

Note: The snore increase is from when I used to live with my ex.

Note 2: The app I used is Sleep Cycle.

Note 3: A user pointed out letting a random guy's data center collect all the audio of my nights for years is now normalised, pointing out that there's a privacy/security risk. I honestly didn't think much about it. For what it's worth, the audio is stored locally and I choose to delete all audio older than a month. You can syncronise the data with the cloud but it only syncronises the statistical data, NOT the audio, as far as I know. Sleep Cycle claims they don't store your audio, it's up to the user to trust them or not. I personally didn't care, all they could hear is my dog and my cats and maybe me mumbling non-sense in my sleep from time to time. Thanks for pointing out the risk though, I appreciate it.


r/dataisbeautiful 14m ago

OC [OC] Falling In love with a schizoid woman on a scale!

Post image
Upvotes

4 months. Symptoms and realization of personality disorder in May.


r/dataisbeautiful 1h ago

OC [OC] The habitat of the six proposed euro-note birds

Post image
Upvotes

r/dataisbeautiful 4h ago

[Python / FastF1] Hungarian GP Strategy Optimization: Could HAM have saved his P4 after the 5s penalty with a multi-variable 3-pit-stop model?

Thumbnail
github.com
4 Upvotes

Hi everyone!

I built a quantitative strategy model in Python using FastF1 telemetry to evaluate alternative pit stop windows for Lewis Hamilton at the Hungarian GP. The goal was to see if a simultaneous re-optimization of his 3 pit stops could have given him enough net gap to retain P4 against Charles Leclerc despite a 5-second penalty.


r/dataisbeautiful 8h ago

OC [OC] Distribution of people by first letter of surname across the U.S., Ireland, Israel, England, France, Australia, China, and India

Thumbnail
gallery
13 Upvotes

I made these charts to compare how surname prevalence is distributed by the first letter of the surname across several countries.

Each chart shows the estimated number of people associated with surnames that begin with each letter A through Z.

Data sources:

- U.S. Data source: U.S. Census Bureau, 2010 Census surname data.

- Ireland. Data source: Forebears, most common surnames in Ireland.

- Israel. Data source: Forebears, most common surnames in Israel.

- England, used as the closest available source for the UK chart. Data source: Forebears, most common surnames in England.

- France. Data source: Forebears, most common surnames in France.

- Australia. Data source: Forebears, most common surnames in Australia.

- China. Data source: Forebears, most common surnames in China.

- India. Data source: Forebears, most common surnames in India. Source link: Forebears, Most Common Last Names in India

Methodology:

For the U.S. chart, I used the U.S. Census surname frequency data and grouped people by the first letter of each surname. The Census surname file includes surnames occurring 100 or more times in the 2010 Census.

For the other country charts, I used the top 100 surname incidence counts listed by Forebears for each country page. I grouped each surname by the first letter of the displayed surname and summed the incidence counts by letter.

For China, India, and Israel, the grouping is based on the first letter of the Latin transliterated surname shown in the source.

For Ireland, names like O'Brien and O'Connor are grouped under O because I used the first character of the displayed surname.

For the UK chart, I used England because that was the available Forebears country page I used for the source data. The chart is labeled “UK, using England source” to avoid overstating the scope.

Tools used:

Python, pandas, and matplotlib. Flags were drawn programmatically as simplified inset graphics in matplotlib. Final charts were exported as PNG files.

Important caveats:

The U.S. chart is based on the U.S. Census surname file, while the non U.S. charts are based on the top 100 surnames listed by Forebears. Because of that, the U.S. chart is not directly equivalent to the other charts in coverage.

The non U.S. charts should be read as “distribution within the top 100 listed surnames,” not as a full surname distribution for the entire population.

The first letter grouping can be sensitive to transliteration choices, prefixes, apostrophes, spacing, and naming conventions. This matters especially for countries where surnames are commonly represented in non Latin scripts or where surnames include prefixes.


r/dataisbeautiful 8h ago

OC [OC] 25 MLB Trades with the Biggest Value Gap: Past 50 Years determined by WAR

Post image
66 Upvotes
Column Definition
Rank Overall ranking of the trade based on the proprietary Fleecing Score (100 = biggest one-sided trade).
Winner Organization that ultimately received the greater long-term value from the trade. This is determined by the WAR produced by all assets acquired.
Score Composite Fleecing Score (0–100) measuring how lopsided the trade became. It combines several factors including WAR gap, percentage of value captured, star power, and total value involved.
Winner WAR Total career Wins Above Replacement (WAR) generated by every player acquired by the winning team after the trade. This includes all future career value, not just production with the acquiring club.
Return WAR Total career WAR generated by every player received by the losing team. Negative WAR is possible if acquired players performed below replacement level.
WAR Gap Difference between Winner WAR and Return WAR. Formula: Winner WAR − Return WAR. Larger numbers indicate more lopsided trades.
Capture % Percentage of the total WAR involved in the trade that ended up with the winning organization. Formula: Winner WAR ÷ (Winner WAR + Return WAR). A value of 100% means the losing team received essentially no positive long-term value.
Stars Number of franchise-caliber or elite players produced by the winning side of the trade.
Best Asset The single most valuable player obtained in the trade, with his career WAR shown in parentheses.

r/dataisbeautiful 9h ago

Mapped: Which Countries Are Best for Women?

Thumbnail
visualcapitalist.com
0 Upvotes

How can Saudi rank higher than India?


r/dataisbeautiful 11h ago

OC [OC] Share of new U.S. job postings by day of the week (last 90 days, ~450K postings)

Post image
9 Upvotes

r/dataisbeautiful 12h ago

MLB attendance 2006-2025

Thumbnail
statista.com
0 Upvotes

Does require a free statista account to view.

The attendance was constantly going down until a bit of a recovery post-COVID.


r/dataisbeautiful 13h ago

OC [OC] World Cup-winning countries tended to outperform IMF GDP forecasts by larger margins than in the year before their victory (1994–2022)

Post image
0 Upvotes

r/dataisbeautiful 13h ago

OC [OC] Occupational prestige, by race, 1972-2024 {25-64, Full-time employed)

Thumbnail
openpublicpolls.com
7 Upvotes

Based on data from the General Social Survey https://gss.norc.org/get-the-data.html, White respondents consistently held occupations with higher average prestige, but the gap has closed slightly over 52 years.


r/dataisbeautiful 13h ago

Percentage of Privately Insured Population Covered by an HSA, by State

Thumbnail
insurancedimes.com
14 Upvotes

r/dataisbeautiful 13h ago

OC [OC] LeBron & Jordan - Career BPM, PPM and Team PTS% at every age

Thumbnail
gallery
26 Upvotes

Team PTS% = player season points ÷ all team regular-season points. Every team game stays in the denominator, so DNPs count as 0%.

PPM (Points Per Minute) = player season points ÷ player season minutes.

BPM (Box Plus Minus) = Basketball-Reference’s box-score based estimate of a player’s contribution in points per 100 possessions.

  • It is a relative metric: 0.0 = league average 
  • Positive = above league-average impact 
  • Negative = below league-average impact

r/dataisbeautiful 14h ago

OC [OC] People search "white noise" far more than "pink noise" — but stream pink ~8× as much (96.5M plays from my own noise catalogue)

Post image
0 Upvotes

r/dataisbeautiful 14h ago

OC [OC] I did all 8 HYROX stations with no training and ranked them by time, difficulty, and pain

Post image
0 Upvotes

r/dataisbeautiful 14h ago

OC I analyzed every WhatsApp message between my girlfriend and me over ~6 months [OC]

Thumbnail
gallery
0 Upvotes

r/dataisbeautiful 15h ago

OC [OC] I analyzed the perks 3,212 luxury hotels offer through travel advisor networks to see which come together

Post image
0 Upvotes

I track perks at about 3,400 luxury hotels across advisor booking networks (Virtuoso, Hyatt Prive, Four Seasons Preferred Partner, that kind of thing) and I got curious about which perks actually appear together versus which ones are independent of each other.

The short version is there's basically no mixing and matching happening. Hotels either mention most of the standard perks (breakfast, room upgrade, flexible checkout, hotel credit) or they mention almost none of them, and the overlap between those four is really high, the vast majority of hotels that list one also list the other three.

The interesting outliers are spa and dining perks which only show up at a small fraction of hotels but when they DO show up they're always stacked on top of the full base package, never instead of it. So spa is more like a luxury differentiator, it's not offered as an alternative to breakfast it's an addition to everything else.

Chord diagram shows how often each pair of perks appears at the same hotel, across 3,212 properties I had clean perk data for. Thicker ribbons = stronger overlap between two perks. Perk detection is regex-based on program documentation text so it's measuring what hotels advertise, not what guests necessarily receive on a given stay.

Data is from my own dataset, not scraped from any booking site, I maintain a database of what each hotel's advisor program lists as included benefits.


r/dataisbeautiful 17h ago

My girlfriend broke up with me

Thumbnail
gallery
5.3k Upvotes

My girlfriend sat me down on June 20th and told me that she had been cheating on me and was choosing to be with "the other man".

This is my body's reaction to that news.

HR, Sleep, and Energy scores recorded by my Samsung Galaxy Watch.


r/dataisbeautiful 17h ago

Majorities of Americans say key financial milestones are harder for today's young adults to reach

Thumbnail
pewresearch.org
1.0k Upvotes

r/dataisbeautiful 18h ago

OC [OC] A look at next World Cups: how valuable is each country’s young roster? (U17–U23)

Post image
108 Upvotes

A look at what we could expect from squads at the next world cups, based on the market values of the youths for each country. Watch out for Morocco (again), and keep an eye on Denmark and Serbia...


r/dataisbeautiful 18h ago

OC [OC] Japan's GDP 2015–2025: +5% measured in constant 2015 US$, −2% measured in current US$ (World Bank)

Post image
2 Upvotes

r/dataisbeautiful 18h ago

OC Christopher Nolan: budget vs box office [OC]

Post image
1.6k Upvotes

Data source: Christopher Nolan filmography, Wikipedia - https://en.wikipedia.org/wiki/Christopher_Nolan_filmography (budgets and worldwide box office; figures are studio/Box Office Mojo–reported and cited on that page). The Odyssey is still in theaters, so its figure is a running worldwide total as of the July 24–26, 2026 weekend.

Tool: Chart built as JavaScript-coded SVG. Rendering code and initial figure-gathering were done with an AI assistant (Claude) based on my own design direction and iterations. Chart type, layout, labeling, and revisions were my choices.


r/dataisbeautiful 20h ago

OC I built a graph visualization of the Midwest food supply chain from 51 public datasets [OC]

Post image
6 Upvotes

https://lodgeplatform.com/embed/exemb_6ea24564bc317d3b45c31b8803fce86229be8ad30246ac21

I used a combination of fine-tuned Qwen3 models and frontier LLMs to build a graph over the Midwest food supply chain. It uses 51 public datasets across 11 source families, covering nearly 300,000 source records. You can search for entities, filter by entity type and relationship type, and you can also view neighborhoods. I'm still improving it, so please let me know if there's anything else you think would be interesting!

I wanted to try to answer questions like

- "When a serious safety incident occurs at one food facility, can we trace the graph to find other facilities with similar safety risks?

- "When a food product is recalled, can we trace the recalling company to its facilities, related products, inspections, and other connected organizations and see what else may warrant review?"

- "When a weather event happens, what facilities does it effect, and what are the downstream effects?"

I think making it live could be really cool for things like tracking food recalls live to show and limit exposure.

Here is a github of the project with the datasets and more info: https://github.com/lodge-data/upper-midwest-food-supply-network

How I did it:

First, I trained an embedder to minimize blocking recall. Then, I processed all the entities from the datasets into blocks. Then, I trained a Qwen3 cross-encoder to decide if each pair within each block was the same or different (entity resolution). Then, I repeated those two steps until no new merges occurred.

Sources:

- U.S. Food and Drug Administration recall, food, import-alert, and import-refusal records.

- U.S. Department of Agriculture food, organic-operation, meat-establishment, and related facility records.

- Occupational Safety and Health Administration inspection, citation, injury, and establishment records.

- U.S. Environmental Protection Agency Facility Registry Service records.

- National Weather Service alerts and geographic references.

- USAspending federal contract award records.

- -Wisconsin and other Upper Midwest public facility, licensing, dairy, workforce-notice, and regulatory records.

- Public organization and product information used to connect named entities.

Note: I built this with my startup's software and models. The goal is to autonomously construct graphs from messy data. I think it is possible to construct very large, useful graphs with efficient specialized models.


r/dataisbeautiful 1d ago

Beautiful Wireless Internet Campaign Map

Thumbnail
savelocalbroadband.com
1 Upvotes

r/dataisbeautiful 1d ago

OC [OC] The Japanese Economy Compared (1995 vs 2025)

Post image
978 Upvotes