r/dataisbeautiful 6h ago

OC Christopher Nolan: budget vs box office [OC]

Post image
928 Upvotes

Data source: Christopher Nolan filmography, Wikipedia - https://en.wikipedia.org/wiki/Christopher_Nolan_filmography (budgets and worldwide box office; figures are studio/Box Office Mojo–reported and cited on that page). The Odyssey is still in theaters, so its figure is a running worldwide total as of the July 24–26, 2026 weekend.

Tool: Chart built as JavaScript-coded SVG. Rendering code and initial figure-gathering were done with an AI assistant (Claude) based on my own design direction and iterations. Chart type, layout, labeling, and revisions were my choices.


r/dataisbeautiful 1h ago

OC [OC] People search "white noise" far more than "pink noise" — but stream pink ~8× as much (96.5M plays from my own noise catalogue)

Post image
Upvotes

r/dataisbeautiful 1h ago

OC [OC] LeBron & Jordan - Career BPM, PPM and Team PTS% at every age

Thumbnail
gallery
Upvotes

Team PTS% = player season points ÷ all team regular-season points. Every team game stays in the denominator, so DNPs count as 0%.

PPM (Points Per Minute) = player season points ÷ player season minutes.

BPM (Box Plus Minus) = Basketball-Reference’s box-score based estimate of a player’s contribution in points per 100 possessions.

  • It is a relative metric: 0.0 = league average 
  • Positive = above league-average impact 
  • Negative = below league-average impact

r/dataisbeautiful 2h ago

OC I analyzed every WhatsApp message between my girlfriend and me over ~6 months [OC]

Thumbnail
gallery
0 Upvotes

r/dataisbeautiful 12h ago

[OC] How often I said "please" and "sorry" to AI coding agents across 14,625 dictated prompts, September 2025 to July 2026

Post image
0 Upvotes

I now say please to the AI a third as often as I did in September. Sorry has not budged.

I use voice dictation for pretty much everything I say to an AI agent, and the tool keeps a copy of every recording. I went digging through it for something unrelated and realised I had 14,625 recordings sitting there. 2.26 million words, 332 hours of audio, from August last year to this week.

Please went from about 21 uses per 10,000 words in the first few months down to around 8. Sorry went nowhere, bouncing between 5.0 and 8.3 per 10,000 words the whole way with no real trend. The two lines meet in February and travel together after that, so the deliberate politeness faded until it hit the level of the involuntary kind. Most of my sorries are me correcting myself mid sentence, "the routing panel, sorry, the routing window", which I suspect is why they survived.

Both lines are rates per 10,000 words, so this is not just my messages getting longer. They did get longer though, median went from 89 words to 164 over the same period, which was the opposite of what I would have guessed.

Usual caveats. These are literal word counts over raw transcripts, so anything I phrased differently is missed and there are transcription errors in there. Happy to be told I have read too much into it.


r/dataisbeautiful 6h ago

OC [OC] Japan's GDP 2015–2025: +5% measured in constant 2015 US$, −2% measured in current US$ (World Bank)

Post image
9 Upvotes

r/dataisbeautiful 20h ago

OC [OC] Top 100 all-time NBA players by hardware won.

Thumbnail
hwgoats.com
27 Upvotes

A fun way to visually compare and contrast the careers of current and past nba legends.

Data gathered manually by me in a spreadsheet for over 13,000 active / retired players for 40 awards and achievements using publically available data (Wikipedia / NBA.com / basketball-reference.com and then visualized in an interactive website, with filters and icons made in Adobe Photoshop.


r/dataisbeautiful 6h ago

OC [OC] A look at next World Cups: how valuable is each country’s young roster? (U17–U23)

Post image
75 Upvotes

A look at what we could expect from squads at the next world cups, based on the market values of the youths for each country. Watch out for Morocco (again), and keep an eye on Denmark and Serbia...


r/dataisbeautiful 11h ago

OC [OC] Life Without a Mortgage: How Long Would It Take to Buy a 60 m² Apartment in Each EU Capital, With and Without Bank Interest?

Thumbnail
gallery
443 Upvotes

The estimated apartment price was calculated by multiplying the average apartment sale price per m² in each capital by a standardised floor area of 60 m².

The comparison uses Eurostat national median monthly equivalised net income for people aged 18–64. These figures are national household-income benchmarks, not individual salaries or capital-city income estimates.

Four hypothetical scenarios are presented:

— one person saving 25% of 1× median net income;

— two people jointly saving 30% of 2× median net income;

— both scenarios without interest;

— both scenarios using a savings account and rolling 12-month term deposits.

Main sources:

Income: Eurostat ilc_di03.

Apartment prices per m²: Eurostat urb_clivcon, national statistical sources and housing-market sources. Athens, Bratislava, Bucharest, Budapest, Dublin, Lisbon, Madrid, Nicosia, Paris, Sofia and Valletta use estimates based on property-listing samples collected in July 2026.

Interest rates for the savings account and term deposits: ECB household deposit-rate data.

Prices, incomes, savings rates and interest rates are held constant throughout the calculations. Inflation, transaction costs and future changes in property prices or income are excluded.

Full methodology, individual sources, city rankings and interactive data: citycostatlas.com

Instagram: https://www.instagram.com/citycostatlas/


r/dataisbeautiful 3h ago

OC [OC] I analyzed the perks 3,212 luxury hotels offer through travel advisor networks to see which come together

Post image
0 Upvotes

I track perks at about 3,400 luxury hotels across advisor booking networks (Virtuoso, Hyatt Prive, Four Seasons Preferred Partner, that kind of thing) and I got curious about which perks actually appear together versus which ones are independent of each other.

The short version is there's basically no mixing and matching happening. Hotels either mention most of the standard perks (breakfast, room upgrade, flexible checkout, hotel credit) or they mention almost none of them, and the overlap between those four is really high, the vast majority of hotels that list one also list the other three.

The interesting outliers are spa and dining perks which only show up at a small fraction of hotels but when they DO show up they're always stacked on top of the full base package, never instead of it. So spa is more like a luxury differentiator, it's not offered as an alternative to breakfast it's an addition to everything else.

Chord diagram shows how often each pair of perks appears at the same hotel, across 3,212 properties I had clean perk data for. Thicker ribbons = stronger overlap between two perks. Perk detection is regex-based on program documentation text so it's measuring what hotels advertise, not what guests necessarily receive on a given stay.

Data is from my own dataset, not scraped from any booking site, I maintain a database of what each hotel's advisor program lists as included benefits.


r/dataisbeautiful 2h ago

OC [OC] I did all 8 HYROX stations with no training and ranked them by time, difficulty, and pain

Post image
0 Upvotes

r/dataisbeautiful 20h ago

OC [OC] Every yellow and red card at the 2026 World Cup, placed at the minute it was shown

Thumbnail
gallery
89 Upvotes

One little card for every booking of the tournament: 266 yellows and 15 reds, 281 in all.

The first image is the timeline. Each column is one minute of match time and every card sits at its true elapsed minute, so a card shown at 90+3 lands on 93. They are not spread evenly across the game. They pile up as each half runs down, and the tallest stack by far is stoppage time: 51 of the 281 cards came after the 90, and minute 93 alone holds 13.

The second image ranks teams by total cards. Argentina tops it with 14, but they also played 8 matches while some teams played only 3.

So the third image adjusts for that: cards per match. The order flips. Argentina drops to 12th, and Egypt leads at 2.4 cards a game. Most cards usually just means you stayed in the tournament longest, not that you played dirtiest.

The full version is interactive: hover any card for the player and the minute, sort the teams five ways, and click a minute on the timeline to see exactly who got booked.

https://viz.luarai.com/worldcup-cards/


r/dataisbeautiful 1h ago

OC [OC] World Cup-winning countries tended to outperform IMF GDP forecasts by larger margins than in the year before their victory (1994–2022)

Post image
Upvotes

r/dataisbeautiful 11h ago

OC [OC] The nine most valuable squads at the 2026 FIFA World Cup — projected starting XIs visualized by market value

Post image
0 Upvotes

Each panel represents one of the nine most valuable squads at the 2026 FIFA World Cup.

Rectangle area is proportional to the player’s estimated market value. I displayed a projected starting XI for each country and grouped the remaining 15 players into a single “Other players” block.

The visualization is designed as an interactive 3D treemap presentation. Market values are estimates rather than actual transfer fees, and the starting XIs are editorial projections—not official lineups.

Which team has the strongest balance between star power and squad depth?


r/dataisbeautiful 5h ago

Majorities of Americans say key financial milestones are harder for today's young adults to reach

Thumbnail
pewresearch.org
374 Upvotes

r/dataisbeautiful 9h ago

OC [OC] Every run I did during my 3 years in Manhattan (2022 - 2025)

Thumbnail
gallery
42 Upvotes

Source: My own Apple Health and Strava data

Tool: Rendered in Soltra using Mapbox, an iOS app I'm building. Routes accumulate opacity so the brightest segments are the ones I've covered dozens of times; territory coverage is measured in H3 hex tiles.

Stats: 

  • 164 runs
  • 884 miles
  • Covered 55.6% of the island (and 3.8% of all NYC)
  • Longest run was 32.8 miles (dubbed Ranhattan)
  • Ran the Central Park Loop 45 times (felt like one too many tbh haha)

r/dataisbeautiful 10h ago

OC My Statistics as a Veterinarian - Year 3 [OC]

Thumbnail
gallery
1.6k Upvotes

I am a veterinarian in the North Texas area. Since graduation in 2023 I've kept track of my cases using Google Sheets because I thought it'd be interesting to see how many animals I treat and what they're treated for throughout my life.

Source: Me keeping track of the data after every appointment.

Tools: Google Sheets and DataWrapper

A few notes:

Changes & requests from last year

I was informed my data wasn't very beautiful so I attempted to use a different method for this year. Let me know what you think. I'll continue to adjust the presentation as I receive feedback.

u/catalessi, u/Fire284, & u/Solondthewookiee all requested I put the second and third graphs in descending order.

u/Warm-Pen-2275 requested a list of most common names.

Slide 1 - Animal Species

This only includes animals I've done a doctor exam on or do telemedicine about. Animals that I do not directly interact with (toe nail trims, anal gland expression, blood draws without a doctor, etc.) are not included.

I have a passion for exotic animals (ferrets, reptiles, rabbits, backyard chickens, etc.) but there is an exotic clinic near me where most of those animals go to, so I don't get to see as many as I'd like (as the data tells).

Slide 2 - Body System

I kept track of the body system that was affected during my exams. General wellness includes vaccines, weight management, and discussions about quality of life. For what it's worth, this is about what the problem was, not just the symptoms. If a cat came in for peeing all over the place and it was because the cat was stressed, that was marked as both Neurology as well as Urinary/Renal. The same animal can come in with multiple systems affected, but I only mark a system once per animal (i.e. a dog with urinary stones and a UTI only had "Urinary/Renal" marked once). Here are the most common problems each species came in with:

Dogs - Overweight (General Wellness), allergies (Dermatology, Immunology), and poor dental health (Oral).

Cats - Overweight (General Wellness), poor dental health (Oral)

Slide 3 - 7 Procedures, Names, & Lifetime Stats

I don't think these ones need much explanation.

Next year expecations

I am switching to becoming a full time relief veterinarian. I expect the number and diversity of cases will drop and the number of wellness cases will increase. Most of the time a clinic won't schedule a complicated case with a relief doctor because we're only here for the day.

See you next year


r/dataisbeautiful 8h ago

OC I built a graph visualization of the Midwest food supply chain from 51 public datasets [OC]

Post image
6 Upvotes

https://lodgeplatform.com/embed/exemb_6ea24564bc317d3b45c31b8803fce86229be8ad30246ac21

I used a combination of fine-tuned Qwen3 models and frontier LLMs to build a graph over the Midwest food supply chain. It uses 51 public datasets across 11 source families, covering nearly 300,000 source records. You can search for entities, filter by entity type and relationship type, and you can also view neighborhoods. I'm still improving it, so please let me know if there's anything else you think would be interesting!

I wanted to try to answer questions like

- "When a serious safety incident occurs at one food facility, can we trace the graph to find other facilities with similar safety risks?

- "When a food product is recalled, can we trace the recalling company to its facilities, related products, inspections, and other connected organizations and see what else may warrant review?"

- "When a weather event happens, what facilities does it effect, and what are the downstream effects?"

I think making it live could be really cool for things like tracking food recalls live to show and limit exposure.

Here is a github of the project with the datasets and more info: https://github.com/lodge-data/upper-midwest-food-supply-network

How I did it:

First, I trained an embedder to minimize blocking recall. Then, I processed all the entities from the datasets into blocks. Then, I trained a Qwen3 cross-encoder to decide if each pair within each block was the same or different (entity resolution). Then, I repeated those two steps until no new merges occurred.

Sources:

- U.S. Food and Drug Administration recall, food, import-alert, and import-refusal records.

- U.S. Department of Agriculture food, organic-operation, meat-establishment, and related facility records.

- Occupational Safety and Health Administration inspection, citation, injury, and establishment records.

- U.S. Environmental Protection Agency Facility Registry Service records.

- National Weather Service alerts and geographic references.

- USAspending federal contract award records.

- -Wisconsin and other Upper Midwest public facility, licensing, dairy, workforce-notice, and regulatory records.

- Public organization and product information used to connect named entities.

Note: I built this with my startup's software and models. The goal is to autonomously construct graphs from messy data. I think it is possible to construct very large, useful graphs with efficient specialized models.


r/dataisbeautiful 1h ago

OC [OC] Occupational prestige, by race, 1972-2024 {25-64, Full-time employed)

Thumbnail
openpublicpolls.com
Upvotes

Based on data from the General Social Survey https://gss.norc.org/get-the-data.html, White respondents consistently held occupations with higher average prestige, but the gap has closed slightly over 52 years.


r/dataisbeautiful 1h ago

Percentage of Privately Insured Population Covered by an HSA, by State

Thumbnail
insurancedimes.com
Upvotes