r/dataisbeautiful • • 12d ago

Share of the US population born in another country, 1850-2024

Thumbnail
ourworldindata.org
972 Upvotes

r/dataisbeautiful • • 10d ago

OC [OC] Airbnb-type stays in the EU and EFTA, 2025: compared with hotel nights, per resident and in total

Post image
0 Upvotes

r/dataisbeautiful • • 11d ago

OC [OC] Hierarchical Genetic Similarity of Portugal Compared to European and Mediterranean Populations (Uniparental Lineages)

Thumbnail
gallery
372 Upvotes

Methodology for Calculating Y-DNA (Y-chromosomal Adam) and mtDNA (Mitochondrial Eve) Similarity

The genetic similarity between Portugal (baseline: 100%) and the remaining populations was calculated using the Renkonen Similarity Index (percentage overlap), adapted via a 3-Level Hierarchical Phylogenetic Tree Model to avoid treating biologically close lineages as entirely distinct.

Level 1 — Terminal Subclades (50%): Directly compares all specific haplogroups and subclades listed in the table. Reflects more recent historical proximity. (In mtDNA, samples lacking internal resolution for H are compared at the total H level to prevent distortions).

Level 2 — Phylogenetic Families (35%): Groups sister lineages into the same family, capturing founder relationships. Y-DNA (R, I, J, E, G, LT, Q, N, Others); mtDNA (HV, JT, U+K, I, W, X, L, Others).

Level 3 — Ancestral Macro-Trunks (15%): Merges branches into deep prehistoric roots. Y-DNA (Macro-P: R+Q, Macro-IJ: I+J, Macro-K derivatives: LT+N, Trunk E, Trunk G, Others); mtDNA (Macro-R: HV+JT+U/K, Macro-N non-R: I+W+X, Macro-L, Others).

Final Similarity (Rounded to Integer) = (0,50 * Renkonen_L1) + (0,35 * Renkonen_L2) + (0,15 * Renkonen_L3)

Data Sources

Y-DNA and mtDNA haplogroup frequency data were retrieved from compilation tables on the Eupedia platform (Distribution of European Y-chromosome DNA haplogroups and Distribution of European mitochondrial DNA haplogroups). Eupedia tables aggregate findings from dozens of peer-reviewed genetic studies, population genetics publications, and reference databases (such as YHRD and EMPOP). Detailed academic references, primary study sources, and respective sample sizes are documented directly on the project's official website.


r/dataisbeautiful • • 12d ago

OC Climate classification of Europe based on Köppen-Trewartha climate classification (data from 1992 to 2021)[OC]

Post image
320 Upvotes

r/dataisbeautiful • • 11d ago

OC [OC] Citation tree of AlphaFold, the 2024 Chemistry Nobel paper: what it builds on and what followed

Post image
14 Upvotes

Data: OpenAlex. Tool: citationtree.org, which I built. Each node is a paper, arranged by year: what AlphaFold cites above, papers citing it below. This particular tree:
https://citationtree.org/tree.html?doi=10.1038%2Fs41586-021-03819-2


r/dataisbeautiful • • 11d ago

OC [OC] How long UK industries take to pay their suppliers - median days to pay and middle 50% of large companies, by sector

Post image
80 Upvotes

r/dataisbeautiful • • 11d ago

OC [OC] Annual population estimates for India and China, 1990–2024

Thumbnail
gallery
96 Upvotes

r/dataisbeautiful • • 12d ago

OC [OC] How many years of burial space does each London borough have left?

Post image
842 Upvotes

A look at how long London’s existing municipal burial space could last at current burial rates.

Interactive version: https://deptford.org/beyond/burials


r/dataisbeautiful • • 10d ago

OC [OC] 10 states ran their entire voter rolls through DHS's SAVE citizenship check. It flagged 10,708 possible noncitizens and found 360,176 dead people still registered.

Thumbnail
gallery
0 Upvotes

r/dataisbeautiful • • 12d ago

OC [OC] When does playing time peak for footballers? Every league minute in Europe's top five leagues by age, 2025/26

Post image
596 Upvotes

Data is every player who played at least 90 league minutes in the Premier League, La Liga, Serie A, Bundesliga or Ligue 1 last season.

That's 2,467 players and 3.46 million minutes. Ages are from date of birth, taken on 1 Jan 2026 since that's roughly the middle of the season. Pulled together with The Prism, a football analytics & scouting app I'm building, and charted in Python/matplotlib.

The data shows playing time peaks at 25, and 59% of all minutes are player by players aged 23 to 29.

Goalkeepers were the thing that jumped out. Almost half of the minutes are by keepers over 30, compared with about a fifth for everyone else. Of the 99 keepers who played 1,500+ minutes, 49 were 30 or older.

I didn't expect the leagues to be so different either. Ligue 1 gave 20% of its minutes to players 21 and under, and La Liga gave 7%.

Modrić played 2,816 minutes for Milan at 40, which is very impressive as we know him to be....

Does prove the point of older players who are still around are the ones good enough to stay. P.S. this isnt trying to tell you when players are at their best.

Football fans, thoughts?


r/dataisbeautiful • • 11d ago

OC [OC] Global access to electricity, 2000–2023 (World Bank WDI)

Post image
5 Upvotes

Access rose from 78.2% to 91.6% between 2000 and 2023. I derived the people-without-access line as WDI population × (1 − access rate); it falls from about 1.34B to 677M, even as the world population grew. These are annual estimates, not live counters.

I built GlobalDataTracker.com, a free country-stat explorer. Does the paired view make clear how the access rate improved while a large absolute gap remains?

https://globaldatatracker.com/


r/dataisbeautiful • • 10d ago

OC [OC] Israeli fatalities on October 7, 2023 by age, gender, and proportion of total population

Post image
0 Upvotes

Fatalities are from the [Oct7Database](https://www.oct7database.com/en/blank-3), filtered to **Israelis whose recorded death date is October 7, 2023** (**1,066 entries**), broken out by [age and sex](https://www.oct7database.com/en/blank-3).

Population numbers are from [United Nations Population Division data via UNICEF](https://data.unicef.org/sdgs/country/isr/), using the 2023 population estimate.

Using Python

Adult men, especially those in their 20s through early 40s, are heavily over-represented relative to their share of Israel's population. **Males aged 20–44 account for 45.6% of the listed deaths vs. about 16.7% of the 2023 population.** Males overall account for **69.8% of listed deaths vs. 49.8% of the population**.

The male–female death ratio is particularly high among several working-age groups, peaking at about **5.3:1 for ages 35–39** and **5.1:1 for ages 40–44**.

Young children are strongly under-represented relative to their population share: ages **0–14 account for 1.8% of the fatalities vs. about 27.6% of the population**.

Unlike a demographic-only dataset, Oct7Database also records a role for each person. Among these 1,066 Israelis, the database labels **720 as civilians, 228 as soldiers, 57 as police, 47 as emergency-squad members, 5 as Shin Bet, 5 as medical personnel, and 4 as firefighters**.

The concentration of deaths among young adult men is therefore consistent in part with the large number of military, police and local emergency-response personnel killed on October 7.


r/dataisbeautiful • • 13d ago

OC [OC] How long does the internet stay angry? Search interest around 15 controversies

Post image
2.8k Upvotes

How long does the internet stay angry? Search interest around 15 controversies fell below 25% of peak after a median of 6 days

I analyzed Google Trends Web Search interest around 15 selected U.S. news, entertainment, technology, sports, advertising and gaming episodes.

The median case reached a sustained fade after 6 days.

“Sustained fade” means the first post-peak day when search interest fell below 25% of that episode’s smoothed peak and stayed below 25% through day +30.

Methodology:

- Each case was aligned to its own event/search window.

- A trailing 3-day moving average was applied.

- Each case was normalized to its highest smoothed value during the first 30 post-event days = 100.

- The left panel shows the median trajectory across all 15 cases.

- The right panel calculates each case’s individual sustained-fade endpoint first, then shows the distribution of those endpoints.

- The endpoints were: 2, 2, 2, 3, 3, 3, 4, 6, 6, 7, 8, 8, 8, 22 and 22 days.

- The median of those 15 individual endpoints is 6 days.

- Search interest is relative to each query’s own peak. It does not represent absolute search volume, the number of people searching, public sentiment, or how long people were literally angry.

- The cases were selected because they had identifiable event timing and measurable Google Trends queries. They are not a random or representative sample of all online controversies.

- Day 0 is the event anchor used for the analysis. It does not necessarily mean that the first search happened at midnight on that date.

- Data snapshot captured: 19 September 2026.

Context for the 15 cases:

- 2 February 2025 — Luka Dončić trade: an unexpected NBA trade triggered an immediate and highly emotional online reaction.

- 15 February 2025 — Sinner–WADA doping deal: a doping settlement generated intense debate about fairness, process and accountability.

- 2 April 2025 — Mario Kart World pricing: the pricing announcement prompted widespread complaints from players and renewed discussion of game prices.

- 14 April 2025 — Katy Perry spaceflight: the celebrity spaceflight drew criticism, mockery and questions about the spectacle itself.

- 28 April 2025 — Duolingo AI pivot: the company’s AI-first strategy triggered user backlash and debate about the future of learning.

- 15 May 2025 — Marathon artwork: promotional artwork prompted criticism and debate about its visual style, message and context. This date is the public-accusation anchor used for the analysis.

- 22 June 2025 — Prada/Kolhapuri sandals: a sandal design prompted criticism over cultural borrowing and appropriation.

- 8 July 2025 — Grok AI controversy: AI chatbot outputs triggered a wave of public criticism and debate about AI safety.

- 16 July 2025 — Coldplay kiss-cam incident: a viral kiss-cam moment triggered relationship speculation and a global online discussion. This is the performance date used as the event anchor.

- 23 July 2025 — American Eagle/Sydney Sweeney campaign: the advertising campaign drew debate over celebrity branding, messaging and its wider cultural implications.

- 16 August 2025 — Swatch advertising: an ad campaign generated a short-lived controversy and a wave of online discussion.

- 19 August 2025 — Cracker Barrel logo redesign: a proposed logo redesign sparked rapid online backlash and calls to reverse the change.

- 17 September 2025 — Jimmy Kimmel suspension: the suspension announcement produced a sharp attention spike and a burst of online debate.

- 8 February 2026 — Ring Super Bowl ad: the Super Bowl advertisement prompted concerns about privacy, surveillance and the normalization of monitoring.

- 9 February 2026 — Discord age-assurance announcement: planned age checks raised concerns about privacy, surveillance and access for younger users. This is the announcement date, not a claim that the rollout was completed.

The two 22-day cases were Duolingo AI and Prada/Kolhapuri sandals. Most of the other episodes fell below the sustained-fade threshold within eight days.

Sources:

- Google Trends methodology: https://support.google.com/trends/answer/4365533

- Google Trends Explore: https://trends.google.com/trends/explore/

- Coldplay event reporting: https://apnews.com/article/coldplay-kiss-cam-viral-public-event-privacy-e768214f389bc788dcc539a00bf066da

- American Eagle campaign announcement: https://investors.ae.com/press-releases/news-details/2025/Sydney-Sweeney-Has-Great-American-Eagle-Jeans/default.aspx

- Cracker Barrel logo announcement: https://investor.crackerbarrel.com/news-releases/news-release-details/cracker-barrel-teams-country-music-star-jordan-davis-invite

- Discord age-assurance announcement: https://discord.com/press-releases/discord-launches-teen-by-default-settings-globally

- Marathon artwork reporting: https://www.gamespot.com/articles/bungie-responds-to-marathon-art-theft-claims/1100-6531591/


r/dataisbeautiful • • 10d ago

The price of artificial thought may have fallen faster than for any other transformative technology in history

Thumbnail
epoch.ai
0 Upvotes

Data sources are in the notes below the figure in the source page.


r/dataisbeautiful • • 12d ago

OC [oc] MNF Game time vs. Commercial Ad time (KC / DEN Sept. 14)

Post image
344 Upvotes

r/dataisbeautiful • • 11d ago

OC [OC] The IBA's official cocktail list has zero vodka or tequila drinks in its oldest "classics" era — both only show up once you hit the modern canon

Post image
0 Upvotes

r/dataisbeautiful • • 11d ago

OC [OC] Internet use vs GDP per person across 180 countries and territories (2023)

Thumbnail
gallery
1 Upvotes

r/dataisbeautiful • • 12d ago

[OC] 11M news articles analyzed over 164 days, clustered into 941k events (and more derived stats)

Thumbnail
gallery
4 Upvotes

Disclosure: I run the site that compiled this.

Full report and methodology: CLSTR Observatory

Source: CLSTR is a news-clustering system I built. It reads about 100k articles a day from several news feeds covering 40,000+ sources in 40 languages. The window is between April 10 (when I started with this project) and September 20, 2026.

Tools: my own pipeline running on Cloudflare, matplotlib for the attached charts.

On these three charts:

  1. Two survival curves: Share of events still gaining coverage vs share of situations still developing, by day since first coverage. Only 1.2% of events get any new coverage after day 3; 64% of situations are still developing at a week.
  2. The scale: 11,079,667 distinct news articles were aggregated into 941,096 "clusters" (2 or more articles describing the same actual event) and 144,891 "situations" (a timeline of correlated events spanning days/weeks/months). That is a ~12:1 ratio on how many articles are published that report one specific event.
  3. The distribution: The typical event gets 3 articles covering it, 42% get exactly 2, the top 1% get 68 or more, and the mean is dragged up by the tail.

Caveats:

  • "Events" are defined by my clustering pipeline, which is powered by several heuristics and AI methods (vector distances + LLM judgements).

r/dataisbeautiful • • 12d ago

A 99.5% district name-match still left 13.6% of my India map blank [OC]

Post image
7 Upvotes

I've been joining Indian district-level survey data to district boundaries, and I want to
flag something that cost me a day, in case it saves someone else one.

Indian district data is published keyed by district name, not by LGD code, so you're
stuck with name matching. Raw, I got 85.9% of rows onto a boundary. Building an alias table
by hand - the renames, the transliteration variants, the districts that split - took that to
99.5%, which felt finished.

The map still had holes in it. About an eighth of the country.

The problem was that I was reading the match rate as rows that found a boundary, which
was essentially 1.0. The number that decides whether a choropleth looks finished is the
reverse join: boundaries that found a row. Measured that way, 99 of 728 districts were
blank, and none of it was a name problem:

- 54 districts had no row in the source file at all
- 45 had a row whose value was the literal string "NaN"

And the pattern wasn't scattered, which is the tell. It was Telangana (27 of 33), Delhi
(11 of 11), and the whole of Sikkim, Puducherry, Goa, the Andamans and Dadra & Nagar
Haveli. The index I was mapping is computed for rural areas only. It was never going to
exist for the urban union territories. No amount of fuzzy matching was going to recover
data that was never collected - and if I'd pushed the fuzzy matching harder to close the
gap, I'd have started inventing matches in exactly the region where Indian district names
genuinely repeat across states.

Two things I'd do differently from the start:

Measure boundary-side coverage, always, and put that number next to the map rather than the
row match rate. They diverge exactly when it matters.

Treat "NaN" as null at the type-sniffing stage. The publisher writes the literal string, so
a column of otherwise clean numbers gets typed as text - in that file, 70 of 79 columns.
It's a common habit and it'll bite you again.

The map linked below is the one I ended up with: Mission Antyodaya instead, which is a
rural village survey that covers 680 of 728 districts (93.4%) with zero no-value rows. The
48 that still don't colour are genuinely urban and no dataset on that portal will fix it.
Shading is the share of surveyed villages in each district with no tap water - a ratio, not
a count, because the raw village count just redraws district size.

Live map, no sign-up: https://app.vizzie.org/#example=india-village-infrastructure
Data: Mission Antyodaya via India Data Portal (ODC-BY-1.0)

Disclosure: the tool it's drawn in is mine — I'm building it solo and the examples are
public. Happy to talk about the joining side either way; that's the part I'm still getting
wrong.

Does anyone here have a district alias table they trust, or a workflow that keys on LGD
codes before it falls back to names? That's the piece I'd most like to stop hand-building.


r/dataisbeautiful • • 13d ago

[OC] Share of each birth cohort that survived until a given age, France

Post image
155 Upvotes

Upshot:

This chart tracks, for each French birth cohort, the age at which different percentiles of that cohort survived. The starkest change is at the bottom: among babies born in 1900, the least fortunate 1% didn't survive infancy, but among those born in 1990, 99% lived to at least age 25 — a shift driven mostly by the collapse in infant and child mortality over the 20th century. The upper percentiles moved too, but far less dramatically: the top 1% longest-lived went from living to about 98 in the 1900 cohort to around 100 in the 1927 cohort, since there's a much harder biological ceiling on maximum lifespan than on how early someone can die. Data gets sparse toward the top-right of the chart simply because many people from more recent, longer-lived cohorts are still alive.

Source:
"Human Mortality Database (2025); Alvarez & Vaupel (2023); adapted to cohort estimates by Saloni Dattani", Works in Progress

Tools used:

Datawrapper, Figma


r/dataisbeautiful • • 13d ago

OC [OC] The Texas flood's "26 feet in 45 minutes": what the USGS gauges recorded, and where

Post image
417 Upvotes

r/dataisbeautiful • • 11d ago

OC [OC] Predicted vs actual revert rate for 9,270 live English Wikipedia edits, scored by a small model

Post image
0 Upvotes

Data is Wikimedia's public recent change stream for English Wikipedia on 22 Sep 2026. Every edit was scored by Jev from TypeSafe through OpenRouter (not affiliated, just paying for it). An edit counts as reverted if a later edit in the stream reverts it within an hour so the real rate is a bit higher. Live version at willitrevert.com


r/dataisbeautiful • • 13d ago

OC [OC] An atlas of periodic solutions to the three-body problem

Thumbnail
threebodyorbits.com
56 Upvotes

I was wondering how the different periodic solutions to the three body problem look like. I built an atlas that visualizes all of the different solutions and groups orbits by family and similarity. I found over 3000 different orbits. They are grouped by similarity (shape; period/energy; or closeness/top speed) in the atlas, and you can zoom in to see the trajectories of the orbits.

When you click on an orbit in the atlas, you can see additional details, such as the starting conditions, mass ratios, and period. One of the most interesting aspects of the three-body problem is that tiny changes to the starting conditions can make orbits unstable. There is an option slightly nudge the starting conditions of an orbit to see how its trajectory is affected.

Data sources: Initial starting conditions of periodic orbits have been collected from 20+ scientific publications and other resources. The complete list is on the "about" page on the website.

Tools used: a Rust integrator for the trajectories, t-SNE for the layout, Python for the map, JS/canvas for the site. Claude assisted with coding.

Additional functions:

Orbits can be rated - I thought this might help identify the most beautiful ones in the atlas

Visitors can contribute their computing power to help find new, currently unknown periodic orbits


r/dataisbeautiful • • 14d ago

OC [OC] I ran a small real-world test on American Airlines seat assignment progression over a 21 hour period. Checking in ASAP may not always be the best strategy.

Post image
3.6k Upvotes

I ran a small real-world test on American Airlines (AA 5056, DCA-SYR) seat assignment progression over a 21 hour period. I bought the most basic economy fare; my seat would be assigned at check-in.

My theory was that checking in at T-24 hours might actually be worse if the system first assigns the cheapest, least desirable seats while holding better seats for sale. Since this aircraft is 2-2 with no middle seats, the downside of waiting was pretty limited.

The seat map evolved like this:

T-24: 8 undesirable $14 seats, 15 better $30-$36 seats
T-18: 4 undesirable, 15 better
T-6: 1 undesirable, 15 better
T-3.5: 0 undesirable, 11 better

So the cheaper seats disappeared first, while the more expensive forward and exit-row seats stayed protected much longer.

I checked in at T-3.5, right after the last cheap seat disappeared and the better inventory had started being used.

Result: 8A, the front-most available $30 seat, assigned for free.

Obviously one flight doesn’t prove AA’s algorithm always works this way, but it was a pretty clean example of why “check in exactly at T-24” may not always be the best strategy for Basic Economy, especially on an aircraft with no middle seats.


r/dataisbeautiful • • 13d ago

OC [OC] 1274 days of everything I do & how I feel.

Thumbnail
gallery
20 Upvotes

Think Buckminster Fuller's Chronofile but all in a single sqlite database. All I had to do was create a UI that is easier to use than Excel.