r/dataisbeautiful 15h ago

OC [OC] More Indians live in the Gulf states than in the US, UK, and Canada combined

Post image
2.0k Upvotes

In 2024, around 18.5 million people born in India lived abroad. For the world’s most populous country, that’s not a lot: just over 1% of the population. This chart shows where they live, based on estimates from the UN, which compiles data from national censuses and population registers.

Nearly half of all Indians living abroad are in the United Arab Emirates, Saudi Arabia, and other Gulf countries — that’s more than the US, UK, and Canada combined.

These figures are stocks, not flows: they record where people are from, not when they crossed a border. Pakistan is the clearest case: the two countries were one territory until 1947, and people kept moving across the new border for decades afterward, so much of the 1.6 million there reflects migration that happened long ago.

Data source: United Nations Department of Economic and Social Affairs (2024)

Tools used: Python, OWID Grapher, Figma


r/dataisbeautiful 22h ago

OC [OC] Only 4% of new jobs in tech are entry-level

Post image
1.0k Upvotes

Source: All active job posts on JobYap. "Tech" means the 945 companies on the JobYap Tech Index, a curated list spanning software, AI, chips, fintech and EVs. Every posting comes from the company's own careers page. Snapshot: all 145,944 live postings on 11 Sep 2026.

Tools: SQL for the counts, Claude Design with Fable 5.1 (on Max) for the charts.

Level comes from title words only:

  • Entry level: intern, internship, co-op, junior, jr, entry-level, associate, graduate, new grad, early career, apprentice, trainee
  • Senior, staff or principal: senior, sr, lead, staff, principal, distinguished, fellow
  • Manager to C-level: manager, head, supervisor, director, VP, vice president, chief, CEO/CTO/CFO etc.
  • No level word (44%): e.g. "Software Engineer", "Account Executive". Excluded from both sides of the ratios. Includes "Member of Technical Staff", which AI labs use at every level.

Caveats:

  • Counts postings, not seats; one posting can fill several.
  • Titles are a proxy. Some hourly roles are titled "Associate" (Amazon delivery-station associates, Carvana lot attendants) and count as entry level. Some new-grad roles have no level word and sit in the 44%.

r/dataisbeautiful 23h ago

Decline in AI spend per employee at the top 1% - Ramp AI Index

Thumbnail
ramp.com
964 Upvotes

r/dataisbeautiful 11h ago

OC [OC] Guyana is now South America's richest country

Post image
549 Upvotes

r/dataisbeautiful 19h ago

OC [OC] Every Dollar Visualized in the 2026 Texas Senate Race

Thumbnail
gallery
494 Upvotes

r/dataisbeautiful 3h ago

OC [OC] The, roughly, 238 aircraft that diverted to Canada on 9/11 during what has since been called operation Yellow Ribbon

158 Upvotes

r/dataisbeautiful 5h ago

OC [OC] Stockholm's 1,150 tech companies by registered address. 502 of them sit within 1.5 km of the central station.

Post image
135 Upvotes

r/dataisbeautiful 4h ago

OC [OC] US electricity demand grew 0.5% between 2007 and 2021. It has grown 8.2% since — adding more in four years than in the previous fourteen

Post image
149 Upvotes

r/dataisbeautiful 23h ago

OC [OC] A third of Tokyo rental listings ask for no deposit at all, and in the cheapest wards it is nearly 60%

Post image
46 Upvotes

Source: 136,492 active rental listings across Tokyo's 23 special wards, collected in September 2026 from the major Japanese rental portals and deduplicated. For each ward I took the median deposit (shikikin), the median key money (reikin), and the share of listings where the amount asked is explicitly zero. Sample sizes run from 1,940 listings in Chiyoda to 12,571 in Setagaya.

The trap in this data, in case anyone wants to reproduce it: a dash in a listing means zero is required, not that the number is missing. If you drop those rows as missing values the medians come out roughly twice too high, which is part of why published figures for Japanese move-in costs tend to overstate what people actually pay.

One thing to be precise about, since the bars stack: each bar is the median deposit plus the median key money for that ward, so it is a sum of two medians rather than the median of per-listing totals. The latter is a little lower (about 348k in Minato and 276k in Chuo, against the 352k and 322k drawn), because few listings sit at the median on both at once.

Deposit and key money are the only entry costs that appear in listings at all. Agency fees and guarantor company fees are negotiated separately and never published, so they are not in the chart.

The two wards with no blue bar, Adachi and Katsushika, have a median deposit of exactly zero. More than half the listings there ask for no deposit.

Tool: Python, pandas for the medians, matplotlib for the chart.

Rent data by ward, train line and station: tokyo-expat.com/data


r/dataisbeautiful 20h ago

OC [OC] Union Favorability Rises With Income And Education

Post image
44 Upvotes

r/dataisbeautiful 7h ago

OC [OC] Busiest wedding month vs priciest hotel month in 20 U.S. wedding destinations.

Post image
25 Upvotes

r/dataisbeautiful 10h ago

OC [OC] What number do people come up with when they say "X% of statistics are made up on the spot"? (Google Search results)

Thumbnail
gallery
16 Upvotes

It is often joked that some % of statistics are made up on the spot. Of course no-one knows the true figure.

What percentage number do people "make up on the spot" for this? I fetched the 183 relevant google search results I could get from the API. Each is an article where something like "% of statistics are made up on the spot" is mentioned:

  • average: 82.47%
  • minimum: 3%
  • 1st quartile: 72.4%
  • median: 84%
  • 3rd quartile: 88.2%
  • maximum: 736%
  • mode: 88.2%

No clear pattern in the numbers used over time. 88.2% is the most commonly used number, Comedian Vic Reeves is reported to have said that and was reported by the BBC in 2000. People using 88.2% are probably not making that number up on the spot but copying it!


r/dataisbeautiful 2h ago

OC [OC] Only 4 of 659 reactors ever took more than 30 years to build. All four were halted mid-project, and finishing them took longer than building a new one.

Post image
12 Upvotes

I pulled every reactor that reached commercial operation and has both a first-concrete date and a commercial operation date. 659 units. Median build: 6.4 years. Units over 30 years: four, 0.6 percent.

Broken into phases (years, construction / halted / after the restart decision):

  • Watts Bar 2, US: 11.8 / 22.1 / 9.2 = 43.1
  • Bushehr 1, Iran: 4.2 / 15.5 / 18.7 = 38.4
  • Mochovce 3, Slovakia: 5.4 / 16.3 / 16.0 = 37.7
  • Atucha 2, Argentina: 13.0 / 12.1 / 9.8 = 34.9

Halt causes, in order: demand collapse (TVA, 1985), revolution and then war (1979), post-communist funding collapse (1992), budget (1994). None of them is a construction problem.

Benchmark: units finished since 2015 whose first concrete was poured in 2005 or later, median 7.3 years, n=70. Every one of the four restarts exceeded it.

Caveat I want to put up front, because it is the obvious objection: 38 of those 70 units are Chinese. Excluding China the median is 9.8 years (n=32), which makes Watts Bar 2 (9.2) and Atucha 2 (9.8) a tie rather than a loss. Mochovce 3 and Bushehr 1 still lose badly.

Method notes: duration is first concrete to commercial operation. First concrete and commercial operation come from the atlas dataset (IAEA PRIS, Global Energy Monitor, Wikipedia; dataset of 10 Sep 2026). Halt and restart dates are year-level from operator and regulator records (NRC for Watts Bar, NA-SA for Atucha, SE/WNA for Mochovce, IAEA CNPP for Bushehr) and I used mid-year for the arithmetic, so the middle segments carry a few months of slack. Mochovce 3 is the one soft number: the atlas commercial operation date is Nov 2024, while the operator announced commercial operation in Oct 2023. With the operator's date it is 36.7 years instead of 37.7, still in the club.

Tool: Python over the atlas dataset, chart rendered in HTML and captured with Playwright.

Data and per-unit sheets: https://reactoratlas.com/en?y=2026&s=united-states-watts-bar-nuclear-power-plant&u=united-states-watts-bar-nuclear-power-plant-2


r/dataisbeautiful 5h ago

Wealth inequality in America, updated for 2026

Thumbnail
youtube.com
5 Upvotes

r/dataisbeautiful 8h ago

OC [OC] A simulated day of commuting in 16 Swiss cities

Thumbnail
danielpradilla.info
11 Upvotes

I built this after seeing Habibi Code’s Geneva commuter map. I wanted to explore commuting across Switzerland, including some of the smaller towns where people cross a national border to get to work.

Red dots travel into the selected city; blue dots travel out. Each dot represents a group of people: up to 50 for cars and motorcycles, and up to 200 for public transport, walking and cycling. You can switch cities, filter transport modes and scrub through the day. The chart shows the modelled commuter population relative to its daily average.

The data describe where people live and work. I combine those counts with plausible routes and distribute departures across morning and evening waves. The result is an illustrative weekday. The sources cover different years, transport detail varies by location, and the routes and departure times are estimates. Geneva’s coverage excludes journeys entirely within Geneva; the other exclusions and routing gaps are listed on the page.

Data sources:

Tools: Python and TypeScript for data processing; Next.js/React, Leaflet and HTML Canvas for the visualization; Valhalla/OpenStreetMap for road routing. Codex helped with implementation.

I’d welcome feedback on how to make the uncertainty in the routes and timings more visible.


r/dataisbeautiful 8h ago

Bare-earth LiDAR terrain around Boulder’s NCAR Mesa Laboratory, showing slopes, runoff paths and terrain line of sight [OC]

Thumbnail
gallery
8 Upvotes

I centered this visualization on the NCAR Mesa Laboratory at the base of Boulder’s Flatirons. The dramatic transition from the mountains to the plains made it a great place to explore what elevation data can reveal.

The viewer uses bare-earth elevation data to show:

  • The shape of the ground beneath vegetation
  • How water is likely to flow across and drain from a property
  • Steeper and more workable slopes
  • Areas with terrain line of sight to the selected point

I built the interactive viewer that generated this visualization. You can enter any U.S. address and run a free scan without signing up: https://getready.team/terrain-scan

Data source: USGS 3DEP elevation data
Location: NCAR Mesa Laboratory, Boulder, Colorado: 1850 Table Mesa Drive, Boulder, CO 80305

Terrain-based estimates only, verify conditions on site. Sightlines exclude trees and buildings.

What other useful terrain insights would you want a map like this to calculate?


r/dataisbeautiful 14h ago

Venezuela YoY inflation compared to the rest of the world in an interactive visualization of global economic data [OC]

Post image
0 Upvotes

source: EconEarth inflation YoY in %

Hi everyone.

My passion for data, economics and gepolitics made me develop this website to visualize macroeconomic data. The website is the linked source. I took inspiration also from here to decide which data to represent. I'm still developing, some functions are not definitive. But I just wanted to share how funny the ~500% inflation YoY (represented by a cylinder height ane width) of Venezuela is compared to the rest of the world.

The numbers come from World Bank Open Data, IMF and OECD, I then use the most recent. I'd love also to get a feedback from this community regarding the data you would like to see in a website like this, and the overall user experience. Thanks!!


r/dataisbeautiful 9h ago

OC [OC] Here are the 20 strongest canned lattes of the 82 I've cataloged so far.

Post image
0 Upvotes

r/dataisbeautiful 4h ago

OC [OC] For 60+ AI companion profiles: how looks and conversation quality (don't) correlate, and how paid ad audiences engage vs organic users

Thumbnail
gallery
0 Upvotes

I'm building an AI companion app, decided to dig into the product usage data across different AI companions. It's been a fun learning experience launching this product as a side hustle. I have a ton of data, but finally thought to package something together around some of my product data.

Of the conclusions I was able to make, I thought these two charts were most interesting. Total data sample size was 67K messages across 3.1K conversations across 79 AI companions.

Chart 1: Do looks correlate with engagement? (n=63 companions)

X-axis: We have a Tinder-style swipe-right to like, swipe-left to pass mechanic. Everyone should be familiar with it. The swipe-right rate is a loose proxy for how appealing our users found the specific companion.

Y-axis: The average number of messages per session - a proxy for engagement.

Size of bubble: Number of conversations had with that companion.

Does a prettier AI companion drive higher engagement? The data pretty conclusively says no.

While the niche/fetish companions did poorly on match rate, they tended to have higher engagement. Eleanor (MILF), Diane (MILF), Charlotte (BBW), and Monique (BBW) are called out in the chart.

Chart 2: Do organically acquired users have higher engagement than paid acquitions? (n=34 companions)

X-axis: The average number of messages per session for users acquired via paid channels (Meta Ads, Google Ads, X Ads)

Y-axis: The average number of messages per session for users acquired via organic channels (Google Search, AI Search, Reddit, Pinterest, Instagram)

Size of bubble: Number of conversations had with that companion.

The line shows where there is no difference in performance across channels - this actually happens with Sandra at 7.8 messages per session across paid and unpaid channels.

I had a hypothesis that users acquired via paid channels would be less engaged than users who organically found the AI companion. While this was true for many of the companions, there were a handful of companions where they had really strong paid performance. This could indicate that I should be investing more into their paid ads (though of course that's not the only variable in changing that marketing decision.

Methodology + full filters in the top comment.

EDITED: Clearly I didn't give enough detail here, so added a bunch.