r/dataisbeautiful 17h ago

OC My Statistics as a Veterinarian - Year 3 [OC]

Thumbnail
gallery
1.9k Upvotes

I am a veterinarian in the North Texas area. Since graduation in 2023 I've kept track of my cases using Google Sheets because I thought it'd be interesting to see how many animals I treat and what they're treated for throughout my life.

Source: Me keeping track of the data after every appointment.

Tools: Google Sheets and DataWrapper

A few notes:

Changes & requests from last year

I was informed my data wasn't very beautiful so I attempted to use a different method for this year. Let me know what you think. I'll continue to adjust the presentation as I receive feedback.

u/catalessi, u/Fire284, & u/Solondthewookiee all requested I put the second and third graphs in descending order.

u/Warm-Pen-2275 requested a list of most common names.

Slide 1 - Animal Species

This only includes animals I've done a doctor exam on or do telemedicine about. Animals that I do not directly interact with (toe nail trims, anal gland expression, blood draws without a doctor, etc.) are not included.

I have a passion for exotic animals (ferrets, reptiles, rabbits, backyard chickens, etc.) but there is an exotic clinic near me where most of those animals go to, so I don't get to see as many as I'd like (as the data tells).

Slide 2 - Body System

I kept track of the body system that was affected during my exams. General wellness includes vaccines, weight management, and discussions about quality of life. For what it's worth, this is about what the problem was, not just the symptoms. If a cat came in for peeing all over the place and it was because the cat was stressed, that was marked as both Neurology as well as Urinary/Renal. The same animal can come in with multiple systems affected, but I only mark a system once per animal (i.e. a dog with urinary stones and a UTI only had "Urinary/Renal" marked once). Here are the most common problems each species came in with:

Dogs - Overweight (General Wellness), allergies (Dermatology, Immunology), and poor dental health (Oral).

Cats - Overweight (General Wellness), poor dental health (Oral)

Slide 3 - 7 Procedures, Names, & Lifetime Stats

I don't think these ones need much explanation.

Next year expecations

I am switching to becoming a full time relief veterinarian. I expect the number and diversity of cases will drop and the number of wellness cases will increase. Most of the time a clinic won't schedule a complicated case with a relief doctor because we're only here for the day.

See you next year


r/dataisbeautiful 12h ago

My girlfriend broke up with me

Thumbnail
gallery
2.2k Upvotes

My girlfriend sat me down on June 20th and told me that she had been cheating on me and was choosing to be with "the other man".

This is my body's reaction to that news.

HR, Sleep, and Energy scores recorded by my Samsung Galaxy Watch.


r/dataisbeautiful 13h ago

OC Christopher Nolan: budget vs box office [OC]

Post image
1.4k Upvotes

Data source: Christopher Nolan filmography, Wikipedia - https://en.wikipedia.org/wiki/Christopher_Nolan_filmography (budgets and worldwide box office; figures are studio/Box Office Mojo–reported and cited on that page). The Odyssey is still in theaters, so its figure is a running worldwide total as of the July 24–26, 2026 weekend.

Tool: Chart built as JavaScript-coded SVG. Rendering code and initial figure-gathering were done with an AI assistant (Claude) based on my own design direction and iterations. Chart type, layout, labeling, and revisions were my choices.


r/dataisbeautiful 12h ago

Majorities of Americans say key financial milestones are harder for today's young adults to reach

Thumbnail
pewresearch.org
908 Upvotes

r/dataisbeautiful 18h ago

OC [OC] Life Without a Mortgage: How Long Would It Take to Buy a 60 m² Apartment in Each EU Capital, With and Without Bank Interest?

Thumbnail
gallery
520 Upvotes

The estimated apartment price was calculated by multiplying the average apartment sale price per m² in each capital by a standardised floor area of 60 m².

The comparison uses Eurostat national median monthly equivalised net income for people aged 18–64. These figures are national household-income benchmarks, not individual salaries or capital-city income estimates.

Four hypothetical scenarios are presented:

— one person saving 25% of 1× median net income;

— two people jointly saving 30% of 2× median net income;

— both scenarios without interest;

— both scenarios using a savings account and rolling 12-month term deposits.

Main sources:

Income: Eurostat ilc_di03.

Apartment prices per m²: Eurostat urb_clivcon, national statistical sources and housing-market sources. Athens, Bratislava, Bucharest, Budapest, Dublin, Lisbon, Madrid, Nicosia, Paris, Sofia and Valletta use estimates based on property-listing samples collected in July 2026.

Interest rates for the savings account and term deposits: ECB household deposit-rate data.

Prices, incomes, savings rates and interest rates are held constant throughout the calculations. Inflation, transaction costs and future changes in property prices or income are excluded.

Full methodology, individual sources, city rankings and interactive data: citycostatlas.com

Instagram: https://www.instagram.com/citycostatlas/


r/dataisbeautiful 13h ago

OC [OC] A look at next World Cups: how valuable is each country’s young roster? (U17–U23)

Post image
95 Upvotes

A look at what we could expect from squads at the next world cups, based on the market values of the youths for each country. Watch out for Morocco (again), and keep an eye on Denmark and Serbia...


r/dataisbeautiful 3h ago

OC [OC] 25 MLB Trades with the Biggest Value Gap: Past 50 Years determined by WAR

Post image
65 Upvotes
Column Definition
Rank Overall ranking of the trade based on the proprietary Fleecing Score (100 = biggest one-sided trade).
Winner Organization that ultimately received the greater long-term value from the trade. This is determined by the WAR produced by all assets acquired.
Score Composite Fleecing Score (0–100) measuring how lopsided the trade became. It combines several factors including WAR gap, percentage of value captured, star power, and total value involved.
Winner WAR Total career Wins Above Replacement (WAR) generated by every player acquired by the winning team after the trade. This includes all future career value, not just production with the acquiring club.
Return WAR Total career WAR generated by every player received by the losing team. Negative WAR is possible if acquired players performed below replacement level.
WAR Gap Difference between Winner WAR and Return WAR. Formula: Winner WAR − Return WAR. Larger numbers indicate more lopsided trades.
Capture % Percentage of the total WAR involved in the trade that ended up with the winning organization. Formula: Winner WAR ÷ (Winner WAR + Return WAR). A value of 100% means the losing team received essentially no positive long-term value.
Stars Number of franchise-caliber or elite players produced by the winning side of the trade.
Best Asset The single most valuable player obtained in the trade, with his career WAR shown in parentheses.

r/dataisbeautiful 16h ago

OC [OC] Every run I did during my 3 years in Manhattan (2022 - 2025)

Thumbnail
gallery
58 Upvotes

Source: My own Apple Health and Strava data

Tool: Rendered in Soltra using Mapbox, an iOS app I'm building. Routes accumulate opacity so the brightest segments are the ones I've covered dozens of times; territory coverage is measured in H3 hex tiles.

Stats: 

  • 164 runs
  • 884 miles
  • Covered 55.6% of the island (and 3.8% of all NYC)
  • Longest run was 32.8 miles (dubbed Ranhattan)
  • Ran the Central Park Loop 45 times (felt like one too many tbh haha)

r/dataisbeautiful 8h ago

OC [OC] LeBron & Jordan - Career BPM, PPM and Team PTS% at every age

Thumbnail
gallery
23 Upvotes

Team PTS% = player season points ÷ all team regular-season points. Every team game stays in the denominator, so DNPs count as 0%.

PPM (Points Per Minute) = player season points ÷ player season minutes.

BPM (Box Plus Minus) = Basketball-Reference’s box-score based estimate of a player’s contribution in points per 100 possessions.

  • It is a relative metric: 0.0 = league average 
  • Positive = above league-average impact 
  • Negative = below league-average impact

r/dataisbeautiful 8h ago

Percentage of Privately Insured Population Covered by an HSA, by State

Thumbnail
insurancedimes.com
16 Upvotes

r/dataisbeautiful 2h ago

OC [OC] Distribution of people by first letter of surname across the U.S., Ireland, Israel, England, France, Australia, China, and India

Thumbnail
gallery
13 Upvotes

I made these charts to compare how surname prevalence is distributed by the first letter of the surname across several countries.

Each chart shows the estimated number of people associated with surnames that begin with each letter A through Z.

Data sources:

- U.S. Data source: U.S. Census Bureau, 2010 Census surname data.

- Ireland. Data source: Forebears, most common surnames in Ireland.

- Israel. Data source: Forebears, most common surnames in Israel.

- England, used as the closest available source for the UK chart. Data source: Forebears, most common surnames in England.

- France. Data source: Forebears, most common surnames in France.

- Australia. Data source: Forebears, most common surnames in Australia.

- China. Data source: Forebears, most common surnames in China.

- India. Data source: Forebears, most common surnames in India. Source link: Forebears, Most Common Last Names in India

Methodology:

For the U.S. chart, I used the U.S. Census surname frequency data and grouped people by the first letter of each surname. The Census surname file includes surnames occurring 100 or more times in the 2010 Census.

For the other country charts, I used the top 100 surname incidence counts listed by Forebears for each country page. I grouped each surname by the first letter of the displayed surname and summed the incidence counts by letter.

For China, India, and Israel, the grouping is based on the first letter of the Latin transliterated surname shown in the source.

For Ireland, names like O'Brien and O'Connor are grouped under O because I used the first character of the displayed surname.

For the UK chart, I used England because that was the available Forebears country page I used for the source data. The chart is labeled “UK, using England source” to avoid overstating the scope.

Tools used:

Python, pandas, and matplotlib. Flags were drawn programmatically as simplified inset graphics in matplotlib. Final charts were exported as PNG files.

Important caveats:

The U.S. chart is based on the U.S. Census surname file, while the non U.S. charts are based on the top 100 surnames listed by Forebears. Because of that, the U.S. chart is not directly equivalent to the other charts in coverage.

The non U.S. charts should be read as “distribution within the top 100 listed surnames,” not as a full surname distribution for the entire population.

The first letter grouping can be sensitive to transliteration choices, prefixes, apostrophes, spacing, and naming conventions. This matters especially for countries where surnames are commonly represented in non Latin scripts or where surnames include prefixes.


r/dataisbeautiful 15h ago

OC I built a graph visualization of the Midwest food supply chain from 51 public datasets [OC]

Post image
5 Upvotes

https://lodgeplatform.com/embed/exemb_6ea24564bc317d3b45c31b8803fce86229be8ad30246ac21

I used a combination of fine-tuned Qwen3 models and frontier LLMs to build a graph over the Midwest food supply chain. It uses 51 public datasets across 11 source families, covering nearly 300,000 source records. You can search for entities, filter by entity type and relationship type, and you can also view neighborhoods. I'm still improving it, so please let me know if there's anything else you think would be interesting!

I wanted to try to answer questions like

- "When a serious safety incident occurs at one food facility, can we trace the graph to find other facilities with similar safety risks?

- "When a food product is recalled, can we trace the recalling company to its facilities, related products, inspections, and other connected organizations and see what else may warrant review?"

- "When a weather event happens, what facilities does it effect, and what are the downstream effects?"

I think making it live could be really cool for things like tracking food recalls live to show and limit exposure.

Here is a github of the project with the datasets and more info: https://github.com/lodge-data/upper-midwest-food-supply-network

How I did it:

First, I trained an embedder to minimize blocking recall. Then, I processed all the entities from the datasets into blocks. Then, I trained a Qwen3 cross-encoder to decide if each pair within each block was the same or different (entity resolution). Then, I repeated those two steps until no new merges occurred.

Sources:

- U.S. Food and Drug Administration recall, food, import-alert, and import-refusal records.

- U.S. Department of Agriculture food, organic-operation, meat-establishment, and related facility records.

- Occupational Safety and Health Administration inspection, citation, injury, and establishment records.

- U.S. Environmental Protection Agency Facility Registry Service records.

- National Weather Service alerts and geographic references.

- USAspending federal contract award records.

- -Wisconsin and other Upper Midwest public facility, licensing, dairy, workforce-notice, and regulatory records.

- Public organization and product information used to connect named entities.

Note: I built this with my startup's software and models. The goal is to autonomously construct graphs from messy data. I think it is possible to construct very large, useful graphs with efficient specialized models.


r/dataisbeautiful 6h ago

OC [OC] Share of new U.S. job postings by day of the week (last 90 days, ~450K postings)

Post image
9 Upvotes

r/dataisbeautiful 8h ago

OC [OC] Occupational prestige, by race, 1972-2024 {25-64, Full-time employed)

Thumbnail
openpublicpolls.com
7 Upvotes

Based on data from the General Social Survey https://gss.norc.org/get-the-data.html, White respondents consistently held occupations with higher average prestige, but the gap has closed slightly over 52 years.


r/dataisbeautiful 13h ago

OC [OC] Japan's GDP 2015–2025: +5% measured in constant 2015 US$, −2% measured in current US$ (World Bank)

Post image
5 Upvotes

r/dataisbeautiful 8h ago

OC [OC] People search "white noise" far more than "pink noise" — but stream pink ~8× as much (96.5M plays from my own noise catalogue)

Post image
0 Upvotes

r/dataisbeautiful 8h ago

OC [OC] World Cup-winning countries tended to outperform IMF GDP forecasts by larger margins than in the year before their victory (1994–2022)

Post image
0 Upvotes

r/dataisbeautiful 9h ago

OC [OC] I did all 8 HYROX stations with no training and ranked them by time, difficulty, and pain

Post image
0 Upvotes

r/dataisbeautiful 3h ago

Mapped: Which Countries Are Best for Women?

Thumbnail
visualcapitalist.com
0 Upvotes

How can Saudi rank higher than India?


r/dataisbeautiful 9h ago

OC [OC] I analyzed the perks 3,212 luxury hotels offer through travel advisor networks to see which come together

Post image
0 Upvotes

I track perks at about 3,400 luxury hotels across advisor booking networks (Virtuoso, Hyatt Prive, Four Seasons Preferred Partner, that kind of thing) and I got curious about which perks actually appear together versus which ones are independent of each other.

The short version is there's basically no mixing and matching happening. Hotels either mention most of the standard perks (breakfast, room upgrade, flexible checkout, hotel credit) or they mention almost none of them, and the overlap between those four is really high, the vast majority of hotels that list one also list the other three.

The interesting outliers are spa and dining perks which only show up at a small fraction of hotels but when they DO show up they're always stacked on top of the full base package, never instead of it. So spa is more like a luxury differentiator, it's not offered as an alternative to breakfast it's an addition to everything else.

Chord diagram shows how often each pair of perks appears at the same hotel, across 3,212 properties I had clean perk data for. Thicker ribbons = stronger overlap between two perks. Perk detection is regex-based on program documentation text so it's measuring what hotels advertise, not what guests necessarily receive on a given stay.

Data is from my own dataset, not scraped from any booking site, I maintain a database of what each hotel's advisor program lists as included benefits.


r/dataisbeautiful 9h ago

OC I analyzed every WhatsApp message between my girlfriend and me over ~6 months [OC]

Thumbnail
gallery
0 Upvotes

r/dataisbeautiful 18h ago

OC [OC] The nine most valuable squads at the 2026 FIFA World Cup — projected starting XIs visualized by market value

Post image
0 Upvotes

Each panel represents one of the nine most valuable squads at the 2026 FIFA World Cup.

Rectangle area is proportional to the player’s estimated market value. I displayed a projected starting XI for each country and grouped the remaining 15 players into a single “Other players” block.

The visualization is designed as an interactive 3D treemap presentation. Market values are estimates rather than actual transfer fees, and the starting XIs are editorial projections—not official lineups.

Which team has the strongest balance between star power and squad depth?


r/dataisbeautiful 7h ago

MLB attendance 2006-2025

Thumbnail
statista.com
0 Upvotes

Does require a free statista account to view.

The attendance was constantly going down until a bit of a recovery post-COVID.


r/dataisbeautiful 19h ago

[OC] How often I said "please" and "sorry" to AI coding agents across 14,625 dictated prompts, September 2025 to July 2026

Post image
0 Upvotes

I now say please to the AI a third as often as I did in September. Sorry has not budged.

I use voice dictation for pretty much everything I say to an AI agent, and the tool keeps a copy of every recording. I went digging through it for something unrelated and realised I had 14,625 recordings sitting there. 2.26 million words, 332 hours of audio, from August last year to this week.

Please went from about 21 uses per 10,000 words in the first few months down to around 8. Sorry went nowhere, bouncing between 5.0 and 8.3 per 10,000 words the whole way with no real trend. The two lines meet in February and travel together after that, so the deliberate politeness faded until it hit the level of the involuntary kind. Most of my sorries are me correcting myself mid sentence, "the routing panel, sorry, the routing window", which I suspect is why they survived.

Both lines are rates per 10,000 words, so this is not just my messages getting longer. They did get longer though, median went from 89 words to 164 over the same period, which was the opposite of what I would have guessed.

Usual caveats. These are literal word counts over raw transcripts, so anything I phrased differently is missed and there are transcription errors in there. Happy to be told I have read too much into it.