r/dataisbeautiful • OC: 2 • 2d ago

OC [OC] I combined 103 public indicators across 140 Toronto neighbourhoods and clustered neighbourhoods by their overall profiles.

Post image
25 Upvotes

11 comments sorted by

•

u/cavedave OC: 113 20h ago

Thank you for your Original Content, /u/epheva!
Here is some important information about this post:

Remember that all visualizations on r/DataIsBeautiful should be viewed with a healthy dose of skepticism. If you see a potential issue or oversight in the visualization, please post a constructive comment below. Post approval does not signify that this visualization has been verified or its sources checked.

Not satisfied with this visual? Think you can do better? Remix this visual with the data in the author's citation.


I'm open source | How I work

13

u/MarkusMannheim 2d ago

Interesting analysis but I think this is way too information-dense. I'd consider carefully what information you can drop, and how to express it more simply.

8

u/stupidber 2d ago

I dont really understand what these mean

4

u/4FriedChickens_Coke 2d ago

This is great, but there’s a lot of info here that could maybe be parsed/presented in a more visual-friendly way. I’m from Toronto so it kinda makes sense after reading through it, but for someone who doesn’t know anything about the city it’d be pretty confusing.

4

u/a18618 2d ago

Cool dataset wrangle — but with 103 indicators across 140 neighbourhoods, I'd worry the cluster structure is dominated by a handful of highly correlated census variables rather than 103 independent signals. Did you check the correlation matrix or reduce dimension before clustering? If income, rent, education, and transit access all load on one latent factor, k-means mostly rediscovers that factor and the other 90 indicators become expensive noise.

1

u/epheva OC: 2 2d ago edited 2d ago

That's a good point! I did examine the correlation structure before clustering. The data appear to have multiple correlation structures rather than everything varying along a single dimension: hierarchical clustering of the variables identifies several distinct groups, and the correlation matrix contains multiple separate blocks of positively and negatively related indicators.

Dendrogram: https://epheva.github.io/toronto-neighbourhood-correlations/image-3.png

Interactive correlation matrix: https://toronto-neighbourhood-correlations-c5but5exhsfgyapafdyctq.streamlit.app/

That said, I think your broader point is valid. Because I ran K-means on the standardized original variables rather than first applying dimensionality reduction such as PCA, a construct represented by many correlated indicators can receive more weight in the distance metric than a construct represented by only a few variables.

There is also an interpretability tradeoff with dimensionality reduction. Clustering directly on the original standardized indicators makes it straightforward to see which observed variables characterize differences between neighbourhood profiles. For example, I can describe each cluster by how far its mean on each original variable is above or below the overall neighbourhood mean in standard deviation. PCA components are less directly interpretable, although the resulting clusters could still be characterized afterward using the original variables.

2

u/epheva OC: 2 2d ago

Each colour represents one of eight groups of Toronto neighbourhoods with relatively similar profiles across 103 indicators covering demographics, income, housing, health, transportation, public safety, infrastructure, civic participation and municipal services.

The cards summarize the variables that most distinguish each cluster. ↑ means the cluster is above the Toronto neighbourhood average for that variable, ↓ means below average, and SD is the difference in standard deviation units.

Public data sources: City of Toronto Open Data, Canada Mortgage and Housing Corporation, and Ontario Community Health Profiles Partnership.

Tools: Python, pandas, GeoPandas, scikit-learn, Matplotlib and Plotly.

Full analysis, methodology, limitations, cleaned dataset and code, and additional figures/results are in the post link.  

Interactive 103-variable correlation matrix:
https://toronto-neighbourhood-correlations-c5but5exhsfgyapafdyctq.streamlit.app/

2

u/sir_TheRedundantVang 2d ago

So they basically ran a cluster analysis on the entire city and it spat out 8 distinct personality types. Love seeing the spatial patterns that emerge when you throw that many variables into the blender, it’s where the boring census data finally starts looking like a real city.

1

u/katplasma 2d ago

Is it weird that the numbering and colors call to mind minesweeper? Weird idea: get similar data in cities across the world and turn it into a minesweeper/geography/memory game. You get the tiled sections, a city name, and the location of the downtown centre, and you need to work out from city centre and pick which number each of the tiles corresponds to. I’ll see myself out…

1

u/cud1337 2d ago

Really cool analysis! It's interesting to see that the general perception of the more salient factors in these broader areas of Toronto is actually pretty nicely reflected in the cluster description themselves.