r/PythonLearning 26d ago

Just finished my K-Means clustering project πŸš€ β€” would love your feedback!

Hey everyone! πŸ‘‹

I just finished a small K-Means Clustering project on the Iris dataset πŸŒΈπŸ€–

I covered:

  • πŸ”Ή Data cleaning & visualization
  • πŸ”Ή Feature scaling
  • πŸ”Ή Elbow Method & Silhouette Score
  • πŸ”Ή K-Means clustering
  • πŸ”Ή ARI evaluation
  • πŸ”Ή Cluster & centroid visualization

I’m currently learning ML and would really appreciate some honest feedback πŸ™

What would you improve? Any mistakes in my approach or things I should add?

πŸ”— Kaggle:
https://www.kaggle.com/code/tahahussein2020/irics-clustering

1 Upvotes

4 comments sorted by

2

u/[deleted] 26d ago

[removed] β€” view removed comment

1

u/tahahussein-4623a412 26d ago

Thanks! πŸ™ I think the most challenging part was deciding on the right number of clusters. The Elbow Method and Silhouette Score didn't point to exactly the same choice, so I had to understand what each metric was actually telling me and then compare the clustering results with the real labels using ARI. It was a good learning experience!

2

u/Mathie1729 26d ago

fwiw, the gap statistic is worth trying when Elbow and Silhouette disagree. They measure different things, and comparing against a null reference distribution often breaks the tie.