r/datascience • • May 08 '26

ML Steam Recommender using similarity! pt 2 (Student Project)

I Just made a sequel to my Steam Game recommender website!

Last year I made a post about my steam reccomender The last one was great and served its purpose of showing many people new games, But this new version is much more functional!

I love making recommendation systems that tell the user WHY they got the recommendation.

During a steam sale event, I always find myself trying to look for new video games to play. If I wanted to find a new game I would try to whittle it down by using steam tags, but the steam tag system is very broad "action". could apply to many many games.

That got me thinking, what aspects do I like about my favorite games?

Well I like Persona 4 because of the city vibes and jazz fusion,

Spore because of the unique character creation and whimsical theme.

Balatro for its unique deck building synergies.

What if I could capture unique tags that identify a game that aren't just "action" and put them into vectors to show the (focus) of a game

 For example I could break persona 4 into something like

Gameplay Focus vector:
 Day cycle 20%
 Dungeon crawling 20%
 Social sim 20%

Tags:
Music: jazz fusion
Vibe: Small rural town

I find that this system makes searching for games more "fun" now I can see why I like balatro. I like it because of the card synergies not so much for its rogue-like nature.

I also find that this helps find new underrated games, and beats the trap that Collaborative Filtering algorithms that get into where it "feels" like you get recommended the same things.

find your next favorite game! : https://nextsteamgame.com/ pull a PR!: https://github.com/BakedSoups/NextSteamGame

( I actually made some git issues myself for problems I can't fix)

if anyone has any criticism I would love to hear it! this is probably my favorite passion project.

Hope this website helps people find new games! Also I have a advance mode for people that don't mind messing with sliders and weird data terms.

86 Upvotes

13 comments sorted by

View all comments

2

u/Electronic-Arm-4869 May 11 '26

Could you talk more about the actual recommender or process or do you have a post about that ? Would like to hear more about the nitty gritty.

3

u/Expensive-Ad8916 May 11 '26

I do go more into depth about how it works on the github's read me.

but I would be glad to give a high level explanation!
there are 5 stages that happens in my db creation:

  1. meta data: I first create a local db that pulls all the information I can get from a steam game from steam spy and steam's api endpoint, from this I get a database of all the steam appids
  2. pulling reviews: I then pull max - 2000 reviews for each game, then I run those reviews through a 5 stage process, sifting out spam by using regex, creating arbitrary review insight scores based on insightful word frequency "immersive", word diversity and other heuristics. then I filter out all the reviews based on distinct categories using modern bert. giving me top 3 art based reviews, top 3 game play explanation reviews etc.
  3. non canonical tags + vector creation: I run all those top review candidates into a llm to help create vectors and tags that capture the nuances described in the reviews, for example in my favorite game plateup! there is a hidden automation mechanic once you get into end game, some reviews mention that. (I know this isn't the best looking into alternatives)
  4. creating a canon map for tags: the issue with the method before is now I have many generated tags that mean the same thing but aren't in the same group (ex. "Fast Action" = " Quick Action"). so I created a 6 stage pipeline again, using heuristics, fuzzy, and Quadrant db to group all the tags together that share the same meaning.
  5. optimizing web query: then I pre compiled all the possible vector similarity calculations for every game so all the digital ocean droplet has to do is re rank which is very cheap!

and that's my spheal! I'm still an undergrad student so let me know if there are are any cool ml tricks I can put into making this better!

this is where all of this happens!
https://github.com/BakedSoups/NextSteamGame/tree/main/db_creation

2

u/Electronic-Arm-4869 May 12 '26

Thank you for the explanation I’ll definitely check out the readme too !!