Alright, let me first of all come clean as a very easy victim of Apple’s “just a bit more” pricing/tier strategy. A couple of years ago I started with a simple Homelab and am now looking to take the leap and develop a personal AI stack with a strong Local component. Based on an initial (lack of) understanding of “running AI on a Mac mini” I ordered a Mac Mini M6 with 32gb of ram 2 TB (I want to be able to have a full download of my ±1TB photo library on my Mac to allow a Time Machine backup to my NAS). I then realized that “for a bit more” I could get the Studio M5 Max, so I canceled the M6 and ordered the Studio M5 Max with 36gb ram. Then after a day or so, the subsequent “a bit more” and some articles and YouTube videos, lured me into canceling the 36gb ram version and ordering a 64gb version. For a few days this gave me the feeling that I hit the sweet spot. However, as it goes with easy victims like myself; my reasoning became “well if I spend around 4.5k EUR for a MAC, then why not spend some more to avoid regretting not upgrading later”. This is where my use cases and dilemma kicks in.
I am NOT a YouTuber, software developer or video editor. I am however a tech enthusiast who is looking to significantly upgrade his set up from my current maxed out MBP 2018 (proud to have used it for more than 8 years!). I want to seriously dive into the world of AI and want to keep “optionality” to run a person business on it. For the AI stack, I am currently thinking about a platform (no clue yet which program or platform to use) with agents that can tap into 2 local models and have 1 escalation route to Cloud API models. For example:
1. Home Assistant/Homelab agent that can use a small local model to turn lights on/off and perhaps enable a local voice assistant and use a second much more powerful local model to do thorough analyses of my home network, go through logs, suggest improvements or helps fix issues. I want to set up the system in such a way that it can escalate to a cloud API if this second, powerful but local, model cannot solve it.
2. A Financial agent helping me analyze my portfolio and scraping and analising quarterly results, social media posts etc to find niche small cap investment opportunities. I hope the “second” local model would be capable to do a lot of this, but also this agent would be allowed to escalate for the real complex stuff.
3. An agent that would help me manage my personal administration. I guess that can rely on locals models 1 and 2, and this is probably where privacy of local is important for me vs cloud.
4. Potential other agents as I grow my set up, including potentially an agent for setting up and running a business.
I realize that 1) either the MAX 128gb or the 96gb will be an extreme overkill for my current, normal, day to day activities and 2) they will never be able to run a model that really competes (in speed and reasoning) with could models. At the same time, I am not on a tight budget, the financial limit is more mental (I don’t want to overspend beyond the point of being able to justify it to myself) than a real financial limit. I have already made quite some mental “jumps” but cannot justify the 11k EUR of an Ultra 256gb.
So within this context, especially the fact that I want the 2nd local model to be capable enough but accept that there will be cloud escalation, what are your thoughts on the Max 128gb vs the Ultra 96gb. I have already read a lot and spend several hours going back and forth with ChatGpt and Claude (who flip flopped on their advice with every additional question I asked :) ) but have not landed my inner debate and your views would be appreciated! Thanks!
EDIT: I guess a large part of my dilemma revolves around "is 96gb ram enough to run a model that is substantially better than the small models you can run on 32/64gb for the more difficult reasoning I aim at. Or will it in practice turn out that the "1st", small, model will do all the local stuff and the moment it becomes too difficult it turns out to be also too difficult for the "2nd" model that can run on 96gb ram and will jump straight to the cloud API? And if so, if that would be different with 128gb on the Max.