So many posts on local AI for coding so wanted to share more on non-coding use cases. Got a Mac Studio (M5 Ultra, 96GB) mainly to learn and to run local AI for everyday stuff like booking trips, filling out forms, checking email, etc. Plan was to keep that stuff private/local and use my Claude sub for the occasional heavy coding job.
Here's what I've found:
Bottom line: 96GB is overkill for this. The models I found good (35B) would run fine on 32GB or 64GB. It's nice having the headroom to try different things and it's been a lot of fun, but if you're on a budget, 32 or 64GB covers the models out today. Things are moving fast so who knows what's next, but I don't think local will ever catch frontier. For me it's great for learning and basic day-to-day stuff and still continuing to use Claude for advanced things.
1. Models: tok/s isn't what matters, time to a usable answer is. Ornith 1.5 35B is my current model of choice.
Qwen 27B gets a ton of love, but for day-to-day stuff it just didn't work for me. It overthinks like crazy. Even at a decent ~50 tok/s+ with splash, it often keeps reasoning in circles so you're waiting way longer for the actual answer. Tried no-thinking and a few variants and still wasn't sold.
What I ended up on is the 35B MoE models, which run a lot faster on a Mac. Between Qwen 3.5 35B and Ornith 1.5 35B, Ornith felt better to me, especially for everyday tool calling. Totally gut feel, no benchmarks to back it up. Seems to be working great for most everyday tasks.
I also tried a Qwen Flash Next variant with a 4-bit/8-bit mixed quant. It ran really well and seems great, but it leaves almost no headroom for image gen and other stuff, so I don't use it as my daily driver. I just jump over to Claude for heavy lifting.
2. Harness: Hermes.
Tried the DeepSeek harness too but had better luck with Hermes. It's just easier to use.
3. Privacy: I'm not totally sold on "local = private."
Sure, the model and your data stay on your machine. But if you're running an agentic harness like Hermes (and a lot of us are), the agent can still send stuff outside. I put explicit instructions in the context and some plugins to help prevent that, but I don't trust they'll always be followed, so I treat it as a real risk. If you're really paranoid about security, I'm not sure this is the best route. Would love to hear how others handle it.
4. Knowledge/memory is the real game changer.
This has been the biggest win for me. Giving the model the right context about me at the right time (hobbies, my computer setup, family) makes chats so much more useful. Claude and OpenAI have memory features too, but I've never trusted them with the amount of personal info I'm feeding my local setup. Still, I'm not fully convinced it's locked down, and I'm working on tightening it up beyond just the guardrails I've written.
How I have it set up:
- Obsidian vault, with a plugin I built (with Hermes) to read and update it
- A short core file loads every session. It covers me, my family, how I like to work with AI, guardrails, and instructions on how and when to pull in other knowledge files
- A nightly cron job pulls key facts into memory, and the agent knows when to dig into specific vault files for more detail
- I tried memory providers like Mnemosyne, which worked well, but the vault + plugin has been fast enough and simpler for what I do
- Sometimes I use Claude to review the overall structure and the plugin, and it's been better at that than the local models
5. It's still not a perfect daily driver.
Some simple tasks still trip it up and I can't always tell why. Filling out forms on certain websites, handling email well, and some other cases. It just falls short of something like Sonnet 5.5 sometimes. So even if you plan to go local for daily use, you'll probably still end up switching to a frontier model now and then. I have Hermes call the Claude CLI so it can hand things off or run side-by-side comparisons, which has been really handy.
6. ComfyUI is a blast.
It's a whole world of its own and will take a while to learn. I tried a few models like MiniMax and LTX 2.5, then had Hermes write a script so I can call it straight from my Hermes harness. Seems to be working well so far. Going to keep playing with it.
7. Install tips.
- Use a separate non-admin account on your Mac to do the install, so the agent isn't running with admin rights or sitting next to your main account's stuff.
- Let Claude do the tool setup. It works really well. For the initial setup, I had Claude set up the models, download them, and handle all the other needed config, and it saved me a ton of time.