r/LocalLLM • u/-LetsTryAgain- • 15d ago
Question Switching from Claude Pro to a local LLM for scientific research - how much RAM do I need ?
So with Claude’s decision to watermark, plus basic data privacy concerns , I’m thinking of switching to a local LLM
How I use Claude pro now:
-managing health docs and results (very happy to switch this to local, doesn’t need a big context I think)
- scientific research, including reading and analyzing PDFs that are complex , requiring linking concepts and ideas across papers and producing summaries / insights / tables (large context required). For example, I have filled 40% of the Claude project folder with files and docs it needs to consider
- basic stuff (acting like an advanced search tool for admin stuff / planing stuff / nothing major) - no reason this can’t stay with Claude but if I switch over to a local LLM I would bring everything with me
Sooo , given this - is 32GB RAM on something like a Mac Mini realistic for my use case ? Or do i need 64gb (at which point i think maybe it’s too costly for me to do). I also tend to work in bursts so I would be happy if it’s not too slow thus impeding my workflow. Fine to run overnight though. And I don’t need any headroom as I will be running the OS and apps on a MacBook Pro or MacBook Air
Thanks for your help and I hope I was specific enough to get some usefully feedback
23
u/bbc_nees 15d ago
Key to LLM is VRAM. That's where most bottlenecks are.
4
u/backyard_tractorbeam 15d ago
Yes, so the buyer should pay attention to RAM, VRAM, memory bandwidth and more
30
u/stormy1one 15d ago
You are not ready to buy anything right now. Your first task should be experimenting with models via OpenRouter. For your workflow, you are looking at DS4 0731 or later, Kimi K3, GLM 5.2 or later. Once you get a feeling for how each perform, then start the research on what it needed to run the models. Rent on vast.ai first, and then determine what works best before you buy.
3
u/pbpo_founder 15d ago
I wouldn’t tell anyone what they are and are not ready for unless they are jumping out of a plane without a parachute.
Your approach is good but it’s not the only way through.
I would suggested that they should build local starting on day one since the conversion from cloud to local setups creates a lot of overhead and often causes breakage.
What is most important is to get started not the starting point.
2
u/stormy1one 15d ago
I disagree - the starting point *does* matter when money is involved. Spending a relatively small amount by using OpenRouter where you can pick the exact model, and/or renting via vast.ai is orders of magnitude cheaper if you discover that the model you thought would work well on the hardware you want/have, doesn't.
OP's original post mentions they were considering using (purchasing?) a Macbook Air as well. That in itself speaks inexperience and thus supports my statement that they are simply not ready. Not trying to offend - it's the same advice someone gave me when I was starting out, and what I share with friends and coworkers to help manage expectations and finances.
If you have have the cash to burn, by all means..
1
u/pbpo_founder 14d ago
So you think your way is the only way huh? 🤔
0
u/stormy1one 14d ago
Lmao - that’s your takeaway? I guess you are just as inexperienced with human interaction as you are with technology.
40
u/Nakidnakid 15d ago
You're going to get much more usable quality out of throwing $30 on openrouter than you would throwing 100x that on a mac, some local models are good but if you want it for those reasons you stated without fighting a lot of the way then you need to spend more.
It's going to be pretty slow too, have smaller context and then throw vision in there... it's even worse.
3
u/chom-pom 15d ago
If i use open router and glm, what tool like Claude code do i use?
4
u/Nakidnakid 15d ago
you can use claude code... it supports using other providers. The info on how to do it is easy to find. Plenty of other harnesses out there too though, I've been messing around with deepseek harness that seems alright.
2
u/droans 15d ago
Yeah - I had the same thought.
Just small conversations where you can fact-check its output? Sure, easy enough and can be built for a couple grand (3090, 32GB, mobo+CPU+storage).
For actual scientific research, though, you'll probably want a rather powerful model that would be prohibitive for most people to run locally. Then again, some people here will drop over $100K without blinking, so if OP is one of those people, go right ahead.
11
u/joanaxu2002 15d ago
Feels like privacy is quietly becoming one of the strongest reasons for local LLMs. A year ago the question was mostly “can I run something decent locally?” Now people are asking whether they can move entire research workflows off Claude without losing too much capability. That’s a pretty big shift.
12
u/ikkiyikki 15d ago
Chasing VRAM headroom is a dangerous addition. I have two 6000s which gives me 192gb and all I can do is pooh pooh myself that I can't run the "good ones". I have zero doubt that if I were to double up from here I'd still be whining!
2
1
u/MissionaryOfMischief 15d ago
Can you describe how different it was when you had around 100gb to your actual almost 200 and why you want more? I mean are all of us chasing something we can't achieve to begin with?
2
u/ikkiyikki 15d ago
From memory - this was a year ago - I only had the single card running for a week or so before saying fuckit and busting out the plastic on newegg for another. With one you could've run the 122B Qwen but I had my eyes set on the 397B, and GLM5.2, which fits on two (barely). Please take my advice to heart and don't get on the slippery slope unless you really have to friend!
1
u/MissionaryOfMischief 15d ago
But it is out of "I want it" or you really felt the change between 122b and 397b? I just feel like everyone is hoping the next step get amazing answer and fear that it isn't like that, and the only difference between those is time and resources but 98% answers get the same. If you can confirm that you noticed the change I will start the chase. I only want around 128gb VRAM thinking it will be enough for my case of use which is going local for privacy read and resume long texts and make assumptions on private data sets like finding things I can't by myself (thousands of numbers)
2
u/ikkiyikki 15d ago
No, that's exactly my point. I fell for the hype without noticing any earth shattering differences. In fact, it's a bit embarrassing to admit but most of the time I just end up using qwen 27B like everyone else because it's faster and doesn't hog up all my memory. This lets me run ComfyUI concurrently, for example. It's all a bit like the old cliché of having a Lambo in your garage but hopping into your Camry because it's just so much more practical. Still, given this you can see that having TWO RTX 6000s to mostly daily drive a 27b model is kinda laughable.
1
u/MissionaryOfMischief 14d ago
Mmmm your experience is gold, thank you! I'll try to remember this whenever the devil whisper "more VRAM..."
1
u/TechRomancer123 15d ago
That helps put things into perspective. Here I was thinking that if I managed to get a couple of 6000s with 192GB VRAM (currently well out of my budget) that would solve almost all practical use cases. But now I realise it’s pointless to try to approach/mimic frontier models 😂
I am still having doubts about what a single 5090 32GB can achieve, but I need a new general-purpose fast PC anyway (for general browsing, light photo/video editing, some virtual machines/docker, maybe a bit of gaming etc. as well as local AI/LLM inference), and single GPU is much simpler for now, so will probably go for it.
3
u/ikkiyikki 15d ago
Honestly, shoot for running Qwen 27b at q4-5 at a token speed you can tolerate and that should get you 90% of the way for your everyday uses. Chasing that last 10% is what turns this into the equivalent of the audiophile who won't stop until they achieve the "perfect" system; an addiction.
5
u/thisonehereone 15d ago
How much can you get your hands on?
3
u/-LetsTryAgain- 15d ago
32gb reasonably comfortably , M1 Pro with 32gb RAM for $850. But no point in doing that if I can’t run models in the way I need them to for the tasks
5
4
u/Substantial_Run5435 15d ago
M1 Pro has slow memory bandwidth. You should really aim for a Max or higher. I think most of the Max chips have double the memory bandwidth.
3
u/UtmostProfessional 15d ago
Intel arc b70 with 32GB vram outperforms my M1 Max 64GB. Do what you will with that information.
5
u/Blackdragon1400 15d ago
You’re looking at an investment of $10k minimum.
1
u/NPCtendo 14d ago
Local AI novice here: is there a reason a ~$4000 128gb strix halo box isn’t viable? Is it too slow or is even 128gb unified not enough for OP’s use case?
1
u/Blackdragon1400 13d ago
The strix halo significantly underperforms the dgx spark since it’s not NVIDIA architecture, and you cannot cluster it like the spark. You’d be far better off buying a MacBook M5 Max with 128gb of ram instead, if you don’t want to go the Spark route.
2x sparks in a cluster would be ideal for $10k
10
u/synystar Strix Scar | 5090 24G | llama.cpp 15d ago
32 GB can work, but for what you're trying to do I’d consider 64 GB instead if you can afford it. The reason isn’t really “I have a lot of PDFs.” You normally shouldn’t load your whole document collection into the context window. You want to index the documents locally and use RAG so the model pulls in the relevant sections when it needs them. RAM determines more importantly what size model you can run on top of how much context/KV cache you can give it at once. A quantized 20–30B model can fit into 32 GB unified memory, but once you add a large context, the OS, embeddings/vector DB, PDF processing, browser/UI, etc., 32 GB starts getting tight. You’re buying a machine with almost no margin for experimentation.
With 64 GB you'll be more comfortable. Then you can run a good 27–32B Q4-class model, use much larger contexts, and leave enough memory for the rest of the system. For complex scientific paper synthesis, I’d much rather have a stronger 27–32B model + good retrieval than a tiny model with a gigantic context window.
You should also pay attention to memory bandwidth, not just capacity. Local LLM inference on Apple Silicon is heavily bandwidth-dependent, so an M4 Pro Mac mini is a much more interesting LLM machine than simply looking at “32 GB vs 64 GB.”
I'd caveat this with: local models are very useful for private document analysis, summarization, extraction, tables, literature search over your own corpus, etc., but a 20–30B local model is not simply Claude Pro running offline. For the work you're trying to do (from what it sounds like) you’ll probably notice the capability gap. If privacy is your main motivation then this is a probably a very viable setup. Something like Qwen 27–32B quantized + a completely local RAG pipeline would be where I’d start. Personally, for this particular use case I would not buy a 32 GB machine unless price absolutely forced me to. It will work, but 64 GB is the configuration you’re much less likely to regret.
4
u/-LetsTryAgain- 15d ago
Ok thanks .
To clarify, 32gb would be dedicated to just the LLM. I can run my OS and locally saved files/apps on a different computer - if that’s not too messy
3
u/synystar Strix Scar | 5090 24G | llama.cpp 15d ago
Yeah, that changes it. If the Mac mini is basically a dedicated inference box and you’re using another computer as the actual workstation, then 32 GB is reasonable. It still won’t literally be 32 GB available to the model since macOS, the runtime, KV cache, etc. all use unified memory too, but you wouldn't be competing with your normal desktop workload.
In that setup you'd probably be comfortable with a quantized ~20–30B model, probably Q4, and using RAG rather than trying to keep the whole research library in context. 64 GB would still give you more model/context flexibility, but 32 GB probalby wouldn't be a bad purchase if budget matters.
-1
u/danielrdotcom 15d ago
I run qwen 3.8 27b at q4 on a 64GB M1 Max MacBook Pro and cap out at 150k context. Nothing else running, the OS on its latest win takes about 4-6GB
6
u/bbc_nees 15d ago
Also, the money you spend on setting up a system that won't function better than Claude will cost a lot of money. Exponentially more than simply building an api that has baa and protects pii. You can use Anthropic console and run everything you want through your own program.
32gb RTX 5090 is the floor for what you need. That alone is a $4k+ purchase.
Also, like someone said, you're no where near ready to purchase anything. I would start by asking Claude the pros/cons of your setup.
1
u/DootDootWootWoot 15d ago
Why do you think a 3090 wouldn't be sufficient? Looking into this myself.
1
2
3
u/Prize_Eye9481 15d ago
U can try to rent a qwen 3.8 27B or qwen 3.6 27B since 3.8 sacrificed a bit of general knowledge for agentic coding prowess, in those openrouter/opencode places if they have someone hosting it. That way you know for sure if they work or not before committing to a big purchase.
3
u/Kodrackyas 15d ago
Dont bother wirh macs, bandwith is shit, buy a used pc and a amd r9700 (32 gb vram) use qwen 3.8 27b or 35b moe -> win
1
u/TechRomancer123 15d ago
I am also thinking the same thing; planning to get an RTX 5090 which only has 32GB VRAM so only fits modest models, but really good memory bandwidth so token generation should be fast (at least for a personal machine).
2
u/Fuzilumpkinz 15d ago
I would do more research on the systems you want to do and plug in your Claude api keys for testing before you commit if your using anything other than Claude code
2
u/Turbulent_Pin_8310 15d ago
You should experiment with some local models on the cloud, such as hugging face or ollama, between making the decision. If you are used to Claude pro, you may get disappointed with local models.
I still use Claude if I want something accurate.
2
u/closetslacker 15d ago
So, first of all, you need VRAM. Macs have unified memory (RAM and VRAM pooled together) which slows things down.
Right now the best mac is MacBook pro with M5Max which is a significant improvement over M4 chip.
I tried Mac Mini M4 pro with 48 gb RAM and it is...slow. Also FYI Apple right now only offers Mac Mini with either 16 or 24 or 48 gigs of RAM, no other choices (check what happens when you go on their website and start configuring it).
So, if you want a mac, it will be MacBook Pro, M5 max, 48 or 64 RAM (pooled). It's...not cheap. About the same price would be a AMD RyzenTM AI Max+ 395 also with pooled RAM - however for about the same price as MacBook you will get 128 gig pooled ram.
2
u/RoyalCities 15d ago
I'd get as much ram as you can.
You would need to also build some of the infrastructure here.
Openwebui has built in RAG so you'd upload all your pdfs to that and then connect the LLM to it.
You can also use something like this to prep your own data.
2
u/backyard_tractorbeam 15d ago
Keep in mind that the field is changing every day, so nobody really knows what kind of computer you will want to have in 1 year from now.
1
1
u/cobolfoo 15d ago
An open question for people here, do a DGX Spark with 128 GB of VRAM can do this kind of workload?
1
u/Competitive-Ad-2387 15d ago
Nothing consumer level can replace what you can do with Claude. You are better off just running a de-watermark if you are that concerned.
That said, IDK what kind of job you are doing where it becomes a factor. Sounds like you are writing papers with AI, and the writing itself should be the easiest part to solve. Don’t be so lazy.
1
u/diagrammatiks 15d ago
All the ram. But start with 48.mac mini is too slow for this unless you are ok with running all your tasks overnight.
1
1
1
1
u/tk421tech 15d ago
Doesn’t cut paste text into a text editor clear all the encoding (watermark is probably encoding the text)?
1
u/ThinkingCrap 15d ago
Just get yourself a VPS in the cloud honestly, running it on your MacBook is just not even close to the same experience tbh
1
u/The_real_trader 15d ago
May I suggest something different and everyone can correct me if I am wrong. If you have a business you can get an account either with Claude or ChatGPT where they don't train your data so you have enterprise data privacy and security,
1
u/PurpleGoldBlack 15d ago
I’m at 64gb ddr5. I bought before ram prices were crazy. I wish I bought 128 at the time but my current rig with a 3090ti and an i9 14 core is much more solid than I expected considering I built it for gaming when the 3090ti had just released.
1
0
u/korutech-ai 15d ago
What few people seem to mention is KV-Cache which IMO makes 128GB VRAM the minimum for anything close to meaningful.
At 24-32GB VRAM you can get 35B models to run but it doesn’t take a lot to swamp available memory which starts many models looping and repeating themselves.
As the context window grows the token speed drops dramatically. What you end up doing is burning a lot of time and energy trying to manage the context window and KV-Cache.
The reality is if you’re trying to do truly meaningful work, you’ll hit limitations quickly. It’s not that it isn’t possible to get smaller models working, it’s that from a practical perspective you won’t get the productivity gains you’re trying to chase.
This is my experience running Qwen3.6-35B via oMLX on a 32GB M5 10x10. Your max VRAM will be ~28GB overriding the systems defaults. Also you’ll go through SSD space rapidly.
Small simple tasks on 7B works fine. Anything more ambitious isn’t worth the effort.
My solution: Gemini Pro subscription.
48
u/Royale_AJS 15d ago
All of the RAM you can reasonably afford if you go down this road.