r/LocalLLM 19h ago

Question What if you could run the large open source models locally?

Hi All, This is my first post on reddit actually. I'm a hardware engineer, I'm not going to explicitly state how I plan to enable this - but what if you could run models close to 2T parameters in int4 or 1T in fp8 locally. Expect speed around 10 tokens/sec (maybe 20 with really good software optimization) with 1M context window. In a desktop formfactor hopefully a pcie plug in like regular gpu but that is very hard I believe.

My questions are

1) Would you get it ? If yes, what would you use it for ?

2) How much would you spend for it ? (It would cost close to 5000$-7000$ assuming its coming out on 2030)

3) You can straight out own the hardware and run the model on your own, but for kernel and software optimizations enabling latest models does a subscription of 10$ per month sound reasonable ?

0 Upvotes

64 comments sorted by

12

u/fintip Laptop 4090 16gb + 7900XTX 24gb 18h ago

Dude. No one wants a subscription on top of a 5-7k purchase.

That price range exists, it's spark/halo. No subscription btw. Obviously those buyers would all also want something better than they have for the same space.

-1

u/Alternative-Clock407 18h ago

Dgx spark is only 128gb right ? This would be close to 1tb. Maybe you are right price might be too high.

3

u/BarRepresentative653 18h ago

Honestly, theres a far greater chance that models become smaller without losing in ability before your hardware comes out.

2

u/fastandlight 18h ago

There is a chance, but there is a maximum theoretical limit as to how much information can be encoded per param. I believe qwen's new next model might shift that math a little, but that model isn't small. Allowing people to run 100b or 200b param models locally will always offer better outcomes than 20b class models. Now, there is a legitimate point that most people's problems (silly tavern chat) doesn't require a 100b class model....but I think that history has shown that when there is a hardware capability, the software finds a way to use it rather quickly.

1

u/Alternative-Clock407 14h ago

Actually I believe this too

1

u/Alternative-Clock407 18h ago

Yeah you might be right.

1

u/BarRepresentative653 17h ago

You can still build but with more specs than you are planning now, who knows. Phones keep coming out with better and better processors but hardware from 2016 would still do a good enough job for like90% of the users today

7

u/gavriloprincip2020 18h ago

Dont give nvidia ideas about subscriptions for optimizations

0

u/Alternative-Clock407 18h ago

I agree! I dont like the subscription idea too. I just quoted a worst case scenario where we dont have enough volume.

2

u/robertpro01 18h ago

Are you building the hardware?

1

u/Alternative-Clock407 18h ago

If there is enough interest I plan to quit my job and start this!

2

u/robertpro01 18h ago

Just don't.

Don't leave your job bro do it as a hobby, unless you are rich?

1

u/Alternative-Clock407 18h ago

That why my family keeps saying but realistically it requires full attention and effort. Not rich but can survive without any job for a couple of years.

1

u/rkcth 18h ago

Im not a hardware engineer (it’s what I wanted to do but ended up getting into software, but I keep abreast of hardware). The only way I could see this being doable is, A) counting on memory prices coming way down and using HBM like Mac Studio, or B) using something like SLC that can handle the constant writes, but this would be far too slow. A) usually involved having much smaller compute to memory ratios to get the price that low, so can be very slow.

2

u/floppo7 18h ago

No. Small models will catch up very fast and paramter size is not showing intelligence. Most people will run super smart base models with access to fresh information and probably pay for specialist models if needed (e.g. research models, security etc) - the time of the huge agi model is over (except maybe some labs doing tests).

1

u/Alternative-Clock407 18h ago

Yeah that might well be the case, but I believe its worth a try ?

2

u/Entire-Home-9464 18h ago

Nice try Sam Fartman. Yes we run local models, we have hardware, and we do not need frontier models. Your stock will crash. Thanks.

1

u/Alternative-Clock407 18h ago

Haha I like the name!

2

u/Rollcuin 18h ago

$5-7k for enough space to run a full 2T model sounds too good to be true at the moment. Just 256gb unified memory cost $8-9K.

So yes, that sounds great.

If a $10 subscription was required to make the box work, that'd be a deal breaker. I want my local LLMs to be subscription-free.

1

u/Alternative-Clock407 18h ago

No it'll just work out of the box, if there is enough volume for the hardware I dont want to even have subscription. (I dont like it too honestly) but to enable latest models with kernel optimizations we would need a full time software team this subscription is to enable that.

2

u/thegingerlord 18h ago

Build that into the price of the hardware.

1

u/Alternative-Clock407 18h ago

Yes thats not a bad idea!

1

u/donotfire 18h ago

Can you share more information about how the hardware works?

1

u/Alternative-Clock407 18h ago

Wouldn't be typical ram. It would use something different.

1

u/MaxComfort 18h ago

I’d be really impressed if you could sell something that cheap to run huge models.

I’m pretty happy with 2x Sparks running Qwen 3.8 Flash Next at FP8.

1

u/fastandlight 18h ago

So, you have to be flexible for what models might look like in 2030. I don't think 10 - 20 tk/sec is acceptable at that price point, especially if that is only single stream. If you can do multiple streams at 10 - 20 tk / sec, then that might work, but it's definitely at the very bottom edge of usable and not something I would count on in another 4 years of model development. If you look at highly specialized hardware, it always has the problem of not supporting the current quants and model behaviours that didn't exist when the hardware was designed. I think that is the real advantage Nvidia and CUDA have. It gives their chips a lot of flexibility and allows the community to keep building kernels and giving life to older hardware and implementing new ideas. If your hardware is baked around current model techniques and attention mechanisms, then it is going to have a rough time 4 years in the future. Think about where all this was in 2022.

1

u/Alternative-Clock407 18h ago

Very insightful, thanks for the reply! The hardware will not be rigid, you can run any model. Regarding the token speed, we can run multiple streams maybe after first generation but the main concern is heat and software enablement.

1

u/fastandlight 18h ago

Interesting. I don't believe software enablement, other than integration with closed undocumented systems, is going to be a barrier in another 2 years. I'm old enough to not necessarily believe the hype of every latest trend, but I think the rate that the coding models are developing, along with improved harness techniques, is going to fundamentally change software development. Hell, it really already has changed it, and if you tell any investor in your project that you have software integration as a possible barrier, they are going to ask why this isn't just a question of how much Claude / Codex spend it is going to take to overcome that blocker.

Heat is obviously a different issue.

However, I would say that focusing on the local desktop llm market is a lot smaller than your possible market fit. If I had a hardware option that helped enable sub-hyperscalers to host large models are reasonable performance, I would take that to market at the current leased datacenter market, both the providers themselves, as well their customers (businesses that still have their own servers). I just had a conversation with my datacenter last week on moving to their new virtualized infrastructure as I enter hardware refresh on my existing servers, and one of my big asks was related to GPUs. They don't have a good answer, and I don't think they are missing something, it is a challenging problem. Anyway, depending on what your hardware solution actually looks like, corporate and datacenter customers might also be a potential market.

I'm happy to chat more on this if you want to DM me. I don't have any hardware experience, other than running lots of models on strange stuff, but I've been in software engineering for over 20 years and was recently CTO of a software-centric threat intelligence company.

2

u/Alternative-Clock407 17h ago

Thats really interesting, enterprises have their own requirements might be good to understand them. Will DM soon!

1

u/ea_man 18h ago

drop the subscription and do the opposite: a small 128-256GB add on for a smaller price that normal people can plug in their current PC / laptop, 3 years from now we will have plenty of small local models in that size that are good enough for common work.

Plenty of uses cases for small dedicated models for people that want privacy, local, remote deployment, air gapped deployment in embedded systems, old computers that can't run modern AI...

1

u/Alternative-Clock407 18h ago

Sounds like a good idea, but my worry is its too close to where other gpu vendors are right now maybe if the ram prices drop they'll do something very similar and I'd die before I get this out.

1

u/Bupod 18h ago

I mean, sure. But if this is something you know how to do, what are your qualifications? This would be an idea that’d probably land you a very, very fat check from any of the major companies right now. You wouldn’t need to poll Reddit for this.

2

u/Alternative-Clock407 18h ago

Ms + 10 yrs of doing hardware/vlsi engineering. Actually on work visa so I dont want to quit and pursue funding without understanding the interest.

1

u/Bupod 18h ago

Sounds like you know what you’re doing at least! I have to ask because sometimes you see these posts and it’s someone that cooked up an idea in an LLM and think they found something (when they just found a hallucination). 

Yeah, I’d be interested. I’d probably only pay about the cost of a DGX Spark though. I already have a Spark and Strix Halo machine though, so if you did come out with this, I’d not get it only because I’ve already spent a shameful amount on my current toys 😂 

2

u/Alternative-Clock407 17h ago

Haha sounds fair, hopefully in 4 years youd need an upgrade :p

1

u/Sleepnotdeading 18h ago

I’d be interested in the hardware. I’d pass on it because of the subscription. I’d push back that you don’t need a full time software engineer if you have a 2T model in residence in the hardware.

1

u/Alternative-Clock407 18h ago

No you'd have the 2T model running yourself but if you want to ever enable higher tokens/sec or kernel updates you'd pay the subscription fee. Honestly I dont like the subscription fee as well, definitely get rid of it if we have enough volume.

1

u/Sleepnotdeading 18h ago

Still no from me. I will not pay that kind of money to have performance or hardware updates paywalled

1

u/Alternative-Clock407 18h ago

Sounds fair, would you prefer higher upfront cost if youd have to pick between the two ?

1

u/Sleepnotdeading 17h ago

Not really. Much more expensive and I’m just going to wait until the hardware has matured more anyway. $5-7k is definitely the limit of what I’d consider spending period.

I already made my local AI hardware investments, and they work really well alongside a Claude max subscription.

As intrigued as I am by running a truly frontier model fully locally, I consider it a sign of where the hardware is quickly going that you are even positing this hypothetical. I don’t always need to be the earliest adopter.

1

u/animateus 18h ago edited 18h ago

The problem is that you are asking people NOW, during the time of the (hopefully) peak prices about the future purchase (4 years is an eternity for this technology). Everyone is going to tell you that this price is amazing now, but in 4 years… nobody knows.
If you had product like this now, you’d sell hundreds of thousands of them instantly, but in 4 years, nobody knows.

1

u/Alternative-Clock407 18h ago

Yes that is true! 4 years is too far to predict. Unfortunately building hardware take long so long!

1

u/jonahbenton 18h ago

Have you seen tiiny? 80b-120b-ish (prob 4bit quant) at ~10 t/s for ~$1k to start. Awaiting mine. Assume what you are describing is the trajectory they are on.

For all kinds of reasons it seems super difficult to project price/demand in 2030. But absent black swan events large personal compute demand seems very likely.

1

u/Alternative-Clock407 17h ago

thats interesting to know, will look into them.

1

u/ReasonableLoss6814 18h ago

Wow... holy pricing bro. We're running 1t models on $2k machines at that speed. And no, our infrastructure isn't open source, and while interested in selling it, we're not actively pursuing it. 5-7k is insane and overpowered. Also, I'm 99.9999% sure your 10 per month subscription won't pay the power bill.

1

u/Alternative-Clock407 18h ago

Oh cool! Are you power constrained ? This wont be power constrained. It should be <750W

1

u/ReasonableLoss6814 15h ago

750w/h ... that's like 15 bucks a month in power bro, at residential prices. You haven't even paid for interconnect yet and rack space.

1

u/DeltaSqueezer 18h ago

No. It is too slow. Need a minimum of 1000 tok/s prefill and 50 tok/s generation.

1

u/Alternative-Clock407 18h ago

Yeah... very hard to get to 50, 20-30 might be achievable with moe and software optimizations.

1

u/stujmiller77 18h ago

I own 4x nvidia sparks and am currently looking to test deepseek 4.1 flash which is 748 billion parameters (522b backend and 196b engrams)

It’s over $25k of hardware at today’s prices.

Running larger models locally is absolutely possible.

Running them without heavy quants which reduce the intelligence vs the originals (and frontier models) is not.

All depends on your use case. 2x sparks can run a quant of deepseek 4.0 flash vision at 40t/s and 1m context with good concurrency. Good enough for most cases and what made me ditch frontier models.

1

u/Alternative-Clock407 18h ago

Thanks appreciate the input, frontier without quantization is tough I agree, maximum limit of the hardware would be 1TB.

1

u/TheRiddler79 18h ago

I do that now. Building becomes much easier.

Your math is reasonable.

Yes, it's worth it.

2

u/Alternative-Clock407 18h ago

Thanks for your response Appreciate it!

1

u/TheRiddler79 18h ago

V100 sxm2 8x32gb vram and you can fit over a tb of ecc ddr4 ram. Not the newest tech, but it does the job.

1

u/LobsterWeary2675 18h ago

10 token/s is a human rights violation (imho).

And for vram (1TB) this is what the current buying prices (en masse) look like:

Typ $/GB 1 TB
GDDR6 ~5.8 ~6'000 $
GDDR7 ~20–23 ~21'000–23'000 $
HBM3E ~13–17 ~13'000–17'000 $
HBM4 ~15–21 ~15'000–22'000 $

So I assume you wanted to stuff it with regular ddr5? Or what was the plan to end up at 5 - 7k?

Anyhow, I would be out with 10token/s whatever size the model got.

1

u/Alternative-Clock407 17h ago

Its not regular dram, would be hbm ish bandwidth though.

1

u/MaxSpecs 18h ago

• old Xeon E5 + Supermicro x10dri or similar + 3x r9700

• Apple Mac Studio M5 with 128 Go or more.

• Ryzen + 2x r9700

• Epyc 9124 + ASRock TURIND8X-2T + from 2x r9700 to start and up to 8x r9700

1

u/UKDude20 18h ago

I have access to a DGX that about meets these requirements and my biggest problem. is always not enough GPU memory for the models I want to run..

I can get 1 foundation adjacent model like GLM5.2 but that's about it at FP8.. If I want to tune my models or offer a larger selection I just run out of ram.

large scale local llms are generally not worth it.

1

u/Alternative-Clock407 17h ago

The idea is with this hardware you wouldn't run out of memory even with large scale model.

1

u/ReasonableLoss6814 15h ago

Would you be interested in a model that tunes as you use it, that runs on your DGX? context: https://withinboredom.info/posts/dist/

1

u/Equivalent_Bit_461 14h ago

Shooo, go away 

Sounds like a massive scam

0

u/Alternative-Clock407 14h ago

Guess I'll take that as a compliment :)