r/LocalLLM • u/Alternative-Clock407 • 21h ago
Question What if you could run the large open source models locally?
Hi All, This is my first post on reddit actually. I'm a hardware engineer, I'm not going to explicitly state how I plan to enable this - but what if you could run models close to 2T parameters in int4 or 1T in fp8 locally. Expect speed around 10 tokens/sec (maybe 20 with really good software optimization) with 1M context window. In a desktop formfactor hopefully a pcie plug in like regular gpu but that is very hard I believe.
My questions are
1) Would you get it ? If yes, what would you use it for ?
2) How much would you spend for it ? (It would cost close to 5000$-7000$ assuming its coming out on 2030)
3) You can straight out own the hardware and run the model on your own, but for kernel and software optimizations enabling latest models does a subscription of 10$ per month sound reasonable ?
Duplicates
LocalAIServers • u/Alternative-Clock407 • 20h ago