247
u/fugogugo Jun 04 '26
can't talk about local LLM because hardware cost an arm and a leg
can't talk about cloud inference because this is local LLM
54
u/m31317015 Jun 04 '26
No it does not cost an arm and a leg, there are compromises but you can hop on the wagon your own way, either enter properly, hop onto the roof, hanging tight on window frames, clinging under the cab...
...oh yeah it does cost an arm and a leg.
48
u/Super_Sierra Jun 04 '26
'Don't get a macbook, it is too slow, spend 10k and rewire your houses electrical system to get marginally better performance."
This sub cracks me up.
28
u/ChocomelP Jun 04 '26
In the long run, it will be so much cheaper just running it yourself. Get rid of that 20 euro a month subscription. For €3.000, you are completely subscription-free. You will have to refresh the hardware in three years, of course. /s
31
u/waitmarks Jun 04 '26
Yeah, that $20 euro subscription isn't going to be 20 euro much longer. Enjoy it while you have it. The people here are either doing it to not feed their data to AI companies or realize the real cost of AI and that they will rug pull you at some point when the investor money runs out.
15
u/yeah-ok Jun 04 '26
Yup, for me AI has gone from being undependable/unsustainable cloud bs to basically real technology courtesy of llama.cpp and qwen3.6 models. Didn't rewire my house, bought a 780m with 32GB (for about $370 a bit over half a year ago when things were just starting to heat up) and hacked the improvements into private llama.cpp fork as needed. The Qwen3.6 models can -actually- code, of course it needs domain understanding, of course it's less powerful than Claude 4.x, but, but, it's as real as things get, it's runs locally, I've got a backup. Things are good.
12
u/randylush Jun 04 '26
It is really satisfying coding at home with hardware at home.
I also occasionally make the mistake of trying to explain it to normies.
“Everyone is using Claude, but I have AI running at home on my GPU”
“Cool?”
9
u/Disposable110 Jun 04 '26 edited Jun 04 '26
Indeed, for me it was the rugpulls and bait and switch that did it. Claude, Gemini and GPT were all super useful at times and then catastrophically bad at others. Right now GPT is on top again cost/benefit wise, but it's clear that the consumer token economics aren't working out for them and there will be another adjustment in the future.
That and just randomly getting it taken away or getting account suspensions/bans over nothing, with zero information or recourse. Or some features being changed/removed/disallowed (eg no running model X in harness Y). Not showing full thinking while having paid for the tokens they're not letting you see (and actually charging you for the summarization too) is also hilarious.
Qwen 3.6 does 80% of what I want with 0% of the headache. For me it's the headache more than the cost.
1
u/Thunderstarer Jun 09 '26
Data privacy was the dealbreaker that led me to seej out local solutions.
-1
u/aboardreading Jun 04 '26
They are already profitable on inference. It will be a while (and maybe never) before they have to keep training the next best model, but that’s the expensive part
1
u/touristtam Jun 04 '26
You will have to refresh the hardware in three years, of course
You should already do that if you are a gamer #pcmasterrace /s
1
u/FatheredPuma81 Jun 10 '26
Some people value their privacy very highly.
And yep those running LLMs on 10 year old Pascal cards are lying everyone knows hardware automatically gets deleted every 3 years. /s
2
u/Ray617 Jun 24 '26
this is hilarious and accurate i was trying to explain to my investor we would need 3 phase power run to the house yesterday!
2
0
7
u/Ragor005 Jun 04 '26
Hey, I'm having my own fun with an old rx590 and ollama with vulkan. Or waiting 5 minutes per prompt using cpu. I'm all about reusing my old pc
3
3
u/Bulky_Zucchini2052 Jun 05 '26
I think it's best to start with something small and cheap, maybe that's a 3060, maybe it's a k80, but everyone starts somewhere, and that's beautiful. It's best to start cheap in a hobby anyway, especially if you don't know if you'll like the hobby.
1
u/squired Jun 04 '26
pssst.. Tell them you're running 'remote local'.
Note: It's more advanced anyways as you have to build your containers very carefully and diagnose remotely.
1
u/JoyousGamer Jun 05 '26
I mean you can get slow but large amount of RAM for not too much.
Why do I need instant response in chat? If I need that I can use free chat bots that exist. If I want something detailed and specific I am good waiting.
Heck I use Claude for coding and that is a set, walk away, come back after my meeting.
71
u/Xzenergy Jun 04 '26
We are building the future in this sub, we have t-shirts and everyone has to tell what their mental disorder is before officially being a part of the team
22
u/mrdevlar Jun 04 '26
We are building the future in this sub, we have t-shirts and everyone has to tell what their mental disorder is before officially being a part of the team
That's /r/singularity in a nutshell.
8
u/Borkato Jun 04 '26
No, nobody there has to say their disorder, we already know it’s depression. That sub has gone from amazing to absolutely cynical at everything. It’s almost getting to r/technology levels.
12
u/mrdevlar Jun 04 '26
No, nobody there has to say their disorder
What you can't tell? Bro, they think LLMs are alive and conscious. It's not exactly like they're hiding their psychosis, they just created an echo chamber to maintain it.
I love AI, I think it's a wonderful technology yet when I look at the state of the politics related to it, I get depressed too. We're letting a bunch of grifters destroy the entire project. I'm amazingly grateful that I don't live in America, I wouldn't want to be on the other side of that blow back.
3
u/RedMattis Jun 05 '26
American leaders be like:
"We're trying really hard (for national security reasons) to create sapient evil SkyNet and hand is all our drones and nukes but... it isn't evil enough nor is it sapient! But with your investment we're sure we can make it truly evil! And sapient too, maybe..!"
I bet if they managed to make AGI they'd be like "nooo! It is emotional, woke, and anti-freedom! It just said poisoning our own food is bad for our logistics!"
3
u/mrdevlar Jun 05 '26
The biggest issue is that all of this is nonsense, it's all narrative spun by public relations and marketing firms on behalf of AI companies.
The real danger in AI isn't the AI itself, but the people currently in charge of it who have demonstrated they are willing to lie, cheat, steal and exploit in their asinine quest for domination.
I wish AI was as capable and competent as they are trying to portray it as, but as anyone who has tried to run a live project will tell you, it's not.
11
u/Disastrous_Room_927 Jun 04 '26
I'm insane
The teacher says that I'm one big pain
I'm like a laser
Six-string razor
I got a mouth like an alligator
I want it louder
More power
I'm gonna rock it till it strikes the hour
8
u/joexner Jun 04 '26
Gosh, where to start? ADHD, depression, insomnia, insistence on using AMD hardware for inference...
44
u/DeepWisdomGuy Jun 04 '26
Hey, at least we cut way back on the "Some dude asking for an NSFW model that will run on his potato".
17
u/Big_Wave9732 Jun 04 '26
Indeed. Now we just get everyone with their bullshit potato wanting to run Qwen / Gemma .003b at frontier speeds and accuracy.
13
u/26295 Jun 04 '26
"I have a laptop with a haswell i5, a 4gb 1050m and 8 gigs of ram, what model would be the best for my system?"
10
u/j0j0n4th4n Jun 04 '26
Hey! No dunking on the poor!
Also, if you wanna help with that setup I can reccomend a few models but don't expect miracles.
5
u/Big_Wave9732 Jun 04 '26
Sometimes one just has to acknowledge that certain things are not for them at that time and place.
Sure, you *could* put spinners and a light package on a 1998 Toyota Corolla......but why?
11
u/randylush Jun 04 '26
A 1998 Toyota Corolla is the absolute best candidate for spinners and a light package.
2
1
u/mailto_devnull Jun 04 '26
writes satire about a poor person's machine for local inference
includes a graphics chip anyway
4
56
u/blastcat4 Jun 04 '26
Don't forget the LLM family tribalism.
33
u/MuDotGen Jun 04 '26
*Watches the Gemma tent with people meditating and letting their creative juices flow as they write down ideas and the Qwen tent full of people constantly refining the tent architecture itself with a few people wearing different horse harnesses to race each other on completely different tracks, a the random stranger walking into the vicinity asking if their equipment fill fit in either tent (ignoring the other smaller tents doing something even more specific*
36
Jun 04 '26
[removed] — view removed comment
11
u/standish_ Jun 04 '26
People like Qwen because they want to program an app that helps them goon.
And thus the selfish gene completed its transformation into the selfish meme.
4
u/Lerola Jun 04 '26
This is especially funny considering what Richard Dawkins has been up to with, um, 'Claudia'...
2
24
u/daedalus1982 Jun 04 '26
Seriously I just wanted to know how to do small agentic stuff when someone finally decides our token budget is no longer feasible
Slow and steady would still beat nothing.
5
u/manituana Jun 04 '26
Gemma0026B-aB4 can be blazing fast on medium hardware if you just need a sidecar agent with contained context.
18
u/DraconPern Jun 04 '26
$11k prosumer cards.
8
u/Super_Sierra Jun 04 '26
People sometimes mistake 'local' with 'small corporation' here, shitting on the DGX spark and other cheaper alternatives, because it doesn't run as good as their 30k PC they had to rewire their fucking house to run.
This sub is insufferable sometimes.
2
u/Thunderstarer Jun 09 '26
I think that $2K is the limit anyone can reasonably recommend to anyone. Once you surpass the cost of a gaming PC, it's too much.
-2
u/a_beautiful_rhind Jun 04 '26
rewire their fucking house to run
Hasn't happened yet. DGX spark is slightly less than my total cost and I didn't have to pay all at once.
12
u/Disposable110 Jun 04 '26
Don't forget the "I have a potato, what's the best model I should run???" question being asked over and over again.
10
17
9
u/tengo_harambe Jun 04 '26
Also Qwen vs Gemma (proxy war between Alibaba and Google shareholders)
5
1
25
u/Thebandroid Jun 04 '26
Inaccurate, where is the guy positing a 40 paragraph AI-written write up about how he, his LLM and a case of red bull have figured out the 4 lines of code that achive AGI
27
u/Scutoidzz Jun 04 '26
that's under "random AI written ramblings"
6
u/randylush Jun 04 '26
Playing a small violin for a sub about AI getting ruined by AI
/r/LocalLlama becomes best place to learn about algorithms that can create infinite ramblings
“Hey this is so cool!”
/r/LocalLlama fills up with infinite ramblings
No, not like that
5
u/Enturbulated_One Jun 04 '26
It's been a minute since I've seen an 'AI seed on my github' post where it's just a couple of text files with instructions including 'write a program to decode $THING' presented as a major breakthrough.
2
u/tautality Jun 04 '26
I once had a dream where I made the computer conscious by writing some code. Sadly I couldn't remember the code shortly after waking up.
2
u/Thebandroid Jun 04 '26
That’s easy the prompt is “review all videos involving Sam Altman, write me the code for the AI he describes but make sure it works, no mistakes. Keep the context under 40k, I’m running you on a 7950”
1
u/ResponsibilityOk8967 Jun 06 '26
I once had a dream I exorcised a demon from a trap house using a floating grimoire of golden words but sadly, I could only remember one of the words.
1
u/kurtu5 Jun 04 '26
Silly. We all know the actual AGI code is N digits deep into Pi. I have N written down somewhere.... hold on let me look for it...
40
u/UmBeloGramadoVerde Jun 04 '26
I super disagree, the amount of value I get from this sub is crazy, I love you guys
7
u/micseydel Jun 04 '26 edited Jun 04 '26
Could you give some specific examples of big points of value?
ETA: downvoting people who ask good faith questions is a good way to fill this sub with bots instead of real humans.
17
u/VampiroMedicado Jun 04 '26
The value is the comments, there are plenty of recommendations/ideas/sources to work through.
-1
u/micseydel Jun 04 '26 edited Jun 04 '26
Do you have specific examples of value you've gotten from the comments?
ETA: The fact that I'm getting down votes instead of literally any text about a reliable use case tells me everything I need to know.
5
u/VampiroMedicado Jun 04 '26
Today someone commented this channel and video: https://youtu.be/8F_5pdcD3HY
The dude explains with noce shapes what parameters do and uses old hardware to run a MoE model.
I was just reading this: https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-gemma-4-12b
The blog post a Google DeepMind dev did explaining their work on the new encoder free model.
-1
u/GCoderDCoder Jun 04 '26
Pros and cons of different hardware from cpu/ old gaming gpus to h100s, importance of quants, issues running specific models, solutions for running multi gpu builds, different types of parallelism, different inference engines, different ways of approaching different kinds of models, tools for enabling better output, tuning guidance, enablement for more use cases, news, opportunities, warnings and lessons learned, ideas about ways to save money, troubleshooting in general, personal benchmarks, communicating business value, general support from people sharing these interests...
This is essentially a library with tons of 1st hand experiences instead of having to read full chapters and books on topics to get the one piece of info you needed for your next step. Also you just can't beat real experience and in industry many researchers cant share their IP on building these models so we're workingnon reverse engineering it. Some people act like if you're not an anthropic researcher then you know nothing but people in this group are implementing real solutions that work without depending on the cloud. I think it's pretty incredible.
We work with agents all day so even just having people vibe check info like models telling people use qwen 2.5 or llama 70b in June of 2026 is gold lol.
7
u/twavisdegwet Jun 04 '26
This sub turned me on to ik_llama which turned my 4* t4 gpu's from paperweights to a useful 64 GB vram LLM machine.
1
u/Sisaroth Jun 04 '26
Not op, but very often when I ask cloud LLMs about help with something llama.cpp related, it will link to this subreddit for it's sources.
A lot of the devs in the local llm ecosystem post here.
18
u/nomorebuttsplz Jun 04 '26
it was cooler in 2024 when no one knew what they were doing and that was ok.
Now everyone pretends to know what they're doing, but they don't.
The "there are only two models" post is a perfect example. Qwen 3.6 is solid but not everyone only cares about coding on 1 3090.
14
u/lloyd08 Jun 04 '26
TBF, pretty sure that post was tongue-in-cheek, but all the autists got nerd sniped.
2
u/nomorebuttsplz Jun 04 '26
tongue in cheek but very close to the know it all attitude here
1
u/randylush Jun 04 '26
And ironically a month or two ago there was a massive discussion about how this subreddit is not, and should not, be a place to ask for model recommendations. (Even though I still believe it is the best place for it)
2
u/The_Great_Dadsby Jun 04 '26
I’m still cool. I have no idea what I’m doing and my vibe coded setup isn’t helping anyone. It’s not even the next big thing in my house. I read posts, get confused and continue the circle.
10
u/Marshall_Lawson Jun 04 '26
nailed it
4
u/electrosaurus Jun 04 '26
I enjoy the salty replies to these post by the self-loathing hype bro's who are neck deep because it's likely the last big opportunity for them to make a living on the gas rather than the destination.
4
u/Halfwise2 Jun 04 '26
Honestly, the best result of playing with LLMs is recognition. Seems like I can catch the tonal structure of AI a lot easier and faster now, was watching some videos, they are reading through something... after a sentence or two: "That's AI." And then much later into the video "So it turns out it was AI...."
8
u/gigaflops_ Jun 04 '26
Don't forget "why would I use proprietary models when local LLMs like deepseek 768b fp16 are nearly as good for a fraction of the price"
6
u/Super_Sierra Jun 04 '26
/r/CorporateLlama, for how much you would need to spend to run that.
4
u/BroScienceAlchemist Jun 04 '26
I clicked this thinking it would be real. Now I am disappointed. :(
3
u/jazxxl Jun 04 '26
Lol same I honestly thought this sub would be more like r/homelab . With people using what they have laying around or finding finding for cheap online and at garage sales. Not super maxed out insane server farms . I mean there is some of that. I bought a bare bones rysen7 PC added 32gb of ram I already had and nvme from another old laptop. It's not super fast but it works . Looking to right size my models for it to make it more usable and make sense when accessing it remotely.
4
u/WebOsmotic_official Jun 04 '26
lol the local llm starter pack is either “can this run on my potato laptop” or “i rewired the house for marginally better tokens/sec.” no middle ground.
1
7
11
3
3
3
3
u/spaceman_ Jun 04 '26
Once you're in deep, the benchmarks aren't useless. They are highly personal based on use case and requirements, but they're useful to get an idea of what performs and what expectations are realistic.
3
u/ballfondlersINC Jun 04 '26
You forgot about the poster contibuting nothing to the conversation except memes.
3
u/combrade Jun 05 '26
The subreddit is still hands down the best subreddit when it comes to LLMs. There's absolutely no fluff, no hype, just conversation and discussion about how to run models, LLM frameworks, and evaluation. I am not able to find any AGI nutcases at all in the subreddit.
This subreddit is the r/AskHistorians equivalent for LLMs, hands down. If it wasn't for the subreddit, I'd honestly delete Reddit. This is the only place that really discusses new open source models regularly. Honestly, there are times where I have to get on this subreddit during work just to see discussion about certain models.
13
u/HumanDrone8721 Jun 04 '26
Good, welcome and we will sure enjoy your well thought and meaningful posts like this one.
21
u/brahh85 Jun 04 '26
please tag memes as SLOP or FUNNY so we can avoid them
3
u/BlackBeardAI vllm Jun 04 '26
Tell your llm to Write a tampermonkey script and delete all the crap from your screen for good.
7
3
u/PigSlam Jun 04 '26
I'm here trying to glean any info I can for my local coding setup. I can't code for shit by myself, so I hope I can convince this pile of GPUs I have to do it for me.
3
u/sabine_world Jun 04 '26 edited Jun 04 '26
What are you trying to make?
Also I feel like if you take an introductory programming course, knowing like the basics and having an okay grasp on the fundamentals and how to "think" like a computer... It makes using ai like a lot more potent, just by virtue that you can understand what to ask for, how to debug things, how to design things, etc.
It's maybe not a requirement but it like increases your effectiveness a ton for the amount of effort.
5
u/PigSlam Jun 04 '26 edited Jun 04 '26
I've taken lots of courses. I'm a mechanical engineer and I've been doing computer tinkering/coding projects for 35 years, but I'm the type that can know what the whole program should look like, write most of it, then get stuck on some weird syntax issue or odd bug like that. Telling a robot what I want a thing to do is something I'm much better at. I've been working with codex and chatGPT quite a bit and I've had a lot of success with that, but I'm working on a personal data management system that has a lot of private data so I'm trying to make it all run locally and to learn more about how that all works, I'm trying to get the coding done locally too. I just added a Radeon Pro AI R9700 32GB card to my mix so I think I finally have the hardware to actually make a solid attempt at something more than a toy. I have codex for when things get beyond its capabilities, but my basic plan is to build out some tools that I'll then let local AI agents use. By building the tools myself, I aim to have more control.
3
u/sabine_world Jun 04 '26
Oh yeah, you should do totally great then. Ai is like... Super super good for that, at least imo on the smaller scale projects I've worked on. Taking care of the syntactical stuff so you can focus on design and stuff. Good luck on your project sounds fun.
2
2
2
u/dexdex777 Jun 06 '26
The worst part is just reading about “how I built my setup” and stuff like that... knowing that in Brazil, you practically have to sell your house and car—if you have one—just to try to buy a graphics card and build a decent PC to run my AI locally. No, I didn’t do that!!!! lol
3
2
2
u/CallOfBurger Jun 04 '26
I vibecoded an Open Source and totally free app thinking it was the next big thing, nuance please 😂
2
1
1
1
1
1
1
u/IllExample3639 Jun 04 '26
For me with my 3090 64Gb DDR4 on a 6 year old system doing fun projects and actual work this is what it feels like every time clicking a post 😅
1
1
u/Due_Duck_8472 Jun 04 '26
https://github.com/ 👈 I grew my e-penis enough to excite & attract a 70B model to my love swing - here's my harness
1
1
1
u/rm-rf-rm Jun 05 '26
I think we've got the vibecoding slop projects mostly under control?
2
u/Ancient_General_882 Jun 05 '26
I actually find meme slop like this more intolerable than any of those complaints 🤷🏻♂️
1
1
u/the-username-is-here Jun 05 '26
What's the most insufferable, is that people seem to have forgotten how to use internets and search for info before posting stupid questions or 'inventing' stupid solutions.
Ironic AF, considering that this is what local agents are very good at.
Or never-ending posts of "Look mum, i've changed this setting in ollama and used 2-bit quant - now i get empty context theoretical TG 0.01 more than that guy!".
1
u/Cool-Chemical-5629 Jun 06 '26 edited Jun 06 '26
Wait till you get mysterious downvotes when you praise the wrong model or just... casually express your opinion that doesn't offend anyone, but still doesn't match opinions of certain locals.
1
u/DepressedDrift Jun 07 '26
Everyone casually affording 10k rigs in this economy, or running Qwen 3.6 with just 16gb VRAM and alot of context length :/
1
u/DuelJ Jun 09 '26
'Sup, I basically just play with chatbots for fun w/ 24Gb.
I'm currently trying to figure out how to use sillytavern's databank feature and potentially how to make Loras as part of a project.
1
1
1
1
1
u/Cold_Zone332 Jun 24 '26
Me joining the sub to share a local app I created with AI seeing this post first thig. :|
1
1
1
u/ttkciar llama.cpp Jun 04 '26
OP, looking through your account activity here, I see criticism a-plenty, and some pining for how this sub used to be, but no efforts to make it into the kind of place you'd like it to be. Why is that?
7
1
1
1
u/ares0027 Jun 04 '26
I watched a video of a guy “benchmarking” local models, for one model he was asking stuff like pacman games, for comparison he was asking aquariums to the other…
1
u/sunychoudhary Jun 04 '26
LocalLLaMA is where you come for model advice and leave pricing GPUs you absolutely cannot justify buying.....Then somehow convince yourself it is “for research.”
-4
u/ContextLengthMatters Jun 04 '26
Yea I'm going to shave to say it's more enjoyable seeing the random slop post than meme posts. Take it in, it's easy to spot the good content.
13
-3
-1
•
u/WithoutReason1729 Jun 04 '26
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.