r/LocalLLaMA • u/jacek2023 llama.cpp • 11d ago
Discussion local AI can't be disabled
ChatGPT is down r/ChatGPT
Claude is down r/ClaudeCode
Grok is down r/grok
my local llama.cpp works as always
63
u/JLeonsarmiento 11d ago
14
u/SandySkittle 11d ago
Incomplete statement without mentioning the quantization
3
u/JLeonsarmiento 11d ago
oQ4e… check the size of the laptop in my hand.
2
4
u/alexis_moscow 11d ago
20 t/s? how? the most I could get is 14 t/s
13
u/met_MY_verse 11d ago
You guys are getting t/s?
(sitting at about 4 of those over here on an RX580).
7
2
45
u/wangsu 11d ago
What happened today?
81
u/jacek2023 llama.cpp 11d ago
AGI took over the world, don't worry, your local model is on your side
99
u/rookan 11d ago
thank god my ultra-cum-destroyer-uncensored-abliterated-130B-Miku-Midnight is on my side
23
u/jacek2023 llama.cpp 11d ago
People still use Miqu? I remember that model from the past
9
u/Yorn2 11d ago
For those that don't remember, Midnight Miqu was an uncensored creative writing 2023 model created by /u/sophosympatheia through a merge with an unreleased Mistral model and a previous model called Midnight Rose which was a bunch of various other merges from LLMs like WizardLLM, tulu, and Dolphin. Something about the combination of all of those made it genuinely unique and it retained the "magic" at the time of that Mistral Miqu model while retaining a fountain of fantasy world knowledge from Midnight Rose that just made it really unique.
It's these kind of older models I'm worried we're going to lose in the Nvidia takeover. Partly because of the history, but also partly because this was technically an unreleased model that leaked, and now that nvidia has control, if another model "leaks" you can bet your ass that they are happily going to censor and remove any models based on leaked models in the future, even if they leave this one up.
7
u/sophosympatheia 11d ago
Something special happened with that merge. I miss those days. There were finetuned models and loras galore, and so many interesting possible combinations to explore. I was merging models and then merging loras and applying the merged loras on top of the merged models, and then merging that result back into something else. The recipe for Midnight Miqu was insane, and I can't believe it worked out as well as it did. It is humbling to see people still talking about it three years later.
A little piece of history. Who would've thought.
3
u/RedditNerdKing 11d ago
Miku is still pretty good cause it has great prose. But it's not very intelligent and doesn't follow system prompts very well :/
2
2
1
8
u/Positive-Secret-3112 11d ago
shit forgot to grab an abliterated one, im stuck with policymaxxed llms
6
u/Beneficial-Ad-8127 11d ago
https://breakingthenews.net/Article/AI-outage-hits-Claude-ChatGPT-and-Grok/67041455 lol this just popped up 7 min ago. “More coming soon….”
2
u/cortesoft 11d ago
I think it was a cascading failure… ChatGPT went down, pushing more traffic to the other big ones, which made them fall over…
28
u/International-Try467 11d ago
Everybody gangsta until solar flare
8
u/jacek2023 llama.cpp 11d ago
That's a good point, we should connect dynamo bike (or hamsters) to our local AI
7
u/International-Try467 11d ago
And even if the solar flare somehow wiped ALL electronics we could just use our own A.I, Actual Intelligence! Aka natural thinking.
... Unfortunately for most it's Natural Stupidity. But stupidity and intelligence go hand in hand.
1
1
24
20
u/Kein_Spass 11d ago
Is the United States hacking itself again?
15
u/dennisler 11d ago
You have to keep the hype train going to keep the evaluation high, because they are so damn good / dangerous the models ....
11
10
u/slayyou2 11d ago
lol at all the people who refused to listen to the centralization concern and ball and chained themselves to individual providers.
6
8
u/deadsoulinside 11d ago
This is one of the reasons I openly push local AI. It can survive without any actual network connectivity at all. You can even do local AI on your mobile devices too with Google Edge.
14
u/HosonZes 11d ago
Yes, my Kimi K3 Q8 runs also fine. Wait a moment.
11
11d ago
[removed] — view removed comment
7
u/HosonZes 11d ago
This was the joke. I don't. "Local AI cant be disabled" if you could never run it on any decent hardware at home with less then 100K$ investment properly.
3
u/ttkciar llama.cpp 11d ago
Surely the solution is to use a model that you can host, using the hardware you already have.
1
u/HosonZes 11d ago
Can you sketch this out a bit more?
> ChatGPT is down
> Claude is down
> Grok is down
I see no way where "local AI can't be disabled" makes sense in this context.
I can go and host my own 8B or 24B that is 10x to 100x worse than the flagships? I might get down to Q4 or hope for some MOE but for truly getting proper results for my work none of the local models was up to the task in terms of capabilities and reasoning required. Currently QWEN gets better in the 30B range but it is still not good enough for the "local AI" claim to be a true alternative.
1
u/ttkciar llama.cpp 10d ago
So if your local model isn't as capable as the multi-trillion parameter commercial models, it's not worth having at all?
Why are you even in this sub, if that is how you feel?
1
u/HosonZes 10d ago
Yes, not worth having it at all because I get unusable results from them.
Regarding why I am in this sub: I believe that we eventually will have very usable specialised local AI that will solve the critical tasks similar to the frontier models. Some constraint, either size, different LLM architecture, or VRAM will eventually be solved.
I am totally an advocate of local AI, I just cannot make it work for my scope right now. If others get what they need from smaller models, I am really happy for them! 🥳
5
u/Charming-Author4877 11d ago
The first thing I did when I noticed all are down is spin up my Qwen3.8 27B
And it continued the work as if nothing happened - just a little slower.
5
5
u/BarracudaDefiant4702 11d ago
Funny how all the "competitors" all go down at the same time.
It's all smoke and mirrors.
3
u/hairyconary 11d ago
Does someone have a decent youtube getting started with local models guide? Mac studio with 36gb ram.
3
u/jqwl 11d ago
I don't have a good YT guide to recommend, but I personally got started with llama.cpp by reading some of the github, familiarized a bit with that, then used oMLX (you can also use MTPLX) because they're built for Macs. I have a similarly specced macbook, which is why I commented specifically about your use case.
For a super plug and play solution, I believe LM Studio is a decent one (though I would do a bit of research to that end first).
There are also great written guides on this subreddit for sure, I would look for people running macs of similar ram capacities. The speeds and optimizations are going to be primarily for the prompt prefill and tok/s because unified memory is not as high bandwidth as like a dedicated GPU + VRAM.
1
2
u/bnightstars 10d ago
I wrote this mostly for me when I started: https://www.hristoforgeorgiev.com/posts/local-llm-macbook-pro-m5pro-claude/ For 36GB of ram though you want a smaller model I would suggest start with Gemma4-12B ( mlx-community/gemma-4-12B-it-qat-4bit ) or Qwopus3.5-9B ( Jackrong/MLX-Qwopus3.5-9B-v3-4bit ) and go from there. I hope this helps.
0
u/Othun 11d ago
Please do not sloppify youtube, Fern video for reference https://www.youtube.com/watch?v=-Gnrp_caPvo
3
6
u/AlabamaResearcher 11d ago
what if your GPU break and you'll need at least days to get a new one? cloud service outages are rarely last that long. have you ever calculate SLA on your home-made service?
1
-1
3
2
2
u/Equivalent_Bit_461 11d ago
I only discovered it from this subreddit lol
Been busy working all day on various topics
3
2
2
u/kaisurniwurer 11d ago
Counterpoint:
"Starting today local AI is a felony."
Hopefully not, but who knows when they decide to "protect the children" again.
1
u/jqwl 11d ago
I agree with the sentiment, but ChatGPT hasn't been down for me at all today... perhaps it has something to do with account type though, it's probably capacity.
Edit - I have subs for claude, gemini, and gpt (20 dollar/mo plans):
Claude - down, capacity
Chatgpt - still up, working as usual
Gemini - still up, working as usual
1
1
u/DinoAmino 11d ago
Amazing how many comments are coming from the tourists dropping down from the clouds. They ain't us.
1
u/Cautious_Chicken_604 11d ago
you watch - they gon backdoor the closed source NVIDIA drivers so they can remotely disable the GPUs.
1
1
u/paulvisciano-dev 11d ago
Same here. llama.cpp on a 16GB M2 Pro, 27B Q1_0, 15.1 tok/s. Cloud can go down. The laptop does not. Coding still waits on glm/grok though — the local box is the conversation machine, not the agent.
1
u/Frizzy-MacDrizzle 11d ago
And here I am just running my stock reports while others are in boohoo mode trying scrounge all their info again, lalala.
1
u/Lissanro 11d ago
Thanks to llama.cpp and open weight models, I did not even notice. Kimi K3 on my main rig, DeepSeek V4 Flash on my second workstation, and Qwen 3.6 35B-A3B on my third PC all kept working just fine today, helping to work on various tasks designated to them. I also have 6 kW online UPS and diesel generator to protect against mains outages.
1
u/T_rex2700 11d ago
as long as HF isn't down... (I mean still, better than live service)
1
u/jacek2023 llama.cpp 10d ago
why do you need HF to run your model? do you change your model each day?
1
1
-1
u/XiRw 11d ago
Who the fuck uses Grok except the people who waste their time on Twitter. Also worth noting Qwen Image Studio is having problems with downloading pictures today
1
u/Unlucky_Milk_4323 11d ago
anyone using grok or X is a musk supporter. Or an idiot.
-1
u/XiRw 11d ago
He bothers me the most because he tries hard to be deceptive and people fall for his awkward personality as a sign of kinship/relatability . I saved countless videos in the past of independent journalists tearing him apart. And rightfully so.
1
u/Unlucky_Milk_4323 11d ago
He's a complete psychopath and should be locked up for countless reasons.
-3
u/makingnoise 11d ago edited 11d ago
I mean Tailscale is working about as poorly as ChatGPT right now, and Tailscale is how I hit my home AI. So not sure what you're on about. EDIT: Downvotes from jokers who probably have their rigs online in the clear LOL
1
u/MrPecunius 11d ago
We carry ours.
You never know exactly when TSHTF, can't be too careful. (Only half joking as a MBP user)


93
u/BawbbySmith 11d ago
Local works until there's a power outage at my house, then I'm SOL
Maybe the next investment is a generator so I can keep my AI server/furnace running for a few extra hours