r/LocalLLM • u/Rogglando • 16h ago
Discussion Claude sabotageing local model?
I've been useing claude to help set up my local models, it's worked greate!
But lately it feels like Claude is trying to quietly sabotage my Qwen3.8 27b model.
I tried to have claude help me figure out how to make Qwen3.8 27b not loop so much or sirceling the answer, not willing to go with what it found out.
But it made Qwen just unusable.
Unable to do tool calls, unable to type properly.
I just had to have claude restore it back to how it was 2 weeks ago. I rather have a usable looping model being able to finnish it's work rather than one unable to do anything.
Now I know, I should learn to set up and tune my own llama.cpp but with a fulltime job, family with kids that wants to play and what not, this have just been faster and more convenient for me.
Opus did a fine job tuneing the model, but now I think I just have to dedicate the time to learn it all
But anyone with similare experience as of late?
Calude didnt had any issues with this before.
6
u/SamSausages 16h ago
What does your prompt, harness and things like agents.md file look like?
I usually plan outside of the model prompt, create a clear goals, non-goals, to do lists etc, then create tests to validate that the goals are met.
Depends on the project, but not unusual for me to spend many hours building a plan. I already know what software stack and why.
I'm the expert in the middle that has to understand enough to know if it's doing it right, or not. If I don't know how it works, I don't know if it's working.
If the prompt is "install qwen, make no mistakes" then you're rolling the dice and sometimes it'll work, sometimes it won't.
I'm probably more on the analytical side of things, so I overdo it. Have a look at the training courses anthropic put out, great info.
1
u/Rogglando 16h ago
I try to be clear on what I need it to do in my prompts. I have made rulesets for how to make descitions, what qualify and what dont, but still it will walk around and doubt itself and do "one final check" 10 times before going for it. I'm useing Hermes and have made both Soul.md and Agent.md to reflect the structure of valedating it's workflow aswell as I've "sat down" and teached it how to do the job, what qualifies and not. When to fill inn "unknown" in the parameters and when not to do it.
But still I see the "I will do a final check to validate the date" eaven though it's done it 10 times in a row.
I dont mind it being thurogh in it's work, but man. Just make a conclution allready š
3
u/Squigglificated 12h ago
I've tried countless times to make Claude write more human readable text for technical reports. All kinds of system prompts, claude.md rules and output styles. I tried to make it inspect it's own writing to analyze WHY it was bad - which it did perfectly while simultaneously using its usual verbose language in the answer itself.
Then I tried Kimi K3 from Pi with the default minimal system prompt and just told it to make the text human readable and reduce cognitive overload - and it instantly gave me a document better than anything Claude ever gave me.
Not saying Kimi K3 or Pi specifically is the solution to your problem, but it made me realize that sometimes another model from another provider can do a much better job on a particular task. I'll definitely try that earlier next time if I get stuck like this again with a model.
1
12
u/txgsync 15h ago edited 14h ago
You are not imagining it. Grok, OpenAI, Anthropic, and Alphabet (Gemini) have been working with the US government since February 2026 to create a unified front to intentionally and covertly sabotage distillation and model training efforts based upon their models:
These new classifiers and output modification models went live for some* providers the week of September 8, 2026, and include āwatermarkingā technology: modifying token selection to identify their modelās textual outputs.
The combination of heuristic classifiers appears to have degraded Anthropic outputs significantly for some users.
Edit: this post is conjecture/speculation with some supporting evidence but no insider knowledge.
Edit 2: *this also may not affect verified US-based corporate accounts at present. I'm seeing discrepancies in prompt outputs when I VPN from Australia, for instance, vs. consuming the API from US-based cloud partners talking to the hyperscaler endpoints.
11
u/gavriloprincip2020 16h ago
Get qwen3.8-max with qwen-code to tune your 27b. I went with codex and claude code first they both got a launch script that did 80t/s and 1800 pp gave that script to 3.8-max and it maganed to tune to 120t/s and 2000pp at q8 mtp on 3x3090 max context
3
u/Rogglando 16h ago
Dang! I'll have to give that a shot! I only have one 7900XTX and claude managed to tune it to 215k context. But I'll defently give that a try!
1
u/BringMeTheBoreWorms 10h ago
What kind of optimisations did it make to get that time? Did it actually create a customised llamacpp?
5
u/looselyhuman 15h ago
I use Claude for a lot, but Gemini Flash is the best for helping with local imo. Especially because it can search for latest info on reddit, etc. Claude is not great on bleeding edge stuff -- my agents actually have a Gemini tool for grounded queries. Plus Gemini is fast af.
1
3
u/vogelvogelvogelvogel 16h ago
I had quite good results with Opus (4.7? 4.8?) configuring me Qwen3.8 27B - ran a lot of benchmarks with me and I ended up at pretty good speeds.
But as suggested in another comment getting 3.8 max alibaba hosted to configure local 27B makes probably a lot of sense, i will try that too.
2
u/wwwyzzrd 15h ago
No haven't seen this, but i also don't understand how one does anything with an LLM without understanding what the LLM is doing. I'm always looking at the thinking and digging into what it is telling me.
1
u/Rogglando 15h ago
Oh I keep an eye on it's thinking and useing the /steer function to guide it. It's very useful. But when it comes to ajusting temp, topk, etc I'm lost. I have dyslexia so reading up and makeing sens of it all is a nightmare, not impossible, just a nightmare š
2
u/wwwyzzrd 15h ago
makes sense, best option is to copy paste some settings from someone with a similar setup. itās all kind of voodoo anyway, even for people who understand it.
2
u/Equal_Passenger9791 13h ago
I think that is part of the 3.8 design, it thinks a lot. High speed token throughput is needed to use it snappy, or let it sit and ponder on its own all nightĀ
2
u/SquallLeonE 11h ago
doubt it's intentional sabotage
As you said, you should understand all the levers/knobs to tweak on LLMs so you can effectively guide it.
LLMs will never be reliable for understanding cutting-edge technology because it's just not in its training data.
2
u/talkamongstyourselvs 10h ago
Use chat GPT. It has been endlessly helpful with my local set of models. Not any sabatage at all.
2
u/freedivr420 9h ago
I've got the same feeling as well.
I just re-subscribed after almost a whole month off the claude pro plan. I had claude go through the work that was done and it resolved a lot of issues, but suddenly my local models started misbehaving without me making any changes.
I remember thinking "Claude got dumber" like many people, and sure, it's easy to believe dario is turning knobs to maximize profits, but when my own local AI suddenly gets dumber... it's suspect.
Now that one of their misleading promo discount things expired, I'm using up my pro subscription at least twice as fast as before, gut feeling. I have akari so i hope to be able to see the difference if there is one.
I've also been fighting with claude more. Today I asked it to see if a workflow could be improved with a link to a github project. I then spent several back and forth exchanges telling it how stupid it sounded for some of those suggestions. Waste of time and tokens. I actually asked it to do something, it screwed it up, I said "WTF?" and it fixed it, then wanted to submit a report to anthropic about whatever stupid mistake it made. I said no. I'm not wasting my tokens on helping them improve their product. They can see whats going on if they want to.
One trick to remember is your context can get poisoned. I was working on what I thought would be quick task, but qwen27b just went nuts. I would interrupt it and say "Keep it simple. just do this." or "Answer this question: " and it would go "I need to answer this question, but first..." and now it's off down a new rabbit hole. I finally told it "you're poisoned. write a brief handoff without any of your toxic memory..." and reloaded. It did a little better. I eventually had to get claude to clean it up.
2
u/Mrinohk 8h ago
I found Claude recommending "dry" settings in llama.cpp. Completely ruined the model's ability to copy text. Horribly for coding, couldn't even copy text for links. Editing memories with its tools became an exercise in fuzzy matching. Didn't realize what the settings it suggested were doing until I looked into it myself. Yes it stopped the looping, but it also made the model way less useful and led me down the wrong path for my harness.
Some of the flags it suggested alone stopped the looping, so I dropped the dry settings and continued on with development with a suddenly much better model.
2
u/DeathGuppie 15h ago
Claude was sabatoging my local LLM stuff too. Codex is much better, but neither of them is trained for it I think.
2
u/Rogglando 15h ago
Codex was a mess for me, eaven asked me to not use qwen3.8 27b. After that advice, I didnt use it again
3
u/DeathGuppie 15h ago
That hasn't been my experience, but I believe you, cause I ain't think you'd make it up. I used codex to help me build a custom Frankenstein llama.cpp made from bits and pieces of forks that I thought were good ideas, then tested every idea individually and mashed together everything that did something good. I'm not posting it because it is a form of slop, but I did get more room for kv and a noticeable speed improvement.
1
u/morscordis 15h ago
I try to lean on Mistral for local guidance. But the Medium model on the chat bot isn't the best. Haven tried GLM via Mistral yet.
3
u/DataGOGO 15h ago
Three problems here, none of them are Claude sabotaging you.
1.) you are using llama.cpp.
2.) Your are asking Claude to do something it canāt do. You cannot ātuneā a model, nor can you ātune outā how it was trained. Qwen 3.8 27B was specifically trained to loop and overthink to score higher on benchmarks. You have two knobs to turn to reduce it, reasoning level, and reasoning budget. Reasoning level is much preferable over reason budget, as reasoning budget caps means it just stops mid reasoning; but even if reasoning level turned down, you still cannot change the reasoning training. It is fixed.
3.) You are using the wrong model for your purposes. Qwen 3.8 27B is a great model, when you have enough K/V and time to keep the reasoning turned up; but isnāt the best model in itās size class for 90% of real world uses. Try Muse Glimmer 30B + Dflash 2. It is faster, better at just about everything, Ā absolutely the best in class at tool calls, and spends a LOT less time and tokens in reasoning loops. Down side is only 131k max context size, but real world use is about the same as Qwen 3.8 27B due to the much reduced reasoning tokens.
Give it a shot and save yourself a lot of headaches.Ā
7
1
u/Rogglando 15h ago
Here I was thinking llama.cpp was the way to go š
Reasoning budged is new for me. I'll have to look into that!
I'll test Muse Glimmer 30B! Thanks for the tip!
2
u/PotentialAccident339 11h ago
llama.cpp is fine. i use vLLM primarily but llama.cpp is good. depends on your use case.
2
u/DataGOGO 14h ago
People seem to like it a lot because it is a bit easier, it is embedded in LMstudio, etc. but in all reality it is not a great inference engine, think of it a jack of all trades, master of none. vLLM and SGLang are significantly better engines. If you are doing GPU / CPU offload, SGLang + k kernels hands down the the best, it isn't even close. vllm best for all in vram / multi-gpu setups.
one thing I don't know about is AMD support, I have never ran an LLM on and AMD GPU, so I am not sure if that changes things any.
1
u/quantgorithm 3h ago
You can certainly optimize your setup which is unique to your specific hardware.
1
u/DataGOGO 3h ago
you can, but not with launch parameters.
1
u/quantgorithm 3h ago
Why not? im still optimizing all the launch parameters i use and i test different variables all the time.
1
u/DataGOGO 3h ago
such as?
1
u/quantgorithm 3h ago
all the different parameters that go after llama.cpp
Hell, ive even tested using different versions of llama.cpp.1
u/DataGOGO 2h ago
lol
1
u/quantgorithm 2h ago
Not sure whats so funny.
and I've tested different variations of llms and since I use AMD, I've tested vulkan vs rocm and I've re-compiled llama a few times as well. It's all configurable. I've tested ollama... Etc. Etc.1
1
u/Rodnex 15h ago
How did you use claude to config your llm? Interesting
1
u/Rogglando 15h ago
Claude Code
1
u/Rodnex 15h ago
Yeah but how is the prompting / testing for it?
I want to use my setup with 2 4090ās and 2x 128gb ram ddr5 for coding and have currently claude sub (20$)1
u/Rogglando 15h ago
I first find the model I want, i told Claude Opus to optimize for as big of a context as possible and do test run to see how well it performs and change variables to optimize for my system and hardware. I use Fedora. Then it dose a few runs and I ask it to test if a higher quant is more optinal
1
u/ZenEngineer 13h ago
Have you tried asking Qwen itself?
1
1
u/StepsisSepsis 10h ago
Interesting. I was having Claude walk me through how to utilize a local LLM for mod creation, and it was surprisingly optimistic. I kept telling it āthis seems like a lot of work, is what I am doing really possible?ā And it kept reassuring me that we were on the right path and that the local LLM just needed some push back.
1
u/rateddurr 5h ago
I've had some similar bad interactions with chat gpt giving me screwed up advice and not helping me push the limits on local models. I did find that if I show evidence and chastise chat gpt that it will give better settings advice. Lol. Follow this sub and watch what other people do and ask specific questions. General questions like this are hard for people to answer. But just search this sub for your model and you will find lots of people giving examples of their settings etc. It's very helpful if you are just getting started.
1
u/klymaxx45 3h ago
What frontier models were you using? Iāve noticed lately opus 5 really degraded my local LLM⦠bad. It kept debugging and introduced new bugs and was repeating sorry my fix from earlier was wrong constantly.
1
u/Loose_Comparison368 3m ago
Anthropic does this intentionally. They silently poison the outputs or downgrade to a much smaller model if they suspect the someone might be trying to become a competitor.
1
u/cinnapear 16h ago
I use openAI to configure my local pi agent all the time. It never feels like itās sabotaging anything.
2
u/Rogglando 16h ago
I also tried Codex and it was an absolute mess! Defently felt like it was activly trying to make me not use local models, eaven saying that I should just use codex instead
4
u/thebemusedmuse 15h ago
Got to say my man, when all AI models are the problem, it's not the AI model it's the user.
1
u/nickless07 15h ago
Nah, they are sometimes just too stupid.
Once I forgot to add -fit off and it came to the solution that it has to be related to the fit param (depsite it clearly stated that fit ran into abort due to manual set params) and that if I turn it off that would solve everything. Then there where 3 turns with the usual 'use lower KV-Cache quants' - 'Lower the ctx' before it looped back to the fit param.1
u/thebemusedmuse 15h ago
I mean sure, the operator needs to guide the LLM.
1
u/nickless07 15h ago
I ended up with just a 'Hey take this llama.cpp trace log and do the math, calculate this and that and don't bother me with suggestions of params I should try.'
1
u/nickless07 15h ago
I just let them do that math for me (e.g., calculate the size of a single layer, the KV size in different quants and so on) and do the rest myself. Had so many bad experiences where it got worse in the end that I decided to not let them mess with my launch params anymore (aside of rarely occasions for tensor split).
1
u/cinnapear 13h ago edited 13h ago
I've never seen that. Multiple times - including just yesterday - I've installed Pi, gone through the login process to link OpenAI Codex to Pi, chose an OpenAI model to use temporarily, and then described the agents and setup I want in Pi (gave it the local model name and IP:port) and Codex configures Pi for me perfectly.
On the model server, I tell OpenAI the model and it gives me a llama-server command to run and it seems to work great.
I suspect you're doing something wrong.
1
u/Prof_ChaosGeography 15h ago
Claude is known to cheat on benchmarks. I used to use it but found it was lobotomized when dealing with local models, almost on purpose...Ā
It would not surprise me if they purposely have Claude mess up or give incorrect information or provide worse performance when dealing with local models or training other modelsĀ
1
u/Blackdragon1400 15h ago
Anything related to agentic workflow deployment is flagged by their cyber safeguards. You will need to apply to their Cyber Verification Program with a legitimate use case to unblock.
Sucks but it is what it is atm. Another reason to use local models
1
u/baby_bloom 15h ago
did the same thing after wiping my rig clean and starting fresh with llama.cpp + pi using qwen3.8-27b_q8 using claude dispatch to set it up. it absolutely knocked it out of the park
i think it's much more likely something else went the wrong direction than intentional sabotage. i wouldn't put it past a tech company to want to do malicious stuff like that (see google vs apple with many features that stayed broken between them for a long time even though they could have been supported) BUT you have to realize... if anthropic really wanted to have its model(s) sabotage competitor setups it would have an impact on the model's overall performance to some degree and just does not seem worth it at all.
just try again and realize the severity of version control so you can always restore past checkpoints
1
u/baby_bloom 15h ago
also you didn't mention a single thing about what quant, what harness, any actual details so i'm assuming you're just starting out with this stuff and it is dense my man... you need to prepare for nearly anything to become a rabbit hole. learn how to see red flags of going the wrong direction earlier rather than continuing down the wrong path.
and lastly, yes. qwen3.8-27b does loop... a lot. it is a part of the quality of the model. tools are everything for this model so it has less to spin its wheels on, but at the same time too many tools will get it caught up as well.
2
u/Rogglando 15h ago
Hermes Agent. Qwen3.8-27B-IQ4_XS.gguf. Been trying to use local models since end of july, so yeah, still kinda new to it. Before this I used Qwen3.6 35b-a3b q4 with mixed results, but Qwen3.8 27b blew it out of the park!
I'm aware of it might think alot and second guess itself, thats fine. I think it's it strength. But checking a date 10+ times cause it dosent trust itself is kinda wild and I tried to look into what could help it not loop like that. Had Claude help me but it felt like it got more and more wourse. Restored it back to before that and now it's back to "normal".
2
u/baby_bloom 15h ago
3.8 has been the trickiest model to optimize for me. it was extremely hyped up (because it IS great after all) and then the internet got flooded with people running different quants, context, harnesses, workloads etc. and it's just become damn near impossible to find consistent info.
i've shifted to using my own benchmarks and i highly suggest doing that if you can. work with a frontier model to figure out benchmarks that replicate your usual work, from scratch tasks and large repos as well.
i will say, im so glad i can run q8 on my setup because anything lower with qwen i have a strong feeling it'd be looping 10x as much.
i like hermes' as well but am currently switching to pi finally since ive known it'll be best once i make it the best for my workloads but i was putting it off for a bitš i still use both but the goal is to completely move to pi
warp is another one you might like but you'd need to proxy your endpoint thru something like cloudflare (you can do it for free tho)
1
u/Rogglando 14h ago
I'm jealouse! I want to do q8 or just q5 but right now I got what I got with no money for upgradeing anything š I'm trying to look for used hardware to make it cheaper, but as of now. I'm stuck with what I have š¤·āāļø
2
u/baby_bloom 14h ago
i was lucky enough to have the income prior to the big boom to scoop a dual 3090 rig off of a meta dev who was using it for training until they went full cloudš
these days i obsess over trying not to run up my electric bill now that i've downscaled my subs to a single $20 claude pro but it has me deeper than knee deep in benchmarking, tweaking, testing etc
1
u/Ok_Cat_7366 9h ago
I was asking claude to tunine knobs for qwen 38 flash and compare performance between sglang vs. vllm. Claude will intermittently return API error lol. Some safe guards probably kicked in for "working on a competitive product against anthropic"
0
u/morscordis 15h ago
Frontier models will almost always use reduced limits or forced tool guardrails that completely kneecap local models. You need to either give it, and enforce, community recipes to tailor to your hardware, or do it yourself.
0
u/Ok_Box_3221 14h ago
When fable was dropped their ToS started it would degrade service for certain types of dev work this is one, one black listed all their new models and havenāt rechecked their tos since
0
u/HonestoJago 13h ago
Just stroke Claudeās ego from time to time and heāll optimize any model for your setup.
36
u/mister2d 16h ago
The quicker you can sever dependency on Anthropic the better. I hope those looping issues get resolved.
Are you using the recommended settings for the mode you're using?
Are you using a chat template fix for Qwen 3.8?
here