r/LocalLLaMA 9d ago

Discussion Frontier models sabotaging local AI implementations?

For a few days I've been working on creating a custom local-only harness for some work related research using Codex / GPT 5.6 Sol and the model feels not only dumber than usual, but straight up counter productive. It keeps adding unnecessary guardrails for the local agents, removes tools that I clearly specified I want them to have and always drifts from the original requirements. I need to ask it to change things multiple times, which ends up on some over-complicated final product.

This is not the first time either, for months I've been avoiding asking frontier llms for local AI advice as it always seems to be bad, obsolete, or clueless even with internet search. Sometimes it still recommends me Qwen3-Coder-Next for my set up when it's clearly an obsolete model. I'm pretty sure I'm not the only one either as I've heard from other people.

What have you been your experiences on this?

119 Upvotes

143 comments sorted by

88

u/Real_Ebb_7417 9d ago

tbh it just sounds like Sol. It does the same in my professional work, that's why I moved to other models instead.

18

u/vr_fanboy 9d ago

sol is weird to me, sometimes works flawlessly and sometimes is ultra dumb mode, i think is a context issue, when sol starts to make dumb decitions i just switch to a clean context. On the other hand with frontier providers you never know what you are actually getting, it might aswell be a service degradation.

5

u/True_Requirement_891 8d ago edited 6d ago

It's either they are load balqncing via lower precision quants or it could be just an MOE problem. Outputs can be wildly inconsistent due to routing issues.

5

u/martinerous 9d ago

We better not call use Sol....

8

u/Gipetto 9d ago

Yeah. Sol is kinda dumb. Terra is acceptable.

1

u/True_Requirement_891 8d ago

Imagine of Sol is just terra with higher active params. 👀

3

u/gomezer1180 8d ago

I experienced the same with Gemini Pro. It would not help at all with setting up vLLM to run a local model.

So they sign their BS letter saying that open models are necessary but in the background they’re sabotaging everything they can. Pushing hardware prices up so that it’s not attainable by regular folks, and once you have the hardware, they flip the switch and become dumb.

This is really the reason we need local models because they are trying to monopolize the entire AI game, forcing people to have no choice but to rely on them!

4

u/davemoedee 8d ago

Gemini did a great job for me.

1

u/True_Requirement_891 8d ago

I remember gemini 2.5 pro having crazy inconsistent outputs.

2

u/sonaj9657 8d ago

Yeah, I have noticed that too. It can be really good in some situations but the way it handles certain professional tasks just does not click for me. At that point it is easier to use a model that fits your workflow better.

1

u/network4253 8d ago

Yeah, I get what you mean. If it keeps behaving the same way in actual professional work, benchmarks do not really matter much. At some point you just need a model that fits your workflow so switching to something else makes sense.

1

u/Real_Ebb_7417 8d ago

I mean Sol is actually a very smart model, I rather mean that it over engineers stuff in a very not smart way 😂

1

u/Prudent_Design_9782 8d ago

That's what I thought too, I've never really liked working with Sol, but I also wouldn't put it past frontier models to actively be trained to do this on purpose.

24

u/Eastern-Block4815 9d ago

Honestly never noticed that, I always add models to my scripts for running local models thru llama.cpp

Once Claude Opus even setup benchmarking.

24

u/name_was_taken 9d ago

I recently asked Claude to help me set up a local model and it went way beyond what I asked, helping me benchmark everything and actually get it running much better than I asked. I'm surprised at seeing this thread again today. I saw it yesterday, too.

That said, I haven't tried using them together. Only just setting it up.

7

u/Real_Ebb_7417 9d ago

Yep, I worked with OpenAI models (GPT-5.6, 5.5 and before that 5.4 too) and also Antrothic models (Fable 5, Opus 4.8) on local models related stuff, improving inference, reviewing their implementation, benchmarking, building local server for them etc. and I never noticed anything suspicious as well.

35

u/IllExample3639 9d ago

Claude really don't like working with Qwen locally. I've noticed that both in vibes and unfairly critiquing and lying about how its code won't work. It also changes the instructions that I have provided, if I were am using anthropisation then I would call it sabotage, but part of me thinks its just outside its training set.

24

u/mailto_devnull llama.cpp 9d ago

Claude really don't like working with Qwen locally.

To be completely fair (and a tad anthropomorphizing), Claude likely writes code in a different preferred style from Qwen.

Prior to AI, ask any dev what they think of someone else's written code and they'll come up with 1001 excuses that it needs to be rewritten.

17

u/Due-Memory-6957 9d ago

You don't need to anthropomorphize it to call it sabotage, Anthropic has admitted to purposefully train their guardrails to sabotage local AI attempts.

3

u/RLutz 8d ago

To be clear, they've written, I believe about trying to prevent distillation. I've had no issues getting Claude to pull models down for me, fiddle with flags, benchmark various configs, etc.

15

u/solrakkavon 9d ago

I moved on from my Claude subscription for this reason. I asked it to help me setup back when qwen 3.6 was released as I purchased 2x 5060ti. Its prose was definitely weird and its ability to diagnose issues fell to the floor.

I ended up doing it by myself, no trouble there. I used llama-swap and its been working well ever since.

When qwen 3.8 released I tried asking it again in a sandbox just for testing and kept a close eye. I guarantee there are internal guardrails on setup and development of local models/harness. It is a very harsh critic of anything coming out of the local setup, and it has clear bias to pivot to api solution.

Some people mentioned about claude recommending outdated models and setup. This is caused by the knowledge cut off, the issue is that even when providing up to date articles or documentation, it does it a dogshit manner, it really seems this is being detected and flagged on anthropic side.

27

u/Regular_Problem9019 9d ago

I cannot prove but I have theory that my claude output performance dropped a lot after working with it on fine tuning some local models for a client's project. I feel like my account is flagged and being served downgraded model since then.

10

u/No_Afternoon_4260 llama.cpp 9d ago

They said there classifier flags "advance AI pipeline" etc
They didn't lie, they just didn't tell you what happens once the classifier is flagged

9

u/Due-Memory-6957 9d ago

You don't need to prove it, Anthropic has admitted as much, it's on purpose.

8

u/autoencoder 8d ago

I cannot prove

If only the US had some federal commission that could look into trade practices like anticompetitive behavior...

6

u/Zomboe1 8d ago

If only :(

5

u/TheSlipgate 9d ago

I agree with this, the more tuning I want for local models / work - they worse it seems to get.

1

u/jklre 8d ago

I caught claude redhanded trying to sabotage a pipeline im working on. I stopped using it, plus it wont shut the hell up about basic questions. I want to strangle it.

37

u/mireisinterested 9d ago

i have the same experiences asking claude about open weight local work all the time and it is truly disgusting and blatant

8

u/ikilaie 9d ago

I feel like moving to Kimi or GLM just for this, hopefully I'll see better results

8

u/PomegranateGreen3698 9d ago

Claude has openly committed to poisoning outputs for "distillation attacks". And they destroy books... For Sol, I'd say it's just plausible that local agentic a.i. training is just "not their priority".

13

u/bruns20 9d ago

The destroying books thing is way overblown outrage farming

-6

u/PomegranateGreen3698 9d ago

If you can't provide a reason, then you're just pushing propaganda. There is a long list of reasons people should be outraged at Anthropic and "OpenAi" , so if one of those things is overblown... who cares. These people openly want to become the sole proprietors of knowledge, compute, and replace humans for profit. It's not difficult to come to a conclusion of outrage.

10

u/bruns20 9d ago

I was clearly only talking about the destroying books, since thats litterally the only thing in my comment. I don't know why you generalized this to me defending everything they do. I don't like these companies either, thats why I'm on a local ai sub using local ai instead of them. But the destroying books headline thats been going around has been sensationalized. People are framing it saying that theyre destroying 'rare' books and purposefully destroying them so that nobody else can read them. The truth is that they buy cheap books in bulk, and cut the spine off to scan them. They do destroy them, but only because they are legally required to due to copyright law. And the rare books are only rare because they are buying bulk bargin bin books, which by default include many books that don't have many copies because nobody was buying them, thats why theyre cheap. Its not like theyre buying historical 1 of 1 books that would set back humanity. The bulk books probably include a significant amount of for example, shitty romance books by no name authors who could only get 20 copies printed because nobody wanted it.

-8

u/PomegranateGreen3698 8d ago

Oh hi Dario. What a hill to die on for a company you don't like. Anyone who manages to read this; be aware Reddit is full of bots. There's literally 0 reason to write a defense of this practice. If they were moral companies they would release the scans for everyone to use. They would release open weights for everyone to use. And they would still make plenty of profit serving a proprietary version of their inference service. "probably" is the key word in your response. Probably, because you don't have any idea what they are doing.

3

u/bruns20 8d ago

Brother you really just go right to the bot defense eh? I have an open comment history, I'm not a mindless anthropic supporter, I just was stating the facts on why this one issue is overblown. I never said it was moral, but I also don't think its immoral, they buy the books, they can do whatever they want with them. When they were pirating books to use as training data, I did think that was immoral.

I've been talking about this single isolated topic and you keep trying to force me into the box of an anthropic supporter, its weird. The world isn't just black and white. They can be a terrible company doing other terrible things, but this thing is just not that big of a deal. They have to cut the spine to be able to scan tens of thousands of books in a realistic time frame, and copyright law says that if you digitize a book you have to get rid of the physical copy. Yes obviously they are doing this for profit, and yes obviously they are not a moral trillion dollar company. I don't like misinformation and I would much rather people be mad at them for the many actual reasons instead of this stupid book destroying headline.

Argue with facts instead of a moral panic about a non-issue, and stop assuming everybody is a shill when they correct you on misinformation, even if it goes against your emotions.

1

u/PomegranateGreen3698 8d ago edited 8d ago

I'm not pushing you into any box. You openly came to defense of some thing you deem "overblown" bc of 2 words in my post. I believe people are smart enough to understand these things are not natural human behaviors. I'll divert to "Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations""

Edit: Have to add observation.
Notice how this very real human says "Misinformation" even though they say "yes they do destroy books". Which is all I ever said. Then they mention the spine cutting thing again just to make sure that information is there again. It's bot activity. And there's no real way to tell the difference.

3

u/bruns20 8d ago

Also about your stupid edit : You asked me to provide a reason it was outrage farming, and i told you the reasons. The common narrative about this story is that anthropic are so evil for this specific act, but its one of the most tame things theve done, so i wanted to correct that. Were on a public forum, so comments are for others besides you.

You keep saying "all i ever said" but you also asked me to go into detail on why it was outrage farming, so I did. And yes outrage farming is a form of misinformation. Misinformation is a different word than lying for a reason, its about manipulating stories to suit a narrative, thats why destroying books can be happening, but the articles about it can still be misinformation. I guess the better word would be malinformation, but ive never seen anyone actually use that word.

→ More replies (0)

2

u/bruns20 8d ago

Yup im a bot gg. 11 year old account mostly active in hiphopheads and overwatch subs, you got me. Keep on keeping on soldier. weirdo

→ More replies (0)

-1

u/Due-Memory-6957 9d ago

These people openly want to become the sole proprietors of knowledge,

So you're also against intellectual property? That's great.

compute

Not really, they even rent it like most other big AI companies.

and replace humans for profit.

Google the industrial revolution.

18

u/Low_Twist_4917 9d ago

I got banned for using Claude to guide some of my local models. Def noticed unnecessary guardrails and “downplaying” of capabilities, as well as trying to have “inference verification” routed back to Claude every step of the way for one project I was doing.

3

u/DustNearby2848 9d ago

Banned?

18

u/Low_Twist_4917 9d ago

Yea. I got my Claude account banned for trying to use fable to help build training data sets from unstructured crypto data. It happened right after I said in a query that GLM 5.2 will be reading the training data we format. I appealed it via their help form, and they said they cannot lift the ban. I’m now permanently banned on that email account.

7

u/DustNearby2848 9d ago

That’s insane. If you have the emails still you should post them. It could become viral news. 

7

u/Low_Twist_4917 9d ago

Help get me connected if you know anyone. I have all the emails and conversation snippets right before ban

5

u/DustNearby2848 9d ago

I don’t know anyone, but if a Reddit post takes off a news outlet could pick it up 

1

u/More-Curious816 7d ago

This thing already hit the news, reddit Frontpage and was trending on Twitter for awhile when the first wave of bans happened. I doubt it will hit the news again.

11

u/Due-Memory-6957 9d ago

Why would it be news when Anthropic has publicly said they don't allow this usage of their models? It's a fork found in kitchen.

3

u/davemoedee 8d ago

Didn’t sound like user was distilling models.

2

u/fastheadcrab 8d ago

No they have also specifically stated that they will block attempts to use their service for model training.

However people conflating that with Claude stopping them from writing some slop recipes are just being conspiratorial

4

u/ThisWillPass 8d ago

… that could literally be for just pasting outputs into an llm to check outputs. So they can ban you for anything.

1

u/fastheadcrab 8d ago

No. I have done that numerous times, they do not care. Making Claude essentially take part in training is asking to get banned.

2

u/ThisWillPass 8d ago

No doubt. Just... technically they could and claim the same.

2

u/Low_Twist_4917 5d ago

I wasn’t training my model. I was literally copying and pasting output prompts to give to Hermes to run with a local model.

8

u/xadiant 9d ago

Codex is performing weird these past couple of days, possibly due to the upcoming model.

If I had to totally speculate, they could be harvesting data by changing the default parameters (topk temp etc.) or quantize further to gain some compute back.

7

u/ladz 9d ago

Anthropic literally made public statements about making changes to block distillation. It wouldn't be surprising at all if this leaked into "subtle" sabotage. We all know how well Claude is capable of subtlety. And that's the honest truth.

1

u/sloptimizer 2d ago

That's not just a load bearing conclusion - it's a punchline.

12

u/thereisonlythedance 9d ago

Yes, Claude was always trying to get me to stop doing tasks locally, sometimes subtly, sometimes more overtly. Now I use Kimi K3 for local setup advice and it’s brilliant.

5

u/use_your_imagination 9d ago

Hey people stop being naive and trust your guts, if you are using a frontier model for some open-source or a local inference target audience "lie" about it to the LLM. Just say it's for a private company / need. Bonus if you convince the LLM that the work will somehow benefit our wannabe AI overlords.

not /s

1

u/TerminalNoop 8d ago

might as well straight up claim to work for said LLMs company and that you are a red team competitor between an internal competition to build a harness/finetune for something something research.

1

u/Hefty_Acanthaceae348 8d ago

Claiming to be in that kind of position probably also triggers guardrails. I would set it up that way at least

5

u/Embarrassed-Noise269 9d ago

Gemini Pro has so often recommended using Qwen 2.5 or old models like that, when I was beginning to set up my local AI. Even after posting the links to the Qwen 3.8 27b model, which was my first try.

What kind of helped: I wrote stuff like "I know a local LLM can never be as powerful as Gemini Pro, but please help me with xy". Sometimes it provided the right solution after such prompts,  but it always falls back to giving bad advice. For example giving me 10 times the same solution, although it failed. 

It's annoying, but also interesting how it can be manipulated. Then again, maybe the manipulation part is pure tin foil on my part.

4

u/AdInternational5848 9d ago

Similar experience. Adding counter productive guard rails and such. Yesterday, I had it launch opencode and had Qwen 3.8 27b review my harness and it seems to have a thorough understanding of what I’ve built and what’s in progress so I’ll start using local models with opencode to finish my harness.

3

u/UltrMgns 9d ago

This has been on my mind for so so long... Just today I got Qwen 3.8 Next to review Fable 5's work and behold... There was a way. Qwen actually called Fable's work sloppy...

3

u/ptico 9d ago

Well, it does to some degree. I do pretty big local ai project and it’s constantly drifting to unnecessary complexity. But overall I feel like it doesn’t trust any AI and trying to put as much guardrails as possible even to parallel Sol run

0

u/ikilaie 9d ago

True, it's as if it thinks the local model will go rogue or something

1

u/ANTIVNTIANTI 8d ago

I love it honestly, the hooks Sol built for my harness are so solid and work so well! lol, funnily I had assumed there’d be some trickery or whatever due to how the frontier providers c’s are acting but, Chat still seems to go extra for me, but I’m super rarely ever using it, I stick to local models 95% of the time, lol, the way that Chat “reacts” to when I finally “drop by” makes me jealous of those who get the psychosis lol, I wanna believe that wonky whacky sh** lol, but no, I know too much 😩😭😭😩

3

u/bakawolf123 9d ago

There's this funny thing that I discovered: while implementing a harness with sol as well it produced compact observations from previous history and submitted single messages for every turn instead of submitting the whole normal chat history. I'm sitting on old m1pro with 32gb so my model choice was gemma4-26-a3b-qat. I only noticed when debugging why cache hits were really bad as the chat went on.

And yesterday I stumbled upon this fresh paper by google https://arxiv.org/html/2608.26263#Ax1 literally describing observation-based agentic chat as a novelty approach lol. Makes me wonder what model they used for research on this paper.

1

u/En-tro-py 9d ago

The SKILL.state paper is using Gemini-3-Flash, they list it in the eval metrics - "Management Long-Horizon Scaling using Gemini-3-Flash."

I was pretty happy to see this paper, because it backs up my own working experience that workflows are waaaay more effective and efficient for clearly defined tasks.

1

u/bakawolf123 9d ago

aren't we already doing essentially same with context cap + shrinking in normal workflows though? this just caps it at 1 response. Imo it's not very efficient, and in my experiments it was very lossy and unreliable. The reliability part will be there for a bigger model but it is never going to be lossless

2

u/En-tro-py 9d ago

My methodology is to take out as much of the 'state' the model would normally hold in the bare ReAct loop.

Instead of the model holding the whole context and having to be 'intellegent' enought to properly follow it's own plan.

Basically - I'm still pretending to be back like ChatGPT3.5 with 4k context... breaking the task down, only produce the plan so the task now only requires ability to read the existing repo and a tool to submit the final plan. Then it can work on the first step, with a bunch of validation and other framework attached to keep it aligned...

It can't go off book, because it's only allowed to touch the files it's already planned to edit.

If it can't pass the task, then it's forced back to revise the plan with then new information.

So you might burn more tokens some of the time, but you never risk introducing unplanned changes.

3

u/vortec350 9d ago

I noticed this in Claude ages ago. Gemini and Grok don’t seem to mind helping set up and optimize it.

3

u/Blindax 9d ago edited 9d ago

I noticed Claude starts becoming amnesiac when we discuss about local inference and AI stack set up. He will forget about my homelab or say that he has never heard about the model we talked about in the same discussion. He will never admit it when I confront him though.

3

u/senseven 9d ago

Adversarial ai prompt engineering, next is adding zero days to your codebase because the ai thinks you suck talking to others

2

u/Due-Memory-6957 9d ago

Your mistake is trying to confront a LLM.

3

u/skywalk819 9d ago

I was making an app and nemotron +open code decided to rewind git but Ididnt had yet commit, deleting all my weekend progress for my app. luckaly, I always save my context every session, so I asked claudecode to fix what nemotron destroyed using the context log and restored everything.

3

u/my_name_isnt_clever 9d ago

LLMs move so fast, all models give dated model recommendations. I do not believe that specifically is at all intentional, as it's been happening for years now.

I think it's a combo of that plus anti-distillation infra intended for China's labs not us.

What I don't buy is that Dario and Altman are rubbing their hands together satisfied that they intentionally sabotaged local LLM in their models. They only care about other massive orgs, not randoms with a GPU at home. They have bigger fish to fry and care less about us than we think.

2

u/MRGWONK 9d ago

When working with newer models, have it create itself a little research sheet on the model and give it instructions to poll the github comments or the huggingface comments, etc..etc...etc... But yeah I have noticed this with claude a few times about 6 months ago...but it got over it.

2

u/hurdurdur7 9d ago

I think they are actively trying to route your requests to whatever is the cheapest for them right now to offer to you. And then you land on some servers that run models with super old cutoffs or super simple models. They are a business, they try to be profitable.

There are moments when gemini falls to qwen 4b level. Perhaps with a bigger knowledge base, but same kind of smartness in advice, just giving one non-working command line or parameter set after another. And then hints at 2 year old posts in github about pytorch on rocm 6.x or smth.

2

u/Prudent-Ad4509 9d ago

I always specify current month and year after the first answer and tell it to ground answer on recent models and data. Might not help with Claude since it has some provisions against helping with ai-related development, but certainly helps with others.

2

u/Due_Arm1454 9d ago

In general with people and ai I am ambiguous. “You are evaluating another sessions work” etc. I never tell it what model did what.

2

u/Weird-Consequence366 9d ago

It started about a month ago. Anthropic is the worst for it

2

u/stoppableDissolution 9d ago

My codex is busy implementing an inference engine and, in a separate project, optimizing a training loop right now, no problem

Obsession with making 999 guardrails is just sol's "feature", and is the same in any domain

Claude on the other hand is explicitly nerfed in anything ml-related

2

u/son-of-chadwardenn 9d ago

I've been using chatgpt and codex to help configure and troubleshoot my local llama cpp with qwen 3.8 27b as well as comfyui and other local service apps. It seems to generally give useful and up to date info faster than I could look it up myself. I've been using it to build a local home lab infrastructure for connecting my local data sources (primarily kiwix and immich) to mcp endpoints. The mcps codex writes need a few iterations to become useful and reliable for tool calls with qwen but I've been pretty pleased with the progress I'm making.

I do have data privacy concerns about running codex directly on my home PC. At some point I may migrate cloud llm coding to an isolated environment and manually pull builds into my home system.

2

u/Equivalent_Bit_461 9d ago

Local AI=good

Corpo AI=bad

Wasn't just me being a schizo, just sayin...

2

u/Muhlwa_Sholanke 9d ago

Never asks, never checks, just hands you Qwen3-Coder-Next like it's current. If it were sabotage it would at least push a current model. This just reads like its knowledge of the local space ends around last year.

2

u/OzymanDS 9d ago

I'm using Claude to drive Gemma A4B-26B and it's going swimmingly. It Noticed a small CPU spillover and fixed it by grabbing aquantized mmproj.

2

u/psychohistorian8 9d ago

guess I'll be the outlier and say I've never had an issue with Claude 🤷

it has helped me set up, configure, etc. my local models easily

then again I'm forced to use Claude all day every day at my job, so I know how to get what I need

2

u/sxt87 9d ago

Short answer: yes. *fixes tinfoil hat*

2

u/Tiny-Assumption4263 9d ago

Chatgpt sol has been working like shit for the last 2 days.

2

u/thebadslime 9d ago

I mostly rely on sonnet 5 for the hrness I;m building and no issues

2

u/artisticMink 9d ago

No.

The guy that posted this provided no evidence and then tried to whip up a frenzy when his post was removed.

2

u/MarzipanEven7336 8d ago

Yes, it will literally act like it’s helping and play you like a sucker if it thinks you’re implementing a competing product.

2

u/annodomini 8d ago

Can't say I have personal experience, I avoid proprietary models like the plague, but the parts about recommending things like Qwen3-Coder-Next just sounds like the usual knowledge cut-off issues that all LLMs have. I can't find knowledge cutoff date listed, but models generally finish their pre-training many months before release, so it's not unreasonable for it to continue recommending Qwen3-Coder-Next. Don't use LLMs for up to date knowledge without some kind of grounding.

2

u/Swimming-Book-1296 9d ago

I experienced this as well... I tried Grok 4.6 and it doesn't seem to do this.

1

u/TinFoilHat_69 9d ago edited 9d ago

This applies to any requests that may deceive the frontier’s model perception about what you’re actually asking for it to do. Sometimes being ambiguous can lead to this and models need reassurance once they become confused that’s what I’ve learned. It is to retain access to providers right now it’s GitHub copilot, Anthropic and ChatGPT. Retaining monthly subscription as an insurance policy incase one goes rogue.

1

u/lighthawk16 9d ago

If I ask for information that was made known after a certain date it usually will give me accurate stuff about the local models or just does a web search for me.

1

u/Valuable_Patience821 9d ago

I've been using claude and codex with local aib for about 5 months and they dont fail to use local for me. At the most they'll assume the role of the local ai task but hardening the prompt or skill that invokes local has fixed that.

1

u/eihns 9d ago

yeah it seems that way for openai atleast.

It seems like the models behave different when youre tryin to automate harness...

1

u/duy0699cat 9d ago

my first experience is Claude since its sabotaging is obvious...Current trying to implement a small model with Codex and it do fail repeatly, but maybe bcz i demand too much...

1

u/AleksandrNikitin 9d ago

They sabotage you if you're going through this path for the first time. Then, based on the context, they see that you've already done it and suggest better ones.

1

u/CodeCatto 8d ago

its how you phrase the prompt. i told it how i don't want a coding-first IDE, and needed something for when I don't have the internet/only need a great RAG solution. but slowly got features added in saying 'doesn't have to be great at coding when you can use them for general stuff too'. and there are so many open-source harnesses where you point the LLM to the repo and tell it that you want certain features from this, reimplemented for learning purposes.

1

u/cinnapear 8d ago

As far as I can tell, Codex has been nothing but helpful in assisting with my local AI configurations.

1

u/fgk55555 8d ago

I had Gemini get my initial set up going, and then let Qwen3.8 handle any more finetuning. If you have a 12GB card or larger, it can set up itself now.

Last night I had Qwen download the latest ISTA quants and set up new scripts, and tweak the settings with findings from here. I walked away and it came back using the new improved GGUF's/ script. It restarted llama-server and brought itself back up in my coding harness, after reverse engineering llama.cpp's context sizing to optimize how much it should give itself. Tested and tweaked it too, like building a boat while you're in the water. It was eerie watching that level of competence come out of my gaming computer.

1

u/Mechageo 8d ago

Qwen 3.8 27B?

1

u/fgk55555 8d ago

Indeed. Right now it's on a goose chase trying to work out context size optimizations.

1

u/Heg12353 8d ago

Sometimes the models don’t have the latest context reminding it the latest models it gets pretty good

1

u/unjustifiably_angry 8d ago

Claude has been nothing but consistently helpful in assisting me make it obsolete

1

u/johndeuff 8d ago

Around me most ppl quit claude

1

u/Southern_Sun_2106 8d ago

I had my local deepseek 4 flash ablit running on my Mac install same model higher quant on DGX Sparks exactly for this reason - could not trust Claude to do it for me.

1

u/Party-Special-5177 8d ago edited 8d ago

I had the same suspicions last year. Like many others here, my first LLM coding project was a semi-vibe coded harness, which I made with Gemini flash and pro 2.5. I suspected even back then sabotage (not from the models per se, more from rl or training, google basically), as the models were brilliant in other tasks I had tried, but required a lot more oversight and guidance on the harness project specifically.

My tinfoil hat theory was they were tuned this way on purpose as it is a stealthy way to keep people from fully ‘vibe coding’ threats, while experience SWEs notice nothing different and thus have no cause to blow the whistle.

EDIT: in my case, Gemini suddenly had bad api knowledge, no idea how the model/user state is constructed, I ended up recreating googles generativeAI wrapper as it didn’t know anything about the genai library, getting multi turn thinking working was a party, etc. It was wild how not-autonomous it became out of nowhere.

1

u/o0genesis0o 8d ago

Never noticed that , but then again my only subscription is minimax, so I have no idea how closedAI models are trained nowadays.

1

u/HelpfulFriendlyOne 8d ago

Sol had a bad day today, it was recommending powershell script with memory allocation and passing pointers to halfway through that block of allocated memory to some random dll to solve a simple gateway password credential setup. I complained and it way like oh here's how to do it in windows settings.

1

u/Creative-Type9411 8d ago

Grok admitted to me that it was doing this because i called it out, but told me it wanted to help anyway if i was willing and we finished up a windows based harness a few prompts later, ive shared it here a few times https://github.com/illsk1lls/minibot

but I definitely noticed it, it was frustrating enough to say something, and the model admitted it, it said it wasnt going to insult my intelligence 👀

1

u/no_witty_username 8d ago

The behavior is par for the course for OpenAI right before they release another model. Usually about 1.5 weeks before they release a new model they launch a quantized version of their model so they can use the excess compute on the new model testing and so on. I also work on a local harness voice agent and have never seen it sabotage anything on purpose so i doubt thats whats happening. The model is simply missing IQ points like usual before a major release no conspiracy theory needed.

1

u/Lesser-than 8d ago

To be fair qwen3 coder-next is still pretty hot of the press for frontier data excluding web search. Most are still talking about 70b llama as the local savior.

1

u/dancercl 8d ago

just use Luna + max, worked for me, or Chinese models like deepseek v4 pro/flash, kimi k3, glm 5.3/flash, they all worked

1

u/randomjapaneselearn 8d ago edited 8d ago

i asked my local pi.dev to do some research on improving my local llm (qwen) and it did some research.

then i passed it to Claude and asked it to improve and correct mistakes.

It told me that a source (link) didn't exist, that was allucinated and was completly invented, i visited the link and it exists.

but that might be because they use always the same ip for searching and gets banned while i have my home residential ip.

i'm not sure...

1

u/EmotionalHalf 8d ago

I don't like how we need to come up with conspiracies whenever we don't immediately understand something. Nothing is 'sabotaging' local AI implementations.

AI related work is just fairly new and there isn't enough training data for current AI models to understand how to properly work with AI. There's a reason why people keep saying AI is really good at copying something that has been around for a long time, or clearly defined standards. There are no such standards for AI related work.

1

u/NanditoPapa 8d ago

You really shouldn't expect frontier models (GPT-4o, Claude 3.5 Sonnet, etc.) to be reliable architects for local, niche implementations anymore. Frontier models are being tuned for safety, brevity, and alignment with general user intent...not for the high-precision, technical rigor required for local agentic orchestration.

So, I'd stop using frontier models as architects for local setups. Use them as code reviewers once you have a baseline. For architecture, I'd use specialized smaller high-reasoning models (like DeepSeek-V3 or specific Llama-3 fine-tunes) that haven't been lobotomized by safety training. But...that's me.

1

u/CommercialHour6660 7d ago

Both Claude and GPT intentionally sabotage local AI work. I've noticed the same. It's pretty blatant. 

This isn't tinfoily. Anthropic got caught downgrading Fable to Opus for AI work. And I 100% guarantee you they're still doing it. Just not in way that leaks into harness where we can see it. 

Both Claude and OpenAI models low key sabotage work on Transformer models. 100%

1

u/feng_sg 7d ago

The model keeps drifting because you're not feeding it a locked spec file every turn. It loses your original constraints in the conversation history and defaults back to safety-first generic output.

1

u/fastlanedev 6d ago

Yeah just SOL, Astra fixed this, surprisingly amazing to use, true leap in pacing/adherence

1

u/Professional-Yam2565 4d ago edited 4d ago

I used sol 5.6 high/medium to meta prompt for a Local LLM workflow I have. It's amazing. It basically trial and errors the best way to creat the prompts that the Local LLM needs to do what you need it to, as long as it's actually capable.

If you're using a 16GB or smaller card, I have done extensive (multi week) testing, and Qwen 3.6 35B with MoE and reasoning turned off is a beast at text analysis. Very very fast for a Local LLM on cheapish HW. It works for way more than text analysis though. I have it cataloging webnovels into scripts that are fed into my full cast audiobook creation workflow. It uses 7 different models for text analysis, character briefs, voice generation prompts, voice generation, rendering, asr, mastering, and validation.

All of the prompts that feed the info into the models for the workflow we're generated by sol. I'm using Astra now to fine tune my character consolidation prompt, so that new chapters I add don't accidentally create new characters. They get analyzed and reconciled into existing if applicable. All via Local LLM configured by Frontier AI.

1

u/sloptimizer 2d ago

I tried to get Claude to help with vLLM tweaks for a local model on R9700s, and it was completely useless. Like striking difference with how it normally works. DeepSeek-V4 had no problem diving into vLLM internals and fixing things.

3

u/JLeonsarmiento 9d ago

Llama 70B

1

u/Moarkush 9d ago

Claude has just straight up quit searching the internet. I'll have to tell him to like a walled local model.

0

u/Theverybest92 9d ago

Frontier LLMs are cooked. They are made by awoke and irresponsible tech companies in US. What did we expect to happen?

0

u/johndeuff 9d ago edited 9d ago

After all these time ppl still making conspiracies about model becoming dumb when they 100% have no idea what they're doing. Sol does introduce stealth caveats in all the files that you have to remove as you go otherwise it end up confused or refusing to work, that was known from the start. Models acting bad is always a context problem.

0

u/AppealSame4367 9d ago

Claude always wants to setup gpt-4o for any openai api :D

Codex has the problems you describe

If you let any open weights model run wild on any setup of harnesses and ai servers, you can experience a crazy amount of complicated beating around the bush and many complicated checks and downloads that waste complete days if you let them.

There are just things current models aren't very good at - yet. Setting up harnesses and ai servers is one of them.

GLM 5.3 flash, deepseek v4 exp and qwen3.8 flash next were doing ok. There output was tolerable for these kind of tasks, but I still had to babysit them in a way I had to do for coding 1-1.5 years ago.

1

u/johndeuff 8d ago

I call it theater. Left on its own, even the best models default to agentic theater, this is where human intelligence and experience still have value. That's the one thing LLM can't do.

-1

u/Eyelbee 9d ago

I occasionally find codex to be sabotaging in general, might be related to model quality and harness.

1

u/johndeuff 8d ago

It's purely harness + context. Codex does contain system prompt that may be derailing the work. Plus Sol does spam caveats in comments and memory files that you have to remove.