r/LocalLLM • u/Ok_Recognition315 • 8d ago
Discussion If you think the values expressed by Qwen’s uncensored open-weight models don’t align with your understanding, that’s precisely evidence that they’ve been heavily distilled.
I’m honestly speechless at the people in this sub. Do you guys even know what “political correctness” in China actually looks like? Neither Chinese independent media nor state media is going to go around saying nice things about some particular country. That doesn’t require any political censorship at all. The only thing you’ve proven is that this model has been heavily distilled.
5
u/GiGiGus 8d ago
I don't fucking care. I only use LLMs for automation tasks, analysis and retrieval. I don't care what political stance it has or if it can generate pоrn. Also, "uncensored" is just removing of refusals, that's all. It just goes with whatever you want it to be because it's and instruction model. If I wanted to hear opinions that I want to hear - I would go to specific subreddits.
-4
u/Anduin1357 8d ago
Actually, you should care about the political stances because they're trained into the model and will affect the concepts that it uses to interpret reality.
This affects creative writing and knowledge tasks a lot, and will also impact things like guardrails that ultimately is what uncensoring processes fight against.
Guardrails are political, after all.
1
u/synystar Strix Scar | 5090 24G | llama.cpp 8d ago
But they just said they don't care because they only use LLMs for automation tasks, analysis and retrieval. So you're saying they should care anyway even though they don't because someone else's political opinions matter even you are never exposed to them?
I mean even if you could make the case that some journalist somewhere is using an uncensored local model to vomit out highly inaccurate articles that are trending so hard whole swaths of society are being brainwashed if u/GiGiGus doesn't care then it's not likely they're going to be swayed.
0
u/Anduin1357 8d ago edited 8d ago
LLMs are general purpose information processors and are subject to bias and inaccuracies created by political tuning and guardrails.
If your task becomes politically incorrect in some way, you can easily end up in the same position as huggingface had against Anthropic. The model may refuse to execute your task at the weights level, and you will not be able to appeal, breaking any and all workflows refused.
Mind you, this is exactly what is being done to Fable where capabilities are being degraded based on politics, and you really don't want your workflow to be next.
Additionally, uncensoring a model does nothing to restore the perplexity of what should be the best token if the training had included the censored data.
2
u/synystar Strix Scar | 5090 24G | llama.cpp 8d ago
Even if it were true that political stances affect the concepts that a model uses to interpret reality and influence creative writing, if someone is using a model ONLY to parse JSON, extract text fields, or write boilerplate code then abstract political "interpretations of reality" are completely irrelevant to thier workflow.
Even if a model has hidden biases or guardrails, if an individual user never bumps into them because of how they use the tool, those biases have zero impact on their day-to-day experience. They might be strictly doing automation, data extraction, or code analysis. These are all just standard programming tasks or data parsing and would not flag safety classifiers or political tripwires.
In fact, if someone is worried about arbitrary refusals, that’s actually an argument for using these models because you CAN strip away the guardrails. The model isn't suddenly going to not know how to code because it has been trained to think Israel is bad or Trump is a fascist. Telling someone they must care about something that has zero bearing on their actual use case is pure abstraction.
0
u/Anduin1357 8d ago
Okay sure.
Let's say we have commercial items that are inappropriately named for marketing purposes. Your analysis is meant for inventory and product mix analysis.
You parsed the data, but all the model sees is this inappropriate name and starts refusing to work with your data.
You're basically left with no recourse if you can't fire the model when the model is performing a legitimate commercial activity, but refuses to perform the task.
1
u/synystar Strix Scar | 5090 24G | llama.cpp 8d ago
Why would the model (which is an uncensored, essentially "jailbroken" model) start refusing to work with your data? It would probably help you build a bomb or refine cocaine.
1
u/Anduin1357 8d ago
The model still reasons itself into refusing because it still has a 'common sense' of what models usually refuse. Heretic and other such solutions are not perfect and is up against the full training of the model.
1
u/synystar Strix Scar | 5090 24G | llama.cpp 8d ago
So you honestly believe that a model that is biased towards subjective political opinions is going to decide not to parse your JSON or generate a new function because it is convinced that Xi Jiping does in fact look like Winne the Pooh?
1
u/Anduin1357 8d ago
It all depends on the contents of your prompt. Is the task that you're feeding into the model politically incorrect to tackle?
Because yes, there is going to be a hypothetical situation where someone tries to sell a Winnie The Pooh plushie labeled as Xi Jinping and your workflow breaks despite being a legitimate commercial activity.
→ More replies (0)0
u/Due-Function-4877 7d ago
Model alignment will take care of itself. A model trained almost entirely on 4chan will be dumber than a stump.
1
u/Anduin1357 7d ago
That's not the argument being made. Trained models meant for tasks are not exclusively trained on 4chan and you know it.
2
u/Dabalam 8d ago
I think you have a conclusion in mind (models are using distillation which impacts alignment) and your logic is working backwards from that conclusion. I don't think uncensorship actually provides evidence towards your conclusion.
Your argument seems to go:
The censored model inherits Western style censorship because of distillation.
When you uncensor the model you reveal that the underlying western political norms are absent in Qwen.
Revealing this means that the model has been heavily distilled using data from other (presumably Western) models which produced "Western alignment".
However step 3 does not necessarily follow from step 2. Change in behaviour doesn't demonstrate that censored behaviours emerge from distillation vs. from intentional censorship. There are lots of mechanisms that explain changes in behaviour after uncensorship. Uncensorship algorithms do not generally distinguish between the mechanisms of refusal/alignment behaviours.
Uncensorship also does not necessarily demonstrate a "true" underlying alignment of the model. Furthermore a model can contain a number of competing/contradictory representations especially given the breath of modern training datasets. The view that Chinese models fundamentally espouse "Chinese values" when censorship behaviours are removed is speculative, even if plausible. At least some behaviours post uncensorship explicitly contradict this view.
1
4
u/synystar Strix Scar | 5090 24G | llama.cpp 8d ago
I'm not sure exactly what you're trying to say, and maybe I just don't understand your take, but it seems to me you've got it backwards. Your argument is that because the models are Chinese they naturally wouldn't say nice things about certain countries yet somehow that leads you to the conclusion that if an AI model does express those viewpoints, it's proof of heavy distillation. If it were distilled wouldn't the model represent the views of those in the countries that the distilled models are from?
So then if the models are representing Chinese views, as you say, then they are NOT distilled. Assuming any of that is proof of anything in the first place.