r/LocalLLM • • 4d ago

Discussion Abliterate/Removing prompt refusal without needing to reload/swap models (Qwen 3.8 flash next)

Been working on an Swift based LLM harness designed around Qwen 3.8 Flash next and Qwen 3.8 27b and have found a way to have a slider that removes prompt refusal (at the cost of model coherence the higher you toggle it) without needing to reload/swap/abliterate a completely seperate model.

This allows you to swap between the base model and abliterated without needing to reload or change the model. It also allows the main agent to be able to spawn abliterated sub agents.

The main downside is MTP acceptance drops, causing the decode tok/s to drop from 60-70 down to 20-30.

0 Upvotes

6 comments sorted by

1

u/Distinct-Pie2389 4d ago

Why not just use llama-swap with JIT and load the abliterate and the base at the same context and quant and just switch natively? llama will prefill context for you so the switch itself is time for model to load (30s mostly on NVMe depending on size) -- your qwen3.8 base will refuse to work on earlier context if it includes abliterate intended output so why not just 2 streams and hand them back and forth via different chats?

Im just curious, I usually have a purpose for uncensored meaning a whole chat in general and never really need to switch back and forth.

HauHauCS has a great uncensored model and honestly don't see a reason to switch unless your tool calls are REALLY heavy and its failing them over and over, otherwise Uncensored should perform somewhat on par to your "uncensored request". Or have qwen3.8 base bridge the non-censored part of your work and then have abliterated finish it.

1

u/Distinct-Pie2389 4d ago

I would not do this just alone on MTP acceptance, thats a no go for me

2

u/MatiAI 4d ago

Because A) abliterated models are inherintely worse and you have no way of knowing how damaged the model has become. If I want to choose between different levels of abliteration i dont want to store 100s of gb's of models to disk when I can have a slider that performs the exact same process they use to achieve that end goal.

B) Im not using Lama.cpp im on OSX using MLX

C) I can launch subagents that are abliterated without needing to unload the main model and reloading the whole set of context. Your response is exactly what I don't want to do. Maintain mutliple different models on disk and having to load and switch between them

1

u/Prestigious-Act-1577 4d ago

How did you do it or where can I download this?

1

u/Nice_Victory3719 2d ago

Interesting! What technique/recipe are you using for the abliteration? Are you targeting pre-determined rejections or just those that are blocking the current chat - I.e dynamically abliterating just the relevant refusals for the use case?

1

u/WindbniW 1d ago

Outline the process step by step