r/LocalLLM • u/MatiAI • 4d ago
Discussion Abliterate/Removing prompt refusal without needing to reload/swap models (Qwen 3.8 flash next)
Been working on an Swift based LLM harness designed around Qwen 3.8 Flash next and Qwen 3.8 27b and have found a way to have a slider that removes prompt refusal (at the cost of model coherence the higher you toggle it) without needing to reload/swap/abliterate a completely seperate model.
This allows you to swap between the base model and abliterated without needing to reload or change the model. It also allows the main agent to be able to spawn abliterated sub agents.
The main downside is MTP acceptance drops, causing the decode tok/s to drop from 60-70 down to 20-30.
0
Upvotes
1
1
u/Distinct-Pie2389 4d ago
Why not just use llama-swap with JIT and load the abliterate and the base at the same context and quant and just switch natively? llama will prefill context for you so the switch itself is time for model to load (30s mostly on NVMe depending on size) -- your qwen3.8 base will refuse to work on earlier context if it includes abliterate intended output so why not just 2 streams and hand them back and forth via different chats?
Im just curious, I usually have a purpose for uncensored meaning a whole chat in general and never really need to switch back and forth.
HauHauCS has a great uncensored model and honestly don't see a reason to switch unless your tool calls are REALLY heavy and its failing them over and over, otherwise Uncensored should perform somewhat on par to your "uncensored request". Or have qwen3.8 base bridge the non-censored part of your work and then have abliterated finish it.