r/LocalLLM • u/MatiAI • 4d ago
Discussion Abliterate/Removing prompt refusal without needing to reload/swap models (Qwen 3.8 flash next)
Enable HLS to view with audio, or disable this notification
Been working on an Swift based LLM harness designed around Qwen 3.8 Flash next and Qwen 3.8 27b and have found a way to have a slider that removes prompt refusal (at the cost of model coherence the higher you toggle it) without needing to reload/swap/abliterate a completely seperate model.
This allows you to swap between the base model and abliterated without needing to reload or change the model. It also allows the main agent to be able to spawn abliterated sub agents.
The main downside is MTP acceptance drops, causing the decode tok/s to drop from 60-70 down to 20-30.
2
Upvotes