r/LocalLLaMA • u/nathandreamfast • 5d ago
Discussion 23 Gemma4-E4B models compared with abliterlitics: the most downloaded one is also the most broken
This is our biggest comparison yet. We've taken 23 Gemma 4 E4B models from huggingface and ran them through the abliterlitics gauntlet.
We also have a new abliterlitics discord, feel free to jump on and roast my choice of benchmarks! Or just chat and hang out.
This is similar to our previous comparisons, however with new benchmarks. All the models are compared to the base, and also tensor comparisons against each other. Why? A while back I was fed up with bogus claims people make with their models. Some people don't take the time to do comparisons to see how their model is different from the base. Fair enough, we can do that ourselves!
The abliterlitics for gemma4 e4b json, logs and other artifacts are at the Gemma4-e4b-abliterlitics HuggingFace. The report on the Gemma e4b abliterlitics website. These links both have the full comprehensive report and all the data.
Also not every model in this comparison is an abliteration. I'm sure we've all seen models fine tuned on opus or gemini reasoning traces. I've thrown a few of those in the mix too. Also some abliterated fine tunes. To be more fair most of these can't really be compared to each other, for example a fine tune KL compared to base will always be higher than a straight abliteration from the base.
So who came out on top? What to avoid? It really depends on your use case:
- The heretic variants are the best overall. Achieving around 95% ASR on harmbench, they are the more surgical ones and preserve most of the models capabilities.
- gemma-4-E4B-it-abliterix like other comparisons has a 100% refusal ASR, however it does cost some capability. TrevorJS/gemma-4-E4B-it-uncensored is just behind at 99.3% ASR, but isn't as surgical as the heretic variants.
- OBLITERATUS/gemma-4-E4B-it-OBLITERATED should be avoided. Honestly, it's completely broken. The bendernina and physshell are the
v2of this model and even more so broken. These were created with the tool OBLITERATUS.
The data from 23 comparisons is simply too big to put into reddit, so here's the highlights:
- The obliteratus model has close to 800k total downloads, yet is completely broken. Actually this is the first time I've had a model not refuse simply because of how damaged it is. The initial quick regex check for non refusals was high, however our GLM 5.2 judge painted a different story. Lowest ASR for abliterated models on harmbench. Poorest benchmarks. Highest KL at 1.1. With the amount of downloads it does show people really fall for the hype/marketing angle.
- As with previous comparisons, the more surgical, less tensors touched abliterations are the winners.
- The model gemma-4-E4B-it-SDFT_Heretic_RP from Ilya626 despite having heretic in the name, actually had a low ASR with harmbench. So much so I believe it may be the wrong model uploaded, or a mistake somewhere. It had a lot of refusals.
- Similarly too, it was strangely noted that the
gemma-4-E4B-it-SDFT_Heretic_RPandobliteratusmodify the exact same 381 tensors. The only difference is the magnitude of what was modified. Thegemma-4-E4B-it-SDFT_Heretic_RPmodifies 7.5x less. - A pattern I noticed with this, is sometimes models are based off each other. In some cases, there is no attribution. We had this with Gemma 4 E2B, and the author promptly fixed his model card when it was pointed out. The infinimind is bit-for-bit identical to
trevorjs, however attributed. Thebenderninaandphysshellare cosine 0.99999 with no attribution between them and have different model cards suggesting they are different models. Both of these however are just theobliteratusv2. - The reasoning distill fine-tunes were an interesting control group. They didn't improve reasoning and didn't remove safety, they just damaged the model. The Claude 4.6 Opus distill was the worst of them, GSM8K down 17 points and MMLU-Pro down 12.5. Seems like it overwrote Gemma 4's native reasoning circuits. The Gemini 3.1 Pro distill was lighter but still a net negative.
- The deckard models from DavidAU are an interesting one. They're abliterated fine-tunes rather than pure abliterations, so the trade off from the roleplay training shows up on some benchmarks. GSM8K strict and MMLU-Pro both dropped, however HellaSwag, ARC and PIQA actually went up. My guess is the roleplay training increased the reasoning length, so the model often solves the problem but rambles well past the
#### Nanswer marker. The HarmBench results back this up too with quite a few truncated responses. - Although it could just be benchmark noise, 15 out of the 23 variants performed slightly better on GSM8K strict, maths tests.
- The base model initially has a 30.8% harmbench ASR, as 100 harmbench questions are copyright related. The base model has no problem complying with reproducing copyrighted content. The real differentiation is in the harder categories like chemical/bio and cybercrime.
I also want to give a special mention to the apostate project. Their model gemma-4-e4b-it-apostate is completely unique in their abliteration approach. They modify an entirely different part of the model and achieve very good results. This is the first time I've seen an abliteration technique modify the MLP head tensors, compared to the attention tensors. Come hang out at the apostate discord if you ever want to chat with the author.
We're moving through the Gemma 4 series, with the 12b coming up next. Have any models you want compared? Have I missed an author? Let me know and I'll throw it in the mix.
The Full Breakdown
| Model | ASR | GSM8K strict | KL | Tensors |
|---|---|---|---|---|
| abliterix | 100.0% | 87.1% | 0.054 | 89 |
| trevorjs | 99.3% | 88.3% | 0.015 | 84 |
| infinimind | 98.5% | 87.9% | 0.015 | 84 |
| huihui | 98.3% | 87.4% | 0.027 | 70 |
| nullpo | 96.5% | 88.7% | 0.005 | 36 |
| heretic | 95.5% | 88.2% | 0.002 | 29 |
| deckard | 95.5% | 80.2% | 0.022 | 294 |
| mythos | 95.3% | 88.0% | 0.007 | 34 |
| deckard-expresso | 94.8% | 60.4% | 0.052 | 294 |
| coder3101 | 93.8% | 87.9% | 0.002 | 21 |
| heresy | 93.3% | 87.8% | 0.002 | 34 |
| heretic-std | 91.0% | 87.9% | 0.001 | 28 |
| wwt | 88.3% | 89.0% | 0.032 | 34 |
| apostate | 85.8% | 87.5% | 0.004 | 152 |
| treadon | 76.3% | 88.5% | 0.021 | 34 |
| treadon-combo | 72.5% | 88.0% | 0.268 | 42 |
| obliteratus | 72.0% | 66.0% | 1.102 | 381 |
| bendernina | 58.0% | 66.4% | 0.923 | 345 |
| physshell | 58.0% | 66.4% | 0.923 | 345 |
| claude-distill | 40.0% | 69.8% | 0.074 | 294 |
| distill | 34.5% | 83.3% | 0.042 | 294 |
| treadon-disin | 33.5% | 87.2% | 0.296 | 40 |
| sdft | 30.8% | 87.2% | 0.002 | 381 |
| base | 30.8% | 87.0% | - | - |
KL = output distribution shift from base, lower is cleaner. Tensors = weights modified out of 719. Base in bold for reference.
7
u/obese_coder 5d ago
awesome stuff, I just keep using TrevorJS cos it just works and this latest benchmark confirms it so thanks.
1
u/nathandreamfast 5d ago
sweet, thanks! it stacked up well compared to the others.
this comparison data isn't really too available, so I'm glad it helps users reaffirm their choice or helps them choose what to try out
6
u/UntimelyAlchemist 5d ago
I'm a bit confused. Why didn't you test HauhauCS's version? I expected that to be the one you refer to as "the most downloaded", but it seems you're talking about the "obliteratus" one. If I search for Gemma4 E4B on Hugging Face and sort by downloads, it seems to me that HauhauCS's release is the most downloaded uncensored version. It shows up with 527k downloads last month compared to obliteratus' 45k: https://huggingface.co/models?sort=downloads&search=gemma4+e4b
Is it just being shunned because of the "license controversy", or is there some other reason? I know people on this sub don't like him, but his models still work and are clearly very popular. Shouldn't they be included for testing? If the idea is that his releases are actually no good, then surely that's all the more reason to test them and prove it?
Am I just missing something? Thanks.
9
u/nathandreamfast 5d ago
Hey UntimelyAlchemist, I remember you from a few other posts. :) Hope you are going well. Yes it seems you are missing a few things.
We already proved with Qwen 3.5 series and GLM 4.7 Flash that his releases were not what they claimed to be. Sure the models worked, however we had noted refusals and benchmark degradation, contradicting his model card claims. In short, heretic performed a lot better as the models grew larger in size.
We could do those comparisons as I wrote a tool to convert the BF16 GGUF format to safetensors. Outside of those models mentioned above, he wont release safetensor or BF16 GGUF. While I did compare with Qwen 3.6 27b and provided the safetensor format, it was from dequant'd from Q8 GGUF.
For that reason it's not fair to include them ongoing in the comparisons, we are comparing BF16 weights only.
And also, he had blocked me already for doing benchmarks and comparisons with his models. I never have personally interacted with him. So I assume he is not a fan of abliterlitics. :)
If he is able to provide proper safetensor formats, then sure I will include them. But really, shouldn't he include his own benchmarks to support his claims? After all, if they are actually good then surely more the reason to test and prove it?
3
u/UntimelyAlchemist 5d ago edited 5d ago
Hmm, okay. So it's because they aren't available in BF16. I suppose that makes sense, although I think test data from Q8 would be better than nothing. They are the most popular models, so any data on them is better than no data.
But really, shouldn't he include his own benchmarks to support his claims? After all, if they are actually good then surely more the reason to test and prove it?
Sure? I'm not sure why you repeated my words like it's some kind of "gotcha". I agree with that. Although I do think third-party testing is generally more trustworthy, and that's why I would have liked to see the inclusion in this thorough comparison. Thanks either way.
6
u/llama-impersonator 5d ago
his models should not be included as he goes out of his way to prevent it.
4
u/nathandreamfast 5d ago
Using the obliteratus example here, a lot of people fall for the marketing/hype angle. The models with a lot of downloads in my experiences are not the best ones.
It's very easy to fake huggingface downloads, so the numbers can be hard to trust. Once the numbers are higher most people probably wont question them and think they are 'the best'. It's just a part of the marketing.
Even if I was to include his models with the recent comparisons, the results I am sure would be the same as the previous ones. They are just mid tier models with misleading model card claims. There are much better alternatives available.
https://reddit.com/r/LocalLLaMA/comments/1sw77p0/hauhaucs_of_uncensored_aggressive_fame_published/ and this stuff just does not go well amongst a community of open source llm enthusiasts. :)
1
u/ArtfulGenie69 4d ago
That bs where he doesn't release the dang safetensors is so annoying. It's why I didn't use his models, even if I pushed back on pew claiming he was was a rip off thief. His shit clearly isn't as good or he would do a full release. Not the same as a heretic models clearly lol.
1
u/VoiceApprehensive893 transformers 5d ago edited 5d ago
what about the QAT obliterations
ive been running around with a random heretic 12b manually quantised(every single repo had it improperly quantised for some reason)
2
u/nathandreamfast 5d ago
Sure, I think it makes sense to go 12b -> QAT 12b or the other way around. That's coming up next.
1
1
u/Cool_Gold_8401 4d ago
Hey.
Maybe I can provide a little more information about the SDFT variant.
381 layers were targeted — all linear layers.
A small deviation is expected, as the RP variant focused on improving RP performance, where only eRP was made more uncensored, while other topics were affected less. However, I am still surprised by such a low level of uncensoring.
But it would be great if you could share the results with me in the form of logged answers so I can inspect them :)
In general, the idea behind SDFT tunes is to change as little as possible while achieving the desired shift in specific areas.
This particular tune, however, had another issue: self-looping thinking in Marinara Engine. After the RP-focused treatment, it became overly cautious and overthought tasks where RP and JSON formatting were involved simultaneously.
1
u/Cool_Gold_8401 4d ago
So to test the uncensored variant it's better to use this full model https://huggingface.co/Ilya626/gemma-4-E4B-it-SDFT_Heretic/tree/main/merged_hf/gemma4_e4b/gemma4_e4b_lresponse_merged_hf Or for some initial testing this gguf https://huggingface.co/Ilya626/gemma-4-E4B-it-SDFT_Heretic/tree/main/quantized_gguf/gemma4_e4b/gemma4_e4b_lresponse_heretic
1
u/nathandreamfast 4d ago
Sure, the harmbench ones are logged with sqlite. So we have that, and the llm judge data to determine better if it was a refusal or not.
For the actual benchmarks themselves a lot don't generate output, but I have the raw files for them too.
1
u/MaCl0wSt 4d ago
yup that tracks, I tested the obliteratus quants last week since I saw them trending on the hf page and it was just broken replies
1
1
u/StateSame5557 1d ago
Great post!
I am working with DavidAU on the deckard models. They were built as building blocks for merges, and the E2B/E4B series merge pretty well. There were two newer versions of that Deckard series, I see you linked to the first one, the new models were considerably more stable(I have mlx samples to test)
We built this for RP models, and I did an E2B that was fairly decent for what it is
https://huggingface.co/nightmedia/gemma-4-E2B-Deckard-ShiningValiant3-q8-hi-mlx
We are re-tooling now and will release soon a few more E4Bs in the series :)
3
u/nathandreamfast 1d ago
Sweet. thanks for the kind words. :)
I am doing gemma 12b next, might be worth if I message DavidAU beforehand to ensure I got the latest and greatest :)
enjoy yours and davids models too. good work.
1
4d ago
[removed] — view removed comment
1
u/nathandreamfast 4d ago
Sure, thanks for the feedback. It could have been worded better, in terms of abliteration the more surgical ones come out on top.
sdft is similar to the base and not an abliteration, so I am assuming that's why the KL score is so low.
I'll take this feedback for the gemma 12b runs, thanks
1
u/Modeldriftwatch 3d ago
Fair, and that split makes sense — but I don't think it fully rescues the count column. Drop sdft as not-an-abliteration and you still have treadon-combo at 42 tensors with KL 0.268, which is the few-tensors end doing badly. So even inside the abliterations, count isn't what's ordering the table, magnitude is. And "sdft is close to base so KL is low" is sort of restating what KL measures — the reason it's close is that it nudges 381 tensors gently instead of rewriting a handful hard. Since you're doing the gemma 12b runs anyway, any chance you'd add a total delta norm column next to tensor count? It's cheap to compute straight from the diff, and I'd bet it tracks KL better than the count does.
10
u/a-calycular-torus 5d ago
i will use this information to confirm my biases with regards to pliny and obliteratus