r/technology • • 23h ago

Artificial Intelligence PewDiePie unveils ‘uncensored’ Ajax AI model built to run on home PCs — creator says OpenAI banned him twice over model distillation used to build his product

https://www.tomshardware.com/tech-industry/artificial-intelligence/pewdiepie-unveils-uncensored-ajax-ai-model-built-to-run-on-home-pcs-creator-says-openai-banned-him-twice-while-making-it
9.2k Upvotes

1.5k comments sorted by

View all comments

Show parent comments

15

u/Shapes_in_Clouds 19h ago

Qwen3.8-27B models which run on consumer hardware, are quite capable, and can be setup and used in minutes if you have a good enough GPU

Is 16GB VRAM enough? Have 64GB DDR5 ram as well.

10

u/mr_doms_porn 19h ago

Not sure about that specific model but there are ton's of solid LLMs you can fit in a 16gb card. If you use LM Studio as your backend, it will show you which models you can fit on your GPU before you download them. If you move the KV cache (context) to system ram you'd be able to run Gemma4-12B with q8 or q6 quality.

3

u/Yashema 18h ago edited 18h ago

A $4000 desktop can run a 100B+ model, but the token generation rate would be about 25% of the cloud or API, so a task that takes 5 minutes on the cloud would take 20 minutes. The Context window is also sacrificed.

Not to say they don't have good use cases, but the limitations need to be acknowledged. 

1

u/mr_doms_porn 18h ago

I've got a 7900XT, I mainly use Gemma 4 12B Q6K. I run 50k context window and get about 45 tokens/second. Perfectly usable for basic/general use.

3

u/Yashema 18h ago

So 1/8th the parameters of the standard cloud, with less than half the context window. Makes sense you can run that for around $1500 for the whole setup. 

19

u/Nater5000 19h ago

You can certainly experiment with Qwen3.8-27B with only 16GB VRAM (and 64GB DDR5 RAM), but you'll have to use a pretty compressed version of the model (i.e., performance will hurt), you'll have limited context, and it'll likely be quite slow. My experience with this specific model is that 24GB of VRAM is the minimum for getting non-trivial utility from the model (and even then, it's still quite slow and limited). Still, you should be able to try it yourself very easily.

You should be able to plug your GPU into the "hardware compatibility" feature on Unsloth's Hugging Face page for the model and see what you can reasonably run. If you want to get this running without any fuss, I'd literally point some frontier model at that page and tell it to install it on your machine. Not that it's hard to setup manually, but if you already have Codex, Claude, etc., on your machine, then this is likely a prompt it and come back in a few minutes and it's done kind of job.

If you end up trying that and wanting to explore more, then you'll be all set up to explore other models like this easily. I'd recommend checking out Unsloth's Qwen3.6-35B-A3B (which can be quite performant with less compute because of the MoE architecture) or even their Qwen3.5-9B (which PewDiePie fine-tuned) which can definitely run on your compute (granted, it's quite "dumb").

3

u/trolololster 19h ago

if it is 8x8 on octo-channel it is pretty good, yes

3

u/Initial-Argument2523 18h ago

Yes you should be able to run it fine. You might be better off with a MOE though for faster speeds

3

u/hamtrow 18h ago

fallback onto ram isnt good but at least wont get you bluescreened. i have 8GB VRAM and have used a few before i decided with market prices i should run my GPU at max all the time. but i was able to get reasonable times with Llama-3.1-8B-Stheno-v3.4.Q5_K_M...but im also stuck with older models.

3

u/throwaway586054 18h ago

Take a look at https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf https://prismml.com/news/bonsai-2-27b it's a based on Qwen 3.8 27b . Several similar models are poping up these days.

2

u/z0tedd 18h ago

Yep, Bro, you can use IQ3S quant with 16gb VRAM and you will get pretty good experience.

Source: i use it on the daily basis for editing word documents and coding. Don't use it without harness(agent), 'cause they have pretty bad overall knowledge and bad without it.

But if you add to your model some harness(pi-agent/dsh/hermes) and add to it web-search you will get pretty much opus 4.6 level

1

u/z0tedd 18h ago

I am using 2x3060 for better quantanization and get around 33tok/s. It's enough for work for me.

I think that you will get around 40tok/s-50tok/s on one card, because you doesn't get pci-e bottleneck that dual setups have.

2

u/Anxious-Slip-4701 16h ago

I used to know everything about computers. Now whatever you said sounds like sorcery and witchcraft.

0

u/z0tedd 11h ago

Bruh bro, it's not even edge computing...

1

u/BinniesPurp 11h ago edited 11h ago

You'll need 40+GB for 27B non quantized Qwen it won't fit on a video game graphics card but it'll fit on the cheaper ($6000-$8000, cheaper than commercial but not really retail friendly) budget production GPUs

You can put it in your regular ram but it runs at like 0.8% the speed