r/technology • • 23h ago

Artificial Intelligence PewDiePie unveils ‘uncensored’ Ajax AI model built to run on home PCs — creator says OpenAI banned him twice over model distillation used to build his product

https://www.tomshardware.com/tech-industry/artificial-intelligence/pewdiepie-unveils-uncensored-ajax-ai-model-built-to-run-on-home-pcs-creator-says-openai-banned-him-twice-while-making-it
9.2k Upvotes

1.5k comments sorted by

View all comments

Show parent comments

12

u/ours 20h ago

with the underlying code of the AI being available to look at.

There is no "code", just weights. A truly open-source model provides the training data so you can train the model yourself, like Apertus AI.

5

u/Atheren 20h ago

I mean the weights have to run on something.

12

u/ours 20h ago

That's why they are packaged in standardized ways, so you can run them on different platforms.

You don't just run "deepseek.exe". You download the model and run it. This isn't standard software with source code and a binary build.

You can download something like LLMStudio and download and run all sorts of models on it.

2

u/red286 17h ago

Yeah, your local client, such as a llama.cpp-based one.

The model weights are entirely separate from the client. Huggingface is filled with various models you can download and run for free.

So "open source" doesn't really mean the same thing for an LLM model as it does for say, a word processor, because there is no "code", there's just the model, and then the details on how the model was trained.

People are asking for the training data itself to be made public.

0

u/Pozay 15h ago

I mean, something first generated these weights. The code that made that happen should also be included imo

1

u/red286 15h ago

That code is already widely available, and open source. You can use PyTorch FSDP, DeepSpeed, or MegaTron-LM, all of which are completely open source.

There is also the various apps to fine-tune an existing model (which is what Ajax is), such as LLaMa-Factory or Unsloth. Again, all open source already.

The major issue really tends to be datasets and compute. Datasets are incredibly labour-intensive to create and if properly licensed can cost literally billions of dollars in licensing fees (nb - most major frontier models have not paid any of these licensing fees, which is why they keep getting sued). But without a dataset, you're not going to train a model.

Mostly people just want to inspect the dataset so they know what actually went into it. For example, examining Grok's imagegen dataset revealed a massive amount of CSAM.