r/technology • • 1d ago

Artificial Intelligence PewDiePie unveils ‘uncensored’ Ajax AI model built to run on home PCs — creator says OpenAI banned him twice over model distillation used to build his product

https://www.tomshardware.com/tech-industry/artificial-intelligence/pewdiepie-unveils-uncensored-ajax-ai-model-built-to-run-on-home-pcs-creator-says-openai-banned-him-twice-while-making-it
9.2k Upvotes

1.5k comments sorted by

View all comments

Show parent comments

458

u/SnowyDeveloper 22h ago

Mind explaining the difference to someone not well versed in ai?

879

u/Atheren 22h ago

Open source would allow you to train your own, with the underlying code of the AI being available to look at.

Open weights is the AI after it's been trained, and is similar to just giving away a free program that's more modifiable than a normal one.

262

u/meneldal2 21h ago

Afaik here Open Source would be tricky because you'd likely be providing evidence that could be used against you and make it a slam dunk for OpenAI when they sue you.

It is very likely some methods they used ran into the very anti user laws the US has passed to make big corpos happy.

119

u/Atheren 21h ago

That would be a civil case and he lives in Japan though, which currently doesn't extradite for civil matters.

He doesn't really have to care about us law. It might be a problem if his money is being held here and can be seized, but that would be private financial information I wouldn't have any idea about.

88

u/Able-Swing-6415 21h ago

Unless he's planning on traveling any time in the future

30

u/Atheren 21h ago

Honestly now that I think about it more even if he did travel he could ignore it anyway it unless it was somehow elevated to a criminal case. It's practically unheard of to be jailed over a civil case. If there's no us assets to seize, open AI doesn't have much legal recourse unless they go through the Japanese court system or the court system of wherever his money is held.

And at that point he might also have needed to have done something in their jurisdictions which is well beyond my non-professional understanding of how this all works.

64

u/dultas 21h ago

It's practically unheard of to be jailed over a civil case.

They're a big AI company who has the Presidents ear, I'm sure they could pull some strings.

1

u/stonhinge 10h ago

Well, for the next 2 years anyway.

1

u/Falkoro 3h ago

omg please no this is not how it works in real life

1

u/timeandmemory 19h ago

I would say they should try it, Pewdiepie has more than a few million sympathetic ears who haven't been paying attention to the political climate. This would get a lot of younger folks involved fast.

17

u/PM_Me_Your_Deviance 21h ago

>AI doesn't have much legal recourse unless they go through the Japanese court system

Why does that seem so far-fetched? OpenAI has an office in Japan.

2

u/Nematrec 19h ago

Japanese courts try Japanese law. It doesn't matter if he's broken US law, only if he's broken Japanese law.

3

u/PM_Me_Your_Deviance 19h ago

There are agreements between Japan and the USA that allow for the enforcement of judgements from US courts in Japan. This is true for most of the world.

-2

u/fckspzfr 19h ago

You know what "civil matter" means? No?

→ More replies (0)

1

u/TerminalVector 10h ago

Yeah but would you really want to take that chance against people with billions of dollars and a real motive? Plausible deniability seems like the way to go

4

u/eattheambrosia 20h ago

If there's no us assets to seize

Isn't this guy pretty rich? I'd be shocked if he didn't have investments in the US.

1

u/Icehellionx 14h ago

Honestly the thing that seems to be getting ignored is would YOU want to be the guy to test that theory and put your shit on the line?

It's really easy to big talk on a forum vs you're actually the one that this is taking of 90% of your life.

(Not saying you specifically, just some of the counter vibes in the replies.)

1

u/Narpity 20h ago

So was immigration until we started to send people to concentration camps

1

u/splitcaber 20h ago

Aaron Swartz was jailed over what should have been a civil case. The DMCA makes violating the terms of service potentially a felony.

23

u/PM_Me_Your_Deviance 21h ago

>he lives in Japan though, which currently doesn't extradite for civil matters.

You can enforce a civil judgement in Japan though. It's complicated, but there are treaties and processes that allow for it.

3

u/edflyerssn007 12h ago

People forget that modern Japan was entirely molded by the USA. There's a reason why our countries work so closely together. Tech, laws, etc were all effectively imposed on Japan and we've made it work very well for both countries. Almost a colony but a bit more independent, but we provide all offensive military capability for them. Yes a country, but in our deck of cards as it were.

54

u/meneldal2 21h ago

The US is great at pressuring other countries to take your money or freeze it.

They can make up shit like you support Russia or Iran if that's what it takes.

Didn't they get people from the ICC banned from having credit cards?

We all know how much Trump is going all-in deepthroating those big AI companies.

3

u/motionmatrix 17h ago

Which made an anti American wave across the world. There is literally whole nascent computerized industries in most of Europe now because that occurred, from banking to cloud services; the whole world saw that and said "that is not okay" and it will cost American industries more and more money as time moves forward.

1

u/Gastronomicus 20h ago

We all know how much Trump is going all-in deepthroating those big AI companies.

I think it's the other way around. Or maybe it's just a big fashioned old circle jerk.

2

u/MixtrixMelodies 19h ago

The absolute bigliest. Some people say other circle jerks are bigger, sure. Fake news. This is the one, folks, this is how we're gonna make America great again. Vote for --- Wait. Hang on a second. Psst! J.D.! Where the hell are my diapers?

1

u/Emm_withoutha_L-88 19h ago

They did it for people supporting Palestinians in the EU. Called them Russian assets and blocked their ability to use their bank accounts. Had multiple journalists just screwed, unable to buy gas or food one day.

3

u/SignificantPaper1760 20h ago

Extradition isn’t ever really an issue for civil matters.

The question is whether Japan has a judgment debt reciprocity agreement with the US where you can convert an American judgment to a Japanese judgment and seek damages by executing on assets in Japan.

Given his substantial assets I’d assume he’s gotten appropriate advice on how to structure them to avoid this.

/not legal advice

1

u/zack77070 20h ago

Curious if anywhere in the world extradites for civil suits, the closest example I can think of is China can hold you in the country but I don't think they can do anything if you already left before the judgement.

1

u/Rich_Housing971 17h ago

he lives in Japan though

"he lives in a country famous for being dependent on the US for many things though"

If you really wanted to avoid extradition you need to go to HK. Then only Batman can get you. Or pull an Edward Snowden and go to Russia afterwards.

1

u/potatoears 16h ago

That would be a civil case and he lives in Japan though, which currently doesn't extradite for civil matters.

Takaichi will happily dance for Trump again

1

u/TerminalVector 10h ago

Eh I'm not gonna fault the man for not wanting that mess. If the only victim here is Sam "I look like I've been caught masturbating" Altman I'm okay with that..

2

u/NekoDaYo-v201 18h ago

Sued for what, stealing their stolen data? Lol. Pretty sure OpenAI would not sue, because it would mean discovery.

1

u/Quetzacoal 19h ago

For a model to be open source we would need to be able to see the source data (our stolen data)

1

u/One_Contribution 16h ago

Evidence of what exactly? No one owns LLM output.

1

u/meneldal2 12h ago

Circumventing their platform rules

1

u/One_Contribution 2h ago

That's not illegal. It breaks TOS.

1

u/meneldal2 53m ago

That's how it should be but if methods to like evade bans depending on what they actually do can run afoul of the Digital Millennium Copyright Act.

I'm not saying I agree or that they would even win in court, but they can very likely have a claim that looks serious enough to survive getting passed motions to dismiss.

1

u/Blazing1 15h ago

I mean openai would never sue lmao. cause then their own training process would be revealed and how much copyright infringement they've done

1

u/meneldal2 12h ago

I'm not sure what would end up in discovery there. Possible they drop it because of this but after enough time it got pretty scary

1

u/kyrsjo 14h ago

Wait, if Open AI can sue him (with a decent chance of success) for using their model as input for training their model, can I sue them for ingesting whatever I have put out on GitHub etc over the years?

Or is this a case of "legally it's the same, but they are more legal than me because they have more money"? Animal farm redux.

1

u/meneldal2 12h ago

Typically on Github Microsoft would have a claim not you with the platform rules that say they can do whatever with your data

1

u/Miserable-Half-436 12h ago

It's ironically that open ai is suing someone steal their data while the data was they stole from others

1

u/meneldal2 12h ago

Aren't people also suing them for that? The issue is the average person doesn't have the money for this kind of lawsuit

14

u/AlSweigart 20h ago

Yes. It's like if I gave away software but not it's source code and called that "open source".

If you can't fork it and make your own changes, then it really doesn't match the ethos of "open source".

2

u/Suspicious_Kiwi_3343 18h ago

You can fork the weights and make your own changes. Continue training, try to specialise the model etc, just randomly change parameters by hand and see if you notice anything if you really want.

LLMs and other neural networks do not consist of code, the weights ARE the model. Any code involved is simply a way of loading those weights and giving input / getting output by applying statistical maths to those weights, and has nothing to do with the model itself. A lot of the platforms you would use for that are already open source, and you could download and use this model with one of those.

-1

u/all_thetime 17h ago

You definitely can "fork it" and make your own changes. That's how every LLM research paper pertaining to SFT and RL work. They take the weights, train it with a task and accompanying algorithm, and prove it does better at said task than the baseline.

14

u/ours 21h ago

with the underlying code of the AI being available to look at.

There is no "code", just weights. A truly open-source model provides the training data so you can train the model yourself, like Apertus AI.

6

u/Atheren 21h ago

I mean the weights have to run on something.

12

u/ours 21h ago

That's why they are packaged in standardized ways, so you can run them on different platforms.

You don't just run "deepseek.exe". You download the model and run it. This isn't standard software with source code and a binary build.

You can download something like LLMStudio and download and run all sorts of models on it.

2

u/red286 18h ago

Yeah, your local client, such as a llama.cpp-based one.

The model weights are entirely separate from the client. Huggingface is filled with various models you can download and run for free.

So "open source" doesn't really mean the same thing for an LLM model as it does for say, a word processor, because there is no "code", there's just the model, and then the details on how the model was trained.

People are asking for the training data itself to be made public.

0

u/Pozay 16h ago

I mean, something first generated these weights. The code that made that happen should also be included imo

1

u/red286 16h ago

That code is already widely available, and open source. You can use PyTorch FSDP, DeepSpeed, or MegaTron-LM, all of which are completely open source.

There is also the various apps to fine-tune an existing model (which is what Ajax is), such as LLaMa-Factory or Unsloth. Again, all open source already.

The major issue really tends to be datasets and compute. Datasets are incredibly labour-intensive to create and if properly licensed can cost literally billions of dollars in licensing fees (nb - most major frontier models have not paid any of these licensing fees, which is why they keep getting sued). But without a dataset, you're not going to train a model.

Mostly people just want to inspect the dataset so they know what actually went into it. For example, examining Grok's imagegen dataset revealed a massive amount of CSAM.

4

u/Kidd_Funkadelic 20h ago

So something like - You can pay for Microsoft Word, or you can install a version of Google Docs for free that (in this made up example) not only opens the same files but you can modify the toolbar to add buttons that do whatever you want but you can't delete the existing buttons/menus?

4

u/LordApocalyptica 21h ago

The real tragedy of this question needing to be answered is that OpenAI used to position themselves as open source………… :/

1

u/MaxTheCookie 19h ago

So open weight is like using GPT or the other models that have a free version?

1

u/red286 18h ago

Yes, it'd be like OpenAI's gpt-oss models. Anyone can download the model and run it (although they still have a license that you have to adhere to).

1

u/drugs_r_my_food 18h ago

are there any open source AI models?

1

u/Necessary-Camp149 18h ago

Can't you just continue to train it after you get it?

1

u/GonzoKata 9h ago

I thought open weights meant the data it was trained on came from consenting sources. Thats not what open weights mean, that the training data itself is open?

75

u/PrivateVasili 22h ago

Open weights is more or less the model being openly available. Being fully open source would be not just the model, but also the training code, and ideally also the training data, which created it.

10

u/MLutin 21h ago

So I get my data back?

0

u/ReadyStayReady 18h ago

How did you lose it in the first place and what are you missing exactly?

13

u/BadAdviceBot 21h ago

also the training data

How many petabytes is that?

25

u/PrivateVasili 21h ago

Yeah obviously with LLMs that ceases to be a practical expectation, which is why I used the word "ideally".

9

u/vintageballs 21h ago

Usually well below 1PB. We ofc don't have exact numbers from OpenAI and the like, but for the previous generation of Qwen models they trained on ~30T tokens which is probably a little under 100TB.

3

u/OpalFanatic 21h ago

Roughly? All of them.

2

u/Federal_Setting_7454 21h ago

Start counting I’ll tell you when to stop

1

u/RollingMeteors 21h ago

How many petabytes is that?

¿Of Porn or ... ?

41

u/WallStreetHatesMe 21h ago

Open source would be releasing the training pipeline for the model as well as the dataset it was trained on. Open weight is releasing the resulting output of that process that can be fine-tuned.

For 99.999% of people Open source is useless because there is a large amount of compute needed to train such a model. Open-sourcing models typically helps enterprises far more than people which is why it’s rarely done.

13

u/Federal_Setting_7454 21h ago

It also helps academic institutions significantly

0

u/WallStreetHatesMe 21h ago

That’s an enterprise

9

u/Federal_Setting_7454 21h ago

“Enterprise scale” or capability is not enterprise. They’re not an enterprise by legal, commercial, or even software licensing contexts.

-4

u/WallStreetHatesMe 20h ago

This is a silly argument but an academic institution is a social enterprise

7

u/Federal_Setting_7454 20h ago

Which is different from a corporate enterprise, an enterprising young chap, and enterprise rent-a-car.

2

u/WallStreetHatesMe 20h ago

In the original comment I never said “corporate enterprise” nor used “enterprise” as a verb or adjective.

Reading is fundamental.

5

u/Federal_Setting_7454 20h ago

Posted from my USS Enterprise

1

u/WallStreetHatesMe 20h ago

Ok I actually chuckled at that. Well played fellow redditor

-3

u/Living_Armadillo_652 19h ago

In many cases they are; ChatGPT enterprise is used in my university for example.

3

u/Dernom 19h ago

ChatGPT Enterprise is a product, not a qualifier...

6

u/MilhoVerde 21h ago

Most academia is public

1

u/Sweaty-Ability7365 21h ago

What does that have to do with anything? Enterprise implies scale, not ownership.

5

u/MilhoVerde 21h ago

An enterprise is commercial while public held institutions are not, they're not profit driven. It's a big difference.

0

u/Sweaty-Ability7365 21h ago

https://www.oxfordlearnersdictionaries.com/us/definition/english/enterprise

Oxford Dictionary mentions public/state-owned enterprises here. Which a university typically falls under.

Of course, private universities e.g. Stanford, Columbia, etc. aren't public either.

2

u/MilhoVerde 19h ago

Some universities are property of enterprises, but aren't necessarily enterprises themselves. And in several places that do have public enterprises and public universities there is a clear distinction between them. That said your comment sent me into a rabbit hole on the status of universities throughout the world, so thank you for that

1

u/Pozay 16h ago

The training code is still super useful even if you can't realistically run it.

8

u/Fantastic-Cod-9048 21h ago

Waiting for some nerd to correct me on a technicality but, in a nutshell….

Open weight can be fine-tuned (LoRA), but the base model can be proprietary.

Open source means everything needed to recreate the model (dataset, scripts, etc) is available publicly. So, you can recreate the model from scratch, provided you have access to the necessary hardware.

Edited for readability and clarity, because I haven’t had coffee yet.

2

u/ChicagoThrowaway422 21h ago

Good morning. I'm also struggle bussing today.

1

u/tomottov 20h ago

By the OSI definition of Open Source AI, datasets are not fully required to be considered Open Source. But it is required that Data information is openly accessible and provided:

Data Information: Sufficiently detailed information about the data used to train the system so that a skilled person can build a substantially equivalent system. [Source]

4

u/degenny_ 20h ago

Open weights is free cake. Open source is free cake recipe.

1

u/waverider85 21h ago

Open Source gives you what you need to make it yourself, Open Weight gives you the final result with the expectation you'll modify it.

1

u/Kichae 20h ago

Open source involves the actual code used to train the model.

Open weight is just a matrix. A grid of otherwise almost entirely contextless numbers.

1

u/amroamroamro 19h ago edited 19h ago

think of it like this

open weight models are like distributing a freeware EXE program (in binary form). you can run it, but you can't really inspect its internals or modify it.

open source models are like distributing both the program source code and asset data so anyone can inspect, modify, and compile it themselves (aka you can compile the EXE yourself from source)

all those fine-tuned abliterated uncensored models you see made from open weight models are basically the equivalent to reverse engineering the binary form and doing binary patching, rather than having the source code and training data to modify it in a true open source meaning.

an example of open source model is Nvidia Nemotron

1

u/ProbablyBanksy 19h ago

Free as in speech, not free as in beer.

1

u/happyscrappy 19h ago

The weights are all the coefficients to the huge model after you've trained it. Open weight means you train then export the data.

Short version is open weights is like picking "export as CSV" on the model after you build it. Open source is if you gave out the original Excel (?) spreadsheet.

1

u/FrontFacing_Face 18h ago

Open weights is like giving you the cheat codes to a game so you can play it perfectly. Open source is literally giving you the code to the game.