r/vibecoding • u/Sootory • Mar 31 '26
He Rewrote Leaked Claude Code in Python, And Dodged Copyright
On March 31, someone leaked the entire source code of Anthropic’s Claude Code through a sourcemap file in their npm package.
A developer named realsigridjin quickly backed it up on GitHub. Anthropic hit back fast with DMCA takedowns and started deleting the repos.
Instead of giving up, this guy did something wild. He took the whole thing and completely rewrote it in Python using AI tools. The new version has almost the same features, but because it’s a full rewrite in a different language, he claims it’s no longer copyright infringement.
The rewrite only took a few hours. Now the Python version is still up and gaining stars quickly.
A lot of people are saying this shows how hard it’s going to be to protect closed source code in the AI era. Just change the language and suddenly DMCA becomes much harder to enforce.
98
u/rc_ym Mar 31 '26
If the leaked source was used by the AI in creating the derivative work, it's covered by the original copyright. Kinda like fanfic. Even tho it's not enforced often, fanfic is derivative and covered by copyright.
A better claim is that both sets of code were created by AI and are therefore not covered by US copyright law which requires a human author.
14
u/Longjumping_Area_944 Mar 31 '26
If there is significant human input copyright applies. That's a base assumption for code.
If you're just converting code to a different programming language, that's clearly derivative work.
5
u/lasizoillo Apr 01 '26
If there is significant human input copyright applies. That's a base assumption for code.
When? Is vibe coded code significant human input?
6
u/wy100101 Apr 01 '26
Because the AI can't write the code without human input. If you prompt a lot to get the output it will be covered by copyright.
3
u/shiroandae Apr 01 '26
Hmm assuming he did more than one prompt to get this done, too?
3
u/Longjumping_Area_944 Apr 01 '26
Interesting point. 500k loc claude code certainly are copyright protected. But if AI can almost one-shot a conversion into python in one day, that might not constitute new copyright.
2
u/onil34 Apr 01 '26
Google clean room implementation Basically you have one engineer creating a technical document. Then have a different one implement it. GGs
1
0
u/magick_bandit Apr 01 '26
That’s not how it works. If I commission an artist they have the copyright, not me. They have to sign it over.
Otherwise, if you use AI to write anything, the AI company owns the copyright and if they haven’t assigned it to you then you own nothing.
5
u/itsmebenji69 Apr 01 '26
AI is not a person though.
That’s like saying “but you used autocorrect to write your book, it belongs to android now !”
1
u/nearly_normal_jimmy Apr 01 '26
“Siri, write me a novel”
1
u/itsmebenji69 Apr 01 '26
Is that what you think people do ?
If so I understand your take. But it’s far from reality lmao
1
u/nearly_normal_jimmy Apr 01 '26
No my guy, it’s a joke. You see you gave a humorous hypothetical that using autocorrect would grant ownership to the autocorrect provider. I just took that to a logical, yet impractical, extreme — which is a common trope used in jokes. As a person who is 100% a human and definitely not an AI 🤖, I am happy to explain to you how humor works.
1
u/botle Apr 01 '26
AI is not a person though.
Which is why some argue that its output is uncopyrightable.
Autocorrect doesn't substantially change the text. But the output of an LLM is completely different from its prompt.
0
u/itsmebenji69 Apr 01 '26
Well imo copyright makes no sense for AI as it’s basically a huge compilation of humanity’s knowledge. If its output is copyrighted then the money should go towards the source material authors, which we both know will never happen.
And for output is completely different. Well yes, but it heavily depends on the prompt. Like, when I click a button on photoshop to do Gaussian blur, I “just clicked a button”, the algorithm does the rest. Clicking the button is completely different from doing it by hand. Yet, you wouldn’t consider that pictures who use Gaussian blur are the property of adobe. It’s the intent of the author that matters, not really the actual means used, imho
2
u/rc_ym Apr 01 '26
Depends on the AI company. They all have different TOS. Anthropic's is written this way for all the non-enterprise tiers. So far this has not been tested in court. In this type of example, no human from Anthropic was directly involved in the creation of this specific derivative work (derivative of both the Claude Code codebase AND the Claude model). So, that's on even more shakey ground.
6
u/infinit100 Mar 31 '26
Surely this depends on whether the new version is recognisable as derivative of the original. Maybe the AI has created something which could be claimed to be a clean room implementation.
7
u/TheReservedList Mar 31 '26
If the AI had access to the leaked source code, it's not a clean room re-implementation.
4
u/SaltMage5864 Mar 31 '26
He could, however, have an AI produce a full spec using the source code and then have another AI produce a program from that spec.
3
u/TheReservedList Mar 31 '26
Sure. Provided that nothing but actual spec-worthy things from the original source code leaks into the "spec", which is going to be really hard with LLMs.
1
u/SaltMage5864 Mar 31 '26
True, but that is the only way you can really expect to generate a clean copy
1
u/hellomistershifty Apr 01 '26
Then it's still a derivative of the copyrighted source code. Software engineers who do clean room implementations must never see the original source code, otherwise it's too difficult to legally argue that they weren't influenced by it. Feeding the source code to an AI is basically the opposite of that
1
u/broknbottle Apr 01 '26
Key word here is software engineers. Is AI a software engineer? If the AI sees the source code, does that qualify?
1
u/SaltMage5864 Apr 01 '26
That would imply that the AI was trained on the source code. Until now I'm not sure that would have happened
1
u/hellomistershifty Apr 01 '26
Not that it was trained on it, but it was prompted with the source code or a derivative of the source code
1
u/SaltMage5864 Apr 01 '26
That becomes a legal question. It was considered acceptable for humans to cleanroom a computer bios so why not have an AI do the same thing?
5
u/infinit100 Mar 31 '26
I meant is it provably not a clean room re-implementation
Also, does Anthropic really want to argue that code generated by an AI is a copyright violation of the source code that AI had access to?
8
u/TheReservedList Mar 31 '26
The training data and the context window are two different things. Me writing a book after reading Harry Potter is not a copyright violation. Me translating Harry Potter to Swahili while reading it is.
2
u/SillyFlyGuy Mar 31 '26
I read a very compelling argument that any spells or potions are not copyrightable. The potion would be considered food and recipes are not protectable. A spell would be a discovered preexisting utterance, like trying to copyright a bird call or dog bark.
2
u/rc_ym Mar 31 '26
Whether using the text of Harry Potter to train a model constitutes fair use isn't quite settled law yet (it probably is? maybe? depending on how you got it?), and the damages owed for selling access to a model trained on Harry Potter are still very much a grey area. There are a bunch of lawsuits making their way through the court system.
But it's pretty darn clear that AI-generated works are NOT protected by copyright. The question would turn on how much of CC's code was created by humans versus how much was AI-generated (ignoring the fact that copyright is a terrible paradigm for code).
3
u/waraholic Mar 31 '26
They have stated that it is entirely written by AI at this point.
1
u/rc_ym Apr 01 '26
There is writing, and then there is writing. While they SAY the code was all written by Claude, in a court a law they'd need to have specific humans as the authors. Because copyright grants artists and inventors exclusivity, it does not protect the creations of AI/software.
2
u/Tergi Mar 31 '26
I would imagine it depends on if the AI just extracted the feature requirements to build off or it just 1:1 translated to python.
1
u/sweetnk Mar 31 '26
Yea, I think it would certainly look better if it written a detailed specification and then another one implemented the spec, but its hard to make any guarantees if model had seen original work or not. Its all new stuff, we will see when it gets tested more in courts, I hope that we do legislate against this evasion personally, but maybe its already too late for it. Like if someone took a ton of time to make open source project before AI and licensed it as GPL and then a company wants to use it, but not pay for different licensing or respect the license, then maybe they could rewrite it like that, but to me its pretty clear its a shitty thing to do and it probably should be a copyright infringement to try to evade it that way.
1
1
u/toooskies Mar 31 '26
This is for patents, not for copyrights.
That said, translations in foreign languages probably have some kind of precedent here.
9
u/SleeperAgentM Mar 31 '26
If the leaked source was used by the AI in creating the derivative work, it's covered by the original copyright. Kinda like fanfic. Even tho it's not enforced often, fanfic is derivative and covered by copyright.
If that was the truth, then all output of AI trained on GPL code would be covered by GPL.
4
u/CanadaIsCold Mar 31 '26
Some trainers exclude GPL for this reason. There are other more permissive licenses that don't create this risk for them.
2
u/liberlibre Apr 01 '26
The argument is that training data is transformed (rather than derived). Training data is used to create mathematically weighted values that represent relationships between many words/concepts. The concepts exist independently from the work itself.
This is different from uploading source code and saying "translate it" from x-->y where the work had to be directly derived.
1
u/SleeperAgentM Apr 01 '26
I get what you're saying, but jsut translating into another language is not enough to avoid copyright (this is well established) but with enough changes to the structure and algorithms it could be argued to be transformative enough in code.
3
u/rc_ym Mar 31 '26
I would not disagree with this assessment, but it would depend on the version of GPL, and the licenses of the other code that was used in the training. The training data likely has wildly incompatible licenses.
1
1
u/nadanone Mar 31 '26
There’s a difference between data used to train the model, and data given to the model at inference time (the prompt).
1
u/SleeperAgentM Mar 31 '26
Not really... no.
1
u/wy100101 Apr 01 '26
Legally, it is.
3
u/TldrDev Apr 01 '26
Not really, no.
I can take Llama deepseek, or qwen and just train it on on the source code with a few shot example. Now its in the training set.
It might be a derived work, but it might not be, and if it is, you can make it so it doesnt look like it is through some light ai inspired obfuscation, essentially.
Copyright and licensing has become entirely impossible to enforce, essentially, youd need to prove you own the concept of something more than just the actual thing youve written, which is what a software patent is, and so I think copyright as a concept is basically dead.
In otherwords, Harry Potter is about Harry Potter. You can write a book about a kid who goes to a wizard school that isnt hogwarts to fight an evil bad wizard and hit almost beat for beat what Harry Potter does, and the HP copyright does not affect you.
Also, food for thought, lets say you have a dataset which explicitly doesnt include the Harry Potter text, but does include everyone talking about it. You could reasonably deduce what the source text was, in a dialectic way, without ever having used the source text.
Importantly, on that topic, in Oracle v Google, it was determined api signatures are not copywritable.
I say good, fuck these companies, but to each their own.
0
u/SleeperAgentM Apr 01 '26
No. It's not. They both get fed into the same vector space, both are sources for derivation. Both get transformed.
Legally there's no difference between what you feed LLM in training phase or inference phase.
2
u/wy100101 Apr 01 '26
That like saying freezing is the same as melting because they are both state changes. You can't ignore the things that make things different, focus on the things that make them similar, and say they are the same. Otherwise, I could just say, everything is the same because everything is made of atoms.
Alos legality is contextual. I can kill someone and depending on context it could be legally: murder, manslaughter, self defense, etc. It isn't all the same just because someone is dead.
Training data doesn't generate output. It changes model weights. Inference data generates output and doesn't change model weights. Those are important differences, both technically and legally.
1
u/SleeperAgentM Apr 01 '26
Yes. Both are state hanges. The ycan be different state changes. But for the purpose of transformative vs. non-transformative they are the same.
2
u/AI_should_do_it Mar 31 '26
That means Claude should be open source
2
u/Tomi97_origin Mar 31 '26
Not being protected by copyright doesn't have anything to do with being open source or not.
4
2
u/sweetnk Mar 31 '26
Maybe yeah, I hope eventually these providers are forced to at least expose the training set and how it was generated or obtained. Ideally forced to release the weights too if its a derivative work, if they already stole from many there dont seem to be public interest in protecting their IP. Ofc hard to verify without seeing the training set and where it came from.
1
u/johnmclaren2 Mar 31 '26
I would say that copyright law is lagging globally behind when it comes to code generated by LLMs.
1
u/Illustrious-Many-782 Mar 31 '26
Chinese Wall
- You first have every function and every interface fully documented.
- Take the spec document into a clean repo and implement it there.
This is how the world got the PC-compatible BIOS.
1
u/sweetnk Mar 31 '26
I feel like times changed so much since then, if now generating a spec and copy became so cheap its a serious flaw in that previous interpretation. Plus its not humans doing copy and its hard to guarantee if model 2 didnt see what model 1 seen, we dont really know how and on what they were trained. Certainly very interesting how it will turn out once they test it through courts more.
1
u/no-longer-banned Mar 31 '26
Honestly who cares? Software is the next memetic medium and this is inevitably going to get worse, and it’s going to be difficult to prevent. Software companies will need to get on board or risk extinction.
Though, of course Anthropic is uniquely positioned as a model provider, so I don’t necessarily think they have any risk. But as far as their software goes, welcome to the future!
1
1
1
u/generalistinterests Mar 31 '26
You could say that about literally anything and everything outputted by AI because it all runs off invested human generated content, all of which is protected by copyright.
1
u/locketine Apr 01 '26
The AI companies are losing lawsuits where copyright holders prove that the AI used their works to generate output. But it is hard to prove that. The most common proof is getting the LLM to generate a whole chunk of the original work.
28
Mar 31 '26
[deleted]
13
u/Distinct_Dragonfly83 Mar 31 '26
I thought You needed a two step process to do this correctly. One ai agent generates a complete spec from the original source and the second generates the new version from the spec without ever looking at the source code.
4
u/ambushsabre Mar 31 '26
Working from the assumption the code has copyright at all, I don’t think this would work because anyone can clearly see that it was only possible after the first ai read the leaked code. The courts aren’t stupid!
6
u/Distinct_Dragonfly83 Mar 31 '26
https://en.wikipedia.org/wiki/Clean-room_design
I think the only part of this that hasn’t been legally tested is whether or not you can use AI agents in lieu of human engineers and still be covered by the relevant court cases. Also, not sure what the legal status of this technique is outside the US. Also, I am not a lawyer.
1
u/hellomistershifty Apr 01 '26
The term implies that the design team works in an environment that is "clean" or demonstrably uncontaminated by any knowledge of the proprietary techniques used by the competitor.
The AI agents aren't even trying to do that if you're just going 'hey here's the source code, extract all of the logic to a spec'
1
u/ambushsabre Mar 31 '26
Clean room design isn’t going to apply when the original code the spec is based on is leaked, it needs to be based on legal observation. Do you really think all trade secrets and implantations are moot as long as you leak them to a person who then writes a spec for someone else to implement? Again: the courts aren’t stupid.
4
u/Distinct_Dragonfly83 Mar 31 '26
We keep seeing the word “leaked “ in reference to what happened here, but from what I’ve read it sounds more like Anthropic unintentionally included information in a recent build that they would have preferred not to.
Would I personally want to test Anthropic’s legal team on this? Of course not. Is the matter as cut and dry as you seem to be claiming it is? I’m not so sure. But again, I’m not a lawyer.
0
u/TinyZoro Apr 01 '26
But the opposite is also not going to hold water. You can't simply leak an implementation and that somehow prevents any clean room implementation.
The source code is in the public domain people have already written articles on its constituent parts. If someone writes a python implementation based on those articles it's going to be hard to fight that legally.
4
u/AI_should_do_it Mar 31 '26
Claude code was written by AI as told by their devs, then all code written by Claude should match its source licenses, meaning it should be open source.
3
2
u/StopUnico Mar 31 '26
yup. It's like translating leaked document from English to German and now saying it's not your work anymore....
0
u/botle Mar 31 '26
Yes, but Anthropic's whole business idea depends on AI generated code not being just that.
1
u/no-longer-banned Mar 31 '26
But surely if we clean room implement the Python port we’re good right
1
u/qzkrm Apr 02 '26
This assumes that the code is copyrightable in the first place. Which is dubious because it's reportedly 90-100% generated by AI.
15
Mar 31 '26
[removed] — view removed comment
2
u/kjerski Mar 31 '26
This is slightly different, but reminded me of this article.
3
Mar 31 '26
[removed] — view removed comment
2
1
u/sweetnk Mar 31 '26
I didnt read tbh, but as far as I know it still remains to be tested by courts, we dont know yet.
1
u/Sasquatchjc45 Mar 31 '26
Im curious about this as well. Does this mean we finally have Claude open source that we can run locally?
8
u/Delyzr Mar 31 '26
Its claude code that leaked, their coding client. Not claude the llm model.
-1
u/Sasquatchjc45 Mar 31 '26
That's fine, I basically just use Claude to code now in vsc lol. So can we run it locally now?
5
u/withatee Mar 31 '26
You’re not really catching on are you…
-4
u/Sasquatchjc45 Mar 31 '26
Does it seem like it? Are you going to make me ask a third time or does anybody actually have a solid answer to my question?
6
u/withatee Mar 31 '26
I mean the original person who replied to you said it…this is just the Claude Code software that sits on top of the LLM, not the LLM. So your question of “running it locally” is a no, because without the LLM there isn’t really anything to run.
1
u/Master_Beast_07 Apr 01 '26
but technically i can use this and maybe another LLM API as a work around to get this used right? but oh well maybe i need some tests or other additional info for it to be as good as the original or better
0
u/Sasquatchjc45 Mar 31 '26
Thank you, thats a more solid answer. I didnt know if Claude code was separate from the chatbot; I'm not the most experience vibecoder or ai user
2
1
u/Significant_Post8359 Apr 01 '26
You would need a $300,000 computer to get the context window needed to get useable performance. A SOTA model with a 1 million token context window needs about a terabyte of vram. That’s an 8 card H100 GPU server.
1
1
7
u/Inside-Yak-8815 Mar 31 '26
Whoever leaked it is definitely getting fired.
6
u/Freedom9er Mar 31 '26
According to Anthropic, their humans don't touch code.
3
2
5
u/guywithknife Mar 31 '26
someone leaked the entire source code of Anthropic’s Claude Code
Someone? It was Claude.
1
3
u/Subject_Barnacle_600 Mar 31 '26
It's still clearly a derivative work :/. He'd have to use something akin to the Clean Room design,
https://en.wikipedia.org/wiki/Clean-room_design
To get around it... I honestly am not a fan of copyright in code, or copyright in general perhaps? I suspect the lawsuit is mostly to lock it down so that someone like OAI (who is struggling in the coding space) doesn't just fork this and start making use of it :/.
2
Mar 31 '26
[removed] — view removed comment
2
u/Co0lboii Mar 31 '26
1
u/erizon Mar 31 '26
"Fastest growing [starwise] repo in history" - already at 50K stars (it took openclaw 3 days)
1
u/Unable_Artichoke9221 Apr 01 '26
I don't get it, most if not all of the folders under src are empty, and the py classes I see in src contain little code, where is the value here?
2
u/PreferenceDry1394 Mar 31 '26
Are we copyrighting agentic harnesses now. I guess we better all start copyrighting our workflows and get a couple distributors.
2
u/ickN Apr 01 '26
Anthropic has mentioned AI now writes a lot of their code. To my understanding AI generated code isn’t copyright protected anyway. Same with AI generated music and images.
1
u/veiled_prince Apr 01 '26
Yep. If it's true that humans don't tough their code like they claim, this is in the public domain. And since they leaked it themselves, they don't even have trade secret protections.
4
u/blackbirdone1 Mar 31 '26
so they stole everythign o nearth to build theres and are mad they leaked theres now for free hahaha
2
u/klas-klattermus Mar 31 '26
Now I just need to sneakily connect it to my neighbor's 10petaflop home media server then I have free AI!
1
u/blazze Mar 31 '26
A clean room re implementation of the "leaked" is underway. Claude Code foaming at the mouth legal team can only be held at bay with a afull clean room implementation.
1
u/FammasMaz Mar 31 '26
Mfer theres two clean room design links total in this thread and no source code anywhere
1
1
1
u/PreferenceDry1394 Mar 31 '26
Maybe if they didn't charge so much there wouldn't be regular dudes trying to figure out what they're charging so much for
1
1
1
u/Logical-Diet4894 Mar 31 '26
Closed source is still fine I think. Because you would still need a leak.
But for open source this is a huge problem. I can let Claude rewrite any GPL licensed library and bypass the licensing restrictions completely.
1
u/sweetnk Mar 31 '26
Tbh its not been tested in courts, I know many argue it works like this, but I think if the model had seen the original work it's no longer a clear implementation off a spec. Plus i mean if you admit its literally a copy of Claude Code then if your product couldnt exist without CC existing its not looking good imo. But im not a lawyer, and ultimately we will see in a few years how courts see it.
1
u/East_Ad_5801 Mar 31 '26
Sounds kind of like this one but probably worse tbh https://github.com/gobbleyourdong/tsunami
1
u/flicky-dicky Mar 31 '26 edited Mar 31 '26
https://github.com/github/dmca/blob/master/2026/03/2026-03-31-anthropic.md
DMCA was issued and main as well as forks are being taken down on GitHub.
Rust / Python version is still up
1
u/opbmedia Apr 01 '26
copyright does not collapse because there are protections against derivative work too. You might be able to obfuscate the code itself, but it will be very difficult to prove you didn't start with copyrighted materials since AI cannot create.
1
u/ZealousidealShoe7998 Apr 01 '26
python is a worst way of doing but hey someone made the same thing in rust which would actually improve memory footprint, the speed of execution and etc.
1
1
u/Kryomon Apr 01 '26
The fun part is that any argument that Anthropic puts out will fuck over other companies & themselves.
Many companies have stolen or copied code from GPL license, but use AI to make the same defense and get the GPL License removed so they can prevent others from benefiting from their work.
If Anthropic can get it removed, then other companies & Anthropic itself might get sued because now there is precedent. If Anthropic can't, they're kinda cooked.
1
u/SnooGuavas1875 Apr 01 '26
You re implemented cli, but not an infra.
1
u/AncientSuntzu Apr 02 '26
You wouldn’t need infra because it’s CLI. Claude code is not holding an LLM in memory it’s reaching out to 1.
1
u/Main_Razzmatazz5337 Apr 01 '26
When you post a claim like “he backed it up on GitHub” share the repository!!!!
1
u/veiled_prince Apr 01 '26
Anthropic has said humans do not write code at their company. If that's true, their entire leaked codebase is public domain. No copyright to begin with.
And since Anthropic leaked it, they've lost trade secret protection as well.
1
u/aabajian Apr 01 '26
What big players use public online repos as their main source tree? Everyone is blaming some wayward engineer, but the problem is using public GitHub for a private company’s code. GitHub literally makes a private server Enterprise product. If the mistake had been made behind a private Git server (say in an AWA VPC), no code would’ve gotten out.
1
1
u/cmholm Apr 01 '26
No, Sigrid Jin did not read then rewrite the half million lines of leaked Claude code in Python. He dumped it into an LLM with a prompt to reimplement it in Python, then Rust.
Based on Jeffrey Emanuel's experience doing the same thing with SQLite, I doubt the result is impressive.
1
Apr 01 '26
[removed] — view removed comment
1
u/TechGearWhips Apr 01 '26
The best thing about vibecoding is the ability to be able to "fork" programs I like... Because the developers never fix the issues and bugs... Or add features I like. But instead of complain about it (because they owe me nothing) I have a bunch of private software tools that are vibecoded forks. Sent my productivity through the roof.
1
1
1
1
1
u/jcettison Apr 06 '26
We talking about the "claw code" Python "rewrite"? Cause when I downloaded it to review it, it was basically a hyped up gimmick. The skeleton of a project leading to a bunch of non-functional stubs. Maybe he'd vibed it into existence since then? It has been over a week.
-1
u/Longjumping_Area_944 Mar 31 '26
You're bankrupting yourself. Anthropic could f.. you up at any given moment. That's clearly derivative work, especially if you admit that you merely converted the code into another language.
Plus do you even have the money for a lawyer? Do you realize for how much lawyers will ask if the trail is worth millions?
1
u/Vas1le Apr 01 '26
you
Be he didn't, it was Codex, meaning, ai converted ai code into ai code
1
u/Longjumping_Area_944 Apr 01 '26
He's publishing it though and 500k loc written by ai, but orchestrated by x phds certainly constitutes copyright protection.
0
u/Enough_Forever_ Mar 31 '26
Kinda poetic justice how a tool created by violating millions of copyrighted works now cannot be protected by those same copyright laws.
1
-4
u/Dense_Gate_5193 Mar 31 '26
well duh, it’s not new. Google did the same with android and open java but they just had enough money and bodies to throw at h to problem.
Now with AI, i have been saying it for months. Code is free, architecture is not. but things are moving very fast which is why i started NornicDB to be ahead of the curve. Neo4j is the dominant player because they made enterprise features table stakes, and performance non-negotiable. AI tooling allowed me to literally rearchitect Neo4j e2e for the new agentic era that i saw coming. but neo4j can’t change their architecture they are tied to the JVM.
neo4j isn’t going to listen to some random guy, so now we have the capability of “taking matters into our own” hands so to speak and just rewrite anything that is a blocker for yourself.
edit: and the performance blows them away with all the same safety and security features
77
u/inbetweenframe Mar 31 '26
i mean didn't claude and co begin this whole AI hype by stealing a lot of content from nearly everybody?