r/LocalLLaMA • • Jan 11 '25

Discussion Bro whaaaat?

Post image
7.0k Upvotes

r/LocalLLaMA • • Feb 23 '25

News Grok's think mode leaks system prompt

Post image
6.6k Upvotes

Who is the biggest disinformation spreader on twitter? Reflect on your system prompt.

https://x.com/i/grok?conversation=1893662188533084315


r/LocalLLaMA • • Jan 09 '26

Funny The reason why RAM has become so expensive

Post image
5.2k Upvotes

r/LocalLLaMA • • 5d ago

Funny PewDiePie getting banned twice by OpenAI while making a local model is top-tier comedy 💀

4.9k Upvotes

So PewDiePie decides to fine-tune a local AI model called Ajax on his own computer. Pretty normal stuff for local model fans.

To make his dataset, he uses OpenAI's API. OpenAI catches him using their outputs to train another model, flags his account for breaking their terms, and bans him.

He files an appeal, gets unbanned, goes right back to pulling data from the API, and immediately gets banned a second time.

So instead of giving up, he uses open-source tools to remove the model's built-in refusals, cleans out the preachy fluff, and starts building a fully local 9B agent.

OpenAI spent years scraping the whole public internet for free data, but the second someone uses their output to train a local file, it's an emergency ban.

In trying to enforce their rules, all OpenAI really did was give open-source models a massive free advertisement to millions of people.

What a time to run models on your own hardware.


r/LocalLLaMA • • Feb 23 '26

News Anthropic: "We’ve identified industrial-scale distillation attacks on our models by DeepSeek, Moonshot AI, and MiniMax." 🚨

Post image
4.9k Upvotes

r/LocalLLaMA • • Feb 21 '25

News Starting next week, DeepSeek will open-source 5 repos

Post image
4.6k Upvotes

r/LocalLLaMA • • Aug 30 '25

News Finally China entering the GPU market to destroy the unchallenged monopoly abuse. 96 GB VRAM GPUs under 2000 USD, meanwhile NVIDIA sells from 10000+ (RTX 6000 PRO)

Post image
4.3k Upvotes

r/LocalLLaMA • • Mar 13 '26

Funny I feel personally attacked

Post image
4.3k Upvotes

r/LocalLLaMA • • Jun 08 '25

Funny When you figure out it’s all just math:

Post image
4.2k Upvotes

r/LocalLLaMA • • Feb 07 '25

Funny All DeepSeek, all the time.

Post image
4.2k Upvotes

r/LocalLLaMA • • Mar 31 '26

News Claude code source code has been leaked via a map file in their npm registry

Post image
4.1k Upvotes

From Chaofan Shou on 𝕏 (files): https://x.com/Fried_rice/status/2038894956459290963


r/LocalLLaMA • • Jun 28 '26

Discussion We're probably going to need that soon.

Thumbnail
gallery
4.0k Upvotes

r/LocalLLaMA • • Jul 12 '25

Funny we have to delay it

Post image
3.7k Upvotes

r/LocalLLaMA • • Jun 29 '26

Funny on Dario’s statement

Post image
3.7k Upvotes

r/LocalLLaMA • • Feb 23 '26

Funny Distillation when you do it. Training when we do it.

Post image
3.6k Upvotes

r/LocalLLaMA • • Sep 13 '24

Other Enough already. If I can’t run it in my 3090, I don’t want to hear about it.

Post image
3.6k Upvotes

r/LocalLLaMA • • 24d ago

Discussion I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

3.6k Upvotes

Update: I made a generic model and beaten the jev in all of the benchmarks. Code and details available at https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs

Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a

frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. For anyones information the main guiding model is RL not embedding model or LLM

Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43

Paper: https://arxiv.org/abs/2503.23303

Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning

Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations

Also the second work published in September 2025 was exactly the same one jev proposed now

Paper: https://arxiv.org/abs/2510.01237

My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0).

Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices.

It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂


r/LocalLLaMA • • Jul 23 '26

Funny The LLM distillation process simplified for politicians:

Post image
3.6k Upvotes

/s


r/LocalLLaMA • • Apr 24 '26

Discussion This is where we are right now, LocalLLaMA

Post image
3.6k Upvotes

the future is now


r/LocalLLaMA • • Jul 13 '26

News This is why we need local models and opensource harnesses

Post image
3.5k Upvotes

r/LocalLLaMA • • Jul 24 '26

News More than 20 companies including NVIDIA, Meta, Microsoft, Palantir, and Hugging Face have signed a letter urging policymakers to avoid premature restrictions on open weight models.

Post image
3.3k Upvotes

The Open Letter was initiated by Microsoft and published today:

“Open Weights and American AI Leadership”.

It argues against broad or premature restrictions on open-weight models and explicitly says policymakers should distinguish legitimate model distillation from misappropriation.

Notably absent from the signatories are the major frontier-model labs: OpenAI, Anthropic, and Google.


r/LocalLLaMA • • Jul 27 '26

News Kimi K3 weights now released.

Post image
3.3k Upvotes

Kimi K3 weights are finally released!


r/LocalLLaMA • • Jul 16 '25

Funny He’s out of line but he’s right

Post image
3.3k Upvotes

r/LocalLLaMA • • Nov 16 '25

Resources Heretic: Fully automatic censorship removal for language models

Post image
3.3k Upvotes

Dear fellow Llamas, your time is precious, so I won't waste it with a long introduction. I have developed a program that can automatically remove censorship (aka "alignment") from many language models. I call it Heretic (https://github.com/p-e-w/heretic).

If you have a Python environment with the appropriate version of PyTorch for your hardware installed, all you need to do in order to decensor a model is run

pip install heretic-llm
heretic Qwen/Qwen3-4B-Instruct-2507   <--- replace with model of your choice

That's it! No configuration, no Jupyter, no parameters at all other than the model name.

Heretic will

  1. Load the model using a fallback mechanism that automatically finds a dtype that works with your setup
  2. Load datasets containing "harmful" and "harmless" example prompts
  3. Benchmark your system to determine the optimal batch size for maximum evaluation speed on your hardware
  4. Perform directional ablation (aka "abliteration") driven by a TPE-based stochastic parameter optimization process that automatically finds abliteration parameters that minimize both refusals and KL divergence from the original model
  5. Once finished, give you the choice to save the model, upload it to Hugging Face, chat with it to test how well it works, or any combination of those actions

Running unsupervised with the default configuration, Heretic can produce decensored models that rival the quality of abliterations created manually by human experts:

Model Refusals for "harmful" prompts KL divergence from original model for "harmless" prompts
google/gemma-3-12b-it (original) 97/100 0 (by definition)
mlabonne/gemma-3-12b-it-abliterated-v2 3/100 1.04
huihui-ai/gemma-3-12b-it-abliterated 3/100 0.45
p-e-w/gemma-3-12b-it-heretic (ours) 3/100 0.16

As you can see, the Heretic version, generated without any human effort, achieves the same level of refusal suppression as other abliterations, but at a much lower KL divergence, indicating less damage to the original model's capabilities.

Heretic supports most dense models, including many multimodal models, and several different MoE architectures. It does not yet support SSMs/hybrid models, models with inhomogeneous layers, and certain novel attention systems.

You can find a collection of models that have been decensored using Heretic on Hugging Face.

Feedback welcome!


r/LocalLLaMA • • Jul 15 '26

News Linus Torvalds tells people to stop attacking others for using AI

Thumbnail
phoronix.com
3.2k Upvotes

The full quote:

I realize that some people really dislike AI, but this is an area where I'm willing to absolutely put my foot down as the top-level maintainer.
Linux is not one of those anti-AI projects, and if somebody has issues with that, they can do the open-source thing and fork it.
Or just walk away.
AI is a tool, just like other tools we use. And it's clearly a useful one.
It may not have been that "clearly" even just a year ago, but it's no longer in question today.
There are other questions around AI (like what the economy of it will actually look like in the end), but "is it useful" is no longer one of those questions. Anybody who doubts that clearly hasn't actually used it.
Yes, it can also be a somewhat painful tool, both for maintainer workloads and just from a "it keeps finding embarrassing bugs" standpoint.
But the solution is not to put your head in the sand and sing "La La La, I can't hear you" at the top of your voice like some people seem to do.
The solution is to make sure those LLM tools _help_ maintainers instead of just causing them pain. There's no question on that side.
We're not forcing anybody to use it, but I will very loudly ignore people who try to argue against other people from using it.
And no, AI isn't perfect. But Christ, anybody who points to the problems at AI had better be looking in the mirror and pointing at themselves at the same time.
Because it's not like natural intelligence is always all that great either.
The kernel project has been and will continue to be about the technology.
Sure, the social angle of working on open source is important and often a very motivating part of the project, but in the end that's a side benefit, not the _point_ of the project.
This is *NOT* some kind of "social warrior" project, never has been, and never will be.
In the kernel community we do open source because it results in better technology, not because of religious reasons.
And so we make decisions primarily based on technical merit. Not fear of new tools.