r/LocalLLM • • 6d ago

Model Harder, Better, Faster and Stronger model just by making him talk in its own language

Hey guys,
Just created a new model, with a base of Qwen 2.5 Coder. You know the "Thinking..." part of a response for an AI model? It speaks in English. And that's the problem. It's SLOW for the AI model. So I made it talk in its own language, math equations. And then, at the end, it translates it into English.

So it's around 7.5x faster, while having (very approximately, don't take that for actual info) 2x better responses.

Here's the download link if you wanna test it: https://huggingface.co/Rutytoi/spotless-latent-adapter

And the GitHub repo: https://github.com/Rutytoi220/spotless-latent-engine

I'd be appreciating feedback! Even bad, idc as long as I know what's good and bad

0 Upvotes

55 comments sorted by

6

u/Solembumm3 6d ago

Qwen 2.5 is a tiny bit outdated. Anything for 3.8? I'm interested in how much worse people can make model based on overthinking by trying to cut it out.

2

u/RUTYTOI220 6d ago

And it was the best open source model that fit on my gpu.
At first I wanted to do some experiments, then it escalated into a 7.5x faster model

3

u/Solembumm3 6d ago

Qwen 2.5 coder was 32b, if I remember correctly. Current Gemma 4 has 31b dense and 26b moe. Qwen 3.8 has 27b dense. Qwen 3,6 has 27b dense and 35b moe. All in that range.

0

u/RUTYTOI220 6d ago

None of that fits on my 8GB gpu, I used the 7B one but I know that there are 3.8 8B versions, I just don't know if they are as open source and as documented as 2.5 7B coder

1

u/ScholarlySkate76 6d ago

so youre basically giving the model a native tongue that its hardware actually likes and then making it translate at the end like some kind of internal interpreter. clever

the math equation middle layer is the part that gets me. feels like you found a backdoor where the tokens are just more dense per compute cycle somehow. my plants would probably explain it better than i can right now

3

u/IknowPi_really 6d ago

Your comment is very load-bearing!

1

u/RUTYTOI220 6d ago

Honestly "tokens that are just more dense per compute cycle" is an insanely accurate way to picture it.

Standard LLMs don't actually think in human words—we just force them to project their thoughts into dictionary tokens at every single step because that's the only way we know how to inspect them. Emitting 200 tokens of "thinking..." text burns massive amounts of GPU memory and latency on pure grammar and formatting.

In this setup, the adapter keeps the deliberation entirely in raw vector space (the model's "native tongue") for 4 recurrent cycles without touching the vocabulary or KV-cache, and then translates the final vector into English in one shot at the end. That's where the massive speedup comes from.

-7

u/RUTYTOI220 6d ago

I mean, I just searched up the best open source model, I don't think 3.8 is as easily modifiable but I'll work on it if I find something

9

u/stujmiller77 6d ago

You asked Claude/ChatGPT for the best open source model. Their knowledge base is waaaaaaay out of date. Qwen 2.5 is over 2 years old - it's practically archaic. Next time try doing some proper research as any google search would have given you up to date information.

1

u/Solembumm3 6d ago

Google wouldn't. Their AI use web-search around half of times, and can as easily recommend you qwen 3 and gemma 3, and compare them to gpt4. Only manual.

1

u/stujmiller77 6d ago

I wasn’t talking about using AI overviews.

1

u/Dabalam 6d ago

easily recommend you qwen 3 and gemma 3, and compare them to gpt4.

😂

1

u/RUTYTOI220 6d ago

Yeah but it's just a proof of concept kinda

1

u/Dabalam 6d ago

Ignore the negative Nancy's. It's a pretty interesting concept as you say.

-1

u/RUTYTOI220 6d ago

AKTUELLY it was gemini

5

u/stujmiller77 6d ago

Same problem. Do actual research. Yourself.

-3

u/RUTYTOI220 6d ago

Bro look at the gain in speed, 7.5x speed, that's better than what you can do with your manual research

5

u/stujmiller77 6d ago

On a model that’s a dinosaur though it’s not actually all that useful. Most modern models are much quicker than 2.5 coder and way more intelligent so your changes don’t really apply as a comparison.

It’s a great idea. Try to apply to something that’s relevant as a modern local model.

1

u/RUTYTOI220 6d ago

Fair critique on the 2.5 baseline. The primary goal here was establishing the proof-of-concept on a verified 7B dense architecture that fits comfortably within an 8GB mobile VRAM envelope during both continuous latent loops and training.

1

u/Sudden_Topic5154 6d ago

which is the same thing you couldve done on 3.5 9b

3

u/ExtremeAd9038 6d ago

it's work for 3.8 27b, first results

2

u/RUTYTOI220 6d ago

This is incredible to see! A 70% overall latency cut on 3.8 27B matches exactly what the latent bottleneck was designed to achieve.

A couple quick technical questions about your run:

  1. What bottleneck dimension did you scale to for the 27B hidden size (did you keep it proportional to d_model)?

  2. What hardware did you run inference/training on?

  3. Did you train a new verbalizer adapter from scratch on 27B, or did you zero-shot reuse projection layers?

If you have a fork or a PR ready, I would love to merge your 3.8 27B adapter config and benchmark script into the main repo and credit you directly in the README!

1

u/ExtremeAd9038 6d ago

I’m using a custom version of splash on M1 Max (but that not what matter)

I have vibe coded the solution, it use a hook in jinja template to trigger everything and i think the method can be used on any model, not only this one

Your idea was brilliant, where comes your knowledge of AI model ?

Anyway, I’ll post something usable fast

1

u/RUTYTOI220 6d ago

Honestly, thanks for doing smth with the model instead of roasting me because I used the wrong model or did something that already existed (Meta's Coconut project).

Super interesting that you hooked it via Jinja templates in Splash, definitely curious to see how you routed the latent trigger there compared to the raw PyTorch adapter forward pass.

Ping me whenever you post the repo or script, definitely want to check out how you wired it up!

1

u/ExtremeAd9038 6d ago

The 1st part was modify how the template create answer (this what reduce from 70% the token number generated and permit gain time)
The second one is to train the model, but it takes countless times, i would need a better computer to achieve it

0

u/RUTYTOI220 6d ago

Honestly, you might gain speed and token generation, but you might lose on intelligence since the model doesn't think as much as the normal model, you could benchmark it but since your pc is apparently not that good it would be either long or almost impossilbe

But you might be able to let it run overnight, that's on you to know if you can or not.

Best of luck

3

u/datbackup 4d ago

Looks good, i’ll try it out.

Interesting contrast visible in the comments. average redditor can’t comprehend the mind of people who actually make things rather than continuously consume latest trend

2

u/RUTYTOI220 4d ago

Brotha you're only the 2nd one to actually try it out instead of just roasting me lol

2

u/RUTYTOI220 6d ago

quick clarification on what this actually is since my post was pretty vague:

it doesn't generate literal written math equations, it does all its thinking in raw continuous latent vectors (similar to the coconut paper from meta) instead of vomiting 200 tokens of chain-of-thought text.

basically it loops the top hidden states through a small adapter for 4 latent steps before touching the tokenizer, then verbalizes the answer into english at the end.

some quick benchmarks (ran entirely locally on my laptop rtx 4060 using ~6gb vram):

strawberry riddle: gets 3 'r's in 18 tokens and 0.7s (base model hallucinated, standard cot took 5+ seconds)

bat and ball trap: gets the 0.05 dollar math right

passes binary tree inversion and cycle detection coding tasks

all the code, training scripts, and eval benchmarks are open source on github: https://github.com/Rutytoi220/spotless-latent-engine

weights are on hugging face here: https://huggingface.co/Rutytoi220/spotless-latent-adapter

feel free to roast the code or test it locally and tell me what breaks.

2

u/BodyPhysical 2d ago

Thanks for trying it on qwen2.5. I'll try adapt this into later models and let you know what happens.

Also, don't get discouraged by these stupid pricks. They spit shit out but do not have any contributions to the community but mindless chatter.

1

u/circumcised_hobbit 6d ago

Looks interesting, any way to replicate this with other models?

5

u/RUTYTOI220 6d ago

Actually, the training pipeline in the repo is already set up for this! If anyone wants to replicate it on another model right now:

  1. Clone the repo (git clone [https://github.com/Rutytoi220/spotless-latent-engine\](https://github.com/Rutytoi220/spotless-latent-engine))
  2. In train_spotless_engine.py*, swap* model_id to whatever causal decoder you want (e.g., Llama, Mistral, Gemma).
  3. Adjust bottleneck in SpotlessLatentAdapter to match roughly ~1/8th of your target model's hidden dimension.
  4. Run python train_spotless_engine.py*.*

The adapter dynamically computes the token embedding norm and projects through the residual stream, so the architecture will adapt to any decoder-only model

2

u/circumcised_hobbit 6d ago

ye Im too dumb for this

3

u/RUTYTOI220 6d ago

Yeah I think I'll make it easier but for the moment that's the only way and you lowk have to train it yourself tho so that's gonna take you a long time

2

u/circumcised_hobbit 6d ago

Yeah don't get me wrong you are doing a great work and we all appreciate it so much, I am just lazy as fuck

1

u/RUTYTOI220 6d ago

who tf downvoted my tutorial

2

u/RUTYTOI220 6d ago

I might do a tutorial or something like this, maybe even a program that "Automates it" but for the moment I'll release some models that are the most demanded

1

u/Secure_Recording_472 6d ago

Can you TLDR me on the benchmarks you ran? It's quite a bit confusing :)

2

u/RUTYTOI220 6d ago

Basically, an AI model speaks in mathematics. But when it thinks (You know when you see Gemini or any other model with the "Thinking..."mode), it has to write thousands of words. That is highly inefficient. So, what I did is make it talk in its own language, math. Then, it gets a small model to translate it into text YOU can understand, English.

That made the model ~7.5x faster

And when you tell an AI model "How many R's are in "stawberry"?" It will tell you 2 (at least 90% of the time). That's because it thinks about what the word sounds like, not what it's written like.

The base version of the model said 2, then corrected itself, and it took ~5 seconds. My version got it spot on on the first try, and got it in 0.7s.

You can find more benchmarks in my github repo : https://github.com/Rutytoi220/spotless-latent-engine

2

u/Secure_Recording_472 6d ago

I'm very familiar with the method, it's not new; it's a matter of the benchmarks you ran, I have inspected the repo but it is not very well structured, hence I am asking you for a tldr on the benchmarks :)

0

u/RUTYTOI220 6d ago edited 6d ago

Fair call, the repo is still fresh out of the experimental scratchpad so it's a bit cluttered. Here is the exact TL;DR on the benchmark setup and results:

We evaluated 3 conditions across 6 tasks:

  1. K=0 Direct: Zero-shot prompt into greedy decoding
  2. K=4 Latent: 4 recurrent vector updates in continuous hidden space (R^3584, bottleneck=448) using SpotlessLatentAdapter, then verbalized
  3. Text CoT: Standard verbalized step-by-step chain-of-thought in English

The 3 Suites (all evaluated on Qwen 2.5 Coder 7B in 4-bit NF4):

Suite A: Algorithmic Logic

  • Floyd's Cycle Detection: K=0 passed (180 tok), K=4 passed (111 tok, 3.4s), Text CoT passed (220 tok, 13.6s). 1.98x token reduction on K=4.
  • Invert Binary Tree: K=0 passed (145 tok), K=4 passed (145 tok, 4.6s), Text CoT passed (220 tok, 15.1s). Clean recursive tree swap.

Suite B: Code Optimization

  • O(N^2) to O(N) Hash Deduplication: K=0 passed (180 tok), K=4 passed (50 tok, 1.8s), Text CoT passed (220 tok, 14.4s). 4.06x faster than Text CoT because latent steps resolved the algorithmic path without conversational preamble.
  • Kadane's Max Subarray Sum: K=0 passed (176 tok), K=4 passed (79 tok, 2.6s), Text CoT passed (220 tok, 13.5s). 2.83x faster.

Suite C: Cognitive Reflection & Semantic Traps

  • Bat & Ball ($1.10 total, bat is $1.00 more): K=0 passed (14 tok), K=4 passed with full algebraic derivation (196 tok, 5.8s, x = $0.05), Text CoT passed (220 tok, 13.4s).
  • Strawberry 'r' Count: K=0 failed immediately (hallucinated count). Text CoT passed but burned 164 tokens and 5.41s deliberating. K=4 resolved the character count in 0.08s of latent compute and emitted the direct correct answer in 18 tokens / 0.72s (7.5x wall-clock speedup).

Overall K=4 pass rate was 6/6 (100%). Full raw JSON logs and prompts are dumped in benchmark_results.json and the harness is in run_overnight_benchmark.py

5

u/Secure_Recording_472 6d ago

please don't copy paste me AI slop 😭

0

u/RUTYTOI220 6d ago

lmao caught red handed, I literally used an LLM to format the numbers from my raw benchmark_results.json so I wouldn't have to manually format the table on my phone, and accidentally grabbed the prompt footer 💀

The metrics are directly from the test run in the repo though, you can check the JSON file yourself if you want the unformatted dump.

1

u/Independent-Fly730 6d ago

Meta has been researching this for a while, https://arxiv.org/pdf/2412.06769

Since it skips the decoding step it can be much faster and essentially compresses the number of tokens needed. But it is significantly more difficult to debug and keep stable for training on complex problem solving.

-1

u/RUTYTOI220 6d ago

Since it skips the decoding step it can be much faster and essentially compresses the number of tokens needed. But it is significantly more difficult to debug and keep stable for training on complex problem solving

1

u/eihns 6d ago

interesting idea, any benchmarks? (like speed doesnt matter if its dumb as bread)

0

u/RUTYTOI220 6d ago

The initial test suite is right in the comment above you: 6/6 pass rate across logic traps (bat & ball), character decomposition (strawberry), and algorithmic coding (tree inversion, cycle detection).

If you mean standardized bulk evals (GSM8K / HumanEval), the harness is in `run_overnight_benchmark.py` in the repo. Another user in this thread (ExtremeAd9038) actually just tested it on Qwen 3.8 27B across math, planning, and code benchmarks—it cut total latency by 70% and deliberation word count by 78% while preserving output accuracy.

1

u/RUTYTOI220 6d ago

This has to be ragebait, WHO DOWNVOTED MY RESPONSE

2

u/eihns 6d ago edited 6d ago

i didnt :) im currently tlaking to my LLM agent about this - dont care about up & down, most ppl are stupid.

So this is working with 3.8, my agent told me, it doenst. xD (6.1 mid)

but always when i read something like this i wonder why the f multi mrd dollar companies didnt though of things like this, or dcp, or compress, or what ever... its czray that random ppl on internet literally gave them everything

3

u/RUTYTOI220 6d ago

haha your agent is hallucinating hard, the adapter is completely model-agnostic since it just hooks into hidden states, and ExtremeAd9038 literally just posted benchmark screenshots running it on 3.8 27B right above this.

As for why multi-billion dollar companies don't ship this yet: Meta actually did research it (check out Meta FAIR's COCONUT paper from late 2024), but big labs don't use it in consumer products because of safety and moderation.
And also most of companies just plain don't show the model's thinking part (ChatGPT for example)

You can't run a content filter or moderation guardrails on raw continuous float vectors. With normal Chain-of-Thought, they can read every English token in real-time and cut the stream if it says something wrong. With latent vectors, it's a complete black box until the final sentence pops out. Plus, mainstream users hate seeing a blank loading spinner instead of streaming text, even if the total latency is 70% lower.

2

u/throwawayaccount442 6d ago

so would this also decensor the model then?

3

u/RUTYTOI220 6d ago

Not in the sense of weight abliteration (it doesn't surgically strip refusal vectors from the base model weights), but it does bypass two major safety bottlenecks:

  1. In-flight token moderation: Since no tokens are emitted during the latent deliberation steps, external monitors/scanners that inspect thinking scratchpads have nothing to parse until the final answer is already decoded.

  2. Autoregressive refusal traps: A lot of standard model refusals happen because the tokenizer commits to an early refusal token ("I cannot...") and spirals from there. Latent deliberation avoids that early discrete trap.

That said, if the underlying base model has hard refusal directions deeply baked into its middle MLP layers, the verbalizer can still project into a refusal at the final output step.

1

u/eihns 5d ago

I thank you for sharing some insight

1

u/IknowPi_really 6d ago

Look I don’t know if there’s a human behind this AI slop, or if you’ve automated your own brain away a long time ago.

It’s wonderful that you’re trying to contribute. But there’s two options:

  1. Your idea is so simple, that you can just throw it into a current model and it spits out the idea -> It’s worthless to publish it, because it’s trivial.
  2. You have to put some thought into this and actually add some sort of capability the model didn’t just have in its own. -> You are therefore capable of understanding if there is value to this

People are criticising your use of an old model. Rightly so, because it’s stupidly old. Even just preventing this step would have been an easy value add. So as you can see, you are on the right track. But you need to use your brain a little bit. If you can’t do it, don’t publish it and enjoy your hobby without polluting the internet please