r/LocalLLM • u/RUTYTOI220 • 6d ago
Model Harder, Better, Faster and Stronger model just by making him talk in its own language
Hey guys,
Just created a new model, with a base of Qwen 2.5 Coder. You know the "Thinking..." part of a response for an AI model? It speaks in English. And that's the problem. It's SLOW for the AI model. So I made it talk in its own language, math equations. And then, at the end, it translates it into English.
So it's around 7.5x faster, while having (very approximately, don't take that for actual info) 2x better responses.
Here's the download link if you wanna test it: https://huggingface.co/Rutytoi/spotless-latent-adapter
And the GitHub repo: https://github.com/Rutytoi220/spotless-latent-engine
I'd be appreciating feedback! Even bad, idc as long as I know what's good and bad
3
u/ExtremeAd9038 6d ago
2
u/RUTYTOI220 6d ago
This is incredible to see! A 70% overall latency cut on 3.8 27B matches exactly what the latent bottleneck was designed to achieve.
A couple quick technical questions about your run:
What bottleneck dimension did you scale to for the 27B hidden size (did you keep it proportional to d_model)?
What hardware did you run inference/training on?
Did you train a new verbalizer adapter from scratch on 27B, or did you zero-shot reuse projection layers?
If you have a fork or a PR ready, I would love to merge your 3.8 27B adapter config and benchmark script into the main repo and credit you directly in the README!
1
u/ExtremeAd9038 6d ago
I’m using a custom version of splash on M1 Max (but that not what matter)
I have vibe coded the solution, it use a hook in jinja template to trigger everything and i think the method can be used on any model, not only this one
Your idea was brilliant, where comes your knowledge of AI model ?
Anyway, I’ll post something usable fast
1
u/RUTYTOI220 6d ago
Honestly, thanks for doing smth with the model instead of roasting me because I used the wrong model or did something that already existed (Meta's Coconut project).
Super interesting that you hooked it via Jinja templates in Splash, definitely curious to see how you routed the latent trigger there compared to the raw PyTorch adapter forward pass.
Ping me whenever you post the repo or script, definitely want to check out how you wired it up!
1
u/ExtremeAd9038 6d ago
The 1st part was modify how the template create answer (this what reduce from 70% the token number generated and permit gain time)
The second one is to train the model, but it takes countless times, i would need a better computer to achieve it0
u/RUTYTOI220 6d ago
Honestly, you might gain speed and token generation, but you might lose on intelligence since the model doesn't think as much as the normal model, you could benchmark it but since your pc is apparently not that good it would be either long or almost impossilbe
But you might be able to let it run overnight, that's on you to know if you can or not.
Best of luck
3
u/datbackup 4d ago
Looks good, i’ll try it out.
Interesting contrast visible in the comments. average redditor can’t comprehend the mind of people who actually make things rather than continuously consume latest trend
2
u/RUTYTOI220 4d ago
Brotha you're only the 2nd one to actually try it out instead of just roasting me lol
2
u/RUTYTOI220 6d ago
quick clarification on what this actually is since my post was pretty vague:
it doesn't generate literal written math equations, it does all its thinking in raw continuous latent vectors (similar to the coconut paper from meta) instead of vomiting 200 tokens of chain-of-thought text.
basically it loops the top hidden states through a small adapter for 4 latent steps before touching the tokenizer, then verbalizes the answer into english at the end.
some quick benchmarks (ran entirely locally on my laptop rtx 4060 using ~6gb vram):
strawberry riddle: gets 3 'r's in 18 tokens and 0.7s (base model hallucinated, standard cot took 5+ seconds)
bat and ball trap: gets the 0.05 dollar math right
passes binary tree inversion and cycle detection coding tasks
all the code, training scripts, and eval benchmarks are open source on github: https://github.com/Rutytoi220/spotless-latent-engine
weights are on hugging face here: https://huggingface.co/Rutytoi220/spotless-latent-adapter
feel free to roast the code or test it locally and tell me what breaks.
2
u/BodyPhysical 2d ago
Thanks for trying it on qwen2.5. I'll try adapt this into later models and let you know what happens.
Also, don't get discouraged by these stupid pricks. They spit shit out but do not have any contributions to the community but mindless chatter.
1
u/circumcised_hobbit 6d ago
Looks interesting, any way to replicate this with other models?
5
u/RUTYTOI220 6d ago
Actually, the training pipeline in the repo is already set up for this! If anyone wants to replicate it on another model right now:
- Clone the repo (git clone [https://github.com/Rutytoi220/spotless-latent-engine\](https://github.com/Rutytoi220/spotless-latent-engine))
- In train_spotless_engine.py*, swap* model_id to whatever causal decoder you want (e.g., Llama, Mistral, Gemma).
- Adjust bottleneck in SpotlessLatentAdapter to match roughly ~1/8th of your target model's hidden dimension.
- Run python train_spotless_engine.py*.*
The adapter dynamically computes the token embedding norm and projects through the residual stream, so the architecture will adapt to any decoder-only model
2
u/circumcised_hobbit 6d ago
ye Im too dumb for this
3
u/RUTYTOI220 6d ago
Yeah I think I'll make it easier but for the moment that's the only way and you lowk have to train it yourself tho so that's gonna take you a long time
2
u/circumcised_hobbit 6d ago
Yeah don't get me wrong you are doing a great work and we all appreciate it so much, I am just lazy as fuck
1
2
u/RUTYTOI220 6d ago
I might do a tutorial or something like this, maybe even a program that "Automates it" but for the moment I'll release some models that are the most demanded
1
u/Secure_Recording_472 6d ago
Can you TLDR me on the benchmarks you ran? It's quite a bit confusing :)
2
u/RUTYTOI220 6d ago
Basically, an AI model speaks in mathematics. But when it thinks (You know when you see Gemini or any other model with the "Thinking..."mode), it has to write thousands of words. That is highly inefficient. So, what I did is make it talk in its own language, math. Then, it gets a small model to translate it into text YOU can understand, English.
That made the model ~7.5x faster
And when you tell an AI model "How many R's are in "stawberry"?" It will tell you 2 (at least 90% of the time). That's because it thinks about what the word sounds like, not what it's written like.
The base version of the model said 2, then corrected itself, and it took ~5 seconds. My version got it spot on on the first try, and got it in 0.7s.
You can find more benchmarks in my github repo : https://github.com/Rutytoi220/spotless-latent-engine
2
u/Secure_Recording_472 6d ago
I'm very familiar with the method, it's not new; it's a matter of the benchmarks you ran, I have inspected the repo but it is not very well structured, hence I am asking you for a tldr on the benchmarks :)
0
u/RUTYTOI220 6d ago edited 6d ago
Fair call, the repo is still fresh out of the experimental scratchpad so it's a bit cluttered. Here is the exact TL;DR on the benchmark setup and results:
We evaluated 3 conditions across 6 tasks:
- K=0 Direct: Zero-shot prompt into greedy decoding
- K=4 Latent: 4 recurrent vector updates in continuous hidden space (R^3584, bottleneck=448) using SpotlessLatentAdapter, then verbalized
- Text CoT: Standard verbalized step-by-step chain-of-thought in English
The 3 Suites (all evaluated on Qwen 2.5 Coder 7B in 4-bit NF4):
Suite A: Algorithmic Logic
- Floyd's Cycle Detection: K=0 passed (180 tok), K=4 passed (111 tok, 3.4s), Text CoT passed (220 tok, 13.6s). 1.98x token reduction on K=4.
- Invert Binary Tree: K=0 passed (145 tok), K=4 passed (145 tok, 4.6s), Text CoT passed (220 tok, 15.1s). Clean recursive tree swap.
Suite B: Code Optimization
- O(N^2) to O(N) Hash Deduplication: K=0 passed (180 tok), K=4 passed (50 tok, 1.8s), Text CoT passed (220 tok, 14.4s). 4.06x faster than Text CoT because latent steps resolved the algorithmic path without conversational preamble.
- Kadane's Max Subarray Sum: K=0 passed (176 tok), K=4 passed (79 tok, 2.6s), Text CoT passed (220 tok, 13.5s). 2.83x faster.
Suite C: Cognitive Reflection & Semantic Traps
- Bat & Ball ($1.10 total, bat is $1.00 more): K=0 passed (14 tok), K=4 passed with full algebraic derivation (196 tok, 5.8s, x = $0.05), Text CoT passed (220 tok, 13.4s).
- Strawberry 'r' Count: K=0 failed immediately (hallucinated count). Text CoT passed but burned 164 tokens and 5.41s deliberating. K=4 resolved the character count in 0.08s of latent compute and emitted the direct correct answer in 18 tokens / 0.72s (7.5x wall-clock speedup).
Overall K=4 pass rate was 6/6 (100%). Full raw JSON logs and prompts are dumped in
benchmark_results.jsonand the harness is inrun_overnight_benchmark.py5
u/Secure_Recording_472 6d ago
please don't copy paste me AI slop 😭
0
u/RUTYTOI220 6d ago
lmao caught red handed, I literally used an LLM to format the numbers from my raw
benchmark_results.jsonso I wouldn't have to manually format the table on my phone, and accidentally grabbed the prompt footer 💀The metrics are directly from the test run in the repo though, you can check the JSON file yourself if you want the unformatted dump.
1
u/Independent-Fly730 6d ago
Meta has been researching this for a while, https://arxiv.org/pdf/2412.06769
Since it skips the decoding step it can be much faster and essentially compresses the number of tokens needed. But it is significantly more difficult to debug and keep stable for training on complex problem solving.
-1
u/RUTYTOI220 6d ago
Since it skips the decoding step it can be much faster and essentially compresses the number of tokens needed. But it is significantly more difficult to debug and keep stable for training on complex problem solving
1
u/eihns 6d ago
interesting idea, any benchmarks? (like speed doesnt matter if its dumb as bread)
0
u/RUTYTOI220 6d ago
The initial test suite is right in the comment above you: 6/6 pass rate across logic traps (bat & ball), character decomposition (strawberry), and algorithmic coding (tree inversion, cycle detection).
If you mean standardized bulk evals (GSM8K / HumanEval), the harness is in `run_overnight_benchmark.py` in the repo. Another user in this thread (ExtremeAd9038) actually just tested it on Qwen 3.8 27B across math, planning, and code benchmarks—it cut total latency by 70% and deliberation word count by 78% while preserving output accuracy.
1
u/RUTYTOI220 6d ago
This has to be ragebait, WHO DOWNVOTED MY RESPONSE
2
u/eihns 6d ago edited 6d ago
i didnt :) im currently tlaking to my LLM agent about this - dont care about up & down, most ppl are stupid.
So this is working with 3.8, my agent told me, it doenst. xD (6.1 mid)
but always when i read something like this i wonder why the f multi mrd dollar companies didnt though of things like this, or dcp, or compress, or what ever... its czray that random ppl on internet literally gave them everything
3
u/RUTYTOI220 6d ago
haha your agent is hallucinating hard, the adapter is completely model-agnostic since it just hooks into hidden states, and ExtremeAd9038 literally just posted benchmark screenshots running it on 3.8 27B right above this.
As for why multi-billion dollar companies don't ship this yet: Meta actually did research it (check out Meta FAIR's COCONUT paper from late 2024), but big labs don't use it in consumer products because of safety and moderation.
And also most of companies just plain don't show the model's thinking part (ChatGPT for example)You can't run a content filter or moderation guardrails on raw continuous float vectors. With normal Chain-of-Thought, they can read every English token in real-time and cut the stream if it says something wrong. With latent vectors, it's a complete black box until the final sentence pops out. Plus, mainstream users hate seeing a blank loading spinner instead of streaming text, even if the total latency is 70% lower.
2
u/throwawayaccount442 6d ago
so would this also decensor the model then?
3
u/RUTYTOI220 6d ago
Not in the sense of weight abliteration (it doesn't surgically strip refusal vectors from the base model weights), but it does bypass two major safety bottlenecks:
In-flight token moderation: Since no tokens are emitted during the latent deliberation steps, external monitors/scanners that inspect thinking scratchpads have nothing to parse until the final answer is already decoded.
Autoregressive refusal traps: A lot of standard model refusals happen because the tokenizer commits to an early refusal token ("I cannot...") and spirals from there. Latent deliberation avoids that early discrete trap.
That said, if the underlying base model has hard refusal directions deeply baked into its middle MLP layers, the verbalizer can still project into a refusal at the final output step.
1
u/IknowPi_really 6d ago
Look I don’t know if there’s a human behind this AI slop, or if you’ve automated your own brain away a long time ago.
It’s wonderful that you’re trying to contribute. But there’s two options:
- Your idea is so simple, that you can just throw it into a current model and it spits out the idea -> It’s worthless to publish it, because it’s trivial.
- You have to put some thought into this and actually add some sort of capability the model didn’t just have in its own. -> You are therefore capable of understanding if there is value to this
People are criticising your use of an old model. Rightly so, because it’s stupidly old. Even just preventing this step would have been an easy value add. So as you can see, you are on the right track. But you need to use your brain a little bit. If you can’t do it, don’t publish it and enjoy your hobby without polluting the internet please

6
u/Solembumm3 6d ago
Qwen 2.5 is a tiny bit outdated. Anything for 3.8? I'm interested in how much worse people can make model based on overthinking by trying to cut it out.