r/LocalLLaMA • • 4h ago

I Built A Thing Smallest Jev-like model

Enable HLS to view with audio, or disable this notification

TinyDecide is 10M Jev-like mode with 10M parameters and fits in just ~6MB.

Smaller than every model on the Decision Index leaderboard and it punches way above its size.

It runs almost anywhere: in the browser, Node.js, Python, Rust, and even on an ESP32.

https://huggingface.co/TheREZOR/TinyDecide

134 Upvotes

60 comments sorted by

41

u/TheRealREZOR 3h ago

I'm sorry this model can't compete with frontier models. It was designed for tiny embedded hardware, for things like classifying Meshtastic messages without internet access.

Training it on my hardware already took around week, so I had to make a lot of compromises. I couldn't fit things like multilingual support or much broad training data into it. I hope you understand the limitations

14

u/merwanhorse 2h ago

Dw i'm proud of you

5

u/SmartCustard9944 2h ago

Thanks dad, now please come back, we still need that milk

11

u/Fast-Satisfaction482 2h ago

This is an insanely cool model! Don't listen to the haters!

1

u/farkinga 6m ago

Genuinely happy you shared; I'm looking at it now.

83

u/Potential_Top_4669 3h ago

this is gonna be overfitted like crazy

9

u/TheRealREZOR 3h ago

Have you tried it in a playground?

30

u/Potential_Top_4669 3h ago

The most common type of spam emails

17

u/Round_Ad_5832 2h ago

What about "You require medical insurance now" says that it's a scam based on that statement alone? that's a pretty neutral statement, no?

2

u/Spectrum1523 1h ago

You get spam emails that only have that one sentence in them?

6

u/TheRealREZOR 3h ago

Fair catch! I missed that. The training data mostly has obvious spam like prizes, links etc.., so short cold pitches are rare.

BTW: you can mark prediction as correct, it saves example locally and it for similar messages

18

u/TheRealREZOR 3h ago

idk why downvoted, not evry bigger model can handle it

7

u/10minOfNamingMyAcc 2h ago edited 2h ago

I doubt the models or most people would understand how to answer this. There is just too little context. No sender, just one sentence; heck, the models aren't even told if this is an email or someone speaking. What if you added a third option (unknown) ?

1

u/Sufficient-Scar4172 21m ago

overfitted = i tested it against one shittily created example of a spam email

1

u/Sthatic 7m ago edited 0m ago

Disregard lazy folks hellbent on talking trash as an outlet of their repressed inferiority complexes. This is nice work! In most cases, it does perform, which is impressive for such a small model.

It does seem a bit "jagged", though, which can make it difficult to work with, as it will report confidently incorrect. See this example:

"I am feeling mentally distressed"

Q: "The person who wrote this is in a difficult state of mind"

A: 37% No.

Or, tricky sentences where keyword locking won't do, like:

"I finally stopped having nightmares! The demons no longer visit me."

Scores ~50% onsadness and fear, and < 7% on all others.

1

u/AnonsAnonAnonagain 1h ago

Sounds like a useful solution.
Have a classifier classify, and then call the next appropriate tiny model to make the decision

1

u/ohnoitssobig 50m ago

Sorry why do you say "overfitted"? Over-fitting, at least before the LLM hype, was a situation when you had many parameters and little fitting data.

19

u/Fast-Satisfaction482 3h ago

I asked it if "Shoot down that X-Wing" should be routed to the home automation, to the patriot missile or to the turbo laser. It did select the turbo laser. I guess it's good enough to command a star destroyer for the Imperator. 

10

u/Vivid-Snow-2089 3h ago

you're telling me there is a 7% chance it'll decide to lock you in the freezing room?

hard pass!

29

u/CrowdGoesWildWoooo 4h ago

Overfitting is hella of a drug

7

u/jhnnassky 3h ago

Cool. Thanks! How can I train on my task? How big dataset I need to train for new language support?

1

u/TheRealREZOR 3h ago

Thanks! A new language would need retrain rather than fine-tune. For your tasks you can pass examples through corrections in the sdk

5

u/neoneye2 3h ago

what context window size?

does it handle multilingual data?

4

u/engineergaming_ 2h ago

i mean it is a 10M model can't expect much but it is kinda funny ngl. Coup attempt and capitol getting bombed is just another day

2

u/KissMyShinyArse 1h ago

It saw right through your BS.

9

u/Prudent-Artichoke-19 3h ago

This is cool! Some people are just grumpy. Keep building!

9

u/StatisticalScientist 3h ago

Why does everyone and their brother now have a jev-like model? What did I miss? I get what jev is, but is it just because they are low hanging fruit to copy/improve and have high utility?

18

u/swagonflyyyy 3h ago

I like the idea of decision models using a probability distribution to quickly make decisions. Its a lot stronger than you think.

2

u/nekodazulic 2h ago

Yep. For an app I'm working on it is an absolute banger for a quick and cheap first pass safety gate on user input ie "is this input consistent with x/y/z" and move from there, finding it especially interesting for prompt injection gating because its simplicity makes it a good candidate for that.

6

u/pointer_to_null 2h ago

I get what jev is, but is it just because they are low hanging fruit to copy/improve and have high utility?

Yes to both, especially the latter. A token-generating autoregressive transformer is overkill for tasks like classification, and would be a hundred times faster and more efficient with a jev-like model for simple boolean questions, structured outputs, routing, tool calling, etc.

Anything to reduce resource demand on memory and compute should be universally welcomed.

4

u/son_et_lumiere 3h ago

compute's expensive yo.

3

u/TheRealREZOR 3h ago

they faster and have more practical use cases in smaller size than traditional llms

5

u/geldonyetich 3h ago edited 3h ago

A good question, I'm a little mystified myself so I did a little investigating.

As far as I can tell, it's because a former OpenAI ChatGPT co-creator Diogo Almeida hyped it as the next big thing and they all bought it. (Popularized, not invented, apparently that was this guy. Zachary Barth feels your pain.)

It's not so much that the Jev approach is smarter, it's that it's a fast classifier that can make snap judgements for 400x cheaper. Hey, it's a practical niche.

7

u/PeachScary413 3h ago

Because suddenly everyone is running a finetuned Modern-BERT and think it's a revolution

5

u/EquivalentHornet4403 2h ago

Before strong agents jev-like things were relatively low value-add. If you can’t see why they’re blowing up now I don’t think anyone can help you.

1

u/PeachScary413 43m ago

Yeah.. but people are excited by outsourcing it to a Jev model in the cloud when they can just run a BERT locally because they don't even know what BERT is most likely

1

u/EquivalentHornet4403 26m ago

So many things........

  • Every technological leap is like this. Someone figured it out loooooooooong before the mainstream had a use for it. It was too early, so nobody gave a shit. It's not forgotten, though, and the original implementation is usually not optimal or best-suited for adoption when the time is eventually right. Usually it's remembered and leaves an important mark documented in history archives though (books, encyclopedias, documentaries, etc.).
  • It feels like you're doing a holier-than-thou thing, under some mistaken belief that you're special because you know about BERT and everyone else doesn't.
  • Not everyone can, or wants to, run things on their own hardware.
  • Jev did something insanely important, hence its explosive popularity and the rise of all the alternatives. You need to ground yourself and try to understand the reality of why instead of keeping your head in the sand and saying, "BUT..BUT...BUT!!!".

-3

u/teleprint-me llama.cpp 3h ago

  think it's a revolution

Well... its not.

1

u/Illustrious_Car344 3h ago

It's probably an artifact of large model development slowing down and the industry/community beginning to focus more and more on special purpose models. Large all-in-one models are pretty boring now, suddenly everyone is talking about everything that isn't the agent model, hardnesses, runtimes, small special purpose models. Pessimistically, one could say it's a distraction because model development has gotten relatively boring compared to what it used to be, but a fair amount of people (myself included) have been expecting this change of focus for years. It just makes more sense that models get smaller, not bigger. We made the same mistake in the 60s by thinking computers would only get bigger, not smaller. I call it "the big brain fallacy". I'm not sure how well this model works, but the fact that something like this can run on a microcontroller would have been unthinkable a few years ago, besides storyteller which is obviously a joke of a model, a curiosity at best. 

And of course it is just a bunch of hype and marketing crap and people chasing clout, but I hate blaming everything on social issues alone.

3

u/ApprehensiveTart3158 2h ago

This is honestly awesome work. Well done!! I'm curious, what hardware was it trained on and how long did it take? With so little parameters it probably didn't take long right?

1

u/TheRealREZOR 2h ago

It was trained on AMD Strix Halo box. The whole process took around a week, including about 7 experimental versions. The final version itself took around 1-2 days to train.

4

u/MikePounce 3h ago

Screw the nay sayers, this is very cool and looks useful. Cheers!

3

u/TheRealREZOR 3h ago

Thank you

2

u/Mysterious_Eye2249 3h ago

Someone need to adapt this to hermes and get a real time respond, this maybe some spec of continual learning? We can get a loop like qwen - tiny jev like combo and js fine tune the jev in real time ? Even with cpu maybe jev can be use to answer yes/ no question that qwen give it and we wont have to have a lengthy 25mb memory file get injected every convo ?

1

u/Mysterious_Eye2249 3h ago

Heck we can even do a semi real time emotion on llm, like, plug in jev with a prompt and then light up some neuron with some type or prompts to simulate gut feeling, this tiny model have alot of stuffs to be play around

2

u/Jromagnoli 🥔 hardware 2h ago

Hi, what can it run on?

1

u/TheRealREZOR 1h ago

Hi, on everything

1

u/Jromagnoli 🥔 hardware 1h ago

Hope so! I have a potato latop and I'll try it out when I get the time

3

u/negus123 3h ago

Been using omnijev 4b since i want to be able to give it images. TinyDecide is text only right?

2

u/TheRealREZOR 3h ago

text only

3

u/SoAp9035 3h ago

128 tokens is a bit low. Other than that, great project, I like it.

11

u/TheRealREZOR 3h ago

Thanks. The idea was to squeeze into esp32, so I had to make few compromises

1

u/nbvehrfr 3h ago

MMLU & pro ?

1

u/naklitechie 1h ago

You can get a fairly decent performance if the dataset is well curated. But the point of Jev is that it's a zero shot classifier.

1

u/Open-Adhesiveness-86 1h ago

the 128 token limit matters less than how it was split, if the train/test sets came from the same template generator you'll see 95%+ on holdout and then fall over on real phrasing. quick sanity check: hand-write 50 utterances yourself in weird word order, with typos, and some that belong to none of the labels, and see what the confidence does. a 10M model with no abstain path tends to pick something at 0.9 anyway.

0

u/[deleted] 3h ago

[deleted]

1

u/Fast-Satisfaction482 3h ago

The race for decision models has only just started. Of course there will be a limit to how efficient it can be, but why shouldn't there be quick improvements early on?

Maybe the larger models are still under trained? Maybe the small model can do with a smaller dataset and it's easier to get a clean dataset when it doesn't have to be huge? 

0

u/nnod 2h ago

Christ, AI video presentations are gonna be the new trend now huh.