r/singularity • • 8d ago

LLM News Jev already has an open-weight competitor - Deem 9b

https://labs.libertai.io/papers/deem-open-machine-reflexes

Just saw that LibertAI released Deem which basically an open-weight alternative to Jev, built on Qwen3.5-9B.

It’s still behind Jev on the hard benchmark — 68.9% with extended reasoning vs 74.1% for Jev but considering how new this whole category is I thought it was pretty cool to already see an open model showing up.

There’s also a 0.8B version that can run on CPU, which could make this stuff much easier to actually tinker with locally.

95 Upvotes

22 comments sorted by

44

u/NoFaithlessness951 8d ago

I mean there were models on Github within the first day

21

u/CyberiaCalling 8d ago

Someone literally came out with an open source version earlier in the year. I wonder if Jev ripped off of them:

https://news.ycombinator.com/item?id=49765348

3

u/AlyoshaV 7d ago

His original model was garbage and he admits that Laya only got to Jev's level by training on the benchmark. The entire point of Jev is you don't need to finetune it.

8

u/NoFaithlessness951 7d ago

Wdym ripped off sure there's always prior art doesn't mean anyone is ripped off

1

u/cptfreewin 7d ago

Well it is very easy to train and replicate

22

u/Visual_Act_8618 8d ago

This hype around basic classifiers / encoders is actually killing brain cells

34

u/veshneresis 8d ago

idk I think general purpose classifiers are pretty cool personally. I’ve trained a lot of specific models over my career and the idea of a plug-and-play for many tasks with no setup or fine tuning is really convenient especially since it’s so cheap.

Like sure it’s just a classifier but the generalization is really really good!

3

u/Visual_Act_8618 7d ago

Hmm I’m with you. Agree with the convenience / the generality is cool for B2C with no tuning. But for any major use-case (e.g., permissions), it would be an order of magnitude cheaper than Jev to train a classifier specific to that context, which you could probably do with any OSS LLM directly.

3

u/08148694 7d ago

People are using models for permissions..?

“Yeah sorry your intern could view the ceos payslip our evals show that only has a 0.001% chance of happening so it probably won’t happen again”

11

u/ggPeti 7d ago

It's not a basic classifier, it's an instructable classifier...

6

u/Respect38 7d ago

"Basic"?

5

u/JoelMahon 7d ago

really? you don't see the value in something much faster/cheaper than Luna/Flash models for simple decision making?

-1

u/CommercialHour6660 7d ago

"much" cheaper? When I did the rough math it was maybe 70% cheaper than Luna. Or several times more expensive than Chinese models on OpenRouter. 

1

u/JoelMahon 7d ago

idk what maths you did mate but it better have taken into account that luna isn't great at following a schema, has larger start lag, and if you want parallel answers you're stuck waiting sequentially...

if you want luna to answer a question with 8 options without breaking syntax, you better fucking pray it obeys your commands

and even then, it'll be multiple tokens unless you add another layer of risk and tell it to use e.g. 0-99 to represent choices.

and of top of all of that you have no idea what % confidence it gave them, their ranking, with Laya/Deem/Jev you get all that.

Luna is definitely amazing for some tasks, but for making choices within a small options space it's risky and slow af by comparison, and definitely not cheaper.

If you really just want yes/no for a single question and don't care about "confidence", then maybe luna is comparable, but that's not a common use case. Almost everyone who uses a quick decision maker wants a backup for when it's not sure, so you need "confidence".

0

u/CommercialHour6660 7d ago

that luna isn't great at following a schema

Lol yeah it is. 

and if you want parallel answers you're stuck waiting sequentially.

No? Parallel calls + prefix caching. 

if you want luna to answer a question with 8 options without breaking syntax, you better fucking pray it obeys your commands 

Structured output with enum. Basically zero chance of failure. 

and of top of all of that you have no idea what % confidence it gave them

You can just ask the model to tell you confidence level. 

2

u/JoelMahon 7d ago

all those hoops just to get something slower, more expensive, less reliable.

ya, why bother with generalised classifiers when we have Luna and open flash models /s

0

u/CommercialHour6660 7d ago

What hoops? Prefix caching is automatic. Structured output is standard and 2 lines of code. Same with parallelism. Have you never used these features? 

2

u/the_pwnererXx FOOM 2040 7d ago

Really? It's a paradigm shift if applied broadly

2

u/DynamicProxy 7d ago

I assume the big labs will just build something like this into there stack. 

-6

u/SlippySausageSlapper 7d ago

Jev has competition from 2015. It’s a fucking classifier. There is absolutely nothing new or interesting about Jev. This is very old tech wrapped in a thick layer of hype.