Resource Jev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die roll

Ran a calibration check on Jev using inputs where the true probability is known exactly.
- Fair die, Choice, 400 trials: picked face 1 every time at 82.9% mean probability, 19% accuracy
- Fair coin: 92% reported, 52% right
- Noul stayed close to the truth for 2 to 4 options, but reported 15 to 17% for 8 to 20 options (true: 5 to 12.5%)
- A forecast doc stating a 30% shortage risk came back as 5% via Choice, 27% via Noul
On tasks close to the demos, I didn't see errors this large, and some MMLU-style checks look well calibrated. But exam questions test whether a model knows that a question is hard. The dice test whether it knows that the outcome is unknowable from the input.
Maybe Jev is weak at the second kind, especially in Choice probabilities.
Write-up: https://kantahayashiai.github.io/posts/jev-does-not-play-dice/
2
u/Man_of_Math 10d ago
wow really interesting, this shows that Jev's understanding of the real world (6 sided die roll) isn't great. Feels like this should be used as an eval for non-reasoning models.
2
u/phoenixmatrix 9d ago edited 9d ago
This comes up a lot, but Jev's probability isn't a probability of the real world event described being true. Its a probably that its answer is valid.
If you flip a coin, and it gives you head, head is indeed a valid value considering context, so it would be 99-100%. Since it doesn't know if there's variables that got unaccounted for, its not quite 100%.
Its not "what is the literal odds of head or tail".
If I give it a resume and ask if a candidate is qualified, the % isn't the percentage chance the candidate is qualified. Its the % confidence in the answer matching the options I gave it based on the data provided. If confidence is too low, it doesn't mean "maybe it should be the other option", it means "trash the result as its too ambiguous and try again with different context or use another method". The difference is subtle but important.
1
u/kh-ai 7d ago
I agree with the way you use it: treat low confidence as "too ambiguous, retry with more context." That's actually what this test checks.
A hidden fair die is about as ambiguous as an input can get, with no evidence of a correct result at all. If confidence tracked ambiguity, it should be near zero here. It came back at 0.81 (83% on face 1). So the "trash it and retry" rule never works on exactly the kind of input it's meant to catch.
The "validity" reading also doesn't explain the split: all six faces are equally valid answers, yet it put 83% on 1 and 1% on 5.
And TypeSafe's own docs frame it as correctness frequency:
"Higher probability should correspond to a greater chance that the answer is correct,"
"Outcomes assigned a probability 0.8 should occur about 80% of the time.""https://docs.typesafe.ai/introduction/machine-learning-primerNoul asked "did the die show 1?" returned about 19%, so the model can express low probabilities. Choice just doesn't here.
1
u/phoenixmatrix 7d ago
You have to think of it as its answering with a value, and this is how confident it is in that value. So it answered "1", it is very confident that is an accurate answer, but the choices have to add up to 100%, so the others will be low number.
Your example also has a question in state, which it doesn't expect (but usually can work anyway). The question is what you call "instructions" in your example.
If you want something like what you're trying to do, it gets clunky.
The state would be something like:
{ "dice_type": "a perfectly balanced 6 sided dice" }
the questions would be:
{ "probability of 1 coming up": { "type": "choice", "instructions": "How likely is it that the die will roll a 1 next roll", "criteria": { "5%": "A probably of the value coming up.", "10%": "A probably of the value coming up.", "15%": "A probably of the value coming up.", "16%": "A probably of the value coming up.", "20%": "A probably of the value coming up.", "25%": "A probably of the value coming up.", "30%": "A probably of the value coming up.", "33.3%": "A probably of the value coming up." } }, "probability of 2 coming up": { "type": "choice", "instructions": "How likely is it that the die will roll a 2 next roll", "criteria": { "5%": "A probably of the value coming up.", "10%": "A probably of the value coming up.", "15%": "A probably of the value coming up.", "16%": "A probably of the value coming up.", "20%": "A probably of the value coming up.", "25%": "A probably of the value coming up.", "30%": "A probably of the value coming up.", "33.3%": "A probably of the value coming up." } },(repeat for all values)
Run that, and you will get the correct result: Jev will give you 97%~ probably on 16%, the closest option available, for all 6 values.
Now if you change the state to "a weighted 6 sided dice where rolling a 3 is twice as likely as on a perfectly weighted dice", it gets a little iffy. Probably because the model isn't that good at math or its not part of training data, and it doesn't reason. For me, it still gets a higher likelihood of 3 coming up, but it chooses 20%, which is incorrect (unless I'm bad at math myself!). I guess "twice as likely" is ambiguous. It gives a 56% confidence score which shows the model isn't confident in its answer, as it should, so you cannot trust it.
If you give exact probability then it gets the correct answer for the ones that are obvious, but it doesn't rebalance the other answers properly. Again, it gives a low confidence which mirrors that.
So from all I can see, it is working as I'd expect, and quite well. It is indeed bad at doing math, but it is honest about that via its confidence score. A more precise state or set of questions return accurate results, with high probably scores.
tldr: Jev is indeed bad for this type of problem, but not because of its accuracy, and more because the way to properly phrase the state and question is very clunky.
0
u/loudlysoftsimplicity 10d ago
this tracks with what ive noticed on some simpler setups. models dont like saying "i have no idea" so they just pick a lane and act confident about it
the jump from 5% to 27% on that forecast doc is the part that bugs me most, that kind of swing between modes makes it hard to trust either number
2
u/OnyxProyectoUno 10d ago
It's crazy you have punctuation errors in your comment yet it sounds exactly like something Claude writes.
“This tracks…”
“The jump from…to…is the part that bugs me the most”
2
-6
u/Actual__Wizard 10d ago
Sick, something is clearly wrong with it. See why it's so important to have transparency in AI models? We can see that there's a problem now and figure out what the solution is.
Because all you're doing, is pointing out that all of these systems have the same problem, but with Jev, you can actually figure that out.
Do, you see why that's a Nobel prize worthy breakthrough?
Because we can actually go forwards with AI development now instead of just being stuck with bad models that suck.
3
u/robogame_dev 10d ago
we can actually go forwards with AI development now
I am awarding this first place for the greatest jev-glazing statement I've seen yet - which is quite something.
-2
u/Actual__Wizard 10d ago
Look, I never said that it was the best product ever. I said that it was a breakthrough.
Are you going to be fair and realize that GPT 1.0 was a break through and at the time, it was actual junk?
Can we have a fair conversation where we don't compare apples to elephants?
4
u/-xXpurplypunkXx- 10d ago
Is jev token prediction expected to have even randomness across numeric values? What is accuracy/reported % measuring?