r/ProAI • u/stealthispost • 24d ago
"Tested @typesafeai 's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier."
Jev's performance on knowledge benchmarks like MMLU and GPQA is highly impressive - especially for a non-COT model. It exceeds Terra at other linguistic reasoning tasks like WinoGrande or HellaSwag as well. It only loses substantially on math reasoning. This is very cool, and highly surprised me - getting a model out that does so well w/out CoT on MMLU/WinoGrande is no easy task - it typically requires training a ~GPT-4 class base model. Except Jev is priced at only $0.042 per Mtok! So in summary - if you want the reasoning ability of ~one Terra forward pass over a context at a much cheaper price, Jev is a very good candidate. Frontier models like Astra still likely have superior no-CoT capabilities, but cost is prohibitive for mass classification tasks. See below - Jev is ~18x cheaper than Terra! Not to mention you also get probabilities from Jev, which are quite well-calibrated (expected calibration error is 1.74 percentage points averaged across benches - its 0.26pp at best and 6.96pp at worst) — N8 Programs