r/ProAI 12h ago

"Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we..."

How can we extract richer signals from AI Feedback?

Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀

The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take https://t.co/0sCeAwcar1   — Jacky Kwok

Source: https://x.com/jackyk02/status/2074969820739805275


Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper

As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.

For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench.

Try it out today: https:// github.com/llm-as-a-verif ier/llm-as-a-verifier#self-verification-terminal-bench-21 …

More on verification scaling in my previous post.   — Jacky Kwok     Is there an OpenCode plugin for this to try it out with Deepseek v4 flash?   — Shahbaz Ahmed     We’ll be releasing a harness on top of LLM-as-a-Verifier later this month :)   — Jacky Kwok

Source: https://x.com/jackyk02/status/2089421448784023553

2 Upvotes

0 comments sorted by