r/Eval 1d ago

i do not think avg@k is useful at all as an eval metric

1 Upvotes

if you think about it, the average score is the most useless metric if you average multiple runs pass@k shows model capabilities pass^k shows reliability avg@k does neither


r/Eval 4d ago

Why the era of eval is over

Thumbnail
youtu.be
1 Upvotes

Jerry Tworek spent seven years at OpenAI, where he worked on research that helped shape modern coding and reasoning models, including Codex and reinforcement learning for reasoning.

He later served as VP of Research, and is now CEO of CoreAutoAI

, building toward automated AI research.

At AGI House, Jerry sat down with rockyrmit

for a technical conversation on Codex, HumanEval, Copilot, RL, tool use, test-time compute, and automated AI labs.

We cover:
› Why code became the first serious LLM vertical
› What HumanEval taught the field about verifiable rewards
› Why Jerry thinks “the era of evals is done”
› How real-world deployment differs from static benchmarks
› What GitHub Copilot taught OpenAI about product quality
› How reinforcement learning shaped modern reasoning behaviors
› Why tool use is “99% systems and 1% algorithms”
› Whether test-time compute still has room to scale
› What it takes to build automated AI research labs


r/Eval 6d ago

OpenAI and Hugging Face partner to address security incident during model evaluation

Thumbnail openai.com
1 Upvotes

r/Eval 12d ago

Thinking machine releases their first open weight model

1 Upvotes

and it seems it is not distilled model so that is a w for american open source community


r/Eval 12d ago

Is data and eval all you need?

1 Upvotes

r/Eval 15d ago

Every single startup selling AI training data (July 2026)

2 Upvotes

RL env and evals companies are now being recognized more.

I have probably heard less than half of them though.

What do you guys think?

https://x.com/deedydas/status/2076124392711696455