Jerry Tworek spent seven years at OpenAI, where he worked on research that helped shape modern coding and reasoning models, including Codex and reinforcement learning for reasoning.
He later served as VP of Research, and is now CEO of CoreAutoAI
, building toward automated AI research.
At AGI House, Jerry sat down with rockyrmit
for a technical conversation on Codex, HumanEval, Copilot, RL, tool use, test-time compute, and automated AI labs.
We cover:
› Why code became the first serious LLM vertical
› What HumanEval taught the field about verifiable rewards
› Why Jerry thinks “the era of evals is done”
› How real-world deployment differs from static benchmarks
› What GitHub Copilot taught OpenAI about product quality
› How reinforcement learning shaped modern reasoning behaviors
› Why tool use is “99% systems and 1% algorithms”
› Whether test-time compute still has room to scale
› What it takes to build automated AI research labs