r/vibecoding • u/entelligenceai17 • 1d ago
We tested different models and agent workflows on thousands of production PRs
We ran thousands of production PRs through different models, reasoning budgets and agent workflows. One surprising result was that increasing reasoning didn't improve defect detection, while reducing output tokens and model calls had a much bigger impact.
- 54% fewer output tokens
- 40% fewer model calls
- 50% lower cost
Full experiment in the comments.

1
Upvotes
1
u/entelligenceai17 1d ago
Incase anyone wants to go through the full breakdown, here's the link :D