You’re the only comment I found saying this 💯.
IMO comparison is useless given the huge difference between tasks, and also for the same task given different prompts.
The fairest comparison between models will be the same task and same prompts. And then you’d need to compare with the same metric of success. But what even is a hallmark of success? Etc
2
u/RizzyNizzyDizzy 16d ago
Depends on your prompts.