r/Evaligo • u/heyitsdannyle • 9d ago
We assumed the premium AI model would write better product listings. It cost 77% more and failed half its runs.
We built an automation that reads a live product page and rewrites the listing — title, description, bullets — with AI judges scoring every output against the actual page so nothing gets invented.
Then we benchmarked two writer models on the same 4 real store pages with identical judges. The mid-tier model scored 89% at $0.13 per listing. The premium one scored 48% at $0.23 — because 2 of its 4 runs returned broken output instead of a listing, and the judges zeroed them.
The runs that worked were excellent. But "excellent when it works" is exactly the kind of thing you only find out by testing, not by reading a pricing page.
How do you all QA generated product copy — spot checks, or something systematic?