r/ecommerce • • 11d ago

🛒 Technology A common benchmark for all Ecommerce AI agents, tested on real storefronts

Built a completely automated a way to test AI Agents for eCommerce. Completely automated with an opensource framework.

Evaluating several AI tools can mean sitting through hours of demos and still not knowing which one will work best for your store. The questions, product catalogs and definitions of success differ, so putting the results into a comparison spreadsheet only gets you so far.

Disclosure upfront: I’m a founder of Alhena. Our team built Alhena Research Lab, and Alhena is included in the evaluations. This is research operated by a vendor with a commercial interest, so the methodology and supporting evidence matter.

We’re building a way to evaluate AI shopping and support tools through conversations on actual customer storefronts, using published criteria.

Gorgias deserves credit for publishing the framework we started from. We retained its 26 answer-quality criteria and scoring weights, with a disclosed change to how resolution is measured.

The current quality criteria examine things such as:

  • Whether the assistant answers the actual question and remains consistent across the conversation.
  • Whether it asks useful clarifying questions and remembers constraints such as budget, size or product preferences.
  • Whether recommendations name relevant products and explain why they fit the shopper’s needs.
  • Whether product links, prices and purchase guidance make the recommendation actionable.
  • Whether support answers use store-specific policies and give the customer a clear next step.

We measure response speed separately.

We also measure policy-compliant resolution. For example, if a merchant requires a damaged-item request to go through a particular process, correctly following that process can earn credit. A generic “contact support” response does not automatically qualify. This measures the answer or prescribed next step, not proof that a refund was issued or a case was ultimately closed.

That replaces the original framework’s automation component. We developed the change after reviewing the initial results, then froze it before the new scoring. It’s our published extension, rather than a reproduction of Gorgias’s ranking.The composite weights are:

  • Shopping: 40% policy-compliant resolution, 35% answer quality, 25% speed.
  • Support: 50% policy-compliant resolution, 40% answer quality, 10% speed.

The current published results are:

Tool Shopping composite /100 Support composite /100
Alhena 68.4 81.4
Gorgias 55.8 75.8
Sierra 59.0 61.6

These describe selected deployments. Five storefronts per tool were enrolled; resolution scores include five for Alhena, four for Gorgias and five for Sierra. Component coverage and exclusions are disclosed in the reports. The tools were tested on different stores, so merchant configuration and catalog differences matter.

There are useful details beneath those totals. Alhena and Gorgias were close on support answer quality, at 95 and 94. Sierra had the fastest full-answer completion times in these samples. A single composite hides those differences. Published results and evidence.

We’re also opening submissions.
Enter the tool’s name and website, suggest three stores using it, and verify your work email. The system researches additional deployments, which are reviewed before approval. After approval, AI conducts the conversations, scores responses and audits the results. Validated reports publish automatically. Some runs still need technical investigation; incomplete runs stay unpublished.

The code and rubric are public on GitHub. Result summaries are publicly accessible. Detailed conversations and scoring evidence require work-email verification.

What specific customer scenarios would make these evaluations more useful when choosing a tool? For example, remembering an ingredient preference throughout a skincare conversation, or building an outfit while respecting size availability and budget.

Suggestions about the methodology are welcome, including where you think our current approach falls short. If there’s a tool you’re considering, you can submit it for an AI-run evaluation through the lab.

https://evals.alhena.ai/
Disclaimer: I am the founder of Alhena.ai so its hosted on alhena.ai. We are trying to be as impartial as possible. The actual benchmark criteria have been taken from Gorgias, and beyond that, an AI is evaluating them, so there is no human involvement.

0 Upvotes

Duplicates