How we measure answer quality
This page used to publish quality scores for our assistants and for three named competitors. Those numbers were not the result of a measurement we could show you, so we removed them. Here is what is actually true today.
Every answer is already scored, for the customer who owns it
The platform reviews its own answers as they are given: each one is scored for factuality and flagged when it states something the retrieved content does not support. Those results are visible to each customer for their own site, alongside the questions the assistant could not answer at all.
That is a real measurement, and it is the one that matters to a buyer: how the assistant performs on your content, not an average across nine industries.
A published benchmark
A number worth publishing needs a fixed set of questions, a written rubric, and a run that someone outside this company could repeat and check. We are building that. Until it has run, there is nothing here to show, and we would rather show nothing than restate figures we cannot stand behind.
We will not score a competitor's product on our own question set and publish the result as though it were neutral. Where we compare, we compare on facts you can verify against a vendor's own published terms: price, channels, languages, compliance posture.
Judge it on your own content
The free tier gives you 100 conversations a month against your own material. That will tell you more than any scoreboard.