Method
How this works, in full.
Every ranking on this site is arithmetic you can check. This page is the whole method: what the proof levels mean, how the seven factors are weighted, and what each published metric is required to mean. If something here seems wrong to you, it probably is — tell us.
Proof levels
Weaker evidence is weighted lower, not hidden
- Verified
We saw evidence — a document, a recording, or a reference call — and checked it.
×1.00
- Vendor attested
The vendor formally attested to this, with a named person accountable for it.
×0.90
- Vendor claim
The vendor typed this into their profile. Nobody has checked it.
×0.75
- Compiled by us
We wrote this from public sources. The vendor has not seen or confirmed it.
×0.55
The multiplier is applied to every factor score that rests on that claim. A vendor cannot raise their own proof level, and a compiled profile cannot carry verified claims at all.
Ranking
Seven factors, summing to 100
22
language
Weighted highest because language fidelity is where voice deployments in this market actually fail. Proficiency and mid-sentence code-switching are scored as separate questions.
18
capability
How much of what you asked for the vendor actually does, discounted by evidence quality.
16
speed
Scored against your deadline rather than an absolute scale — three weeks is fast for one buyer and fatal for another.
14
effort
Measured as your team's effort, not the vendor's. Needing your call transcripts is penalised heavily because assembling and redacting them is usually the longest task in the project.
12
poc
How cheaply you can find out whether this works. A pilot gated behind an annual contract loses most of the points.
12
evidence
Published metrics with denominators, and sample calls you can listen to.
6
commercial
Pricing model and whether the minimum commitment fits your budget.
A vendor who has not published a field scores as unknown rather than as zero — silence is not evidence of a bad answer. It costs them rank and turns into a question on your first call.
Language support
Two different questions, scored separately
Native
Trained on in-market speech. Holds up against regional accents and background noise.
Accented
Understandable but audibly non-native. Fine for menus and confirmations, risky for sales.
Basic
Demo-grade. Expect to be disappointed on a real call.
Separately from proficiency, we record whether a vendor handles a caller switching language mid-sentence. Supporting Hindi and English is not the same capability as following someone who moves between them inside one sentence, and most vendors mean the first when they say they support both.
Training requirements
What it costs you before go-live
Prompt only
Write instructions in plain language. Changes take minutes. No data assembly.
Documents
Point it at your existing PDFs and help centre. Days, if the documents already exist.
Call transcripts
You must supply real call transcripts before it performs. Budget weeks for collection, redaction and consent review — this is usually the longest pole, and it is your effort, not theirs.
Custom model
A model is fine-tuned for you. Highest ceiling, but retraining is a scheduled project rather than an edit, so every later change is slow too.
The metric dictionary
A number without a definition is marketing
Vendors may only publish numbers that resolve to an entry here, and every number needs a sample size. Below the credibility floor for its metric, a figure is shown struck through.
Containment rate
needs n ≥ 500
Share of conversations fully resolved by the bot with no transfer to a human and no callback within 24 hours.
Denominator — All conversations that reached the bot, including ones abandoned in the first five seconds.
Commonly distorted by — Excluding abandoned calls from the denominator, which can move the figure 15–20 points without changing anything real.
Intent accuracy
needs n ≥ 1,000
Share of utterances where the intent the bot selected matched a human annotator's label, on a held-out sample.
Denominator — Utterances annotated, not calls.
Commonly distorted by — Reporting on the training distribution rather than held-out production traffic. Ask which one, always.
Speech recognition WER
needs n ≥ 10,000
Word error rate of the speech-to-text layer against human transcription, on production audio in the stated language.
Denominator — Words transcribed.
Commonly distorted by — Quoting a benchmark figure on clean studio audio. Telephony audio is 8 kHz and noisy; the real number is usually far worse.
Response latency (p95)
needs n ≥ 5,000
95th percentile of the gap between the caller finishing speaking and the bot beginning to speak, measured end to end.
Denominator — Turns measured.
Commonly distorted by — Quoting median instead of p95, or timing only model inference and excluding telephony transit — where most of the delay lives.
Barge-in handling
needs n ≥ 500
Share of caller interruptions where the bot stopped speaking within 300 ms and responded to what was said.
Denominator — Interruption events.
Commonly distorted by — Counting the stop but not whether the bot then answered the right question.
Connect-to-qualify rate
needs n ≥ 1,000
Share of answered outbound calls that reached a defined qualification outcome, agreed with the client in advance.
Denominator — Answered calls, not dialled calls.
Commonly distorted by — Shifting the qualification definition per client so the number is not comparable across vendors. Read the scope text.
Code-switch accuracy
needs n ≥ 500
Intent accuracy restricted to utterances containing a mid-sentence language switch, such as Hindi-English.
Denominator — Code-switched utterances only.
Commonly distorted by — Not measuring it at all, and quoting the blended figure instead. If a vendor sells into India and has no number here, that is itself the finding.
All-in cost per call
needs n ≥ 1,000
Total monthly spend divided by conversations handled, including platform fees, telephony and overage.
Denominator — Conversations handled in the billing period.
Commonly distorted by — Quoting the per-minute rate alone, which excludes the platform fee and makes low-volume pilots look far cheaper than they bill.
What we do not do
- We do not take placement fees, and ranking cannot be bought.
- We do not sell contact lists or broker introductions for a fee.
- We do not privately recommend a vendor when asked. Every buyer gets the same arithmetic, and the moment we start recommending, our independence is gone and vendors are right to treat us as a competitor.
- We are not a party to any contract you sign. We are the directory, not the supplier.