aivoicebots

Method

How this works, in full.

Every ranking on this site is arithmetic you can check. This page is the whole method: what the proof levels mean, how the seven factors are weighted, and what each published metric is required to mean. If something here seems wrong to you, it probably is — tell us.

Proof levels

Weaker evidence is weighted lower, not hidden

  • Verified

    We saw evidence — a document, a recording, or a reference call — and checked it.

    ×1.00

  • Vendor attested

    The vendor formally attested to this, with a named person accountable for it.

    ×0.90

  • Vendor claim

    The vendor typed this into their profile. Nobody has checked it.

    ×0.75

  • Compiled by us

    We wrote this from public sources. The vendor has not seen or confirmed it.

    ×0.55

The multiplier is applied to every factor score that rests on that claim. A vendor cannot raise their own proof level, and a compiled profile cannot carry verified claims at all.

Ranking

Seven factors, summing to 100

  • 22

    language

    Weighted highest because language fidelity is where voice deployments in this market actually fail. Proficiency and mid-sentence code-switching are scored as separate questions.

  • 18

    capability

    How much of what you asked for the vendor actually does, discounted by evidence quality.

  • 16

    speed

    Scored against your deadline rather than an absolute scale — three weeks is fast for one buyer and fatal for another.

  • 14

    effort

    Measured as your team's effort, not the vendor's. Needing your call transcripts is penalised heavily because assembling and redacting them is usually the longest task in the project.

  • 12

    poc

    How cheaply you can find out whether this works. A pilot gated behind an annual contract loses most of the points.

  • 12

    evidence

    Published metrics with denominators, and sample calls you can listen to.

  • 6

    commercial

    Pricing model and whether the minimum commitment fits your budget.

A vendor who has not published a field scores as unknown rather than as zero — silence is not evidence of a bad answer. It costs them rank and turns into a question on your first call.

Language support

Two different questions, scored separately

  • Native

    Trained on in-market speech. Holds up against regional accents and background noise.

  • Accented

    Understandable but audibly non-native. Fine for menus and confirmations, risky for sales.

  • Basic

    Demo-grade. Expect to be disappointed on a real call.

Separately from proficiency, we record whether a vendor handles a caller switching language mid-sentence. Supporting Hindi and English is not the same capability as following someone who moves between them inside one sentence, and most vendors mean the first when they say they support both.

Training requirements

What it costs you before go-live

  • Prompt only

    Write instructions in plain language. Changes take minutes. No data assembly.

  • Documents

    Point it at your existing PDFs and help centre. Days, if the documents already exist.

  • Call transcripts

    You must supply real call transcripts before it performs. Budget weeks for collection, redaction and consent review — this is usually the longest pole, and it is your effort, not theirs.

  • Custom model

    A model is fine-tuned for you. Highest ceiling, but retraining is a scheduled project rather than an edit, so every later change is slow too.

The metric dictionary

A number without a definition is marketing

Vendors may only publish numbers that resolve to an entry here, and every number needs a sample size. Below the credibility floor for its metric, a figure is shown struck through.

  • Containment rate

    needs n ≥ 500

    Share of conversations fully resolved by the bot with no transfer to a human and no callback within 24 hours.

    DenominatorAll conversations that reached the bot, including ones abandoned in the first five seconds.

    Commonly distorted by Excluding abandoned calls from the denominator, which can move the figure 15–20 points without changing anything real.

  • Intent accuracy

    needs n ≥ 1,000

    Share of utterances where the intent the bot selected matched a human annotator's label, on a held-out sample.

    DenominatorUtterances annotated, not calls.

    Commonly distorted by Reporting on the training distribution rather than held-out production traffic. Ask which one, always.

  • Speech recognition WER

    needs n ≥ 10,000

    Word error rate of the speech-to-text layer against human transcription, on production audio in the stated language.

    DenominatorWords transcribed.

    Commonly distorted by Quoting a benchmark figure on clean studio audio. Telephony audio is 8 kHz and noisy; the real number is usually far worse.

  • Response latency (p95)

    needs n ≥ 5,000

    95th percentile of the gap between the caller finishing speaking and the bot beginning to speak, measured end to end.

    DenominatorTurns measured.

    Commonly distorted by Quoting median instead of p95, or timing only model inference and excluding telephony transit — where most of the delay lives.

  • Barge-in handling

    needs n ≥ 500

    Share of caller interruptions where the bot stopped speaking within 300 ms and responded to what was said.

    DenominatorInterruption events.

    Commonly distorted by Counting the stop but not whether the bot then answered the right question.

  • Connect-to-qualify rate

    needs n ≥ 1,000

    Share of answered outbound calls that reached a defined qualification outcome, agreed with the client in advance.

    DenominatorAnswered calls, not dialled calls.

    Commonly distorted by Shifting the qualification definition per client so the number is not comparable across vendors. Read the scope text.

  • Code-switch accuracy

    needs n ≥ 500

    Intent accuracy restricted to utterances containing a mid-sentence language switch, such as Hindi-English.

    DenominatorCode-switched utterances only.

    Commonly distorted by Not measuring it at all, and quoting the blended figure instead. If a vendor sells into India and has no number here, that is itself the finding.

  • All-in cost per call

    needs n ≥ 1,000

    Total monthly spend divided by conversations handled, including platform fees, telephony and overage.

    DenominatorConversations handled in the billing period.

    Commonly distorted by Quoting the per-minute rate alone, which excludes the platform fee and makes low-volume pilots look far cheaper than they bill.

What we do not do

  • We do not take placement fees, and ranking cannot be bought.
  • We do not sell contact lists or broker introductions for a fee.
  • We do not privately recommend a vendor when asked. Every buyer gets the same arithmetic, and the moment we start recommending, our independence is gone and vendors are right to treat us as a competitor.
  • We are not a party to any contract you sign. We are the directory, not the supplier.