Evaluating AI for an insurance floor.
Most of these evaluations go wrong in the same way. The demo is judged on what the model can say, when the thing that decides the outcome is where the work already happens and what the software is allowed to touch.
Find the work before you look at features
Before any vendor call, sit behind three agents for one hour each. Write down every moment their hands leave the customer and go to a keyboard for the system's benefit rather than the customer's. That list is your evaluation criteria. Everything else is the vendor's list.
The usual entries: re-keying what the caller just said into the CRM, copying between the dialer and the quoting tool, writing the disposition, filling the application, chasing the missing document. Note which of those happen during the call and which happen after it. They are different problems and they buy different things.
Then ask what it runs on
The single most expensive answer in this category is "we replace your stack". You already own the dialer, the CRM and the policy system, your people know them, and the integrations that matter took years. A tool that works alongside them can be removed in an afternoon. A tool that replaces them cannot.
Three questions that separate the two:
- Which of my systems does this read, and which does it write? Get the answer per system, not as a category.
- What happens on the day we turn it off? If the answer involves a migration, it is a replacement wearing a layer's clothes.
- Does it need its own softphone, or does it work with the one on the desk?
Find where the human is, and whether they can be skipped
In a regulated line the interesting question is not whether a person reviews the output. Every vendor says yes. It is whether the person can be bypassed, and under what conditions.
Ask for the specific mechanism. "An agent approves" is a policy. "The record cannot post without an approval event, and here is what happens when someone tries" is a design. Ask to see the second one. Ask what an administrator can turn off.
McKinsey's State of AI reporting puts contact-center and customer-service automation among the most common uses organizations report, so you are not evaluating whether this is done. You are evaluating whether this vendor does it in a way your compliance function can live with.
Ask for the number you will be judged on
Whoever signs this will be asked, six months later, what changed. Decide now what that number is and how it will be read. The honest version has three parts: the measurement before, the measurement after, and who takes it.
If the vendor supplies all three, you have a marketing claim rather than a measurement. Take the baseline yourself, from your own system, before anything is installed. Our guide on measuring after-call work is one way to do that for the most common case.
The short version
- Watch three agents for an hour before you take a demo.
- Judge on what it runs on, not what it can say.
- Ask how the human gate is enforced, and who can switch it off.
- Take your own baseline before anyone installs anything.
- Know how you exit before you sign.
Related: questions to ask an AI vendor · choosing your capture tier
Want this looked at on your own floor?
Twenty minutes on your systems and your numbers. You leave with the map either way.