A record should trace to a moment in the call.
The distinction sounds like positioning. It is actually the whole design, and it decides what happens on the day something is wrong.
The failure everyone has heard about
A hallucination is a confident answer that is not grounded in a true source. IBM describes it as a lack of groundedness, and the important part of that phrase is that the problem is structural rather than a matter of the model trying harder.
Grounding a model in real call data reduces the risk substantially. It does not erase it. Any system that tells you it has eliminated hallucination is describing a hope, and the interesting question is what it does about the part that remains.
Generation and provenance are different jobs
Generation asks: produce a summary of this call. The output is fluent, it is usually right, and when it is wrong there is nothing in the artifact that tells you which part to distrust.
Provenance asks: for each field in this record, which moment in the call is it from. A field that cannot answer does not get filled. It gets flagged.
The second is less impressive in a demo and much easier to live with in an audit, because a reviewer can check a field in seconds rather than re-listening to a call.
What that means a system has to do
- Attach a source to every extracted value: a timestamp in the audio, a line in a document, a field in an email.
- Leave a field empty and flagged rather than guessing it. An empty field is a task. A confidently wrong one is a liability.
- Make the reviewer's job checking rather than rewriting. If a person has to redo the work to trust it, nothing was saved.
- Keep the approval as a record, not just a screen. Two years later, the question will be who approved it and when.
How to test it in a demo
Ask the vendor to run a call where a key detail is genuinely ambiguous. Somebody mumbles a policy number, or says a date and corrects it. Then watch what the system does.
If it fills the field, ask what it based that on. If it flags the field and shows you the moment, you are looking at a system built the second way. That one test tells you more than an hour of feature comparison.
Related: questions to ask an AI vendor · choosing your capture tier
Want this looked at on your own floor?
Twenty minutes on your systems and your numbers. You leave with the map either way.