Customer-service reports become hard to trust when every conversation must fit a success label. An answer may have come from approved knowledge. A person may have taken over. Or the system may not have had a source that supported a response. Those are different events and should stay different in the record.
SayTrue uses three outcomes: grounded, escalated to a human, and no source.
Grounded means approved knowledge supported the answer
A grounded outcome links the conversation to the knowledge used: an approved entry or document context that covered the question. The transcript and source log let a supervisor inspect that chain. Grounded does not mean that a customer loved the answer or that the business result is known; it describes the provenance of the response.
Escalated means a person was needed
An escalation is not disguised as an AI resolution. The conversation records that it was handed toward a human path, with the configured department or request flow where applicable. Working hours matter here: outside them, the agent should record the request and contact details instead of transferring into an empty office or promising an unscheduled callback.
No source means the evidence was absent
A no-source outcome keeps an unsupported question visible. The agent can state that it cannot confirm, then follow the configured next action. The question may also enter the knowledge-gap report, grouped with frequency so the team can see whether the same missing fact returns.
There is no decorative fourth category
We do not add an invented confidence score and call it an outcome. A model’s confidence-like number would not answer the operational question: which approved source supported the response, which person received the handover, or which fact was missing? The three categories map to records a supervisor can inspect.
Honest counting changes the dashboard
Monthly reports count the outcomes separately and exclude test traffic. When there are no conversations or ratings, the interface can show an empty state instead of manufacturing a percentage. The report is useful because its denominator reflects the conversations that actually happened.
The knowledge-gap list closes the loop. A repeated no-source question becomes a candidate for a new approved entry. The team writes it, chooses whether the agent may answer, ticket, or transfer, and watches future conversations. Coverage grows through reviewed facts, while escalations and remaining blanks stay visible on their own terms.