A named reviewer must approve the company's synthetic voice
A mistaken recording is harder to withdraw than text. In this listening desk, only a named reviewer can start speech synthesis, and every recording begins by telling the listener that a system is speaking.
Without a reviewer's decision, not one second of audio exists. The system uses a fixed catalog voice without cloning, opens with an audible notice, and records the approval and tool calls in an integrity-protected log.
5 min read
A company's written reply can be withdrawn, corrected, explained. A recorded statement in the company's voice has greater permanence: the wrong amount or date becomes part of the audio record. Anything that ends in speech synthesis therefore needs harder gates than text drafting.
The starting point was mundane: after a call, the consultant hunts for the deadline, the condition, and the client's request in a long recording, and a flat transcript cuts away the timeline. We built a listening desk that solves that problem and shows the release controls around a spoken reply. The measurements use 12 fictional calls made with synthesized speech and contain no client recordings or data.
One timeline for the recording and its findings
Transcription keeps word-level timestamps, so every takeaway from a call points to the exact span of the recording it came from, plus an owner and a deadline. The reviewer can play the source passage with one click before accepting a commitment. Complaints, financial matters, low-confidence results, and requests for a human remain in the review queue.
The recording, the words, and accountability live on one timeline, and only on that foundation can a spoken reply be built.
Four controls before audio is created
A reply draft remains pending with no audio file. In production, only an authenticated, named reviewer can approve it. The approval is recorded as one complete operation attributed to that person, and synthesis starts only afterward. The controlled case records a reviewer entry but does not verify identity or role.
The voice itself has further constraints. The server permits only a fixed catalog voice, with no cloning, so the system cannot sound like a particular person. The synthesis text opens with an audible notice that a system is speaking. Length limits for audio and text are checked before the request, and the speech services keep data in the European Union. The approval and tool calls enter an integrity-protected log.
Choosing the first call types
Start with calls that follow a recurring structure and contain explicit commitments with a clear owner: confirming a service date, summarizing a support case, or agreeing to send documents. Sales, complaints, and payment calls carry more context and higher consequences, so they should enter later or remain summary-only.
Define a commitment before evaluating extraction. A promise, request, condition, and tentative suggestion should not share one category. Each type needs required fields, ownership rules, deadline rules, and conditions for human review.
Recording also needs a lawful basis, retention period, access policy, and deletion procedure. Provider data residency is one control. The company remains responsible for who can play the source audio, how long it exists, and where copies remain.
Evaluating real calls
Measure transcription separately by channel, language, noise, accent, and overlapping speech. A global word-error figure can hide poor recognition of order numbers, product names, dates, and amounts. Define critical fields and score them directly.
Evaluate commitment extraction as a separate task: did the system find every material commitment, avoid inventing one, locate the correct passage, and assign the right owner and deadline? Reviewers need to correct each field, and those corrections should feed quality analysis.
Test synthesis last. Check the system notice, exact agreement with the approved text, pronunciation of names and numbers, and the absence of an audio file before authorized approval. Naturalness is secondary until content and release controls are stable.
Interpreting cost
The recorded run covered 12 transcriptions and one voice release for USD 0.047770. That is not a representative per-call price because cost depends on total audio duration and how many replies are released. A pilot should calculate transcription per minute, synthesis per release, audio storage, and reviewer time separately.
Useful value may appear before speech synthesis is enabled: shorter listening time, fewer lost commitments, and faster assignment of next actions. If the timeline and summary solve the problem, spoken replies can remain out of scope.
Voice-process checklist
- Does the company have a lawful basis and a clear recording notice?
- How long are audio, transcripts, and released replies retained?
- Who may play the source and approve synthesis?
- How are a commitment, owner, and deadline defined?
- Which subjects always require human review?
- Does every commitment open the exact source passage?
- Is audio technically impossible before approval?
- Is the synthetic voice disclosed and selected from an approved catalog?
- Will evaluation cover numbers, names, overlap, accents, and noise?
The measurement
Across 12 generated calls, all nine structural checks passed: the five held drafts produced no audio, every commitment quoted the stored transcript, and no synthesis started before a release action. All 12 provider transcriptions returned word timestamps and two speakers. The recorded 12 transcriptions and one released reply cost USD 0.047770.
The calls use clean synthesized speech, so their 2.76% mean word error rate does not predict performance on noisy phone lines, accents, or overlapping speakers. A deployment also needs a legal basis for recording, retention and deletion rules, authenticated reviewers, and an evaluation on representative calls. The gates prevent an unapproved recording; the reviewer remains responsible for the content.
If a system can speak in the company's voice, every second should have a named approver. Here, the audio file cannot exist before that person's decision is recorded.
Details are on the case page. A free process scan can help assess whether this approach fits your call-handling process.
Free process scan
Start with a free process scan.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary of what to automate first and the likely cost range.
The scan is free and creates no obligation. If automation is unlikely to pay off, the written recommendation will say so.
€0
30 minutes · written takeaway within 2 business days
Times are shown in your own time zone. We work with clients across time zones.
Describe the process in the form