Skip to content
Discuss your operation

Home / Blog / Assist before automate

Contact center AI — voice only

Assist before automate: how call centers use LLMs on live phone calls

Real-time agent assist, auto after-call work, and full-call QA first — autonomous voicebots only after you can measure resolution on the phone.

Contact center floor and agent assist — LLMs helping humans on live phone calls

The quiet ROI is still on the agent’s screen

Most call-center AI pitches skip straight to the autonomous voice agent: a bot that answers the phone, resolves the issue, and never rings a human. That demo is real. It is also the wrong first move for most operations — including VICIdial and hybrid floors we run day to day.

On the phone, the highest-return LLM deployments in 2025–2026 still put a human on the line. The model listens, suggests, summarizes, and scores. It does not own the customer until the operation can measure resolution, not just containment.

This article is voice-only: live calls, IVR, agent assist, after-call work. Not chat. Not WhatsApp.

Quick answer: what a contact center LLM should do first

A contact center LLM should usually start behind the agent: live transcript, next-best action, knowledge snippet, compliance prompt, after-call summary, disposition suggestion, and QA scoring. That is the measurable path for an LLM call center rollout before a customer-facing bot speaks.

Put the model on the caller path only after call recording, STT quality, CRM context, warm transfer, redaction, and resolution metrics work on one queue. The sequence is still Analyze → Augment → Automate.

The problem with “automate the phone first”

Contact centers are under real pressure. Labor is still the bulk of cost (often cited up to ~95% of spend), and executives want AI to fix it. Gartner’s well-known forecast put conversational AI on track to cut $80 billion in agent labor costs by 2026. About 88% of contact centers now say they use some form of AI (industry synthesis, 2026).

The fine print is less flattering:

  • Those same Gartner-era forecasts still assumed only about one in ten agent interactions would be fully automated (up from ~1.6%).
  • Only about 25% of centers have fully integrated automation into daily work — not just licensed a tool.
  • McKinsey’s “gen AI paradox” still holds: roughly eight in ten companies use gen AI, and a similar share report no material earnings impact ((McKinsey, agentic AI report).
  • True agentic AI in production is thin. Syntheses of 2025 leadership surveys put it around ~11%, with copilots and RPA often rebranded as “agents” (InflectionCX operator guide).

Phone makes this harder, not easier. Voice adds latency budgets, barge-in, accents, background noise, PCI pauses, and a customer who cannot scroll back to reread a wrong answer. A bad voicebot is not a mildly annoying chat bubble. It is a hold queue that ends in a transfer and a story told twice.

Operator playbooks keep citing Forrester’s blunt line on customer-service gen AI: avoid consumer-facing applications first, at least until agent-facing foundations work (Cresta use-case guide).

What assist means on a live call

On voice, assist is a real-time loop:

Caller ↔ Human agent
           ↑
   Live STT transcript
           ↓
   LLM + knowledge + CRM context
           ↓
   Screen prompts: next best action, policy snippet, compliance line

After hangup:

Full call transcript → LLM summary + disposition → CRM
                    → 100% QA scoring / coaching flags

That is not a chatbot with a microphone. It is augmenting the agent who already owns the conversation. Industry catalogs put the same cluster at the center of contact-center AI: real-time guidance, knowledge surfacing, compliance prompts, auto summaries, and warm handoff context (CX Foundation, CallMiner).

Why assist wins the first pilot

Cresta’s prioritization matrix for first AI use cases scores risk and time-to-value, not demo dazzle. Automated after-call work and real-time agent assist rank above customer-facing voice AI for order status — even when the voice bot looks higher-value on paper — because data readiness and customer exposure drag the score down.

  • A human reviews output before the customer hears it, or before the CRM note is final.
  • Liability is lower. Wrong advice from an automated system is still your advice (Air Canada / BC tribunal line of cases).
  • Time to value is short: recording/STT and a CRM write path — not a full conversational IVR redesign.
  • It attacks minutes you already pay for. After-call work and mid-call knowledge search show up in AHT today.

Cresta reports auto-summarization can save roughly one minute of typing on about 80% of calls. Multiply by daily call volume and you do not need a board narrative about AGI.

Evidence: assist moves the needle on voice work

The strongest productivity studies on gen AI in support are about human + assist, not fully autonomous phone bots (summarized in InflectionCX’s 2026 operator guide):

  • NBER / QJE (Brynjolfsson, Li, Raymond): 5,179 support agents with a gen-AI assistant → +14% issues resolved per hour. Novices gained 34–35%; top performers barely moved. The system mostly spread tacit knowledge from strong agents to everyone else.
  • HBS / Management Science: 250k+ chats → AI-assisted agents ~20% faster, stronger empathy/thoroughness — again concentrated among less experienced agents. Responses that were too fast made customers suspect a bot (a warning that transfers cleanly to phone pacing).
  • Metrigy (697 firms): agent-assist tools cut average handle time about 29.5%.

These are not “replace the floor” results. They are “make the floor faster and more consistent,” especially for new hires — which matters when turnover runs 30–45% a year and replacement costs sit in the $10k–$20k range per agent.

Conversation intelligence compounds the same stack. Full-coverage scoring turns QA from a 5% sample into every call. Cresta’s CVS Health example: scoring coverage moved from 5% to 100% of calls. When operators later add automation on top of assist and intelligence, gains can stack — Snap Finance (Cresta AI Agent + Assist + CI) reported containment 6% → 33%, AHT −40%, CSAT +23%. Note the product set: analyze and assist are not optional decorations around the bot.

The three-layer voice stack (order that works)

Layer Voice job Customer hears AI? First metrics
Analyze Transcribe, score 100% of calls, find contact drivers No QA coverage, compliance hit rate, driver mix
Augment Live assist + auto ACW summary Indirect (human speaks) AHT, ACW minutes, FCR, agent ramp time
Automate Conversational IVR / voice agent resolves or gathers Yes Resolution (not just containment), transfer rate, repeat contact

Leaders sequence Analyze → Augment → Automate, not the reverse. McKinsey’s agentic-AI thesis matches at strategy level: horizontal copilots scale fast but produce diffuse P&L; vertical, process-embedded agents only pay when you redesign the workflow — not bolt a model onto the old one.

Layer 1 — Analyze (instrument the phone floor)

Before you buy a voice agent:

  • Reliable call recording and STT on the queues that matter.
  • Contact-driver taxonomy from real transcripts — not last year’s IVR menu labels.
  • PII redaction before anything leaves your trust boundary.
  • Compliance and disclosure scoring on all calls, not a sample.

If you cannot answer “why do people call us?” from data, an autonomous voicebot will guess. On VICIdial floors, that usually means proving recording paths and disposition truth before any AI layer.

Layer 2 — Augment (assist the human on the line)

Ship these as one program, not three side projects:

  1. Real-time knowledge and next-best action while the call is live.
  2. Compliance prompts (“read this disclosure now”).
  3. Auto summary + disposition the second the call ends.
  4. Coaching from 100% call scoring so assist quality does not silently decay.

This is where the productivity studies show up. It is also where agent stress can rise if you only automate the easy work and leave humans the emotional residue. Omdia’s 2025 signal that 75% of North American contact-center leaders believe AI may be increasing agent stress is a management problem, not a model problem. Pair assist with coaching — not just harder queues.

Layer 3 — Automate (put AI on the caller path)

Only after assist is measured.

  • Pre-queue / conversational IVR. Natural-language intent instead of “press 1.” Gather context in queue; pop it on the agent screen. Some providers pitch material handle-time cuts from this alone (e.g. ~45 seconds per call claims in CX Foundation’s catalog).
  • Tier-1 voice resolution. Status, balance, appointment, password flows with tool access to CRM/billing — not FAQ theater. High-volume, low-judgment intents first.
  • Warm transfer as a product requirement. When the bot fails, the human must receive intent, attempts, authentication state, and sentiment. Forced repeats after AI-to-human handoffs destroy experience scores.
  • Outbound voice. Reminders, collections nudges, appointment setting. Useful, regulated, easy to over-automate. Consent rules are a design input, not a legal afterthought.

Fully multi-step “agentic” phone workflows (authenticate → update billing → issue refund → confirm) are the destination, not the pilot. Reliability that is fine for a chatbot becomes expensive when the model can move money.

Metrics that belong on a voice AI scorecard

Drop vanity containment if it is the only number.

Metric Why it belongs on voice
AHTDirect labor minutes on the phone
ACW minutesWhere auto-summary pays first
FCR / resolution rateDid the issue actually close?
Repeat contact (24–48h)Catches false containment
Transfer rate + warm-transfer qualityBot failure mode on phone
QA / compliance coverage100% vs sample
Agent ramp / new-hire productivityWhere assist gains concentrate
CSAT / CES on voice onlyDo not average away phone pain with chat

Gartner-cited cost benchmarks still frame the economic gap: about $1.84 per self-service contact vs $13.50 agent-assisted. That gap only materializes if self-service resolves. Traditional self-service fully resolving ~14% of issues is the baseline trap. “Handled by AI” is not “solved by AI.”

If containment rises and repeat contact also rises, the bot is hiding failure — not creating capacity.

What still fails on voice

  1. Starting with the autonomous agent because the board saw a demo. Large fractions of GenAI and agentic projects still die after PoC.
  2. No AI-ready call data. Gartner’s prediction that organizations will abandon 60% of AI projects lacking AI-ready data through 2026 is a data program, not a vendor bake-off.
  3. Ignoring telephony reality. STT under noise, barge-in, hold music, and multi-party calls will beat your prompt engineering.
  4. Cost-first headcount cuts. Public “AI replaced N agents” stories that later rehire humans are a warning: replacement thinking fails when complex cases still need people. Only about 20% of service leaders had actually cut staffing because of AI in recent Gartner-cited surveys.
  5. Measuring the bot, not the journey. A contained call that generates a second call is not savings.
  6. Leaving agents with only the hard calls and no coaching investment.
  7. Underestimating TCO. Operator guides flag 40–60% TCO underestimates once integration, redaction, evaluation, and monitoring land — and warn that promotional per-interaction prices may not survive subsidy eras.

A 90-day voice sequence finance can buy

Days 1–30 — Instrument

One high-volume voice queue. Recording + STT. Contact-driver report. Baseline AHT, ACW, FCR, transfer, repeat contact.

Days 31–60 — Assist

Live knowledge assist + compliance prompts for that queue. Auto summary mandatory. Human edits allowed; measure edit rate. Success = ACW minutes down, AHT stable or down, QA flags not worse.

Days 61–90 — Narrow automate

One conversational IVR intent (status, balance, appointment). Hard escalation rules. Warm transfer packet required. Success = resolution rate on that intent — not “bot talked for 90 seconds.”

Only then expand intents or outbound. Prove one metric on one workflow. Reuse transcripts and CRM connectors for the next step. Do not re-platform between layers.

For how a production VICIdial floor behaves after the stack is healthy, see our fintech contact-center case study. For the cloud dialer install path behind many of those floors, see ViciBox 12 on GCE.

FAQ

Should we deploy an autonomous voicebot first?

Usually no. Instrument and assist first; automate one measured intent later. See three-layer stack.

What is LLM assist on a live phone call?

STT → LLM prompts on the agent screen while a human talks; auto CRM summary after hangup. See assist loop.

Which metrics beat containment?

AHT, ACW, FCR, 24–48h repeat contact, warm-transfer quality, QA coverage, voice-only CSAT. See scorecard.

What is a sensible 90-day plan?

Instrument → assist + ACW → one conversational IVR intent with hard escalation. See 90-day sequence.

Does this apply to VICIdial?

Yes. Same order on dialer floors: recording truth, agent assist and wrap reduction, then narrow automation in front of queues with warm transfer into the agent desktop.

Bottom line

Call centers do not need AGI on the phone. They need a boring sequence that compounds:

  1. Hear every call (analyze)
  2. Help the human while they talk (assist)
  3. Write the CRM note for them (ACW)
  4. Automate only the intents you can resolve and measure (voice agent)

The industry’s loudest product is still the autonomous voice agent. The industry’s quiet ROI is still LLM assist on live calls. Ship the quiet one first.