- What: What a voice agent on Plivo Audio Streaming can do beyond talking.
- Watch out: A tool call that runs longer than a second is silence to the caller unless you cover it.
Written against pipecat-ai 1.11.0. Pipecat moves quickly and import paths change between releases — check the version you have installed before copying a snippet.
Call Your Own Systems Mid-Conversation
The agent decides, during the call, that it needs real data — an order status, an available slot, a customer record — and calls your code to get it.Cover the Silence While It Runs
There is no spinner on a phone call. A lookup that takes three seconds is three seconds of dead air, which is when callers say “hello? are you there?” and start hanging up.Define the Call Flow as Data
A single system prompt holding a whole conversation together stops working the moment the call has real steps. Pipecat Flows models the call as named nodes, each with its own instructions and its own tools, and transitions between them. The flow is a YAML file the agent loads at runtime.transition_only entry in the YAML.
Because the flow is data, one deployed agent can run whichever flow a call needs, and changing what the agent says or where a step leads is a YAML edit rather than a deploy. Write the flow in Python instead when a tool needs JSON Schema constraints an ordinary signature cannot express, or when a node’s shape depends on the conversation.
Remember a Returning Caller
The agent’s conversation history is a list you can read and write, so a caller who phones back can be greeted with what happened last time rather than starting from nothing. Save it when the call ends:start event does not carry it — it carries callId, streamId, accountId, tracks, mediaFormat and extra_headers, and nothing else. The caller’s number reaches you as From on the call webhook. Pass it into the stream yourself:
start event as extra_headers, and the agent can open with context instead of “how can I help you?”.
Reloading a whole prior conversation grows the context every call. For a caller who rings often, summarize the old history before restoring it rather than replaying it verbatim.
Give the Agent Your Existing Tools
If your systems already speak MCP, the agent can use those servers directly rather than having every integration rewritten as a voice-specific function.tools_filter matters on a phone call. An MCP server exposing forty tools gives the model forty things to consider before every reply, and that shows up as latency the caller hears. Register the handful the agent actually needs.
Add Background Audio
Customers describe an agent with complete silence behind it as sounding artificial. Real people are never in silence — there is always a room around them. Faint ambience under the speech makes the call sound like a person at a desk rather than a recording. Pipecat mixes it in the output transport, so it plays for the whole call, including the gaps when the agent is not speaking.PipelineParams line is not optional here. The mixer loads the file only if its sample rate equals the pipeline’s output rate, and Pipecat defaults to 24 kHz out — so an 8 kHz ambience file is silently skipped unless you set the rate to match.
Change it during the call without rebuilding the pipeline — turn the volume down while the agent reads out a reference number, or switch the ambience off entirely.
Tune Barge-In
By default the agent stops the moment it hears the caller. On a phone line that is too eager: a cough, an “mm-hmm”, or someone talking in the background all cut the agent off mid-sentence. A browser would have removed most of that with local echo cancellation. A phone line does not. Require a minimum number of real words before treating speech as an interruption.clearAudio event, which flushes audio already queued for playback so the caller hears the agent stop rather than talk over them.
Tuning for a Phone Line
These are settings rather than features, but they are the ones that decide whether the agent sounds right on a call.SileroVADAnalyzer supports 8 kHz natively, one of the few audio components built for telephony rates rather than adapted to them. It accepts only 8000 or 16000.Related
Build with Pipecat
Connect Pipecat to Plivo Audio Streaming
Audio Streaming Reference
Every WebSocket event, with schemas and types
Transfer to Human Agent
Hand a live call to a person
Best Practices
Voicemail detection and connection handling