Agent Skills: Pipecat

Pipecat realtime voice/multimodal bots. Covers pipelines/frames, transports, RTVI, Pipecat Cloud deploy. Use when building real-time voice bots (STT/LLM/TTS pipelines), multimodal AI agents, WebRTC/WebSocket transports, or deploying to Pipecat Cloud. Keywords: pipecat, pipecat-ai, RTVI, WebRTC, voice bot.

UncategorizedID: itechmeat/llm-code/pipecat

Install this agent skill to your local

pnpm dlx add-skill https://github.com/itechmeat/llm-code/tree/HEAD/skills/pipecat

Skill Files

Browse the full folder contents for pipecat.

Download Skill

Loading file tree…

skills/pipecat/SKILL.md

Skill Metadata

Name
pipecat
Description
"Pipecat realtime voice/multimodal bots. Covers pipelines/frames, transports, RTVI, Pipecat Cloud deploy. Use when building real-time voice bots (STT/LLM/TTS pipelines), multimodal AI agents, WebRTC/WebSocket transports, or deploying to Pipecat Cloud. Keywords: pipecat, pipecat-ai, RTVI, WebRTC, voice bot."

Pipecat

Pipecat is an open-source Python framework for building real-time voice and multimodal bots. It composes streaming speech/LLM/TTS services into a low-latency pipeline, connected via transports (WebRTC/WebSocket) and client SDKs using the RTVI message standard.

Links

Quick navigation

  • Installation (packages/extras/CLI): references/installation.md
  • Migration to 1.0: references/migration-1-0.md
  • Concepts & architecture: references/core-concepts.md
  • Session initialization (runner/bot/client): references/session-initialization.md
  • Pipeline & frames: references/pipeline-and-frames.md
  • Transports: references/transports.md
  • Speech input & turn detection: references/speech-input-and-turn-detection.md
  • Client SDKs + RTVI messaging: references/client-sdks-rtvi.md
  • CLI (init/tail/cloud): references/cli.md
  • Function calling (server): references/function-calling.md
  • Context management: references/context-management.md
  • LLM inference: references/llm-inference.md
  • Text to speech (TTS): references/text-to-speech.md
  • Deployment (pattern/platforms): references/deployment.md
  • Server APIs (supported services): references/server-services.md
  • Server Utilities (runner): references/server-runner.md
  • Server APIs (pipeline/task/params): references/server-pipeline-apis.md
  • Pipecat Cloud ops: references/pipecat-cloud.md
  • Troubleshooting: references/troubleshooting.md

Mental model (cheat sheet)

  • Pipeline: ordered processors that consume/emit frames.
  • Frames: the streaming units (audio/text/video/context/events) flowing through the pipeline.
  • Transport: connectivity + media IO + session state (WebRTC/WebSocket/provider realtime).
  • Runner: HTTP service that starts sessions and spawns a bot process with transport credentials.
  • Client SDK: starts the bot, connects transport, sends messages/requests, receives events.

Recipes

1) Keep secrets server-side

  • Put provider API keys (LLM/STT/TTS) only on the server/bot container.
  • The client should call a server start endpoint (startBot / startBotAndConnect) to receive transport credentials (e.g., a room URL + token), not provider keys.

2) Use WebRTC for production voice

  • Prefer a WebRTC transport (e.g., Daily) for resilience and media quality.
  • Use a WebSocket transport mostly for server↔server, prototypes, or constrained environments.

2b) Design for streaming + overlap

  • Keep the pipeline fully streaming (avoid batching whole turns when you can).
  • If your services support it, start TTS from partial LLM output to reduce perceived latency.

3) Initialize and evolve context via RTVI

  • Initialize the bot’s pipeline context from the server start request payload.
  • For ongoing interaction, prefer a dedicated “send text” style API (when available) instead of deprecated context append methods.

4) Function calling: end-to-end flow

  • LLM requests a function call.
  • Client registers a handler by function name.
  • Client returns a function-call result message back to the bot.

5) Pipecat Cloud deployment basics

  • Build/push an image that matches the expected platform (Pipecat Cloud requires linux/arm64 in the docs).
  • Use a deployment config file for repeatability.
  • Configure pool sizing with min_agents (warm capacity) and max_agents (hard limit).

Critical gotchas / prohibitions

  • Do not embed sensitive API keys in client apps.
  • Expect and handle “at capacity” responses (HTTP 429) when the pool is exhausted.
  • Plan for cold-start latency if min_agents = 0.
  • Ensure secrets and image-pull credentials are created in the same region as the deployed agent.
  • Do not assume deprecated import shims or service-specific context classes still exist in 1.0.0; audit imports before upgrading.
  • Do not keep VAD/turn-detection logic on transport params; current releases route that control through LLMUserAggregator strategies.
  • Do not assume OpenAIResponsesLLMService is HTTP-based anymore; WebSocket is now the default implementation.
  • Do not send a single button field in the RTVI dtmf client message; as of 1.6.0 it requires buttons (a list), and RTVI.PROTOCOL_VERSION is 2.1.0.
  • Do not filter OTel dashboards on old GenAI span attribute names (az.ai.openai, xai, mistral, gen_ai.usage.reasoning_tokens, bare tokens.*); 1.6.0 renamed these to standard gen_ai.* conventions.

Release Highlights (1.7.0 -> 1.8.1)

Turn detection and tools (1.8.0)

  • Turn detection redesign: every in-repo service with built-in turn detection now emits ProposedUserStartedSpeakingFrame / ProposedUserStoppedSpeakingFrame, and ExternalUserTurnStrategies resolves them into real turn frames — one subclassable place that decides turns. Third-party services emitting turn frames directly keep working unchanged.
  • MCP made trivial: MCPClient.tools() now auto-connects, registers tools, and closes the connection at pipeline end (just LLMContext(tools=await mcp.tools())); MCPClient(tools_arguments=...) injects fixed arguments into every call of a tool, hidden from the model's schema. New KeenableWebSearch service (keenable extra) adds live web search + page reading via a hosted MCP server.
  • MoQ client mode: dial a shared relay instead of serving your own socket — works behind NAT (--moq-connect <relay>, plus MOQParams.response_path/request_path).

Workers and error handling (1.8.0)

  • Workers: JobParams/JobGroupParams bundle job dispatch metadata, BaseUIWorker surfaces jobs on a client UI without an LLM, WorkerRunner.get_worker(name) finds peers by name, and request_cancel_job_group() allows external cancellation.
  • Error/usable-state model: ErrorCategory (AUTHENTICATION/SERVER/APPLICATION/UNKNOWN), FrameProcessor.is_usable, and on_usable_changed let handlers tell a briefly-struggling processor from one to retire; PipelineWorker(processor_unusable_policy=CONTINUE|END|CANCEL) decides what happens when a processor goes unusable. AudioVolumeTracker measures rolling 400ms volume.

v1.7.0 additions

  • STT usage metrics: STTUsageMetricsData carries audio_seconds; enable with enable_usage_metrics=True (forwarded to RTVI clients as stt_usage, logged by MetricsLogObserver, attached to OTel stt spans). AWS Nova Sonic LLM adds LLMUsageMetricsData token deltas.
  • New/updated TTS/STT settings: PocketTTSService (local CPU-only TTS, 6 languages + voice cloning); XTTSService deprecated (removal 2.0.0, use Kokoro/Piper); reach_inactive_services on settings frames; plus per-service settings such as Google LLM safety_settings, Azure force_locale, Cartesia keyterm, Deepgram numerals, ElevenLabs filter_background_audio, and Smol endpointing/keywords/format.
  • Breaking-behavior changes: TavusParams.audio_out_faster_than_realtime now defaults to True; GeminiLiveLLMService uses GeminiLiveLLMAdapter (hand-crafted LLMSpecificMessages need llm="gemini-live"); LiveKit transport adds inbound SIP DTMF (on_dtmf_event); LLMSettings.filter_incomplete_user_turns deprecated.

Context Hub (1.8.0)

pipecat context-hub (alias pipecat ch) lets coding agents query a local index of Pipecat APIs instead of hallucinating them; it ships in the cli extra, and pipecat init offers to register it with the coding agents it finds and to build the index. v1.8.1 fixes pipecat eval run to read .yml scenario files and settles Context Hub staleness-warning behavior.

Release Highlights (1.6.0)

New transport and reasoning

  • MOQTransport: a Media over QUIC transport giving bots a bidirectional, low-latency audio + RTVI channel over QUIC instead of WebRTC/WebSocket (pip install "pipecat-ai[moq]"). The bot can run as its own MoQ server for local dev; the development runner gained --moq-serve, --moq-bind, and TLS flags.
  • reasoning in OpenAI Responses services: OpenAIResponsesLLMService / OpenAIResponsesHttpLLMService accept a ReasoningConfig(effort=..., summary=...); encrypted reasoning round-trips across turns and tool calls automatically, and summaries surface as thought frames / on_assistant_thought.

Flows, evals, and new services

  • Pipecat Flows NO_RESPONSE: a function can return (result, NO_RESPONSE) to finish the call without an LLM run or node transition, letting the next user utterance drive the next response.
  • Eval scenarios absent: true: an expectation that passes only if no matching event arrives within within_ms, useful for catching duplicate-output regressions.
  • New services: CrusoeLLMService and BasetenLLMService (OpenAI-compatible LLM services for Crusoe Cloud and Baseten), DeepgramFluxTTSService (websocket TTS for Deepgram Flux, token-streamed by default).
  • Audio token usage: LLMTokenUsage gains input_audio_tokens / output_audio_tokens / cache_read_input_audio_tokens, populated by OpenAIRealtimeLLMService (and Azure realtime) and GeminiLiveLLMService.

Breaking changes and deprecations

  • RTVI dtmf client message uses buttons (a list) instead of button; RTVI.PROTOCOL_VERSION is now 2.1.0.
  • OTel GenAI span attributes were renamed to standard conventions: gen_ai.provider.name values for Azure/xAI/Mistral, and gen_ai.usage.reasoning_tokens -> gen_ai.usage.reasoning.output_tokens; OpenAI Realtime and Gemini Live token attributes are now standardized gen_ai.usage.* names instead of ad hoc tokens.*.
  • ElevenLabs services default to TTS model eleven_flash_v2_5 instead of the now-deprecated eleven_turbo_v2_5.
  • PronunciationDictionaryLocator is deprecated in favor of text_transforms / replace_text (removal in 2.0.0).
  • Turn-strategy reset() is deprecated in favor of handle_user_turn_started() / handle_user_turn_stopped() lifecycle callbacks (removal in 2.0.0).

Release Highlights (0.0.109 -> 1.2.0)

Runtime and service additions

  • OpenAIResponsesLLMService now defaults to a persistent WebSocket connection; the prior HTTP behavior moved to OpenAIResponsesHttpLLMService.
  • Inworld Realtime LLM adds a WebSocket cascade STT/LLM/TTS path with semantic VAD and function calling.
  • MistralTTSService adds streaming Voxtral TTS, and TTS/STT services gained more runtime-update and sample-rate options.
  • The development runner now exports a module-level FastAPI app for custom routes before main().

Tooling and context changes

  • Function calling now supports grouped parallel tool batches, async tool completion after interruption, and streaming intermediate tool results.
  • Context editing now has LLMMessagesTransformFrame, and the framework standardizes on universal LLMContext / LLMContextAggregatorPair.
  • OpenAI tool schemas can now include provider-specific custom_tools.
  • 1.2.0 adds add_tool_change_messages for LLM aggregators, widens tool_resources into deprecated app_resources, and extends async-tool compatibility across more realtime providers.

Turn-taking and client protocol

  • 1.2.0 adds explicit inference/finalization turn hooks (on_user_turn_inference_triggered, LLMTurnCompletionUserTurnStopStrategy, FilterIncompleteUserTurnStrategies) for smarter end-of-turn gating.
  • RTVI grows first-class UI Agent Protocol support with ui-event, ui-snapshot, ui-cancel-task, ui-command, and ui-task, bumping the protocol to 1.3.0.
  • The development runner and runner arguments now carry a stable session_id, which is useful for per-session tracing across local and cloud-like flows.

Breaking migrations

  • Deprecated service-specific context classes, transport params, RTVI shims, frame aliases, and interruption/VAD helpers were removed across the stack.
  • Turn detection and mute behavior moved toward LLMUserAggregator strategies instead of transport-level configuration.
  • Some legacy providers and helpers were removed entirely (OpenPipeLLMService, TTSService.say(), FrameProcessor.wait_for_task(), older beta/alias modules).

Release Highlights (1.3.0 -> 1.5.0)

  • Workers and multi-agent pipelines (1.3.0): PipelineTask / PipelineRunner are renamed toward PipelineWorker / WorkerRunner, and pipecat.workers makes pipelines peers on a typed-message bus for @job dispatch, handoffs, sidecars, UI workers, and distributed Redis/PGMQ patterns.
  • UIWorker and RTVI UI protocol (1.3.0): UIWorker can observe client accessibility snapshots and drive UI commands over RTVI; the UI worker vocabulary moves from task/agent to job/worker (ui-task -> ui-job-group, cancelUITask -> cancelUIJobGroup).
  • Development runner (1.3.0): one runner can serve WebRTC, Daily, telephony, and plain WebSocket clients; /start accepts a transport field, /ws-client supports protobuf WebSocket clients, /status reports enabled transports, and the Daily redirect moved to /daily.
  • Service surface (1.3.0): adds Vonage Video Connector transport, Inception Mercury 2 LLM service, Cartesia turn-based STT, Rime coda TTS defaults, Soniox endpoint-delay settings, LLMService.append_system_instruction(), and STTService.supports_ttfs. Optional service/transport extras now raise ImportError when their dependency is missing, and transformers is no longer a base dependency.
  • Behavioral eval framework (1.4.0): a testing framework for evaluating bot behavior, plus on_user_turn_message_added event handlers and realtime LLM-service metadata frames.
  • Pipecat Flows in core (1.5.0): Pipecat Flows is integrated into the main package, so structured conversation flows no longer require a separate install. New services (1.5.0): Together AI STT/TTS services and NVIDIA per-sentence synthesis, with TTFA (time-to-first-audio) metrics for latency tracking.