Repository navigation
Add Pipecat voice traces and opt-in audio #836
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
David Elner (delner)
wants to merge
18
commits into
main
Choose a base branch
from
delner/pipecat-voice-recording
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
18 commits
Select commit
Hold shift + click to select a range
ede702f
Added: Common audio recording with progressive export
delner e5737d4
Added: Pipecat voice tracing and opt-in recording
delner 27ec415
Added: turn_metrics to Pipecat
delner 56c94aa
Added: TTFB routing to Pipecat
delner 4cb5a2b
Fixed: Preserve LLM usage and timing metrics in native Pipecat traces
delner 7d0dfd0
Fixed: timing-dependent Pipecat attachment test
delner 4308a3e
Fixed: Missing llm_tool_calls in Pipecat
delner 30bfb41
Changed: Pipecat instrumentation fields to generic names
delner 20f1011
Changed: Pipecat tests to use VCR and recording behavior during shutdown
delner 622399e
Added: Pipecat runnable examples
delner 8104157
Fixed: Realtime audio test cassette mismatch
delner 6295eeb
Changed: Restructure Pipecat events, move metadata to contrib.pipecat.*
delner ae0a630
Fixed: TTS early arrival, buffer limits
delner b9257d8
Added: Pipecat state and timeline mappings
delner 20b9392
Refactored: Recording-job cleanup, synthesis recording, alignment
delner bd7df8d
Refactored: Simplify audio recording lifecycle and consolidate Pipeca…
delner 8ce91f4
Fixed: Audio segments stay pending until uploaded
delner 6d51557
Refactored: Audio export and alignment ownership
delner File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,28 @@ | ||
| # Pipecat voice tracing | ||
|
|
||
| Two small voice agents using real OpenAI services and the local Braintrust SDK: | ||
|
|
||
| | Example | Pipeline | | ||
| | --- | --- | | ||
| | `cascade.py` | Speech → OpenAI transcription → LLM + order lookup → synthesized speech | | ||
| | `realtime.py` | Speech → OpenAI Realtime + order lookup → speech | | ||
|
|
||
| Requires Python 3.11–3.13, `uv`, `OPENAI_API_KEY`, and `BRAINTRUST_API_KEY` in your environment. From the repository root, run either command: | ||
|
|
||
| ```sh | ||
| uv run --project examples/pipecat python examples/pipecat/cascade.py | ||
| uv run --project examples/pipecat python examples/pipecat/realtime.py | ||
| ``` | ||
|
|
||
| Each command installs its dependencies, runs one conversation, and prints a trace link in the `example-pipecat` project. OpenAI usage is billable. | ||
|
|
||
| The bundled `order.wav` (about 82 KiB) asks “Where is my order number one zero four two?” `common.py` supplies file input and silent output so no microphone, speaker, browser, or phone setup is needed. The local `lookup_order` tool returns a fictional delivery date. Listen to the captured audio in the trace. | ||
|
|
||
| Instrumentation is enabled with: | ||
|
|
||
| ```python | ||
| logger = braintrust.init_logger(project="example-pipecat") | ||
| setup_pipecat(capture_audio_attachments=True) | ||
| ``` | ||
|
|
||
| Pipecat metrics are enabled in `PipelineParams`. The trace includes turns, model calls, tool execution, metrics, and Ogg audio attachments. Set `capture_audio_attachments=False` to trace without recording audio. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,62 @@ | ||
| """Audio → transcription → model/tool → synthesized speech, traced by Braintrust.""" | ||
|
|
||
| import asyncio | ||
| import os | ||
|
|
||
| import braintrust | ||
| from braintrust.integrations.pipecat import setup_pipecat | ||
| from common import INSTRUCTIONS, RecordedInput, SilentOutput, order_context, run_conversation | ||
|
|
||
| # Pipecat requires Python 3.11+ and is installed by this example, not the shared lint environment. | ||
| # pylint: disable=import-error | ||
| from pipecat.audio.vad.silero import SileroVADAnalyzer | ||
| from pipecat.pipeline.pipeline import Pipeline | ||
| from pipecat.processors.aggregators.llm_response_universal import LLMContextAggregatorPair, LLMUserAggregatorParams | ||
| from pipecat.services.openai.llm import OpenAILLMService | ||
| from pipecat.services.openai.stt import OpenAISTTService | ||
| from pipecat.services.openai.tts import OpenAITTSService | ||
| from pipecat.transports.base_transport import TransportParams | ||
|
|
||
|
|
||
| async def main(): | ||
| logger = braintrust.init_logger(project="example-pipecat") | ||
| setup_pipecat(capture_audio_attachments=True) | ||
|
|
||
| api_key = os.environ["OPENAI_API_KEY"] | ||
| transport = TransportParams(audio_in_enabled=True, audio_out_enabled=True) | ||
| source, output = RecordedInput(transport), SilentOutput(transport) | ||
| aggregators = LLMContextAggregatorPair( | ||
| order_context(), | ||
| user_params=LLMUserAggregatorParams(vad_analyzer=SileroVADAnalyzer()), | ||
| ) | ||
| pipeline = Pipeline( | ||
| [ | ||
| source, | ||
| OpenAISTTService(api_key=api_key), | ||
| aggregators.user(), | ||
| OpenAILLMService( | ||
| api_key=api_key, | ||
| settings=OpenAILLMService.Settings( | ||
| model="gpt-4.1-mini", | ||
| system_instruction=INSTRUCTIONS, | ||
| ), | ||
| ), | ||
| OpenAITTSService( | ||
| api_key=api_key, | ||
| settings=OpenAITTSService.Settings( | ||
| model="gpt-4o-mini-tts", | ||
| voice="alloy", | ||
| ), | ||
| ), | ||
| output, | ||
| aggregators.assistant(), | ||
| ] | ||
| ) | ||
| with logger.start_span(name="cascade") as trace: | ||
| await run_conversation(pipeline, aggregators, source) | ||
| await asyncio.to_thread(logger.flush) | ||
| print(f"Trace: {trace.link()}") | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| asyncio.run(main()) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,115 @@ | ||
| """Example application plumbing: play a WAV instead of opening a microphone.""" | ||
|
|
||
| import asyncio | ||
| import wave | ||
| from pathlib import Path | ||
|
|
||
| # Pipecat requires Python 3.11+ and is installed by this example, not the shared lint environment. | ||
| # pylint: disable=import-error | ||
| from pipecat.adapters.schemas.function_schema import FunctionSchema | ||
| from pipecat.adapters.schemas.tools_schema import ToolsSchema | ||
| from pipecat.audio.utils import create_stream_resampler | ||
| from pipecat.frames.frames import InputAudioRawFrame, LLMRunFrame | ||
| from pipecat.pipeline.worker import PipelineParams, PipelineWorker | ||
| from pipecat.processors.aggregators.llm_context import LLMContext | ||
| from pipecat.transports.base_input import BaseInputTransport | ||
| from pipecat.transports.base_output import BaseOutputTransport | ||
| from pipecat.workers.runner import WorkerRunner | ||
|
|
||
|
|
||
| async def lookup_order(params): | ||
| """A fictional local order database; replace with your application's tool.""" | ||
| await params.result_callback({"order_id": params.arguments["order_id"], "delivery": "Friday"}) | ||
|
|
||
|
|
||
| def order_context(): | ||
| return LLMContext( | ||
| tools=ToolsSchema( | ||
| standard_tools=[ | ||
| FunctionSchema( | ||
| name="lookup_order", | ||
| description="Look up an order's delivery date", | ||
| properties={"order_id": {"type": "string"}}, | ||
| required=["order_id"], | ||
| handler=lookup_order, | ||
| ) | ||
| ] | ||
| ), | ||
| ) | ||
|
|
||
|
|
||
| INSTRUCTIONS = ( | ||
| "Speak English. Greet the caller briefly and ask for their order number. " | ||
| "When they supply it, always call lookup_order. " | ||
| "After its result, say only: Your order arrives Friday." | ||
| ) | ||
|
|
||
|
|
||
| class RecordedInput(BaseInputTransport): | ||
| async def start(self, frame): | ||
| await super().start(frame) | ||
| await self.set_transport_ready(frame) | ||
|
|
||
| async def play(self, sample_rate): | ||
| with wave.open(str(Path(__file__).with_name("order.wav"))) as audio: | ||
| pcm = audio.readframes(audio.getnframes()) | ||
| pcm = await create_stream_resampler().resample(pcm, 16000, sample_rate) | ||
| # Include silence around the request so native VAD determines its boundaries. | ||
| pcm = b"\0" * sample_rate + pcm + b"\0" * (sample_rate * 2) | ||
| frame_bytes = sample_rate // 50 * 2 | ||
| for offset in range(0, len(pcm), frame_bytes): | ||
| await self.push_audio_frame(InputAudioRawFrame(pcm[offset : offset + frame_bytes], sample_rate, 1)) | ||
| await asyncio.sleep(0.02) | ||
|
|
||
|
|
||
| class SilentOutput(BaseOutputTransport): | ||
| """Accept agent audio without requiring an audio device; listen in the trace.""" | ||
|
|
||
| async def start(self, frame): | ||
| await super().start(frame) | ||
| await self.set_transport_ready(frame) | ||
|
|
||
| async def write_audio_frame(self, frame): | ||
| return True | ||
|
|
||
|
|
||
| async def run_conversation(pipeline, aggregators, source, *, realtime=False): | ||
| sample_rate = 24000 if realtime else 16000 | ||
| worker = PipelineWorker( | ||
| pipeline, | ||
| params=PipelineParams( | ||
| audio_in_sample_rate=sample_rate, | ||
| audio_out_sample_rate=24000, | ||
| enable_metrics=True, | ||
| enable_usage_metrics=True, | ||
| ), | ||
| ) | ||
| greeted, finished = asyncio.Event(), asyncio.Event() | ||
|
|
||
| @aggregators.assistant().event_handler("on_assistant_turn_stopped") | ||
| async def assistant_turn(aggregator, message): | ||
| print(f"Agent: {message.content}", flush=True) | ||
| if "friday" in message.content.lower(): | ||
| finished.set() | ||
| else: | ||
| greeted.set() | ||
|
|
||
| @aggregators.user().event_handler("on_user_turn_message_added") | ||
| async def user_turn(aggregator, message): | ||
| print(f"Caller: {message.content}", flush=True) | ||
|
|
||
| @worker.event_handler("on_pipeline_started") | ||
| async def started(worker, frame): | ||
| await worker.queue_frame(LLMRunFrame()) | ||
| await asyncio.wait_for(greeted.wait(), 30) | ||
| await source.play(sample_rate) | ||
|
|
||
| task = asyncio.create_task(WorkerRunner(handle_sigint=False).run(worker)) | ||
| try: | ||
| await asyncio.wait_for(finished.wait(), 60) | ||
| await worker.stop_when_done() | ||
| await task | ||
| finally: | ||
| if not task.done(): | ||
| await worker.cancel() | ||
| await task |
Binary file not shown.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,12 @@ | ||
| [project] | ||
| name = "braintrust-pipecat-example" | ||
| version = "0.1.0" | ||
| description = "Trace cascade and realtime Pipecat voice agents" | ||
| requires-python = ">=3.11,<3.14" | ||
| dependencies = [ | ||
| "braintrust[audio]", | ||
| "pipecat-ai[openai,silero]==1.12.0", | ||
| ] | ||
|
|
||
| [tool.uv.sources] | ||
| braintrust = { path = "../../py", editable = true } |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,50 @@ | ||
| """OpenAI Realtime speech-to-speech with a local tool, traced by Braintrust.""" | ||
|
|
||
| import asyncio | ||
| import os | ||
|
|
||
| import braintrust | ||
| from braintrust.integrations.pipecat import setup_pipecat | ||
| from common import INSTRUCTIONS, RecordedInput, SilentOutput, order_context, run_conversation | ||
|
|
||
| # Pipecat requires Python 3.11+ and is installed by this example, not the shared lint environment. | ||
| # pylint: disable=import-error | ||
| from pipecat.pipeline.pipeline import Pipeline | ||
| from pipecat.processors.aggregators.llm_response_universal import LLMContextAggregatorPair | ||
| from pipecat.services.openai.realtime import events | ||
| from pipecat.services.openai.realtime.llm import OpenAIRealtimeLLMService | ||
| from pipecat.transports.base_transport import TransportParams | ||
|
|
||
|
|
||
| async def main(): | ||
| logger = braintrust.init_logger(project="example-pipecat") | ||
| setup_pipecat(capture_audio_attachments=True) | ||
|
|
||
| transport = TransportParams(audio_in_enabled=True, audio_out_enabled=True) | ||
| source, output = RecordedInput(transport), SilentOutput(transport) | ||
| aggregators = LLMContextAggregatorPair(order_context()) | ||
| model = OpenAIRealtimeLLMService( | ||
| api_key=os.environ["OPENAI_API_KEY"], | ||
| settings=OpenAIRealtimeLLMService.Settings( | ||
| model="gpt-realtime", | ||
| system_instruction=INSTRUCTIONS, | ||
| session_properties=events.SessionProperties( | ||
| audio=events.AudioConfiguration( | ||
| input=events.AudioInput( | ||
| transcription=events.InputAudioTranscription(model="gpt-4o-mini-transcribe", language="en"), | ||
| turn_detection=events.TurnDetection(), | ||
| ), | ||
| output=events.AudioOutput(voice="alloy"), | ||
| ) | ||
| ), | ||
| ), | ||
| ) | ||
| pipeline = Pipeline([source, aggregators.user(), model, output, aggregators.assistant()]) | ||
| with logger.start_span(name="realtime") as trace: | ||
| await run_conversation(pipeline, aggregators, source, realtime=True) | ||
| await asyncio.to_thread(logger.flush) | ||
| print(f"Trace: {trace.link()}") | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| asyncio.run(main()) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,33 @@ | ||
| """Internal audio API shared by integrations; codecs remain optional and lazy.""" | ||
|
|
||
| from .alignment import InputRanges | ||
| from .attachments import UploadFailed, prepare_recording, upload_recording | ||
| from .budget import source_budget | ||
| from .export import AlignmentPublisher, segment_descriptor | ||
| from .jobs import RecordingJobs | ||
| from .options import RecordingOptions | ||
| from .recording import encode_audio | ||
| from .segments import SegmentedRecording | ||
| from .timeline import ClipTimeline, ms_to_samples, pcm_bytes_to_ms, samples_to_ms | ||
| from .worker import RecordingBusy, encode_in_worker | ||
|
|
||
|
|
||
| __all__ = [ | ||
| "AlignmentPublisher", | ||
| "ClipTimeline", | ||
| "InputRanges", | ||
| "RecordingBusy", | ||
| "RecordingJobs", | ||
| "RecordingOptions", | ||
| "SegmentedRecording", | ||
| "UploadFailed", | ||
| "encode_audio", | ||
| "encode_in_worker", | ||
| "ms_to_samples", | ||
| "pcm_bytes_to_ms", | ||
| "prepare_recording", | ||
| "samples_to_ms", | ||
| "segment_descriptor", | ||
| "source_budget", | ||
| "upload_recording", | ||
| ] |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
This adds the
audiooptional extra, but the existingallextra still omits both NumPy and SoundFile. Users following the documentedpip install "braintrust[all]"path and then enabling Pipecat audio recording will hit the encoder's missing-dependency error unless they separately discover and installbraintrust[audio]; add these dependencies toallso it continues to install every optional feature.Useful? React with 👍 / 👎.