Skip to content

Add Pipecat voice traces and opt-in audio - #836

Open
David Elner (delner) wants to merge 18 commits into
mainfrom
delner/pipecat-voice-recording
Open

David Elner (delner) wants to merge 18 commits into
mainfrom
delner/pipecat-voice-recording

Conversation

@delner

Copy link
Copy Markdown
Contributor

Inspect conversations, tool calls, and audio through the existing setup_pipecat() API. Supports cascade and OpenAI Realtime voice pipelines.

  • Readable conversations: transcripts, native turns, model responses, and tool calls.
  • Optional audio: call recordings and speech clips, with sample-based playback selections.
  • Progressive export: upload clips and call segments during conversation; configurable limits preserve completed audio.
  • Opt-in recording: no recording buffers or encoding when disabled; codecs install separately with braintrust[audio].
setup_pipecat(capture_audio_attachments=True)  # Ogg/Opus; keep the existing pipeline

Example trace

pipecat.pipeline                    call recording segments
├─ pipecat.user_turn                 transcript + caller clip
│  └─ pipecat.stt
├─ pipecat.assistant_turn
│  └─ pipecat.llm_response
│     └─ lookup_order                arguments + result
└─ pipecat.assistant_turn             spoken continuation
   ├─ pipecat.llm_response
   └─ pipecat.tts_response            transcript + generated clip

Design

Pipecat lifecycle/media hooks
  ├─ native spans ───────────────┐
  └─ opt-in recorder → worker ───┤
                                 ↓
                       Existing SDK exporter
  • Preserve native turns, IDs, and payloads; map speech to recorded samples.
  • Encode off the voice event loop with bounded buffers/queues; workers share application CPU and memory.
  • Update owning spans as attachments finish; cleanup drains recording work before final flush.
  • Pipecat 1.12.0 native hooks; unsupported versions/pipelines retain existing tracing. Output selections require equal-rate mono audio without mixing.

Performance

53-second cascade replay, 24 kHz audio, 30-second rotation, eight attachments. Means of two fresh-process trials on Apple M5 Pro; baseline is tracing with recording disabled.

Metric Recording off Ogg enabled Added cost
Local response delay¹ 20.92 ms 20.99 ms +0.06 ms
CPU during call 2.84 s 4.16 s +1.31 s (2.48% of one core over the call)
Peak RSS growth² 2.73 MiB 6.29 MiB +3.56 MiB
Final recording drain 5 ms 395 ms +390 ms
Audio / final trace JSON 0 / 30.6 kB 609 kB / 41.6 kB +609 kB / +10.9 kB

¹ Provider-audio schedule to local output write, not caller-heard latency. ² Startup-to-final-drain RSS growth. Offline replay excludes network/upload; production latency and concurrency remain unmeasured. Ogg uses 10.7× fewer audio bytes than WAV in this call.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: babe2621bc

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "Codex (@codex) review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "Codex (@codex) address that feedback".

Comment on lines 164 to +166
async def on_process_frame(self, data: Any) -> None:
if self._voice is not None:
return

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve LLM usage metrics in native voice traces

When a supported Pipecat 1.12 voice pipeline is bound, this early return prevents the legacy _capture_metrics path from seeing any MetricsFrame. NativeObserver.on_push_frame only stores those frames as pipecat.events; it never adds LLMUsageMetricsData or TTFB values to the pipecat.llm_response span. Consequently normal voice calls lose token/cost and time-to-first-token metrics even when Pipecat usage metrics are enabled. Extract and log the metrics in the native observer before bypassing the existing handler.

Useful? React with 👍 / 👎.

Comment thread py/pyproject.toml
braintrust = "braintrust.integrations.harbor:HarborPlugin"

[project.optional-dependencies]
audio = ["numpy>=1.26", "soundfile>=0.13.1"]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Include audio dependencies in the all extra

This adds the audio optional extra, but the existing all extra still omits both NumPy and SoundFile. Users following the documented pip install "braintrust[all]" path and then enabling Pipecat audio recording will hit the encoder's missing-dependency error unless they separately discover and install braintrust[audio]; add these dependencies to all so it continues to install every optional feature.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

most of these tests seem not necessary and use a lot of mocks/fakes. Can we just use vcr tests instead? We can probably combine some redundant tests or remove lower value tests in py/src/braintrust/integrations/pipecat/test_pipecat.py as well

@delner
David Elner (delner) force-pushed the delner/pipecat-voice-recording branch from 8130ef7 to 20b9392 Compare October 7, 2026 02:19
Comment on lines +24 to +25
def test_prepare_recording_preserves_disabled_result():
assert prepare_recording("call", lambda: None) is None

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd take another pass to see if we can remove low value tests like this, or tests that use excess mocks/fakes.

also asking the agent to combine redundant tests is also pretty effective.

a lot of these unit tests don't give me a huge amount of confidence, but maybe that is fine.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I tried to do a couple of passes on this, not sure if it was very effective at reduction. I might have to do some more careful combing.

Comment thread py/src/braintrust/audio/README.md Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we remove these added readmes? The code/file structure should self document itself just fine, these are going to go stale.

Comment thread py/src/braintrust/audio/__init__.py Outdated
from .options import RecordingOptions


__all__ = ["RecordingOptions"]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why is this the only export from this folder? Does everything else get exported directly?

Comment thread py/src/braintrust/audio/attachments.py Outdated
from .worker import RecordingBusy


encoded_budget = ByteBudget(32 * 1024 * 1024)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we don't need a separate budget for audio. We can just use the generic attachment budget.

Comment thread py/src/braintrust/audio/__init__.py Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  1. I find it hard to reason about all the unit versions (sec/ms, byte offsets, etc.). I wonder if we can have some stronger named helpers to reason about this
  2. the classes here feel like they mix reponsibilities too much. class Alignment does both mapping to recording and calling span.log. I'd rather those two be explicitly two different steps / functions / classes

I did some vibing about the class structure, and I think I'd like to see something like so (following is AI GENERATED):

Class Owns / does
RecordingSession Call-level state: accepts capture while active, rotates segments, handles limits, drains on finish. Owns the active segment and tracks pending segment jobs.
AudioSegment One segment’s state and timeline: OPEN → PENDING → READY / OMITTED. When submitted, it holds an immutable PCM snapshot; once the worker finishes, it holds the encoded result or omission reason.
AudioEncoder Converts a segment snapshot into encoded bytes and metadata. Keeps codec details out of recording lifecycle code.
RecordingWorker Runs encoding off the event loop, enforces concurrency limits, and keeps PCM alive until native work actually stops.
AlignmentResolver Purely maps sample ranges to ready segments and produces selection and clip timeline records. It does not know about spans or loggers.
RecordingExporter Turns encoded results into Braintrust attachments and publishes descriptors and selections through the SDK logger. Owns the encoded-byte lease for as long as the exporter retains the attachment.
PipecatRecordingAdapter
  ├─ capture PCM ─────────────→ RecordingSession
  │                                ├─ active AudioSegment
  │                                └─ submitted AudioSegment
  │                                       ↓
  │                                 RecordingWorker
  │                                       ↓
  │                                  AudioEncoder
  │                                       ↓
  │                                 RecordingExporter → SDK logger
  │
  └─ turn sample ranges ──────→ AlignmentResolver ───→ RecordingExporter

from collections import deque
from contextvars import ContextVar

from braintrust.audio.timeline import pcm_bytes_to_ms, samples_to_ms

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So I realized that braintrust.audio is now public API. I think we should make it _audio (the folder name) so that users know it's private/internal to the SDK. Unfortunately we don't have a better way in python to do this.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

also I would really prefer if everything in these files were imported from braintrust._audio, helps me see what exactly the public API of the audio module is in py/src/braintrust/audio/__init__.py

Comment thread py/src/braintrust/audio/export.py Outdated
class Alignment:
"""Publish resolved selections and clip timelines to their owning spans."""

def __init__(self, root, recording):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

feel like all of these should be typed?

Comment thread py/src/braintrust/audio/export.py Outdated
"audio.selection": selections[0] if len(selections) == 1 else None,
}
)
self.root.log(metadata={"braintrust.alignment.ranges_omitted": self.omitted})

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fact that the Alignment class also logs feels confusing. I would like this to be a separate responsibility. But open to discussion.

Comment on lines +24 to +28
def release_once():
nonlocal released
if not released:
released = True
release()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

there feels like a better way we can do this with some of the python concurrency mechanisms built in.

Comment on lines +11 to +14
segment_duration_seconds: float = 60
max_duration_seconds: float = 1800
max_buffer_bytes: int = 32 * 1024 * 1024
flush_fraction: float = 0.5

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why did we pick these? Ditto with some of the other byte numbers we picked.

Comment on lines +65 to +67
max_bytes=8 * 1024 * 1024,
max_duration_ms=120000,
max_packets=16000,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

feel like there are enough of these constants around that we should have a constants.py file and write comments about why they exist.

Comment on lines +82 to +84
def __del__(self):
if getattr(self, "bytes", 0):
self.clear()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we do this?

self.bytes = 0
self.packets.clear()

def capture(self, channel, pcm, sample_rate, channels, *, observed_ns=None, observed_unix_ms=None):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this capture logic is pretty confusing, with a lot of conditionals. I think we need to codify the state machine better here.

I was hoping the usage of AudioSegment would make this better, but maybe not.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants