Repository navigation
test(integrations): consolidate redundant and low-value tests - #898
Merged
Abhijeet Prasad (AbhiPrasad) merged 29 commits intoOct 8, 2026
Merged
Conversation
…ests - Point the sync and async streaming nesting tests at the base streaming cassettes via a vcr_cassette_name parametrization. Their requests match the base tests byte-for-byte (method, URI, body, headers) in latest and 0.32.0. - Remove the sync non-streaming nesting test. Every chat wrapper calls start_span at call time, so the streaming nesting tests already cover the parent capture. Those tests are stricter because their spans finish later. - Delete the 6 cassettes that are no longer used. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Remove test_audio_transcriptions_patchers_target_sdk_surface, which only compared constants and checked hasattr. The setup VCR test has the same gate (cohere.audio.transcriptions.client must import). It now asserts that both TranscriptionsCreatePatcher and AsyncTranscriptionsCreatePatcher report is_patched after setup(), which is stronger than comparing target strings. - Point the setup test at the test_wrap_cohere_audio_transcription_sync cassette. The request is the same except for the random multipart boundary, and VCR matches on method+URL. Delete the duplicate cassette. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Combine the 4 TestHeaderSerialization tests into one round-trip test. It still covers the empty-context no-op, the missing-header None result, keeping existing headers, and the round-trip. - Remove test_plugin_client_context_propagation and its TestWorkflow. test_plugin_activity_after_replay_stays_under_workflow_span[True] already covers a plugin client inside a parent span starting a workflow. It also asserts the workflow and client share root_span_id and that the activities are parented to the workflow span. Removing TestWorkflow also clears its pytest collection warning. - Parametrize the retry, child-workflow and local-activity tests into test_plugin_workflow_tracing. They now assert exact span counts by name instead of >= 1 / > 0. For the retry case they also check that only the first attempt carries the simulated error. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Remove the public run/arun stream early-break tests. The partial-gc cases of test_agno_resume_stream_cleanup run the same _run/_arun_public_dispatch_wrapper and _Traced*Stream code with a real agent and also assert that the span ends. - Remove the public run/arun parent-span nesting tests. Nesting through the same dispatch wrappers is asserted by test_agno_resume_stream_cleanup (sync and async streams), test_agno_workflow_with_agent (non-stream Agent.run) and the async workflow and eval tests. - Remove test_agno_public_arun_non_stream_awaitable_compat. test_agno_resume_session_id[response-async-agent] exercises the same awaited non-stream branch and now also asserts that root output is logged. - Parametrize the three workflow stream aggregation tests into one. - Move the shared cancellable weather-agent setup into a helper. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Merge test_basic_completion_async, test_tool_use_async and test_google_search_grounding_async into their sync counterparts (modes sync/stream/async/async_stream). The async grounding paths now also check usage_metadata token accounting; passes on replay at latest, 1.75.0 and 1.30.0. - Parametrize the sync/async stream parent-preservation tests (#484) over mode, keeping both paths. - Parametrize test_embed_content and test_generate_images over sync/async; merge test_image_input + test_document_input into test_binary_input[image|document]. - test_system_prompt: assert input.config.system_instruction exactly instead of a near-tautological substring check. - Remove test_serialize_content_item_with_content_and_binary_part (end-to-end covered by VCR test_image_input_wrapped_in_content) and test_attachment_with_pydantic_model (bt_safe_deep_copy only; covered by test_bt_json.py attachment identity/pydantic tests). - Cassettes renamed with git mv in every version dir; no re-recording. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Parametrize test_agent_with_binary_content + test_agent_with_document_input into test_agent_with_binary_input[image|pdf] with the union of assertions (media type now checked for images too, chat span metrics checked for both). - Fold test_direct_model_request_stream_complete_output into test_direct_model_request_stream (collects text, asserts first chunk is not dropped; cassette has the same PartStart/delta shape). - Fold test_agent_with_custom_settings into test_agent_with_model_settings_override_in_input (infer_name/usage kwargs + not-leaking assertions + metrics); outbound request unchanged. - Fold the custom tool name check into test_agent_tool_metadata_extraction and drop test_agent_tool_with_custom_name. - Drop test_shape_messages_with_binary_content (covered by the pdf case of test_agent_with_binary_input) and test_agent_with_message_history (test_agent_with_prefill covers message_history capture). - Drop the redundant reasoning_tokens case of test_reasoning_tokens_extraction_provider_keys (covered by test_extract_response_metrics_leaf_fields) and a duplicate pylint enable. - wrap_openai: parametrize test_pydantic_wrapped over stream/completion. - Cassettes renamed with git mv; orphaned ones removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Parametrize sync/async x messages/beta batches.create span tests into one test; existing cassettes are reused via vcr_cassette_name. - Parametrize the sync/async batches.results tests. The sync case still passes the batch id positionally and the async case by keyword, so both extraction paths stay covered. - Collapse the four extract_anthropic_usage unit tests into one parametrized test that checks full metrics/metadata equality. - Fold the system-prompt input assertion into test_anthropic_messages_model_params_inputs, which already sends system=, and drop test_anthropic_messages_system_prompt_inputs. - Parametrize the beta messages create/stream x sync/async tests, keeping each case's original prompt and cassette. Every case now gets the union of the old assertions. - Remove the four legacy streaming tests. Their assertions (cache creation/read metrics, output model/stop_reason, max_tokens, project_id, start/end bounds, get_final_message, synthesized text events) now live in test_anthropic_messages_reasoning_tokens_metrics, checked against the wire at both matrix versions. The text_stream TTFT regression (BT-4702) is still covered by its sync/async text_stream cases. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Remove test_wrapped_tool_handler_keeps_nested_traces_under_stream_tool_span. The same-name tool test exercises the same acquire/nesting path (and now also asserts handler return values), and the VCR calculator test covers it end to end. - Parametrize the ToolSpanTracker lifecycle, MCP metadata, and error tests by tool name, asserting exact metadata. - Remove the subagent overlap check gated on 0.1.11 <= SDK < 0.1.64. It is unreachable with the matrix (0.1.10 returns early, latest is 0.2.x). - Remove the verbatim_prompts=False fallback in the query helper test. The test skips below 0.1.11, and latest (0.2.163) is past 0.2.158. - Remove the non-partial usage elif branch in test_calculator_with_multiple_operations. Instrumentation showed partial usage is always present at both matrix versions; this is now asserted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Fold each sync/async twin into one `is_async`-parametrized test (chat metrics, responses metrics, responses/chat stream helpers, raw-response create, embeddings, client error, and the audio speech/transcription/ translation/streaming tests). Each merged test keeps the union of the twins' assertions, so the async variants now also check provider, span origin, and time_to_first_token where only the sync twin did. - Merge the NOT_GIVEN and Omit filtering tests into one test parametrized over the sentinel; the Omit case keeps its OpenAI >= 2.0 gate. - Drop the chat system-prompt tests: tracing logs `messages` verbatim with no role-specific handling, and test_openai_chat_metrics now asserts the logged input by exact equality. - Parametrize the OpenRouter gateway stream/non-stream is_byok tests. - Cassettes are git-mv'd to the new `[sync]`/`[async]` names in every version dir; the auto-instrument script now reads `test_openai_responses_metrics[sync]`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Delete test_litellm_completion_with_system_prompt[sync/async]: the tracer logs `messages` verbatim with no role-specific handling, and test_litellm_completion_metrics now asserts the logged input by exact equality instead of a substring match. - Remove a stray print() in test_litellm_tool_calls and a time.sleep in test_litellm_async_streaming_with_break whose only assertion is time_to_first_token >= 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Replace the always-true `is not None` checks with assertions that the first setup() installed a FunctionWrapper around each original method, and keep the identity checks proving the second setup() did not re-wrap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Fold test_async_streaming into test_streaming_ttft as a sync/async parametrize; the async path now also asserts TTFT, input and exact output (cassettes git mv'd to test_streaming_ttft[sync|async]). - Trim test_tool_use_with_result to its unique part: it is the only test feeding an AIMessage with tool_calls + ToolMessage through the handler, so it now asserts that logged input instead of re-asserting output/metrics already covered by test_llm_calls/test_tool_usage. - Remove test_async_langchain_invoke (anthropic): the handler runs inline with no async-specific code; anthropic metadata is covered by test_langchain_anthropic_integration and async context by test_consecutive_async_invocations_are_separate_traces. - Remove test_setup_langchain_installs_default_handler: steps 1-3 of auto_test_scripts/test_auto_langchain.py assert the same thing via the same LangChainIntegration.setup() path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Combine test_adk_complex_nested_schema and test_adk_response_json_schema_dict into the parametrized test_adk_structured_output_schema (cassettes git mv'd in all 3 version dirs). - Remove test_adk_input_schema_serialization: input_schema is an agent attribute that never reaches GenerateContentConfig, so its exact-input assertion is a subset of the schema tests. - Assert MAX_TOKENS unconditionally in test_adk_max_tokens_captures_content; every cassette ends with it, so the previously conditional block now always runs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…alizer tests - Remove test_wrap_mistral_agents_complete_tool_spans: agents.complete goes through the same _finalize_completion_response / _log_completion_tool_spans path as test_wrap_mistral_chat_complete_tool_spans, and the agents entry point is covered by test_wrap_mistral_agents_complete_sync/_async. - Parametrize the three test_normalize_mistral_multimodal_value_* tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Combine the three patch_dspy() subprocess tests (patch twice + configure with no callbacks, with a user callback, with an explicit Braintrust callback) into one subprocess script; every assertion is kept and each configure() now checks exactly one BraintrustDSpyCallback. - Run the legacy braintrust.wrappers.dspy import check in-process: the module only re-exports and has no patching side effects. - Tighten test_auto_dspy.py from any(...) to exactly one callback. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Merge test_openai_agents_task_and_turn_span_types into test_openai_agents_integration_setup_creates_spans behind the same TaskSpanData/TurnSpanData hasattr gate. Both ran the same agent and prompt against the same POST /v1/responses request, so the setup test's cassette exercises the same span shapes; drop the orphaned latest cassette. - Merge the two BraintrustTracingProcessor unit tests into one that keeps every assertion and both code paths: a root trace (metadata + origin) and a trace nested under a current span (set/unset current). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Merge test_streaming_outputs_are_not_stringified and test_coroutine_outputs_are_not_stringified into one test parametrized over generator, async generator, and coroutine outputs. Same helper, same assertion, same cleanup of each lazy object. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Fold test_language_specific_translation_name into the test_text2text_task_families parametrize; it now also checks the input shape like the other task families. - Drop test_integration_targets_only_supported_pipeline_classes: it only asserted constants. Each patched class is exercised by its task-family test, and test_unsupported_pipeline_produces_no_span covers that the generic Pipeline.__call__ is not patched. - Drop test_setup_is_idempotent: test_auto_transformers.py already calls setup twice and asserts exactly one span for one pipeline call (the autouse fixture unpatches between in-process tests, so in-process tests never set up twice). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Move the #587 regression assertions (crewai.llm never emits token metrics, has start/end) from test_llm_never_emits_token_metrics into test_kickoff_llm_event_tree_parents_and_shape, which already completes an LLM call with a usage payload; drop the former. - Parametrize test_llm_call_failed_logs_error and test_tool_error_logs_error into test_failure_event_logs_error: identical start -> failure -> error shape, differing only in event builders, span name, and message. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Delete test_auto_instrument_livekit_agents_openai_e2e_voice_turn: it ran the same voice-turn helper and assertions as test_livekit_agents_e2e_parenting_under_custom_span (which also checks parenting). auto_instrument() wiring is covered by the test_auto_livekit_agents.py subprocess script, and the no-outer-parent session path by the agent_speaking and function_tool e2e tests. Remove its two cassettes (~500KB) and the now-dead no-parent helper branch. - test_llm_stream_run_does_not_create_user_turn_on_cancellation asserted nothing; it now checks that only an ended, error-free llm_request_run span is logged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Parametrize the paired/tripled tests in integrations/test_utils.py: _timing_metrics (2), _materialize_attachment from bytes (2), the three _materialize_attachment "returns None" cases, _extract_audio_output (2), _ResolvedAttachment.multimodal_part_payload (2), and _log_and_end_span (2). Same inputs and assertions; the bytes case now applies the full assertion set to both inputs. - Fold the misnamed "preserves_invalid_base64_strings_without_mime_type" test (it asserts None) into the "returns None" parametrize. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- verify_autoinstrument_script now also asserts "SUCCESS" is in stdout, so a script that exits 0 early (or never reaches its checks) fails. - All 28 auto_test_scripts already print SUCCESS on success except test_auto_pipecat.py, which now prints it after asyncio.run(main()). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pytest-vcr forwards vcr marker kwargs to vcrpy as config, where `cassette_name` is silently ignored, so tests meant to share a cassette replayed their own. Override the `vcr_cassette_name` fixture in the root conftest to honor the marker kwarg (default naming is unchanged). - mistral: the setup audio-transcription test now replays the complete_sync cassette at every version; drop its duplicate cassettes. - huggingface_hub: the early-close and context-manager tests make an extra model-mapping GET the base cassette lacks, so drop the no-op kwarg and keep their own cassettes (behavior unchanged). - google_discoveryengine: switch positional marker args to `cassette_name=` and remove its local fixture override. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
31 test cases in test_openai.py repeated the same invariant: make the call on a plain openai.OpenAI()/AsyncOpenAI() client and assert no spans were logged, before repeating it on the wrapped client. Drop those halves and cover the invariant once in test_unwrapped_client_emits_no_spans (sync + async), which replays the test_openai_chat_metrics cassettes. Why each removal is safe: - No unwrapped half compared its result with the wrapped result or checked a return type the wrapped half doesn't. Each wrapped half already asserts the same headers / parse() / stream / helper / output-text properties, so no wrapper-transparency check is lost. - The unwrapped halves that use a different API shape than the wrapped half (async plain create(stream=True) in chat_streaming_async, async with_raw_response.create(stream=True) in the raw-response stream test) only exercised the openai SDK itself, never the wrapper. - Tests where the calls run in order [unwrapped, wrapped] now have the wrapped call replay the first recorded interaction. The request is the same, so the response fits; the extra interaction is unused until the cassettes are re-recorded. - responses_metrics[sync|async] and parallel_tool_calls interleaved different requests ([u create, w create, u parse, w parse] and [u, w, u stream, w stream]), so in-order replay would hand the wrong response to the second call. Their unwrapped interactions (0 and 2) were removed from the cassettes at every version, deleting whole lines only. The remaining interactions are the real wrapped-client recordings. test_auto_openai.py still reads a create response first from test_openai_responses_metrics[sync]. Needs re-record (stale extra interaction only, replay passes now): test_openai_chat_metrics, test_openai_embeddings, test_openai_chat_streaming_sync, test_openai_chat_stream_helper, test_openai_chat_streaming_async, test_openai_chat_async_context_manager, test_openai_response_streaming_async, test_openai_responses_stream_helper, test_openai_responses_with_raw_response_create (incl. _stream, _stream_async), test_openai_responses_with_raw_response_parse, test_openai_images_generate, test_openai_images_edit, test_openai_audio_speech, test_openai_audio_transcription (incl. _text_format), test_openai_audio_translation Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The two tests made the same chat.completions call and differed only in the data-URL content part. One request now sends both an image_url part and a file part, and keeps every assertion from both tests: `image.png` as the default image filename, `test.pdf` kept for the file, content types, attachment keys, metrics, and the response text check. VCR doesn't match on the request body, so the merged test replays the existing image cassette (cassette_name=) at every version. The pdf cassettes are no longer read and were removed. Needs re-record: test_openai_data_urls_convert_to_attachments (its replayed request body is the old image-only one) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add fake-module tests for BaseIntegration.setup() and FunctionWrapperPatcher: setup twice wraps each target once, is_patched reflects state, the patch marker falls back to the root when the target rejects setattr (bound method), superseded_by yields in both setup() and wrap_target(), wrap_target is idempotent and skips missing attributes, and has_patch_marker ignores markers inherited through the MRO. - pydantic_ai: move test_model_classes_patcher_marker_check_is_mro_safe here. ModelClassesPatcher (a ClassScanPatcher) does not override has_patch_marker/mark_patched, so it only tested BasePatcher. - mistral: drop test_mistral_integration_setup_is_idempotent and the _core_method_refs helper (plus the class imports only it used). It only asserted a second setup() leaves methods unchanged, without checking the first setup() patched anything; that generic behavior (including CompositeFunctionWrapperPatcher and superseded_by, which every mistral patcher uses) is now covered here. Real setup() targets are still exercised by the audio/conversations setup tests and the auto-instrument script, which calls auto_instrument() twice and asserts one span. openrouter's setup idempotency test is kept: it is the only test that checks setup() actually wraps Embeddings.generate and Responses.send via their target_module paths (wrap_openrouter only uses the leaf attribute). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vcrpy's `all` record mode appends new interactions after the existing ones instead of replacing them, and replay uses the first match, so the re-recorded cassettes kept replaying the old responses. Keep only the newest recording run in each re-recorded cassette (no interaction is edited; stale ones and earlier runs are removed). google_genai 1.75.0: Google now rejects the legacy Interactions schema for google-genai < 2.0 with HTTP 400, so that re-record captured an error. Restore the original recording; 1.75.0 interactions cassettes can no longer be re-recorded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
pylint on Python 3.10-3.13 can't infer attributes set on a plain types.ModuleType and reports no-member for the fake SDK used by the generic patcher tests. Use a ModuleType subclass that declares them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Abhijeet Prasad (AbhiPrasad)
deleted the
test/consolidate-integration-tests
branch
October 8, 2026 19:39
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR combines redundant tests under
py/src/braintrust/integrations/and removes low-value ones, without losing coverage. Test code shrinks by about 2,300 lines (34 files: +1,783 / −4,096) and the cassette directories get smaller.Each change was checked against the code it tests before being applied. A test was kept whenever it turned out to be the only coverage of a code path, a version-gated branch, or a past bug fix.
What changed
test_utils.py. Both code paths still run, as[sync]/[async]cases. Cassettes were renamed withgit mv, or reused viavcr_cassette_name.input_schematest, whose setting never reaches the request.test_unwrapped_client_emits_no_spanstest. The image and PDF attachment tests are merged into one request.test_base.py: the patching machinery inbase.py(setup-twice, root-marker fallback,superseded_by,wrap_target, MRO-safe markers) now has direct tests that run intest_core. The Mistral setup-twice test is removed. The openrouter one is kept, because it's the only check on two module-path patch targets.>= 1checks are now exact counts, a few tests that asserted nothing now assert something, and async cases gained the stricter checks their sync twins had.SUCCESS, so a script that exits early with code 0 no longer passes.Fixes found along the way
@pytest.mark.vcr(cassette_name=...)was silently ignored. pytest-vcr passes marker kwargs to vcrpy as config, and vcrpy drops this one, so tests meant to share a cassette replayed their own. The rootconftest.pynow overridesvcr_cassette_nameto honor the kwarg; default naming is unchanged. google_discoveryengine's local positional-arg workaround is replaced by the shared fixture.--vcr-record=allappends to existing cassettes instead of replacing them, and replay uses the first match. So the re-record only took effect after removing the stale interactions, which the last commit does (it removes them; no interaction content was edited).test_interactions_deletestays as a separate test.Not included
mainstill containopenai-organization/openai-project/set-cookievalues, recorded before the current scrubbing was added. That fix belongs in a separate PR.Test plan
CI=1) at every matrix version, both before and after. There are no new failures, and the drop in test counts matches what was removed.test_core(982 passed),pylintand pre-commit pass.make check-unused-cassettesfinds no unused cassettes, andcheck-stale-cassettes.pyfinds no stale ones.🤖 Generated with Claude Code
Co-authored by StarfolkAI (@starfolkai)[bot]