Skip to content

Switch explainer to Sonnet 5.5 and upgrade deps (anthropic 1.x) - #48

Merged
mattgodbolt merged 1 commit into
mainfrom
sonnet-5-5
Sep 29, 2026
Merged

mattgodbolt merged 1 commit into
mainfrom
sonnet-5-5

Conversation

@mattgodbolt

Copy link
Copy Markdown
Member

Moves the production explainer from claude-sonnet-5 to claude-sonnet-5-5 (released 2026-09-28) and upgrades all Python deps, including the anthropic SDK 0.120 → 1.9.

Eval: Sonnet 5 vs Sonnet 5.5

Setup: prompt-test run --review with the production prompt, 21 cases, 2 runs per config, Opus 5 reviewer with adaptive thinking. Everything ran on the upgraded deps, so the model is the only thing that changes. Accuracy covers the 16 assembly cases per run. The haiku cases are excluded because the reviewer marks every haiku "incorrect" for not explaining the assembly, and does so for every config.

Config Correct (asm) Errors Warnings Mean latency p50 p95 Max Mean out tok Mean in tok $/call
Sonnet 5 (current prod) 27/32 (84%) 7 50 6.6s 7.6s 10.3s 11.3s 541 2379 $0.0102
Sonnet 5.5, default effort (this PR) 29/32 (91%) 3 34 5.8s 5.8s 9.9s 12.6s 605 2380 $0.0108
Sonnet 5.5, effort medium 29/31 (94%) 2 34 5.9s 6.2s 11.0s 11.8s 600 2380 $0.0108
  • Accuracy: reviewer-flagged errors fell from 7 to 3 and warnings from 50 to 34. Runs are noisy: Sonnet 5 had 1 error in one run and 6 in the other, while Sonnet 5.5 had 2/1. Treat this as "no worse, likely better", not a precise number.
  • Latency: about the same. Mean is slightly lower and max is 12.6s, well under the 30s API Gateway ceiling.
  • Cost: the price is the same ($2/$10). Outputs run about 12% longer, so cost per call is about 6% higher, roughly +$1.30 per fortnight at current traffic. Input token counts are identical because the tokenizer is the same.
  • Effort: medium and high came out indistinguishable, so effort stays unset (API default high).

Changes

Model

  • app/prompt.yaml: claude-sonnet-5-5 with thinking: {type: between_tools}. Sonnet 5.5 returns a 400 for thinking: {type: disabled}. between_tools is its lowest setting, and with no tools it means no thinking. Verified: the response is a single text block. The useThinking per-request path (adaptive) is unchanged.
  • app/model_costs.py: Sonnet 5 goes to $2/$10, since the planned September rise to $3/$15 was cancelled and $2/$10 is now the standard price. Adds Sonnet 5.5 ($2/$10) and Opus 5.5 ($4/$20).
  • CLAUDE.md: updates the thinking, temperature, caching and refusal gotchas for 5.5. Two notable changes: the minimum cacheable prefix drops to 512 tokens, and 5.5 refuses in more categories (bio, reasoning_extraction, general_harms).
  • The reviewer stays on Opus 5 so before and after are graded the same way.

Deps

  • uv lock --upgrade, with pyproject floors raised to the locked versions. Pre-commit ruff goes v0.13.0 → v0.16.9.
  • anthropic 1.x: the only breaking call site was build_api_payload. It put temperature straight into the kwargs, which is a TypeError on 1.x, and it also defaulted to sending 0.0 even when the YAML didn't set one. Now temperature is sent only when configured, via extra_body, and only when thinking is off. New tests cover this. There is no with_raw_response, httpx or output_format usage.
  • httpx2 arrives transitively with anthropic 1.x. It is the SDK's HTTP layer: the maintained fork of httpx, published by Pydantic (github.com/pydantic/httpx2). httpx2-jsfetch in the lock only applies on Emscripten and is never installed here.
  • Newer ruff flags (str, Enum) (UP042), so the two enums move to StrEnum. All uses go through .value or Pydantic, so behaviour is unchanged.
  • The explain log line now shows the thinking type instead of a bool, which would always be True now.

GitHub Actions and Docker base images are left to Dependabot. setup-uv v9 → v10 is a major bump.

Testing

  • uv run pytest: 113 passed. pre-commit run --all-files: clean.
  • Local fastapi dev + ./test-explain.sh against SDK 1.9.0 and Sonnet 5.5: explanation returned, and the log shows thinking=between_tools.
  • After merge, watch ClaudeExplainRefusal because 5.5 refuses in more categories.

🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings September 29, 2026 12:36

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The model migration, SDK adaptation, dependency updates, and associated tests are consistent and complete.

Review effort: Balanced
Findings: None

What changed in this PR

Upgrades the explainer to Sonnet 5.5 and Anthropic SDK 1.x while preserving request, caching, and cost-tracking behavior.

Changes:

  • Migrates production to Sonnet 5.5 with between_tools thinking.
  • Updates SDK payload construction, pricing data, enums, and tests.
  • Refreshes Python dependencies and Ruff configuration.
File Description
app/​prompt.yaml Configures Sonnet 5.5.
app/​prompt.py Adapts payloads for Anthropic SDK 1.x.
app/​explain.py Logs the active thinking type.
app/​model_costs.py Updates Claude 5 pricing.
app/​explanation_types.py Migrates enums to StrEnum.
app/​test_explain.py Tests temperature payload handling.
app/​test_model_costs.py Tests Sonnet 5 pricing.
pyproject.toml Raises dependency floors.
uv.lock Locks upgraded dependencies.
.pre-commit-config.yaml Upgrades Ruff hooks.
CLAUDE.md Documents Sonnet 5.5 behavior.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@mattgodbolt
mattgodbolt merged commit 5d9fd87 into main Sep 29, 2026
3 checks passed
@mattgodbolt
mattgodbolt deleted the sonnet-5-5 branch September 29, 2026 18:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants