Repository navigation
Conversation
claude-opus-5 (the default judge model, config.py DEFAULT_CLAUDE_MODEL) prepends a ThinkingBlock with no .text attribute. Indexing message.content[0].text blindly crashes with AttributeError on the model's first reply. Add failing coverage for analyze_file and analyze_file_with_usage's four call sites (#3884), plus a case asserting a response with no text block raises an actionable error naming the model and block types instead of silently degrading.
claude-opus-5 (DEFAULT_CLAUDE_MODEL) prepends a ThinkingBlock with no .text attribute, so message.content[0].text crashed with AttributeError on the judge's very first reply -- every eval gate that uses the default judge was unrunnable. Add ClaudeClient._first_text(), which scans for the first block that actually has text instead of indexing position 0, and use it at all four call sites in analyze_file / analyze_file_with_usage. A response with no text block at all raises a ValueError naming the model and the block types received, rather than returning None/ and silently turning a broken judge into a wrong eval score.
Verdict: Approve with suggestionsThis fixes a real crash: the eval judge blew up on the very first reply from its own default model, because that model answers with a thinking block first and the code always read block zero. The fix scans for the first block that actually carries text, and fails loudly with a useful message if none does — the right call for a judge, since a quietly empty verdict is worse than a red gate. Two things worth a follow-up, neither blocking:
Real-world evidenceStrong, and it supports the verdict. The no-text-block path was exercised too, and names the real SDK class: CLI ( 🔍 Technical details🟡 Important1. The same
Apply the same shape at 2.
If you deliberately want first-block-only semantics, a one-line WHY comment saying so would stop the next reader from "fixing" it back. Strengths
|
|
Closing in favour of #3885, which fixes the same bug and more. Mine only covered the four sites in It caught a real error in #3884's acceptance criteria too: |
The eval judge crashes the moment it talks to its own default model.
claude-opus-5(the defaultDEFAULT_CLAUDE_MODEL) is an extended-thinking model that prepends aThinkingBlockwith no.textattribute —ClaudeClientindexedmessage.content[0].textblindly, so the very first judged reply threwAttributeErrorand every eval gate using the default judge was unrunnable.ClaudeClient._first_text()scans the response for the first block that actually has text instead of assuming block 0. Used at all fouranalyze_file/analyze_file_with_usagecall sites. A response with no text block at all raises aValueErrornaming the model and the block types received — no silent""/Nonethat would turn a broken judge into a silently wrong score.claude-opus-5sometimes returns['thinking','text']and sometimes['text']. Both shapes are observed; the trigger is unidentified.content[0].textis therefore unsafe on this path regardless of how often each shape occurs. A reviewer who runs one call and sees['text']has observed the benign shape, not disproved the bug.Closes #3884
Test plan
tests/unit/eval/test_claude_judge.py— 34 passed (5 new cases pin_first_text: HTML/binary branches of bothanalyze_fileandanalyze_file_with_usage, plus the no-text-block error path)python -m pytest tests/ -x -k "claude_judge or claude"— 131 passed, 21 skippedpython util/lint.py --all— all quality checks passedANTHROPIC_BASE_URL+ Ocp-Apim key), same model, same call:AttributeError: 'ThinkingBlock' object has no attribute 'text''The sky appears blue because of Rayleigh scattering, in which air molecules scatter shorter (blue) wavelengths of sunlight more strongly than longer ones.'Reviewer note on response-shape variability
Sample from an independent verification run against this branch (same
max_tokens, same model, consecutive calls):A separate call in the same session returned a bare
['text'](no thinking block), which the old code happened to handle. The sample is too small and uncontrolled (gateway routing, thinking budget, and model pinning were not held constant) to support a rate for either shape — the point is only that both shapes occur, not how often. A judge whose pass/fail depends on response shape is worse than one that fails deterministically: it's what trains reviewers to ignore a red eval gate as noise.