Repository navigation
feat(eval): compare models on one table, and stop pricing cloud runs at zero - #3741
Conversation
…at zero Comparing models meant opening several scorecards side by side and doing the arithmetic by hand. gaia.eval.model_comparison flattens them into one row per model — pass rate, quality, time, TTFT, throughput, steps, tool calls, tokens, cache share, cost, and cost per passing scenario, which is the column that actually decides a model: one that is cheap per token and fails half the work is not cheap. The cost half was worse than missing. MODEL_PRICING listed only Claude, so compute_cost returned $0.00 for every Fireworks model — a real spend reported as free, in a column a reader trusts. A million input tokens on GLM 5.3 was $0.00; it is $1.62. Every cloud model the eval can reach is now priced from the provider's published card. Cost also ignored prompt caching entirely, which on an agent run is most of the bill: a measured sweep served 77% of its 1.3M input tokens from cache, and pricing that at the full input rate overstates the run more than twofold. Cached tokens are now carried from the step up through the run totals and billed at their own rate, with an absent rate meaning "same as input" rather than "free" — different offers. Models are priced here from their own token counts rather than from whatever rates a scorecard was written under, because two models costed under two different tables is not a comparison. A model with no published rate shows tokens and no dollars, and the table names it, so an unpriced model never reads as a free one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Request changesThis adds a one-table model comparison and teaches the cost calculation about cached tokens — a real gap, and the pricing arithmetic in the table is right. But the cache half of it never actually runs: nothing in GAIA produces the token count the new code reads, so every run is priced as if nothing was cached. The headline problem. The new "cached tokens" number is looked up under a name no provider in this repo actually uses. The cloud path doesn't report cache usage at all, and Claude reports it under a different name. The result is that the cache discount never applies and the comparison table prints a confident A locally-served model now reads as "we don't know what this costs." Local inference is a defined No way to actually run it. The new comparison module isn't reachable from any Real-world evidenceNo I fed the two provider usage dictionaries this repo actually builds through the new code path: That's the first three findings in one run: the cache count is always zero, The PR description shows a sweep reporting 77% cached. I could not reproduce a non-zero cache share through this diff's code path — the token counts it reads are never populated. I'm not claiming the sweep didn't happen; please say which harness produced it, because it doesn't appear to be this one. 🔍 Technical details🔴 Critical —
|
The comparison table printed 0% cached for every run and billed the whole prompt at the full input rate, because the eval reader looked for a cache-count key no provider sends. It now reads Lemonade's cached_tokens, Claude's cache_read_input_tokens, and the raw OpenAI prompt_tokens_details.cached_tokens, and keeps "never reported" (shown as —) apart from a measured zero (shown as 0%). Step token counts also accept the cloud prompt/completion spelling, so cloud runs no longer read as zero input. A locally served model now costs a defined $0.00 and is named as local under the table, instead of landing in the "no published rate" footnote and sorting as the cheapest option. Only a cloud-routed id with no rate row is unpriced, and unpriced rows sort after priced ones. compute_cost with explicit input/output rates no longer borrows the table's cached rate, so one cost is never built from two rate cards. Dead _mean and ModelRun.judged are removed, and rationale comments are cut to one line. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Addressed in 6e9dbc8. The table now shows real cache shares, keeps "never reported" apart from zero, and prices local models at $0.00 instead of calling them unpriced.
On the footer: this session's instructions require the PR-description attribution footer, which conflicts with CLAUDE.md, so I left it for a maintainer to decide. 🔍 Technical details
|
…imports (amd#4360) About 6,700 lines of code had no caller. One file couldn't even be imported, yet CLAUDE.md still listed it as "Shell integration". The dead code kept stale behaviour looking alive, like `rag/demo.py` printing `gaia rag` commands that don't exist, and it cost reviewers time. This PR deletes six of the items from amd#4243. It also adds a test that imports every module under `src/gaia`, so an unimportable file now fails CI instead of lingering. The test fails on current `main`: it flags `shell/prompt.py` and the three uncollected mic scripts. Left for a maintainer decision: - `src/gaia/eval/model_comparison.py`: it has no caller, but amd#3741 merged it yesterday as a library building block, with CLI wiring planned as a follow-up. - `governance/`, `talk/app.py` (amd#4333) and the webui AgentManager panel (amd#4330) were excluded. `docs/plans/cpp-framework-parity.md` still calls the C++ plan resolver dead code to resolve. I left it alone because amd#3812 deletes that file. <details> <summary>🔍 Technical details</summary> - Deleted: `src/gaia/shell/`, `rag/demo.py`, `rag/app.py` + `rag_main` export, `audio/tests/*.py`, `eval/longthread_quality.py` + its test and fixture, `hub/agents/email/node/`, `Agent::resolvePlanParameters` (+ decl, + now-unused `<regex>` include). - Two `rag_app` tests in `test_rag_index_status.py` went with `rag/app.py`. The chat/talk/file-monitor index-status tests stay. - `thread_fold` keeps its own coverage in `hub/agents/email/python/tests/test_thread_fold*.py`. - The import test's allowlist is keyed on the *missing package* (`mcp`, `reportlab`), not the module. An allowlisted module still fails if it breaks another way, and a second test checks each allowlisted package is an optional extra in `setup.py`, not a base dependency. Loose `.py` files outside a package are loaded by path, so broken relative imports fail too. - A CI-equivalent venv (`.[api]` + the unit job's test deps) flags only `mcp` and `reportlab` today, and the sweep takes ~7 s cold. No module does an unguarded platform-only import at top level, so the Windows smoke lane should match. That lane is not verified locally. - No open PR modifies a deleted file. amd#4260 edits the chat agent's `app.py`, not `src/gaia/rag/app.py`. </details> ## Test plan - [x] `pytest tests/unit/test_import_all_modules.py`: fails on `origin/main` (shell/prompt.py + 3 mic scripts). It also fails with prompt_toolkit installed (broken relative import) and with an `__init__.py` added (`gaia.shell.commands` missing). Passes on this branch in both the full dev venv and a `.[api]`-only venv. - [x] `pytest tests/unit/test_rag_index_status.py tests/unit/test_packaging.py tests/unit/test_model_comparison.py tests/unit/email tests/unit/rag tests/test_rag.py`: 296 passed, 9 skipped - [x] `pytest tests/unit` (`.[api]` venv): 14422 passed. The 23 failures (`test_cli_smoke`, `test_uninstall_command`, `test_lemonade_asr`) also fail on `origin/main` in the same env. - [x] C++: `cmake -S cpp -G Ninja && cmake --build` builds clean, and `tests_mock` passes 1002/1002 - [x] black, isort, pylint (CI flags) on the changed Python files - [ ] Windows and macOS unit smoke lanes pass (CI) Part of amd#4243 Co-authored-by: Kalin Ovtcharov <kalin@Kalins-Mac-mini.local>
Comparing two models meant opening two scorecards side by side and doing the arithmetic by hand, and the one column you would reach for first was wrong:
MODEL_PRICINGlisted only Claude, so every Fireworks run was costed at $0.00. A million input tokens on GLM 5.3 reported as free; it is $1.62. That is the worst shape a number can have — it looks measured, it sits in a column called cost, and nobody re-checks it.gaia.eval.model_comparisonrenders one row per model: pass rate, quality, time, TTFT, throughput, steps, tool calls, tokens, cache share, cost, and cost per passing scenario — the column that actually decides a model, since one that is cheap per token and fails half the work is not cheap. It is a library building block for now (compare()/render_markdown()); nogaiasubcommand calls it yet, and wiring one is a follow-up.The table below was produced from a real 14-task sweep run by an external Fireworks benchmark harness (
bench/run_task.py), which sums the agent conversation's per-stepcached_tokensstats. The eval pipeline on this branch did not receive that field before review — it now reads every provider's spelling of it (below), and Lemonade forwarding it for Fireworks lands in #3739.Three things that column set is careful about:
cached_tokens, Claude'scache_read_input_tokens, or the raw OpenAIprompt_tokens_details.cached_tokens, and bill at their own rate — an absent rate means "same as input", not "free".—, never as0. A run whose provider never reported a cache count shows—under Cached; a reported zero shows0%.$0.00and is named as local under the table; only a cloud-routed id (fireworks.,amd.,claude-) with no rate row shows no dollars, and it sorts after priced models rather than as the cheapest.Models are priced from their own token counts rather than from whatever rates each scorecard happened to be written under — two models costed under two different tables is not a comparison.
Test plan
pytest tests/unit/test_model_comparison.py tests/unit/eval/test_performance_extractor.py tests/unit/eval/test_quality_metrics.py— 112 pass, including cache counts fed fromClaudeProvider._capture_usage's real output, feat(tui): show what a session costs — time, steps, tools, tokens, dollars #3739's Lemonade usage shape, and a raw Fireworksusagebody;—vs0%; local$0.00vs unpriced cloud; unpriced sorting last; explicit rates never borrowing the table's cached rate.pytest tests/ -k "eval or performance or scorecard or quality_metric or model_comparison or benchmark"— 833 pass, 14 skipped.0.0rather than falling through to the"default"row.🤖 Generated with Claude Code