Repository navigation
A hard dollar limit on one GraphRAG indexing run (settings.yaml only) #2584
domondi1
started this conversation in
Community tools
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Indexing cost is one of the most common GraphRAG surprises: entity extraction, gleanings, summarization and community reports add up to thousands of calls, many in flight at once, and the bill for one
graphrag indexdepends on the corpus in ways that are hard to predict up front. Here's one way to put a hard dollar ceiling on a single indexing run using onlysettings.yaml.call_argsare sent with every model request through LiteLLM, so the completion model can carry two extra headers (the run's id and its budget) to a small OpenAI-compatible gateway running next to GraphRAG:Each call reserves its worst-case cost before it's sent, so the concurrent extraction requests can't all spend the same remaining budget, and a call that doesn't fit is refused with a 402 before it reaches OpenAI. Once the budget is spent the run stops with that error instead of continuing to bill.
Caveats: use a new
INDEX_RUN_IDfor each run (a run's budget is fixed the first time its id is seen), setmax_completion_tokens(the reservation is based on it; without it each call reserves a 4,096-token default), and only the completion model goes through the gateway, so embeddings aren't counted. The model needs a price in Inferrail;gpt-4.1-mini,gpt-4.1andgpt-4o-miniare built in. The LiteLLMextra_headerspath is what we tested; I haven't run a full index of a large corpus through it, so reports are welcome.Setup details: guide. I maintain Inferrail (open source, Apache-2.0), so take the suggestion with that in mind.
All reactions