mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
fix(recipes/openai): add max_batch_tokens to embedding touchpoint
OpenAI is the only recipe in the codebase without a max_batch_tokens cap. Every other provider declares one (voyage=120K, azure-openai=8K, dashscope=8K, zhipu=8K, minimax=4K). Without it, gbrain's recursive-halving safety net never engages — batches dispatched purely on the char/4 estimator window will trip OpenAI's 1M-token TPM ceiling on token-dense pages (Discord exports, JSON dumps, code-heavy markdown), then retry storm and block the queue head. Setting cap to 100_000: - gbrain's batcher estimates tokens as chars/4 - Token-dense markdown+JSON tokenizes at ~chars/2.7 - 100K estimated = ~150K real worst-case, safely under OpenAI's 300K per-request hard cap and the 1M/min TPM ceiling - Leaves headroom for recursive-halving on outlier chunks
This commit is contained in:
@@ -17,6 +17,12 @@ export const openai: Recipe = {
|
||||
dims_options: [256, 512, 768, 1024, 1536, 3072],
|
||||
cost_per_1m_tokens_usd: 0.13,
|
||||
price_last_verified: '2026-04-20',
|
||||
// OpenAI per-request hard cap is 300K tokens. Free/Tier-1 TPM is 1M.
|
||||
// Cap batches conservatively at 100K to handle token-dense content
|
||||
// (Discord/Slack markdown+JSON tokenizes at ~chars/2.7, not the chars/4
|
||||
// estimate the batcher uses). 100K estimated = ~150K real tokens worst-case,
|
||||
// safely under both the 300K per-request and 1M TPM ceilings.
|
||||
max_batch_tokens: 100_000,
|
||||
},
|
||||
expansion: {
|
||||
models: ['gpt-5.2', 'gpt-4o-mini'],
|
||||
|
||||
Reference in New Issue
Block a user