mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-28 14:59:47 +00:00
gateway.chat() requested cacheSystem:true but never got a system-prompt cache hit on single-turn callers (page-summary, skillopt, enrich): the call-level providerOptions.anthropic.cacheControl is real (it becomes Anthropic's documented top-level "auto-cache the last cacheable block" shorthand via @ai-sdk/anthropic 3.0.47+), but for a stable system prompt paired with a different user message every call, "the last cacheable block" is that ever-varying tail -- every call writes a fresh cache entry there and never reads a prior one. Fix: pass system as a SystemModelMessage object (ai's documented shape for attaching provider options to the system block) carrying its own providerOptions.anthropic.cacheControl when cacheSystem is requested, and mirror the same marker onto the last tool def (Anthropic caches everything up to and including the last cache_control block it sees). The call-level marker is kept, not removed -- it still gives toolLoop()'s growing multi-turn conversation a rolling cache breakpoint on each turn's tail. All three markers now derive from one canonical cacheControlValue computed after provider_chat_options config merging, so a configured TTL override (e.g. ttl: '1h') applies consistently instead of only reaching the call-level marker. Verified red-before-fix by stashing the gateway.ts diff and confirming the new assertions fail on unfixed code, then restoring. Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>