mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-28 06:23:01 +00:00
Adds a `perplexity` embedding recipe (OpenAI-compatible at https://api.perplexity.ai/v1, auth via PERPLEXITY_API_KEY only — never an OPENAI_API_KEY fallback) covering pplx-embed-v1-0.6b and pplx-embed-v1-4b. Perplexity's /embeddings endpoint diverges from OpenAI's wire shape in two places that break the AI SDK adapter, handled by a new perplexityCompatFetch shim (mirrors the Voyage/ZeroEntropy pattern incl. the two-layer OOM caps): - encoding_format only accepts base64_int8/base64_binary; the SDK's 'float' default is forced to 'base64_int8' outbound. - The response embedding is base64-encoded signed int8 components (natively quantized); decoded to number[] inbound so the SDK's Zod schema validates. Cosine similarity is scale-invariant, so raw int8 components rank correctly. Flexible dims (Matryoshka-style 128..native max: 1024 for 0.6b, 2560 for 4b) validate fail-loud in dims.ts + the init preflight; `dimensions` is Perplexity's native field so no wire translation is needed. default_dims is 1024 (works on a plain vector column for both models); the 4b model's full 2560 width rides the existing halfvec (>2000 dims) storage/ANN path. Pricing entries land in embedding-pricing.ts. Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>