mirror of
https://github.com/open-jarvis/OpenJarvis.git
synced 2026-07-28 14:07:55 +00:00
* feat: add query complexity analyzer with CLI and UI integration Classify incoming queries by difficulty (trivial→very_complex) to suggest appropriate token budgets for local vs. cloud routing. - Add score_complexity() with weighted signals (length, code, math, reasoning, multi-step, creative) and token tier mapping - Wire into `jarvis ask`: auto-suggest max_tokens when not set by user, show complexity in --profile output, log at DEBUG level - Add complexity metadata to /v1/chat/completions API response - Display complexity tier and score in frontend XRayFooter - Extend RoutingContext with complexity_score, suggested_max_tokens, has_reasoning fields - Update HeuristicRouter to route on complexity_score instead of raw query_length - Remove duplicated regex patterns from router.py (use complexity module as single source of truth) - Add 30 unit tests for complexity module - Fix existing router tests for new complexity-based routing Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: pass temperature and max_tokens from UI settings to backend The settings page stores temperature and max_tokens in the frontend store, but these values were never included in the chat API request. The backend Pydantic model defaults max_tokens to 1024 when the field is absent, which is too low for thinking models like qwen3.5 — they consume all tokens on reasoning and return empty content. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: pass UI settings to backend and auto-bump max_tokens from complexity - Pass temperature and max_tokens from frontend Settings store to the backend API (cherry-picked from fix/ui-max-tokens-passthrough) - Server-side: bump max_tokens when the complexity analyzer suggests a higher budget (e.g. for thinking models on complex queries), never reduce below the client-requested value - Fixes empty responses with thinking models (e.g. qwen3.5) that consumed all tokens on internal reasoning Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: raise token budget tiers to prevent empty responses on thinking models Double all tier budgets (trivial: 512→1024, simple: 1024→2048, etc.) so that thinking models like qwen3.5 have enough headroom for internal chain-of-thought plus visible output, even on simple queries. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: CI lint and test failures - Break long lines in routes.py to satisfy 88-char limit (E501) - Add intelligence.max_tokens and temperature to mocked config in test_ask_router.py so complexity analyzer can compare against int instead of MagicMock Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: line too long in test_complexity.py (E501) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>