Files
OpenJarvis/examples/doc_qa/doc_qa.py
T
05f2c02131 feat: Algolia DocSearch + learning subsystem reorganization (#43)
* chore: create learning subdirectory structure (routing, agents, intelligence)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: extract classify_query to routing/_utils.py

Move the classify_query() function and its regex patterns into a shared
utility module so multiple routing policies can import it without
depending on the full trace_policy module.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: move routing files to learning/routing/ subdirectory

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: create LearnedRouterPolicy merging trace-driven + SFT routing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add conditional Algolia DocSearch integration

Add Algolia DocSearch as an optional search upgrade — native lunr.js
search remains the default until credentials are configured. Includes
CDN assets, Jinja2 conditional config injection, init script with
graceful fallback, light/dark theme CSS, improved search tokenization
for snake_case/dotted identifiers, and search boosts for key pages.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: move agent_evolver and skill_discovery to learning/agents/

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: move learning/orchestrator to learning/intelligence/orchestrator

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: delete removed learning policies, rewrite __init__.py, clean up api_routes

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add SFT/GRPO/DSPy/GEPA config dataclasses, update LearningConfig

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add general-purpose SFT trainer (intelligence/sft_trainer.py)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: update stale imports in multi_model_router example

Update imports to use new learning/routing/ paths after the
subdirectory reorganization. Replace BanditRouterPolicy with
LearnedRouterPolicy.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add general-purpose GRPO trainer (intelligence/grpo_trainer.py)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add DSPy agent optimizer (agents/dspy_optimizer.py)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add GEPA agent optimizer (agents/gepa_optimizer.py)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add learning-dspy and learning-gepa optional dependency extras

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: update integration test to check for learned policy instead of grpo

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: clean up stale APIs and unused params in examples

- deep_research: remove system_prompt and max_turns params not accepted
  by Jarvis.ask(), inline system prompt into the query instead
- doc_qa: remove unused --top-k CLI arg that was never passed to the API
- multi_model_router: fix select_model() call to match single-arg signature

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: import SFT/GRPO trainers in intelligence/__init__.py for registry

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: remove .md file changes from PR

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: restore search boost frontmatter for key docs pages

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 21:34:31 -07:00

121 lines
3.2 KiB
Python

#!/usr/bin/env python3
"""Document QA — index documents and answer questions with citations.
Usage:
python examples/doc_qa/doc_qa.py --help
python examples/doc_qa/doc_qa.py --docs-path ./docs \
--query "How does authentication work?"
python examples/doc_qa/doc_qa.py --docs-path ./papers \
--query "What are the main findings?" \
--model gpt-4o --engine cloud
"""
from __future__ import annotations
import argparse
import sys
def main() -> None:
parser = argparse.ArgumentParser(
description=(
"Index documents and answer questions "
"with memory-augmented citations."
),
)
parser.add_argument(
"--docs-path",
type=str,
required=True,
help="Path to the documents directory (or single file) to index.",
)
parser.add_argument(
"--query",
type=str,
required=True,
help="The question to answer based on the indexed documents.",
)
parser.add_argument(
"--model",
type=str,
default="qwen3:8b",
help="Model to use for answering (default: qwen3:8b).",
)
parser.add_argument(
"--engine",
type=str,
default="ollama",
help="Engine backend: ollama, cloud, vllm, etc. (default: ollama).",
)
parser.add_argument(
"--chunk-size",
type=int,
default=512,
help="Chunk size for document indexing (default: 512).",
)
args = parser.parse_args()
try:
from openjarvis import Jarvis
except ImportError:
print(
"Error: openjarvis is not installed. "
"Install it with: uv sync --extra dev",
file=sys.stderr,
)
sys.exit(1)
print(f"Documents: {args.docs_path}")
print(f"Query: {args.query}")
print(f"Model: {args.model} | Engine: {args.engine}")
print("-" * 60)
try:
j = Jarvis(model=args.model, engine_key=args.engine)
except Exception as exc:
print(
f"Error: could not initialize Jarvis -- {exc}\n\n"
"Make sure your engine is running. For Ollama:\n"
" ollama serve\n"
" ollama pull qwen3:8b\n\n"
"For cloud engines, ensure API keys are set in your .env file.",
file=sys.stderr,
)
sys.exit(1)
# Step 1: Index documents into memory
print("Indexing documents...")
try:
result = j.memory.index(args.docs_path, chunk_size=args.chunk_size)
print(f" Indexed {result['chunks']} chunks from {result['path']}")
except Exception as exc:
print(f"Error indexing documents: {exc}", file=sys.stderr)
j.close()
sys.exit(1)
# Step 2: Ask the question with memory context enabled
print("Searching for relevant context...")
try:
response = j.ask(
args.query,
context=True,
)
except Exception as exc:
print(f"Error during QA: {exc}", file=sys.stderr)
sys.exit(1)
finally:
j.close()
print()
print("=" * 60)
print(" Answer")
print("=" * 60)
print()
print(response)
print()
print("=" * 60)
if __name__ == "__main__":
main()