- New provider 'openai' calls any OpenAI-compatible /chat/completions endpoint
(OpenAI, OpenRouter, Together, Groq, NVIDIA, vLLM, Ollama, ...), set via
OPENAI_BASE_URL / OPENAI_API_KEY / OPENAI_MODEL. Now the default; auto-detects.
- Robustness for models that degenerate (e.g. Kimi repetition loops): anti-
repetition sampling (temperature + frequency penalty), and a schema-aware
retry — call_json(require_keys=...) retries a fresh sample on degenerate or
unparseable output (LLMRetryable). Relevance requires 'signals', synth 'briefs'.
- response_format=json_object (toggle via OPENAI_JSON_MODE); base URL accepts a
full /chat/completions URL or a /v1 base.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Relevance filter: evaluate all of an account's signals in one LLM call
(keyed array, chunked at 8) instead of one call per signal.
- Synthesizer: produce all of an account's briefs in one call (one object
per capability) instead of one call per capability. ~2 calls/account.
- llm: wrap JSONDecodeError in LLMError so callers fall back to mock instead
of crashing; disable Gemini thinking (thinkingBudget=0) to stop thinking
tokens truncating JSON; add 429 backoff.
- accounts: add Booz Allen, Dominion Energy, Hope Gas, Markel, GW Medical
Faculty Associates (verified domains; CRM tier captured in crmTier).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>