Skip to content

Model Capabilities Operations (Rerank and Small LLM)

Argus ships two optional external-model capabilities, both disabled by default and both governed by validated configuration, file-based secrets, and a per-source external-model-processing license gate:

Capability Provider and model Where it runs Kill switch
Online filing-fragment rerank Voyage rerank-2.5-lite (POST https://api.voyageai.com/v1/rerank) filing_search and filing_evidence_extract request paths (api/mcp/worker) ARGUS_MODELS__RERANK__ENABLED
Offline small-LLM tasks OpenAI-compatible meta/muse-glimmer-30b (POST https://cliproxy.668618.xyz/v1/chat/completions, schema embedded in the fixed instruction and re-validated locally) Celery worker only, after a filing is ingested and parsed ARGUS_MODELS__SMALL_LLM__ENABLED

Neither capability ever exposes model parameters to CLI/MCP/REST callers, and neither reintroduces question answering: rerank only reorders already authorized fragments, and the small LLM only produces pending-review structured candidates with source positions.

1. Prerequisites

  1. Secrets (one-line, root-owned, mode 0640 files in $ARGUS_SECRETS_DIR_HOST):
  2. voyage_api_key - already required for vector search; also used by rerank.
  3. model_api_key - the small-LLM key; mounted only into the worker service as SMALL_LLM_API_KEY_FILE. The first-version model is not an OpenAI model; only the HTTP protocol is OpenAI-compatible.
  4. External-model-processing license approval. In config/data_licenses.production.yaml, add an external_model_processing entry to the approved SEC EDGAR license (see the commented example in config/data_licenses.yaml). Text without an explicit allowed entry for the provider is never sent - including while generic license enforcement is administratively disabled. Production approvals may only list the reviewed sec_edgar target source. The operator's project-level allowed decision and scope are recorded in the model acceptance decision. The production placeholder grants only this source to Voyage and the reviewed small-LLM capability; excluded third-party material remains out of scope.
  5. Static preflight (scripts/production-preflight.sh) passes. The preflight distinguishes "capability disabled" (fine) from "misconfigured" (deployment stops): it rejects conflicting VOYAGE_RERANK_MODEL vs ARGUS_MODELS__RERANK__MODEL_ID values, non-reviewed model IDs, non-official base URLs, and any attempt to enable the small LLM outside the worker role.

2. Configuration reference

The runtime settings object (models.rerank, models.small_llm) is the single authority; business code never reads these environment variables itself.

Variable Default / range Notes
ARGUS_MODELS__RERANK__ENABLED false Enable only after preflight and an authorized live check.
ARGUS_MODELS__RERANK__MODEL_ID rerank-2.5-lite Must be a reviewed Voyage rerank model (rerank-2.5, rerank-2.5-lite, rerank-2, rerank-2-lite, rerank-1.5, rerank-lite). The legacy VOYAGE_RERANK_MODEL is honored only when the new value is absent; conflicting values fail startup.
ARGUS_MODELS__RERANK__BASE_URL https://api.voyageai.com Fixed; arbitrary hosts are rejected.
ARGUS_MODELS__RERANK__TIMEOUT_SECONDS 3 (1-5) Online call budget.
ARGUS_MODELS__RERANK__CANDIDATE_LIMIT / RESULT_LIMIT 30 / 10 Result limit cannot exceed the candidate limit or the public tool limit.
ARGUS_MODELS__SMALL_LLM__ENABLED false (worker only) Enabling it on api/mcp/beat roles fails startup validation.
ARGUS_MODELS__SMALL_LLM__MODEL_ID meta/muse-glimmer-30b Must be paired with its reviewed base URL. gpt-5.4-mini-2026-03-17 remains configurable only with https://api.openai.com/v1.
ARGUS_MODELS__SMALL_LLM__BASE_URL https://cliproxy.668618.xyz/v1 Reviewed hosts only. The protocol (chat_completions or responses) is derived from the host and is not separately configurable.
ARGUS_MODELS__SMALL_LLM__REASONING_EFFORT low none, minimal, low, or medium. Measured quality needs low; none misses facts in documents with distractor fragments. Streaming keeps the gateway from idle-timing out longer reasoning.
ARGUS_MODELS__SMALL_LLM__TIMEOUT_SECONDS 120 (10-120) Per-call budget. Chat completions stream so the gateway does not idle-timeout a long generation.
ARGUS_MODELS__SMALL_LLM__MAX_INPUT_TOKENS / MAX_OUTPUT_TOKENS / MAX_BATCHES_PER_DOCUMENT 8000 / 2000 / 50 Positive integers with hard upper bounds; exceeding the batch cap marks coverage incomplete instead of silently truncating.
ARGUS_MODELS__SMALL_LLM__FILING_FACTS_ENABLED, ..._EVENTS_ENABLED, ..._EVIDENCE_CLASSIFICATION_ENABLED false Task switches are inert while the master switch is off.

In .env.production set ARGUS_RERANK_ENABLED=true for rerank and ARGUS_SMALL_LLM_ENABLED=true (plus the task switches) for the small LLM, then docker compose ... up -d worker api (and mcp variants). The compose anchor always keeps the small LLM off and keyless; only the worker service receives SMALL_LLM_API_KEY_FILE.

3. What each capability does at runtime

Rerank. filing_search / filing_evidence_extract first apply tenant, company (pushed into the repository query), document, license-status, and as_of filters; if two or more candidates remain and every source document approves external model processing for voyage, the authorized fragment texts are sent once to Voyage and the visible results are reordered by the provider scores (ties keep the keyword order). Any denial, timeout, rate limit, or invalid provider response skips rerank entirely and returns the original order with a machine-readable reason - the user request never fails because of rerank. Provider scores only order candidates; they are never stored as fact confidence.

Small LLM. After a filing is successfully ingested and parsed, the connector dispatches the offline task argus.tasks.model_assisted_filing_processing (Redis-locked per document, idempotent per model snapshot + prompt version). The task skips documents with a deterministic XBRL path, denies unapproved licenses before any network call, batches bounded positioned fragments, and re-validates every candidate deterministically (field enum, excerpt locatability, numeric value presence, currency shape, period consistency, event time sanity). Accepted candidates enter as MODEL_ASSISTED / pending review only; failures are quality issues on the document and never re-mark the connector sync or the ingestion as failed. Fact and event processing each have a 50-batch document cap. Oversized fragments or extra batches remain flagged as incomplete coverage and the offline task reports a partial result for follow-up. When evidence classification is enabled, the same worker builds candidates from the parsed document, calls Voyage once per ambiguous model fact, then uses small-LLM relation classification and deterministic value checks before adding any secondary evidence mapping. These mappings remain pending review.

4. Observing

  • Metrics: argus_model_calls_total{provider,model,operation,status}, argus_model_rejected_candidates_total (candidates discarded by deterministic verification). The provider/model/operation labels on the call and token series attribute provider spend to the exact capability and task for cost accounting, argus_model_call_latency_seconds, argus_model_call_degradations_total, argus_model_call_retries_total, argus_model_tokens_total, and argus_model_pending_review_candidates_total on each role's metrics endpoint (see observability).
  • Structured audit events: event="model_call" records carry model snapshot, prompt version, input SHA-256, source ids, status, degradation reason, and task correlation id - never prompt text, fragment text, or keys.
  • Document-level trail: model_assisted_processing_started, model_assisted_processing_denied, model_assisted_extraction_failed, model_assisted_coverage_incomplete, model_assisted_event_processing_failed, and model_assisted_processing_completed quality issues on the filing document.
  • Live evaluation: uv run python scripts/evaluate_model_quality.py --rerank reports nDCG@10 / Recall@10 versus the keyword baseline (gate: rerank must not be worse); --small-llm reports the joint precision / recall thresholds (facts 90%, events 90%, supports-precision 95%, positive recall 70%). Both need authorized keys (VOYAGE_API_KEY / SMALL_LLM_API_KEY).

Release status (2026-09-23, operator acceptance): The recorded provider smoke calls and switch drill used the reviewed providers (rerank-2.5-lite, voyage-4-lite, meta/muse-glimmer-30b over https://cliproxy.668618.xyz/v1). Sanitized records are in evals/reports/live-integration-2026-09-23.json and evals/reports/switch-drill-2026-09-23.json. The 50-query rerank and 30-per-task small-LLM reports in evals/reports/ measured synthetic samples whose labels are generated by evals/build_datasets.py. The operator's dated decision permits these programmatically labeled samples as this release's evaluation basis and allows the specified SEC EDGAR source for external model processing. The recorded numerical thresholds pass under that policy; no human annotations are claimed. The fact set has 8 negative labels against the original target of 10; the dated decision explicitly accepts this difference. Known limitation: fragments that print values with a scale word ("42 million") are rejected by the deterministic scale check when the model expands the number - candidates from such fragments stay pending review until a prompt revision is re-evaluated. Production enablement on the live server still requires an operator with server access: set ARGUS_RERANK_ENABLED=true and ARGUS_SMALL_LLM_ENABLED=true (plus task switches) in .env.production, install model_api_key as the small-LLM key file, and restart worker/api per this runbook.

5. Disabling and rollback

  • Rerank off: set ARGUS_RERANK_ENABLED=false (or remove the rerank block) and restart api/mcp/worker. filing_search / filing_evidence_extract immediately return the original keyword order; nothing else changes. No key removal is required.
  • Small LLM off: set ARGUS_SMALL_LLM_ENABLED=false and restart the worker. New offline model tasks stop; already-indexed structured events, XBRL facts, history, and audit remain fully readable. Pending-review model candidates simply stop accruing.
  • Full rollback of text egress: additionally remove (or set prohibited) on the external_model_processing entries in the licenses file and restart; every provider call is then blocked before any network access, regardless of the capability switches.
  • Downgrades are safe: both capabilities only add reviewable candidates and ordering; deterministic data and the public tool contracts are untouched.