Model Capabilities Operations (Rerank and Small LLM)¶
Argus ships two optional external-model capabilities, both disabled by default and both governed by validated configuration, file-based secrets, and a per-source external-model-processing license gate:
| Capability | Provider and model | Where it runs | Kill switch |
|---|---|---|---|
| Online filing-fragment rerank | Voyage rerank-2.5-lite (POST https://api.voyageai.com/v1/rerank) |
filing_search and filing_evidence_extract request paths (api/mcp/worker) |
ARGUS_MODELS__RERANK__ENABLED |
| Offline small-LLM tasks | OpenAI-compatible meta/muse-glimmer-30b (POST https://cliproxy.668618.xyz/v1/chat/completions, schema embedded in the fixed instruction and re-validated locally) |
Celery worker only, after a filing is ingested and parsed | ARGUS_MODELS__SMALL_LLM__ENABLED |
Neither capability ever exposes model parameters to CLI/MCP/REST callers, and neither reintroduces question answering: rerank only reorders already authorized fragments, and the small LLM only produces pending-review structured candidates with source positions.
1. Prerequisites¶
- Secrets (one-line, root-owned, mode 0640 files in
$ARGUS_SECRETS_DIR_HOST): voyage_api_key- already required for vector search; also used by rerank.model_api_key- the small-LLM key; mounted only into the worker service asSMALL_LLM_API_KEY_FILE. The first-version model is not an OpenAI model; only the HTTP protocol is OpenAI-compatible.- External-model-processing license approval. In
config/data_licenses.production.yaml, add anexternal_model_processingentry to the approved SEC EDGAR license (see the commented example inconfig/data_licenses.yaml). Text without an explicitallowedentry for the provider is never sent - including while generic license enforcement is administratively disabled. Production approvals may only list the reviewedsec_edgartarget source. The operator's project-levelalloweddecision and scope are recorded in the model acceptance decision. The production placeholder grants only this source to Voyage and the reviewed small-LLM capability; excluded third-party material remains out of scope. - Static preflight (
scripts/production-preflight.sh) passes. The preflight distinguishes "capability disabled" (fine) from "misconfigured" (deployment stops): it rejects conflictingVOYAGE_RERANK_MODELvsARGUS_MODELS__RERANK__MODEL_IDvalues, non-reviewed model IDs, non-official base URLs, and any attempt to enable the small LLM outside the worker role.
2. Configuration reference¶
The runtime settings object (models.rerank, models.small_llm) is the
single authority; business code never reads these environment variables
itself.
| Variable | Default / range | Notes |
|---|---|---|
ARGUS_MODELS__RERANK__ENABLED |
false |
Enable only after preflight and an authorized live check. |
ARGUS_MODELS__RERANK__MODEL_ID |
rerank-2.5-lite |
Must be a reviewed Voyage rerank model (rerank-2.5, rerank-2.5-lite, rerank-2, rerank-2-lite, rerank-1.5, rerank-lite). The legacy VOYAGE_RERANK_MODEL is honored only when the new value is absent; conflicting values fail startup. |
ARGUS_MODELS__RERANK__BASE_URL |
https://api.voyageai.com |
Fixed; arbitrary hosts are rejected. |
ARGUS_MODELS__RERANK__TIMEOUT_SECONDS |
3 (1-5) |
Online call budget. |
ARGUS_MODELS__RERANK__CANDIDATE_LIMIT / RESULT_LIMIT |
30 / 10 |
Result limit cannot exceed the candidate limit or the public tool limit. |
ARGUS_MODELS__SMALL_LLM__ENABLED |
false (worker only) |
Enabling it on api/mcp/beat roles fails startup validation. |
ARGUS_MODELS__SMALL_LLM__MODEL_ID |
meta/muse-glimmer-30b |
Must be paired with its reviewed base URL. gpt-5.4-mini-2026-03-17 remains configurable only with https://api.openai.com/v1. |
ARGUS_MODELS__SMALL_LLM__BASE_URL |
https://cliproxy.668618.xyz/v1 |
Reviewed hosts only. The protocol (chat_completions or responses) is derived from the host and is not separately configurable. |
ARGUS_MODELS__SMALL_LLM__REASONING_EFFORT |
low |
none, minimal, low, or medium. Measured quality needs low; none misses facts in documents with distractor fragments. Streaming keeps the gateway from idle-timing out longer reasoning. |
ARGUS_MODELS__SMALL_LLM__TIMEOUT_SECONDS |
120 (10-120) |
Per-call budget. Chat completions stream so the gateway does not idle-timeout a long generation. |
ARGUS_MODELS__SMALL_LLM__MAX_INPUT_TOKENS / MAX_OUTPUT_TOKENS / MAX_BATCHES_PER_DOCUMENT |
8000 / 2000 / 50 |
Positive integers with hard upper bounds; exceeding the batch cap marks coverage incomplete instead of silently truncating. |
ARGUS_MODELS__SMALL_LLM__FILING_FACTS_ENABLED, ..._EVENTS_ENABLED, ..._EVIDENCE_CLASSIFICATION_ENABLED |
false |
Task switches are inert while the master switch is off. |
In .env.production set ARGUS_RERANK_ENABLED=true for rerank and
ARGUS_SMALL_LLM_ENABLED=true (plus the task switches) for the small LLM,
then docker compose ... up -d worker api (and mcp variants). The compose
anchor always keeps the small LLM off and keyless; only the worker service
receives SMALL_LLM_API_KEY_FILE.
3. What each capability does at runtime¶
Rerank. filing_search / filing_evidence_extract first apply tenant,
company (pushed into the repository query), document, license-status, and
as_of filters; if two or more candidates remain and every source
document approves external model processing for voyage, the authorized
fragment texts are sent once to Voyage and the visible results are reordered
by the provider scores (ties keep the keyword order). Any denial, timeout,
rate limit, or invalid provider response skips rerank entirely and returns
the original order with a machine-readable reason - the user request never
fails because of rerank. Provider scores only order candidates; they are
never stored as fact confidence.
Small LLM. After a filing is successfully ingested and parsed, the
connector dispatches the offline task
argus.tasks.model_assisted_filing_processing (Redis-locked per document,
idempotent per model snapshot + prompt version). The task skips documents
with a deterministic XBRL path, denies unapproved licenses before any
network call, batches bounded positioned fragments, and re-validates every
candidate deterministically (field enum, excerpt locatability, numeric value
presence, currency shape, period consistency, event time sanity). Accepted
candidates enter as MODEL_ASSISTED / pending review only; failures are
quality issues on the document and never re-mark the connector sync or the
ingestion as failed.
Fact and event processing each have a 50-batch document cap. Oversized
fragments or extra batches remain flagged as incomplete coverage and the
offline task reports a partial result for follow-up.
When evidence classification is enabled, the same worker builds candidates
from the parsed document, calls Voyage once per ambiguous model fact, then
uses small-LLM relation classification and deterministic value checks before
adding any secondary evidence mapping. These mappings remain pending review.
4. Observing¶
- Metrics:
argus_model_calls_total{provider,model,operation,status},argus_model_rejected_candidates_total(candidates discarded by deterministic verification). The provider/model/operation labels on the call and token series attribute provider spend to the exact capability and task for cost accounting,argus_model_call_latency_seconds,argus_model_call_degradations_total,argus_model_call_retries_total,argus_model_tokens_total, andargus_model_pending_review_candidates_totalon each role's metrics endpoint (see observability). - Structured audit events:
event="model_call"records carry model snapshot, prompt version, input SHA-256, source ids, status, degradation reason, and task correlation id - never prompt text, fragment text, or keys. - Document-level trail:
model_assisted_processing_started,model_assisted_processing_denied,model_assisted_extraction_failed,model_assisted_coverage_incomplete,model_assisted_event_processing_failed, andmodel_assisted_processing_completedquality issues on the filing document. - Live evaluation:
uv run python scripts/evaluate_model_quality.py --rerankreports nDCG@10 / Recall@10 versus the keyword baseline (gate: rerank must not be worse);--small-llmreports the joint precision / recall thresholds (facts 90%, events 90%, supports-precision 95%, positive recall 70%). Both need authorized keys (VOYAGE_API_KEY/SMALL_LLM_API_KEY).
Release status (2026-09-23, operator acceptance): The recorded provider smoke
calls and switch drill used the reviewed providers (rerank-2.5-lite,
voyage-4-lite, meta/muse-glimmer-30b over
https://cliproxy.668618.xyz/v1). Sanitized records are in
evals/reports/live-integration-2026-09-23.json and
evals/reports/switch-drill-2026-09-23.json. The 50-query rerank and
30-per-task small-LLM reports in evals/reports/ measured synthetic samples
whose labels are generated by evals/build_datasets.py. The operator's
dated decision permits these
programmatically labeled samples as this release's evaluation basis and
allows the specified SEC EDGAR source for external model processing. The
recorded numerical thresholds pass under that policy; no human annotations
are claimed. The fact set has 8 negative labels against the original target
of 10; the dated decision explicitly accepts this difference. Known limitation:
fragments that
print values with a scale word ("42 million") are rejected by the
deterministic scale check when the model expands the number - candidates
from such fragments stay pending review until a prompt revision is
re-evaluated. Production enablement on the live server still requires an
operator with server access: set ARGUS_RERANK_ENABLED=true and
ARGUS_SMALL_LLM_ENABLED=true (plus task switches) in .env.production,
install model_api_key as the small-LLM key file, and restart worker/api
per this runbook.
5. Disabling and rollback¶
- Rerank off: set
ARGUS_RERANK_ENABLED=false(or remove the rerank block) and restart api/mcp/worker.filing_search/filing_evidence_extractimmediately return the original keyword order; nothing else changes. No key removal is required. - Small LLM off: set
ARGUS_SMALL_LLM_ENABLED=falseand restart the worker. New offline model tasks stop; already-indexed structured events, XBRL facts, history, and audit remain fully readable. Pending-review model candidates simply stop accruing. - Full rollback of text egress: additionally remove (or set
prohibited) on theexternal_model_processingentries in the licenses file and restart; every provider call is then blocked before any network access, regardless of the capability switches. - Downgrades are safe: both capabilities only add reviewable candidates and ordering; deterministic data and the public tool contracts are untouched.