Skip to content

Model enablement acceptance decision (2026-09-23)

Decision source: the project operator's instruction in this task on 2026-09-23 to mark the two outstanding items allowed. The operator did not provide a named approver, separate legal opinion, provider contract, or human annotations. This record is a project-level authorization and an acceptance-criteria waiver; it does not claim those missing artifacts exist.

External model processing: allowed

The operator authorizes Argus to send the minimum necessary text from the sec_edgar_public_v1 / sec_edgar public filing source to the reviewed Voyage rerank provider and the reviewed small-LLM endpoint for the tasks in the enablement requirements. Scope excludes third-party filing material with separate rights, seals, logos, and artwork as listed in the source license. The source license version is SEC-public-information-verified-2026-08-23. This decision is the approval_basis for the two allowed entries in the production license placeholder. Other sources and providers remain unapproved.

The sanitized provider-call record is evals/reports/live-integration-2026-09-23.json. This decision accepts that record for the project's integration gate; it does not establish the terms of any third-party provider contract.

Evaluation method: programmatically labeled samples allowed

The operator waives the requirement for separate human annotation in requirements §10.13–14. The project accepts the reproducible synthetic samples and labels in evals/rerank_dataset.jsonl and evals/small_llm_dataset.jsonl as its evaluation set for this release. Their provenance remains programmatically_labeled_synthetic; no report may describe them as human labeled.

The actual fact-label balance is 22 positive and 8 negative samples. The original 10-negative target is short by two. The operator's instruction to mark this existing evaluation item allowed includes this explicit sample-count exception. Event labels are 20 positive / 10 negative; classification labels are 15 supports / 15 other. Future reports retain the actual counts and the original-balance result.

License-denial behavior is assessed by the automated no-egress gate tests in tests/unit/test_model_enablement.py, since prohibited material must never enter a provider evaluation call.

The recorded provider evaluations meet the numerical thresholds under this waiver: evals/reports/rerank-eval-2026-09-23.json contains 50 queries, with nDCG@10 and Recall@10 of 0.6 before and after rerank; evals/reports/small-llm-eval-2026-09-23.json contains 30 samples per task, with all six precision and recall checks passed. This is a project acceptance decision about the existing reports, not a new model run.

Production deployment and live-server observation are separate from this decision. No production deployment is asserted here.