Model enablement acceptance decision (2026-09-23)¶
Decision source: the project operator's instruction in this task on
2026-09-23 to mark the two outstanding items allowed. The operator did not
provide a named approver, separate legal opinion, provider contract, or
human annotations. This record is a project-level authorization and an
acceptance-criteria waiver; it does not claim those missing artifacts exist.
External model processing: allowed¶
The operator authorizes Argus to send the minimum necessary text from the
sec_edgar_public_v1 / sec_edgar public filing source to the reviewed
Voyage rerank provider and the reviewed small-LLM endpoint for the tasks in
the enablement requirements. Scope excludes third-party filing material with
separate rights, seals, logos, and artwork as listed in the source license.
The source license version is
SEC-public-information-verified-2026-08-23. This decision is the
approval_basis for the two allowed entries in the production license
placeholder. Other sources and providers remain unapproved.
The sanitized provider-call record is
evals/reports/live-integration-2026-09-23.json. This decision accepts that
record for the project's integration gate; it does not establish the terms
of any third-party provider contract.
Evaluation method: programmatically labeled samples allowed¶
The operator waives the requirement for separate human annotation in
requirements §10.13–14. The project accepts the reproducible synthetic
samples and labels in evals/rerank_dataset.jsonl and
evals/small_llm_dataset.jsonl as its evaluation set for this release.
Their provenance remains programmatically_labeled_synthetic; no report
may describe them as human labeled.
The actual fact-label balance is 22 positive and 8 negative samples.
The original 10-negative target is short by two. The operator's instruction
to mark this existing evaluation item allowed includes this explicit
sample-count exception. Event labels are 20 positive / 10 negative;
classification labels are 15 supports / 15 other. Future reports retain
the actual counts and the original-balance result.
License-denial behavior is assessed by the automated no-egress gate tests in
tests/unit/test_model_enablement.py, since prohibited material must never
enter a provider evaluation call.
The recorded provider evaluations meet the numerical thresholds under this
waiver: evals/reports/rerank-eval-2026-09-23.json contains 50 queries,
with nDCG@10 and Recall@10 of 0.6 before and after rerank;
evals/reports/small-llm-eval-2026-09-23.json contains 30 samples per
task, with all six precision and recall checks passed. This is a project
acceptance decision about the existing reports, not a new model run.
Production deployment and live-server observation are separate from this decision. No production deployment is asserted here.