Deployment¶
Production deploys mcp-oauth on /mcp and mcp-machine on
/mcp-machine as separate processes, ports, health checks, and private metrics
ports. Production is fixed at ARGUS_MCP_ROLLOUT_PHASE=oauth-cutover: public
/mcp is OAuth-only and machine clients use /mcp-machine. Required Casdoor
environment values are documented in .env.production.example.
The project has two Docker Compose entrypoints:
docker-compose.ymlis for local development only.docker-compose.prod.ymlis the production-oriented Compose template for 1Panel or direct Docker Compose on a Debian server.
Services¶
| Service | Role |
|---|---|
api |
FastAPI REST, readiness, and metrics endpoints. |
mcp |
Transitional machine-authenticated /mcp service; production preflight does not permit this phase. |
mcp-oauth |
OAuth-only /mcp resource server enabled during oauth-cutover. |
mcp-machine |
Permanent legacy-compatible machine endpoint on /mcp-machine. |
worker |
Celery worker for governed asynchronous tasks. |
beat |
Celery Beat scheduler. |
postgres |
PostgreSQL 16 with pgvector image. |
redis |
Redis broker and cache. |
argus_object_storage_data |
Docker volume used for filesystem-backed raw filing storage. |
Local Commands¶
The local Compose file binds all published ports to 127.0.0.1 and explicitly
sets ARGUS_ENABLE_LOCAL_DEMO_CREDENTIALS=true. When running the CLI or API
outside Compose, set that variable only when the fixed demo credentials are
intentionally required.
1Panel / Debian Production Compose¶
Use the production template with an environment file created from .env.production.example.
Edit .env.production on the server or define the same variables in 1Panel. The required production values are:
| Variable | Required value |
|---|---|
ARGUS_SECRETS_DIR_HOST |
Absolute host directory containing the required root-owned, mode-0640 one-line secret files. |
ARGUS_CASDOOR_PREFLIGHT_SECRETS_DIR_HOST |
Separate root-owned directory containing only the Casdoor release-auditor Client ID and secret; it is mounted only into the one-off preflight container. |
ARGUS_SECRETS_GID |
Numeric GID of the dedicated host secret-reader group; defaults to the documented 10002. |
ARGUS_OBJECT_STORAGE__BUCKET_NAME |
Logical bucket directory for filesystem object storage, for example argus-prod. |
ARGUS_OBJECT_STORAGE__SECURE_TRANSPORT |
Must be true when non-local deployments use the S3 backend; the endpoint must use HTTPS. |
ARGUS_API_SECURITY__TRUSTED_HOSTS |
JSON array of production hostnames: ["api.argusfa.com"]. |
ARGUS_API_SECURITY__PUBLIC_BASE_URL |
Public HTTPS API origin: https://api.argusfa.com; onboarding pages use this fixed value rather than a proxy header. |
ARGUS_API_SECURITY__CORS_ALLOWED_ORIGINS |
JSON array of allowed origins, or [] when no browser origin is allowed. |
ARGUS_AUDIT__RETENTION_DAYS |
Number of days to retain complete audit chains; production Compose forwards it to every application service and Celery beat prunes expired chains daily. |
ARGUS_DATA_LICENSE__ENFORCEMENT_ENABLED |
Keep false while administrative license/redistribution enforcement is disabled. Set true only to restore the retained evaluation path. |
ARGUS_CONNECTORS__FAIL_CLOSED_ON_LICENSE_ERROR |
Keep false while enforcement is disabled so license metadata cannot block connector ingestion or Agent calls. |
ARGUS_DATA_LICENSES_FILE_HOST |
Host path to a license YAML. The policy metadata remains mounted for auditability and for a reversible future re-enable; it is not enforced while the switch is off. |
ARGUS_PRIMARY_MACHINE_CREDENTIAL_READY |
Keep NO until an institution:argus-primary credential is stored and deployed to every production caller; set exactly YES before application deployment. |
ARGUS_API_BIND_ADDRESS |
Keep 127.0.0.1; the TLS reverse proxy is the only public entrypoint. |
ARGUS_API_PORT |
Loopback host port for the API service, if not using the default 18000. |
ARGUS_MCP_BIND_ADDRESS |
Keep 127.0.0.1; the TLS reverse proxy is the only public entrypoint. |
ARGUS_MCP_PORT |
Loopback host port for the MCP service, if not using the default 18001. |
ARGUS_MCP_PUBLIC_URL |
Public streamable HTTP MCP URL: https://mcp.argusfa.com/mcp. |
ARGUS_MCP_METRICS_BIND_ADDRESS |
Use 127.0.0.1 for host-local monitoring or the server's RFC1918/100.64.0.0/10 private/VPN IPv4 address for external Prometheus. Public addresses and 0.0.0.0 are rejected. |
ARGUS_MCP_METRICS_PORT |
Authenticated private MCP metrics port, default 18002; permit it only from the monitoring host. |
ARGUS_MCP_MACHINE_BIND_ADDRESS |
Keep 127.0.0.1; the TLS reverse proxy is the only public entrypoint for /mcp-machine. |
ARGUS_MCP_MACHINE_PORT |
Loopback host port for mcp-machine, default 18003. |
ARGUS_MCP_MACHINE_METRICS_BIND_ADDRESS |
Use the same private-only binding policy as the OAuth MCP metrics listener. |
ARGUS_MCP_MACHINE_METRICS_PORT |
Authenticated private metrics port for mcp-machine, default 18004. |
The mcp and mcp-oauth services publish the same ports on purpose — exactly
one authentication variant owns the public MCP port at a time. They carry
mutually exclusive Compose profiles (mcp-machine, mcp-oauth), so a bare
docker compose up -d starts neither; scripts/deploy-production.sh starts the
variant selected by ARGUS_MCP_ROLLOUT_PHASE by naming it explicitly.
Production secrets are not interpolated into Compose environment values or Redis
process arguments. Create /etc/argus/secrets owned by root:ARGUS_SECRETS_GID
with mode 0750 or 0550 and the following one-line files owned by the same
root:ARGUS_SECRETS_GID with mode 0640 or 0440: postgres_password, redis_password,
api_key_signing_secret, jwt_signing_secret, metrics_api_key, stream_cursor_secret, model_api_key,
voyage_api_key, fred_api_key, openfigi_api_key,
bea_api_key, bls_api_key, eia_api_key, and eodhd_api_key. Configure the
non-secret SEC_USER_AGENT as Novai Limited [email protected]. Do not place Casdoor management
credentials in this runtime directory because the entire directory is available
to every Argus application container. The application resolves its runtime *_FILE pointers before
configuration validation, including BEA_API_KEY, BLS_API_KEY, EIA_API_KEY, and EODHD_API_KEY;
PostgreSQL uses POSTGRES_PASSWORD_FILE; Redis generates
a mode-0600 tmpfs configuration so its password is absent from docker inspect
arguments. Passwords may contain URI-reserved characters because application
connection URLs are constructed with percent encoding.
EODHD EOD, Intraday, Financial News, US quote WebSocket, and US trade WebSocket
sources are enabled. Production Compose mounts eodhd_api_key through
EODHD_API_KEY_FILE; the runtime materializer loads it without placing the key in
Compose values, repository files, or logs.
Create ARGUS_CASDOOR_PREFLIGHT_SECRETS_DIR_HOST separately with the same
root ownership, group, directory mode, and file-mode rules. Its resolved path
must be disjoint from, and not nested above or below, the runtime secrets
directory. It must contain only
casdoor_management_client_id and casdoor_management_client_secret for a
dedicated release-auditor application; do not reuse the CLI application secret
or an administrator password. The guarded preflight adds this directory as a
temporary read-only mount to its one-off container. It is absent from API, MCP,
worker, beat, and migration containers.
The Compose services run as non-root users and receive only this numeric supplemental group, so they can read the files without exposing values through container environment inspection. Create the group with the same fixed GID and create each file without placing the value in shell history:
sudo groupadd --system --gid 10002 argus-secrets
sudo install -d -o root -g argus-secrets -m 0750 /etc/argus/secrets
sudo sh -c 'read -r value; printf "%s\n" "$value" > /etc/argus/secrets/postgres_password'
sudo chown root:argus-secrets /etc/argus/secrets/postgres_password
sudo chmod 0640 /etc/argus/secrets/postgres_password
If GID 10002 is already allocated, select another unused numeric GID and update
ARGUS_SECRETS_GID. Repeat the guarded prompt and ownership/mode commands for every
required filename. Do not store these values in
.env.production; that file now contains only paths and non-secret configuration.
Install Cosign from its official release package on the server. After this
repository's Secure Container Release workflow prints the immutable
ARGUS_IMAGE, run the server preflight. It rejects unsafe environment-file
permissions, mutable or foreign images, an invalid GitHub Actions keyless
signature, wrong public domains/bind addresses, missing provider secrets,
disabled vector evidence retrieval, an unexpected license mode, and an invalid Compose
model:
The preflight accepts only ghcr.io/ksahdsambn/argus@sha256:..., verifies that
its Cosign certificate was issued to this repository's release.yml workflow,
requires the secrets directory to be root-owned mode 0750 or 0550, requires
every secret file to be root-owned mode 0640 or 0440, validates the private
MCP metrics binding, reports free space for secrets and Docker data, reports the
backup directory when /etc/argus/backup.env is installed, then pulls the immutable API image and loads AppSettings inside a one-off
container. The image is never executed with production secrets before this
identity check succeeds. A successful result therefore proves the release
identity, Compose interpolation, and application startup validation.
The production Compose template currently keeps generic administrative license
evaluation disabled and mounts config/data_licenses.production.placeholder.yaml
to retain source metadata. Placeholder policies are configuration examples;
they do not prove a provider subscription or production coverage. The separate
external-model processing gate and the record-level checks for market, ETF, and
option data remain active. Do not change the administrative mode to make an
acceptance request pass; verify the actual source rights and coverage separately.
Vector evidence retrieval is enabled in production. Voyage produces and persists
compatible filing-fragment vectors and reranks evidence; Argus returns only the selected
source fragments and their structured evidence identifiers. Celery beat schedules
the bounded filing_vector_index_sweep task every 15 minutes, and Redis prevents overlapping
sweeps. Interactive REST/MCP requests never build indexes; they fall back to keyword
retrieval when no compatible background index is available. ARGUS_VECTOR_INDEX_BATCH_DOCUMENTS
controls the number of documents attempted per sweep and defaults to 5.
Two additional optional model capabilities - online Voyage rerank for
filing_search / filing_evidence_extract and the small-LLM offline
tasks - stay disabled by default. voyage_api_key is shared with rerank;
model_api_key is mounted only into the worker as
SMALL_LLM_API_KEY_FILE. The reviewed small model speaks the OpenAI-compatible
chat completions protocol; it is not an OpenAI model. Both capabilities additionally require an explicit
external_model_processing approval on the target SEC EDGAR license, and
production preflight rejects conflicting or unreviewed model configuration.
The full enable/observe/disable/rollback runbook lives in
Model Capabilities Operations.
Build and publish the release image only with this repository's Secure
Container Release GitHub Actions workflow. The production Compose file
intentionally has no build section: it only accepts the workflow's published,
signed immutable image digest. The workflow output starts with
ARGUS_IMAGE=ghcr.io/ksahdsambn/argus@sha256:; copy the complete output line,
including its 64-character digest.
The release workflow runs only for version tags and cannot be dispatched from a
branch. A release tag must point exactly at the current protected main commit;
protect main, protect semantic version tags, and configure the
production-release GitHub Environment with required reviewers. Before the image
is built or signed, the workflow calls the quality gate from protected main with
only the five connector-smoke secrets it needs, including real connector smoke
tests and the production Compose startup and restore rehearsal. Production
signature verification accepts only release.yml identities rooted at semantic
v* tags.
Paste the workflow's complete value into ARGUS_IMAGE on the server. Do not
substitute another registry and do not deploy a tag, including a
semantic-version tag. The preflight verifies and pulls that exact digest before
the release; an optional explicit pull afterward is safe:
Before the first deployment or an upgrade from credentials issued as
institution:prod, complete the credential migration below. Institution ids are matched
exactly: an existing institution:prod key is not authorized by the production policy and
must not be carried into the new release.
First run preflight, prepare the database schema, and issue the replacement credential without starting the application services:
scripts/production-preflight.sh
docker compose --env-file .env.production -f docker-compose.prod.yml up -d --wait postgres redis
docker compose --env-file .env.production -f docker-compose.prod.yml --profile maintenance run --rm migrate
docker compose --env-file .env.production -f docker-compose.prod.yml run --rm --no-deps api \
python -m argus.cli machine register-subject \
--institution-id institution:argus-primary \
--caller-id service:primary-agent \
--subject-type service_account \
--scopes facts:read,evidence:read,audit:read,licenses:read,connectors:read,connectors:sync \
--tool-whitelist '*'
Store the printed plaintext_api_key in the secret manager; it is shown only once and only
its digest is persisted. Update every production caller to use the replacement key and
verify the stored institution id is institution:argus-primary. Only then set
ARGUS_PRIMARY_MACHINE_CREDENTIAL_READY=YES in .env.production.
Deploy or upgrade through the guarded deployment entrypoint. It refuses to start application services until that exact gate is set, verifies the signed release, starts PostgreSQL and Redis, runs migrations, proves the database Alembic revision exactly matches the image head, then starts API, MCP, worker, and beat with Compose health waiting enabled:
Production /health/ready is an operational gate, not only a process-startup
check. It remains 503 until the database contains persisted business data and
there is a successful, non-dry-run connector synchronization in the previous
seven days that fetched records from a source with persisted raw records.
Unchanged records may be deduplicated on a later run and still count toward this
check. /health/live and container health checks
remain available during an intentional first-deployment backfill. Complete the
initial connector synchronization, poll /v1/connector-tasks/{task_id}, and
require /health/ready to return 200 before exposing the release to callers.
After the new release accepts the replacement credential, revoke the superseded
institution:prod credential. Credentials that need Connector operations must explicitly include connectors:read and
connectors:sync; read-only credentials should omit these scopes.
Do not replace this command with a direct docker compose up: dependency reachability
alone does not prove that the database schema is compatible with the release image.
Before promoting the deployment, complete the Production Release Checklist. It is the production gate for live connector smoke, internal distribution scope, Redis-backed rate limiting, migrations, health checks, credential registration, observability, and rollback evidence.
Public Endpoints¶
Put a TLS-terminating reverse proxy in front of the two public machine services.
Compose binds them to loopback by default; do not change either bind address to
0.0.0.0. The proxy must terminate TLS, forward only the configured hosts, and
be the only host firewall exception for API and MCP traffic:
| Public surface | Container service | Default host port | Typical public URL |
|---|---|---|---|
| REST API | api |
127.0.0.1:18000 |
https://api.argusfa.com |
| OAuth MCP streamable HTTP | mcp-oauth |
127.0.0.1:18001 |
https://mcp.argusfa.com/mcp |
| Machine MCP streamable HTTP | mcp-machine |
127.0.0.1:18003 |
https://mcp.argusfa.com/mcp-machine |
The API service exposes REST under /v1, health checks under
/health/live and /health/ready, and authenticated Prometheus metrics under
/metrics. The API service also exposes public onboarding pages at /cli,
/skill, and /mcp, plus one-command CLI install scripts at /install and
/windows/install. mcp-oauth exposes the OAuth-only streamable HTTP endpoint
at /mcp; mcp-machine exposes the machine-client endpoint at /mcp-machine.
Do not expose PostgreSQL, Redis, or the object-storage volume to the public internet. Keep /metrics
behind monitoring credentials by sending X-Argus-Metrics-Key or
Authorization: Bearer <metrics-token>.
Nginx edge controls¶
The repository provides the production Nginx template at
deploy/nginx/argus.conf.template.
It terminates TLS, applies per-IP request and connection limits, returns 429
when the limit is exceeded, caps requests at 1 MiB, sets client and upstream
timeouts, and disables buffering for streamable MCP responses. Render it with
only the listed variables so Nginx variables are preserved:
export ARGUS_API_HOSTNAME=api.argusfa.com
export ARGUS_MCP_HOSTNAME=mcp.argusfa.com
export ARGUS_API_PORT="${ARGUS_API_PORT:-18000}"
export ARGUS_MCP_PORT="${ARGUS_MCP_PORT:-18001}"
export ARGUS_MCP_MACHINE_PORT="${ARGUS_MCP_MACHINE_PORT:-18003}"
export ARGUS_TLS_CERTIFICATE=/etc/letsencrypt/live/argusfa.com/fullchain.pem
export ARGUS_TLS_CERTIFICATE_KEY=/etc/letsencrypt/live/argusfa.com/privkey.pem
envsubst '$ARGUS_API_HOSTNAME $ARGUS_MCP_HOSTNAME $ARGUS_API_PORT $ARGUS_MCP_PORT $ARGUS_MCP_MACHINE_PORT $ARGUS_TLS_CERTIFICATE $ARGUS_TLS_CERTIFICATE_KEY' \
< deploy/nginx/argus.conf.template > /etc/nginx/conf.d/argus.conf
nginx -t && systemctl reload nginx
Keep only ports 443 (and 80 solely for ACME redirects) open in the host firewall. Tune the displayed limits from measured traffic; do not remove the entry-layer limits because invalid credentials are intentionally audited by the application.
Production monitoring and single-host boundary¶
Docker Compose restarts individual processes but cannot survive loss of the only server, Docker daemon, local PostgreSQL/Redis volume, or local object-storage volume. This topology is documented here but is no longer represented by acknowledgement or minimum-capacity variables, and the production preflight does not review host sizing. For a host-level availability SLA, move PostgreSQL, Redis, and object storage to redundant managed services and run multiple stateless API/MCP/worker replicas on separate hosts or K3s nodes; that is a different deployment topology from this single-server Compose contract.
The repository now ships deploy/prometheus/argus.rules.yml plus reference
Prometheus and blackbox-exporter configurations in deploy/prometheus/. Load the
rules into a Prometheus instance outside this failure domain when possible. The
rules cover public API/MCP availability, API 5xx rate and p95 latency, Celery queue
backlog, internal queue/storage probe failure, object and host disk capacity,
container restart loops, backup age, certificate expiry, audit integrity, and
missing monitoring signals.
The reference Prometheus configuration expects:
- authenticated scraping of
https://api.argusfa.com/metricsusing a credential file containing the same value as the productionmetrics_api_keysecret; - authenticated scraping of
argus-production.internal:18002, where that name resolves to the production server's private/VPN address and the firewall permits only the Prometheus host; - blackbox-exporter for API HTTP and MCP TLS probes;
- node_exporter with the textfile collector enabled; and
- cAdvisor for container restart metrics.
Install the rule and configuration files in the corresponding Prometheus/
blackbox-exporter locations, validate them with promtool check rules and
promtool check config, and configure Alertmanager receivers and escalation routes.
Keep the monitoring credential file mode 0600. After configuring node_exporter,
prepare the backup metric directory before enabling the backup service:
The scheduled backup atomically publishes argus_backup_last_success_unixtime
there only after a successful local backup. Fire a controlled test for each alert
class and confirm Alertmanager delivers it to an attended receiver; merely loading
the rule file is not release acceptance.
For the current single-host deployment, docker-compose.monitoring.yml provides a
separate, hardened monitoring lifecycle without publishing a Prometheus host port.
Prometheus joins the existing production Docker network to scrape mcp:8002
directly and uses the same file-backed credential for the trusted-host-compatible
https://api.argusfa.com/metrics scrape. Its TSDB uses the named
argus-prometheus-data volume with 35-day retention. A separate TLS gateway joins
the private Prometheus backend network and the internal
argus-monitoring-query network; it permits only /api/v1/query, requires a
different Bearer credential, and returns 401/403 for missing/incorrect credentials.
Copy .env.monitoring.example to .env.monitoring, set mode 0600, then generate
the internal CA, gateway certificate, and query token and deploy:
sudo ARGUS_MONITORING_SECRETS_GID=10002 \
bash scripts/generate-monitoring-secrets.sh
bash scripts/deploy-monitoring-compose.sh
The monitoring preflight verifies secret ownership/modes, TLS hostname validation,
credential separation, the production network, and the absence of published
ports. Deployment records protected container identities before and after startup
and fails if API, MCP, worker, beat, PostgreSQL, or Redis changed. Run
scripts/smoke-monitoring-compose.sh after any monitoring change. Keep the CA
private-key file offline-readable only by root; the containers receive only the CA
certificate, gateway certificate/key, and their narrowly required tokens.
Backup, restore, and rehearsal¶
Backups contain the PostgreSQL custom-format dump and the filesystem object
storage archive, plus SHA-256 checksums. Set ARGUS_BACKUP_INCLUDE_REDIS=true
only when the Celery queue must be recoverable; the default is false. PostgreSQL's
dump is transactionally consistent, and filesystem object writes are fsync'd then
atomically published before their database metadata. The backup therefore runs online without
stopping API, MCP, worker, or beat:
Copy the resulting directory off the host using encrypted storage and retain it according to the institution's RPO/RTO policy. Restore is intentionally guarded by an exact confirmation string and must first be rehearsed in an isolated Compose project with different volumes and loopback ports:
export ARGUS_COMPOSE_PROJECT=argus-restore-rehearsal
export ARGUS_COMPOSE_ENV_FILE=.env.restore-rehearsal
export ARGUS_RESTORE_DIR=/srv/argus-backups/argus-YYYYMMDDTHHMMSSZ
export ARGUS_RESTORE_CONFIRM=RESTORE_ARGUS_PRODUCTION
scripts/restore-production.sh
For daily unattended backups, install the supplied systemd service and timer. The wrapper rejects a group/world-readable environment file, runs the online backup, enforces local retention, and optionally sends encrypted incremental copies to a restic repository:
getent group argus-deploy >/dev/null || sudo groupadd --system argus-deploy
id argus-deploy >/dev/null 2>&1 || sudo useradd --system --gid argus-deploy \
--home-dir /srv/argus --shell /usr/sbin/nologin argus-deploy
sudo usermod -aG docker argus-deploy
sudo install -o root -g root -m 0644 deploy/systemd/argus-backup.service /etc/systemd/system/
sudo install -o root -g root -m 0644 deploy/systemd/argus-backup.timer /etc/systemd/system/
sudo install -d -o root -g argus-deploy -m 0750 /etc/argus
sudo install -o root -g argus-deploy -m 0640 deploy/systemd/backup.env.example /etc/argus/backup.env
sudo systemctl daemon-reload
sudo systemctl enable --now argus-backup.timer
systemctl list-timers argus-backup.timer
The service uses systemd StateDirectory=argus-backups, which creates the writable
/var/lib/argus-backups directory before the sandboxed process starts. The deployment
account must retain Docker socket access after group membership changes; start a new login
session before manually testing the unit.
Set ARGUS_RESTIC_ENABLED=true, RESTIC_REPOSITORY, and RESTIC_PASSWORD_FILE
in /etc/argus/backup.env for encrypted off-host replication. Keep the password
file outside the repository with mode 0600. CI performs an isolated restore
rehearsal on every release candidate; production operators must also record a
server-side rehearsal against separate Compose volumes before promotion.
The restore script fails closed unless every existing api, mcp, worker, or
beat container in the selected Compose project is explicitly exited; this
also rejects paused, restarting, created, dead, or removing containers. It validates the object archive before
destructive work and restores PostgreSQL with --exit-on-error in one transaction.
For a real maintenance restore, stop all four application services first and leave
PostgreSQL running; a failed restore leaves the application stopped for operator
inspection rather than automatically exposing partial data.
After every rehearsal, run migrations, start the isolated API, confirm
/health/ready, and verify a known object-storage file is present. Never run a
restore against the live project while application services are accepting writes.
Staging validation with real credentials¶
Run the GitHub Actions Staging validation workflow manually after deploying
the immutable release image to staging. Protect its staging Environment and
configure these environment secrets: ARGUS_STAGING_BASE_URL,
ARGUS_STAGING_API_KEY, ARGUS_STAGING_METRICS_API_KEY,
ARGUS_STAGING_INSTITUTION_ID, ARGUS_STAGING_CALLER_ID,
ARGUS_STAGING_MCP_URL, ARGUS_STAGING_FILING_DOCUMENT_ID,
SEC_USER_AGENT, FRED_API_KEY, OPENFIGI_API_KEY,
BEA_API_KEY, BLS_API_KEY, and EIA_API_KEY.
The workflow fails closed if any is absent, runs the real-provider smoke tests,
then verifies readiness, authenticated metrics, skill delivery, and an
authenticated governed staging request without printing a secret. It also requires
a known parsed and vector-indexed staging filing and proves that structured evidence
retrieval preserves fact-to-evidence identifiers, then performs an
MCP initialize, tool listing, and authenticated tool call through the public MCP URL.
Its live-provider step sets CONNECTOR_LIVE_TESTS_ENABLED=true.
Customer-Side CLI And Skill Setup¶
Customers install the Argus CLI and the appropriate agent skill in their own agent environment. The CLI and skill are client integrations; they do not open a remote shell on the Argus server.
Recommended interactive sign-in flow:
Install the customer-side CLI from the API onboarding endpoint:
Windows PowerShell:
Fetch skill packages from the public API and verify the returned checksum before installing:
curl https://api.argusfa.com/v1/skills/codex
curl https://api.argusfa.com/v1/skills/claude-code
curl https://api.argusfa.com/v1/skills/multi-agent-workflow
An OAuth credential automatically selects the public API. Machine clients can
set ARGUS_BASE_URL explicitly. Service operators can pass --local for
maintenance commands that must execute against the local deployment runtime.
Publish these onboarding URLs for customers:
Production leaves all ARGUS_CLIENT_DISTRIBUTION__* values empty and serves the
CLI wheel built into the signed API image. The production preflight rejects an
external wheel URL so the installed CLI's compiled OAuth client metadata cannot
drift from the API image being deployed.
https://api.argusfa.com/clihttps://api.argusfa.com/skillhttps://api.argusfa.com/mcp
Check service health:
docker compose --env-file .env.production -f docker-compose.prod.yml ps
docker compose --env-file .env.production -f docker-compose.prod.yml logs api --tail=100
The migration command should be run on the production Debian 13 server when it targets the production PostgreSQL service. For safer releases, run the same command first in a staging stack or against a restored production backup, then run it once against production during the deployment window.
K3s¶
Static website operations¶
The website is an independent Compose stack in docker-compose.website.yml; it does not restart or share a lifecycle with API, MCP, workers, databases, or the application production Compose stack. Copy .env.website.example to .env.website, set mode 0600, and configure ARGUS_WEB_PORT, the immutable local release SHA values, the internal https://prometheus-query-gateway:9443 query origin, ARGUS_USAGE_PROMETHEUS_TOKEN_FILE_HOST, ARGUS_USAGE_PROMETHEUS_CA_FILE_HOST, and ARGUS_USAGE_SECRETS_GID. The query token must be root-owned, use root:ARGUS_USAGE_SECRETS_GID, use mode 0640 or 0440, and contain one non-placeholder line. The CA must be a root-owned regular file with mode 0644 or 0444. The deployment entrypoint verifies these files and the internal monitoring-query network before building.
The website container exposes HTTP on ${ARGUS_WEB_BIND_ADDRESS}:${ARGUS_WEB_PORT} (default 127.0.0.1:18080). The same-host reverse proxy owns www.argusfa.com, TLS certificates, and domain routing; the website preflight rejects any non-loopback bind and the Compose stack must never bind 80 or 443 or manage certificates. The container Nginx serves /docs/, gives HTML and discovery files short cache lifetimes, immutable Astro/Pagefind assets long cache lifetimes, serves /data/usage.json from a persistent volume, and returns 404 for unknown routes.
Deploy from a clean, checked-out production repository with bash scripts/deploy-website-compose.sh. It refuses staged, unstaged, or untracked source files and requires any supplied release SHA to equal HEAD, so local image tags remain immutable release identities. It validates Compose inputs, builds both argus-website:<git-sha> and argus-website-usage:<git-sha> from the lockfiles, activates and validates Usage before replacing the website, then runs HTTP, schema, and checksum smoke checks including /data/usage.json. A running pre-state release is adopted as a rollback baseline when both matching images exist. Current/previous SHA state is recorded for both images and at least two rollback-safe releases are retained. If activation or smoke fails, both previous images are restored and rechecked; a first-stage failure with no baseline removes only the failed Usage container and leaves the website untouched. Rehearse bash scripts/rollback-website-compose.sh to switch website and Usage publisher together; a failed rollback restores the current pair and the persistent Usage volume remains untouched.
When upgrading from the legacy state format that recorded only ARGUS_WEB_CURRENT_SHA, leave the old ARGUS_WEB_USAGE_IMAGE=argus-website-usage:local value in .env.website for the first deployment, or set ARGUS_WEB_LEGACY_USAGE_IMAGE explicitly. Before replacing any container, the deploy script verifies the previous website image and retags that actual legacy Usage image as argus-website-usage:<previous-sha>. If either rollback image cannot be proven locally, deployment fails without changing the running release.
The website-usage container is the replacement for the host systemd timer. It immediately publishes then repeats hourly, mounts only the query token, monitoring CA, and writeable Usage volume, and preserves the last valid privacy-filtered snapshot when upstream queries fail. A new TSDB begins with an honest warm-up window: period_start records the actual available-history boundary until 30 days have accumulated, while the existing v1 schema remains readable by the previous release; empty counters publish as zero without inventing earlier history. Its health check rejects a missing, invalid, or checksum-mismatched snapshot; an upstream outage remains healthy only when a previously valid snapshot can be retained and marked stale according to policy. Run bash scripts/smoke-website-compose.sh after local deployment and scripts/smoke-website-public.sh after the external proxy/TLS route is live. To enable IndexNow, set ARGUS_INDEXNOW_KEY_FILE to its public verification-key file; the deploy script validates it, includes <key>.txt in the website image, retries pending URL notifications, and submits canonical HTML changes. Notification failure records only retry state and never rolls back a verified website.
K3s is a valid later deployment target, but it should be generated as a separate manifest set from the same runtime contract: one application image, PostgreSQL with pgvector, Redis, object storage, secrets from Kubernetes Secrets, and a one-shot migration Job before rolling out API, MCP, worker, and beat. Do not convert the local Compose file directly into cluster manifests without preserving the production environment checks.
Configuration¶
Runtime configuration is loaded by argus.config.settings from YAML and ARGUS_ environment variables. Sensitive values must stay in environment variables or secret managers for deployed environments.
The production Compose template uses file-backed *_FILE pointers; direct secret
environment values remain supported for local/test compatibility but are not the
production deployment contract.
Non-local environments require metrics authentication, a metrics token,
non-wildcard trusted hosts and CORS origins, HTTPS MCP, and non-default signing
secrets. The production Compose contract currently keeps license enforcement
disabled through an explicit reversible switch; license metadata remains in the
data contract and audit trail but does not block Agent calls. Set
ARGUS_DATA_LICENSE__ENFORCEMENT_ENABLED=true and
ARGUS_CONNECTORS__FAIL_CLOSED_ON_LICENSE_ERROR=true to restore the previous
institution-scoped fail-closed behavior.
Audit logging and sensitive-value redaction are mandatory invariants. They cannot be disabled by configuration; ARGUS_AUDIT__RETENTION_DAYS controls the daily complete-chain retention job.
ARGUS_MCP_PUBLIC_URL must be a structurally valid URL with an explicit host. Production and staging require the https scheme; credentials, query strings, fragments, whitespace, and invalid ports are rejected during startup validation.
Historical 15.2 completion notes in the engineering log describe contract and local/demo acceptance. Production readiness is declared only after the current production checklist, live data smoke, internal distribution controls, persisted data path, and Compose startup gates pass for the target release.
The production Compose template runs application containers as the non-root argus
image user and applies a read-only root filesystem, dropped Linux capabilities,
no-new-privileges, init: true, tmpfs /tmp, bounded CPU/memory/PIDs, graceful
shutdown windows, and JSON log rotation. Worker health requires a Celery ping;
Beat health requires a current scheduler-loop heartbeat. API health performs a real
loopback HTTP liveness request. Legacy MCP and permanent machine MCP health each
perform an initialize and tool-list protocol round trip against their independent
paths, while OAuth MCP health verifies Protected Resource Metadata. Application
health gates also verify required backing services and schema state where
applicable.
Celery uses persistent messages, late acknowledgement, worker-loss rejection and Redis visibility redelivery. A worker process that exits before completing a task does not acknowledge that message; Redis makes it visible again and another worker can resume the idempotent persisted task. Publish retries and connection-loss cancellation are enabled. Task units are capped at 30 minutes and Redis visibility is 32 minutes, so loss of an entire worker/container delays redelivery by at most that bounded visibility window rather than a day.
The Secure Container Release workflow pushes the image to GHCR, blocks on
high/critical repository and image findings, creates an SPDX JSON SBOM, signs the
immutable digest with keyless Cosign, verifies the digest-bound BuildKit provenance
manifest, and publishes a signed SPDX SBOM attestation through the OCI registry. The
OCI-native attestation path also works for private repositories owned by personal
GitHub accounts, where GitHub's repository attestation API is unavailable.
Copy only the printed digest reference into ARGUS_IMAGE.
Boundary¶
The deployment does not include a web application, dashboard, chat interface, or account self-service flow. Configuration and operation remain machine-oriented through files, commands, services, and institution-owned deployment processes.