Architecture¶
Casdoor is an external OAuth/OIDC authorization server; Argus is only an OAuth resource server. OAuth API and MCP requests use separate audiences, one configured shared institution, and stable caller IDs derived from issuer plus subject. OAuth subjects are not machine accounts. See Casdoor OAuth Resource Server.
Argus separates machine access adapters from governed business logic. CLI + Skill, MCP, and REST translate requests, authenticate the caller, and delegate to the same core services. They do not query databases, object storage, or data vendors directly.
AI agent
-> CLI + Skill | MCP | REST
-> machine authentication
-> CommandLineToolService
-> CoreServiceBoundary
1. write-ahead audit start
2. tenant isolation
3. permission check
4. core operation
5. license check
6. schema, evidence, and output-origin validation
7. terminal audit record
-> shared structured result envelope
Argus does not generate an investment judgment. Investment wording and intended independent analysis do not block otherwise permitted fact access. A client AI agent may use lawfully accessible results for unrestricted independent reasoning and downstream output. Data-access controls such as authentication, tenant isolation, permissions, licenses, restricted fields, rate limits, and audit remain mandatory.
Runtime services¶
| Service | Responsibility | Public exposure |
|---|---|---|
api |
REST, onboarding, skills, health, metrics | API hostname through TLS proxy |
mcp |
MCP tool discovery and invocation | MCP hostname/path through TLS proxy |
worker |
Governed Celery task execution | None |
beat |
Scheduled retention and vector-index maintenance | None |
| PostgreSQL + pgvector | Persistent domain data, audit, task, and vector records | None |
| Redis | Rate limiting, Celery broker/result coordination, locks | None |
| Object storage | Filing source objects and export artifacts | None |
| Static website | Documentation, search, onboarding, aggregate usage snapshot | Website hostname only; no tool execution |
Core domains¶
The src/argus/core package contains the business boundary and services for
company/security master data, filings, evidence mapping, events, corporate
actions, market data, neutral metrics, point-in-time snapshots, data quality,
licenses, trust, async tasks, permissions, and audit. Public Pydantic contracts live in src/argus/schemas; persistence
models and repositories live below src/argus/models and src/argus/db.
Data ingestion path¶
Connectors fetch raw source records under a versioned registry definition. The service validates provider, secret, license, allowed use, rate limit, parser version, and field-mapping version before synchronization. Normalizers convert vendor-specific payloads into neutral domain records and source evidence. Filing content can additionally flow through ingestion, parsing, fragmenting, fact extraction, evidence mapping, and optional vector indexing.
Consistency invariants¶
- Equivalent requests through CLI, MCP, and REST preserve fact, evidence, time, quality, license, restriction, and audit semantics.
- Public schema models reject undeclared fields.
- A fact cannot cite an evidence id absent from its package.
- Historical reads exclude facts whose
known_timeis later thanas_of. - Credentials, secrets, and raw sensitive values are redacted from audit output.
- Connector schema, credential, normalization, persistence, and audit validation remain fail-closed. License and redistribution enforcement is currently disabled by configuration and can be restored without code changes.
- The static website remains independent of the API/MCP runtime lifecycle.
See Core Boundary for the policy sequence and Deployment for concrete process and network boundaries.
P1 contract kernel and generated interfaces¶
argus.contracts contains the shared closed envelope and entity, security, ETF,
market, option, filing, and macro payload families. argus.core.tool_registry selects the
input version, output version, payload family, capability, time semantics,
provider coverage, replacement, and exactly three public bindings for every
tool. scripts/generate_interfaces.py materializes those declarations into
checked-in CLI, MCP, REST, Skill, and documentation artifacts; runtime CLI remote
routing and MCP contract metadata consume the generated authority directly.
The P1 golden contract and live adapter parity test are release gates. Generated
files must pass scripts/generate_interfaces.py --check, and all request/payload
object schemas reject unknown fields.
P3 normalized market domains¶
UsEquityEtfOptionDataService is a read-only domain boundary for US security
master/session data, bitemporal ETF snapshots, market observations and corporate
action adjustments, and option contract/chain data. Persistence is normalized by
revision 0032_us_equity_etf_options. ETF visibility requires effective,
publication, and known-time eligibility; adjustments link to actions and a
versioned policy. Provider coverage, delay, missing fields, and license status are
part of each payload. Provider-supplied greeks retain model and input metadata and
never acquire strategy or transaction semantics.
P4 point-in-time query and export¶
PitCatalogService provides the shared Calendar, Instrument, Feature, and Dataset
point-in-time abstractions. Its closed screener DSL enforces field/type/provider/
license/scan limits and stable cursor pagination. The same selected factual rows
feed deterministic Arrow IPC and Parquet partitions whose manifest records schema,
content, partition, time, source, license, adjustment, cache, and dataset identity.
Exports contain no execution engine or performance semantics.
P5 deterministic metrics and feature DAG¶
FeatureDagService is the only derived-value execution path for new interfaces.
It resolves immutable formula id/version pairs from MetricCatalog, verifies PIT
eligibility and evidence for every fact, validates an acyclic bounded graph, and
executes fixed Decimal operators in a stable topological order. Canonical requests
bind cache identity; derived records retain transitive fact, evidence, provider,
as-of, known-at, and formula-version lineage. Multi-source comparison retains all
values and pairwise differences and has no best-source selection field. Public
outputs are structured values and metadata only.