Skip to content

PIT Screener and Dataset Catalog

P4 adds three read-only data-layer capabilities: universe_screen, dataset_catalog, and point_in_time_dataset_export. They share the same point-in-time row selection and governance rules across CLI, MCP, and REST.

Declarative screening

The screener accepts only a closed predicate structure: field, operator, and typed value. Fields come from an allowlist; operators are eq, ne, lt, lte, gt, gte, and in. Unknown fields, mismatched values, missing field scopes, disallowed providers or licenses, and scans above 10,000 candidate rows fail before results are returned.

Date predicates use ISO dates (YYYY-MM-DD). Datetime predicates must include a UTC offset, for example 2024-06-28T20:00:00Z. Numeric predicates require finite numbers. Invalid predicates are rejected even when no source rows exist. Reuse the returned cursor verbatim; malformed cursors return invalid_argument.

Sorting is deterministic. security_id and session_date are implicit final tie-breakers, and cursors are bound to the canonical query hash. A cursor cannot be reused with changed filters. Results contain canonical conditions, records, reproducible fields, counts, and cost units. They contain no score, recommendation, selection judgment, or ranking explanation.

argus universe-screen \
  --api-key "$ARGUS_MACHINE_API_KEY" \
  --as-of 2024-07-01T00:00:00Z \
  --filters-json '[{"field":"market_cap","operator":"gte","value":800000000}]' \
  --select security_id,ticker,market_cap

Point-in-time abstractions

  • Calendar records bind a session and calendar version to effective_at and known_at.
  • Instrument records use permanent security_id, effective security lifetime, ticker/MIC mapping, listing status, provider, and license.
  • Feature records are append-only revisions keyed by feature/security/session; the visible revision is the last one with effective_at <= as_of and known_at <= as_of.
  • Prepared row selection applies the allowed-provider set to both instruments and features before choosing revisions. A returned row's known_at is the later of the identity mapping and feature capture times. Conflicting values at the latest visible revision and known time require source reconciliation; input order does not select a winner. A later visible correction supersedes an older conflict. These rules also apply to export rows and do not establish a production historical universe or source coverage.
  • Dataset definitions publish primary keys, partition fields, field types, providers, supported formats, and PIT/revision guarantees.

Historical universes are reconstructed from the relevant lifetime and mapping; they are not joined to today's active-security list. The public bindings include delisted securities by default. To exclude them, add an allowlisted predicate {"field":"listing_status","operator":"ne","value":"delisted"}. The internal UniverseDefinition.include_delisted setting is not a public request parameter.

Production coverage

The prepared point_in_time_snapshot reader (point-in-time-db-v2) returns persisted filing financial facts, preserving each fact's period, unit, currency, dimensions, exact numeric representation and source evidence. It selects the latest visible revision of each financial aspect. A document must have been published and captured by as_of; the fact's known, created and revision times must also be at or before that cutoff. Capturing an old filing today does not make its facts visible in a query for yesterday.

snapshot.known_at is metadata containing the requested cutoff; snapshot.record_count counts the selected financial facts, excluding the metadata itself. Both metadata facts reference all contributing source evidence and license scopes. Missing actual facts produce an empty result without fabricated metadata. Historical captures without recorded license terms remain partial and unverified. Missing any supported financial field (revenue, net_income, operating_cash_flow, total_assets, total_liabilities) is reported as partial with a financial_field:<name> coverage gap; missing observations are never replaced by derived or invented supplier values. This prepared reader requires a new deployed release and actual persisted coverage; it does not establish complete company, field or historical coverage for the currently deployed 1.1.7 release.

The built-in PIT repository is a local contract fixture. The current production bindings do not have a persisted PIT dataset reader. universe_screen and point_in_time_dataset_export therefore return unavailable, with fixture_data_disabled and production_source_not_configured metadata, until a production source is integrated. dataset_catalog describes the supported schema and export formats; its presence does not prove data availability.

Reproducible exports

point_in_time_dataset_export writes deterministic Arrow IPC or Parquet partitions. Its manifest contains the dataset identity, schema fingerprint, partition values and object keys, per-partition byte length and SHA-256, aggregate content hash, time window, provider, license, adjustment policy, and canonical parameters.

The cache key always covers tenant, tool version, schema version, provider, as_of, license policy, adjustment policy, and canonical parameters. Repeating an identical request over identical source revisions yields the same partition bytes, manifest values, content hash, and dataset identity.

This point-in-time export does not calculate returns, performance, signals, portfolios, or strategy conclusions. Major-version migration history is available only as structured Registry metadata.