Skip to main content

posture

Runtime-agnostic Python library for CCM (Continuous Control Monitoring) data collection. The entire contract: credentials in, DataFrame out. Runs unchanged in Docker, Airflow, Databricks — the library never knows or cares where it executes.

See docs/ARCHITECTURE.md for the design behind this library — the collect/parse split, locked design decisions, manifest schema, and per-collector implementation notes.

See docs/index.md for every supported collector: its required environment variables, an example query, and the full column schema for each of its tables.

Installation

pip install posture

A few storage backends have extra dependencies not installed by default — install them with the matching extra:

pip install posture[gcs]        # google-cloud-storage, for the "gcs" backend
pip install posture[s3]         # boto3, for the "s3" backend
pip install posture[bigquery]   # google-cloud-bigquery, for the "bigquery" backend
pip install posture[snowflake]  # snowflake-connector-python, for the "snowflake" backend

posture loads a .env file from the current directory (or a parent) automatically on import — no code changes needed. Variables already set in the environment always take precedence over .env values. Each collector's required variables are listed on its page in docs/index.md, e.g.:

# .env
CROWDSTRIKE_CLIENT_ID=xxx
CROWDSTRIKE_CLIENT_SECRET=xxx

Usage

from posture import CCM

ccm = CCM("crowdstrike")                          # creds from CROWDSTRIKE_* env vars
ccm = CCM("crowdstrike", {"client_id": "xxx"})    # partial override, rest from env

df = ccm.collect("hosts")                          # always a complete pandas DataFrame
ccm.flush_cache()                                  # the only cache invalidation

collect() always returns a complete pandas.DataFrame for the requested resource, or raises — there is no such thing as a partial snapshot in this library.

Paginated retrieval, for large resources

For a resource too large to comfortably hold in memory as one DataFrame (e.g. MDE's machine_vulnerabilities), use collect_page() instead — it yields one DataFrame per underlying API page, so peak memory is bounded to a single page rather than the whole resource:

from posture import Storage

store = Storage("sqlite", {"path": "posture.db"})
for df in ccm.collect_page("machine_vulnerabilities"):
    store.write_page(df, "machine_vulnerabilities", mode="append")

Storage("sqlite", ...) mirrors CCM("crowdstrike", ...) — one instance, reused across writes. A concrete class (from posture.storage import SqliteStorage) works identically when the backend is hardcoded rather than a runtime value.

collect() is a thin wrapper over collect_page() — it just concatenates every page into one DataFrame — so both share the same all-or-nothing guarantee: if collection fails partway through, an exception propagates and no partial data is left for the caller to mistake for a complete snapshot.

Discovering what's available

from posture import catalog

catalog()
# {
#   "crowdstrike": {
#     "required_config": {"client_id": "CROWDSTRIKE_CLIENT_ID", "client_secret": "CROWDSTRIKE_CLIENT_SECRET"},
#     "resources": {
#       "hosts": {"derived_from": None, "columns": ["client_id", "device_id", ...]},
#       "vulnerability_remediations": {"derived_from": "vulnerabilities", "columns": [...]},
#       ...
#     },
#   },
#   "knowbe4": {...},
#   ...
# }

catalog() never instantiates a collector, never touches the network, and needs no credentials — it reads sources, required config (as constructor key → env var), and resources (including which are derived, and their declared columns) straight off the registered Collector classes. It only reports required config — optional knobs (e.g. region, base_url) aren't tracked as data, so check a source's page in docs/index.md for those.

runnable_sources() filters catalog() down to sources whose required env vars are all set right now — useful for a universal collector that wants to skip sources with no credentials configured instead of instantiating each one to find out:

from posture import runnable_sources

runnable_sources()
# same shape as catalog(), but only sources ready to run in the current environment

storage_catalog() is the same idea for the storage layer:

from posture import storage_catalog

storage_catalog()
# {
#   "csv":      {"class_name": "CsvStorage", "required_config": {"path": "POSTURE_CSV_PATH"}, "optional_config": {}},
#   "postgres": {"class_name": "PostgresStorage", "required_config": {}, "optional_config": {"dsn": "POSTURE_POSTGRES_DSN", "host": "POSTURE_POSTGRES_HOST", ...}},
#   ...
# }

Same guarantees — no instantiation, no writes, no credentials needed. Postgres's config keys all show up as optional here even though one specific combination (dsn alone, or all of host/dbname/user/password) is actually required — that either/or logic lives in PostgresStorage.__init__, not in a flat required/optional key list.

Example: export Crowdstrike hosts to local JSON

from posture import CCM, write_storage

# CROWDSTRIKE_CLIENT_ID / CROWDSTRIKE_CLIENT_SECRET must be set in the environment
ccm = CCM("crowdstrike")
df = ccm.collect("hosts")

write_storage(df, "json", "hosts", config={"path": "output"}, mode="truncate")

print(f"Wrote {len(df)} hosts to output/default/hosts.json")

Storage: writing a DataFrame somewhere durable

from posture import write_storage

write_storage(df, "csv", "hosts", config={"path": "output"})                 # output/<tenant>/hosts.csv
write_storage(df, "parquet", "hosts", config={"path": "output"})             # output/<tenant>/hosts.parquet
write_storage(df, "sqlite", "hosts", config={"path": "output/posture.db"})   # table "hosts"
write_storage(df, "duckdb", "hosts", config={"path": "output/posture.duckdb"})  # table "hosts"
write_storage(df, "postgres", "hosts", config={"dsn": "postgresql://..."})   # table "hosts"
write_storage(                                                               # same, discrete keys
    df, "postgres", "hosts",
    config={"host": "...", "dbname": "...", "user": "...", "password": "..."},
)
write_storage(df, "gcs", "hosts", config={"bucket": "my-bucket"})            # gs://my-bucket/hosts/<tenant>.parquet
write_storage(df, "s3", "hosts", config={"bucket": "my-bucket"})             # s3://my-bucket/hosts/<tenant>.parquet
write_storage(df, "bigquery", "hosts", config={"project_id": "...", "dataset_id": "..."})  # table "hosts"
write_storage(                                                               # snowflake
    df, "snowflake", "hosts",
    config={
        "account": "...", "database": "...", "schema": "...",
        "authenticator": "SNOWFLAKE", "user": "...", "password": "...",
    },
)

storage is one of "csv", "json", "parquet", "sqlite", "duckdb", "postgres", "gcs", "s3", "bigquery", "snowflake". Postgres accepts either a single dsn or discrete host/port/dbname/user/password keys (same convention every collector uses for its own credentials, resolved from POSTURE_POSTGRES_HOST etc. if not passed explicitly) — dsn takes precedence if both are given.

gcs, s3, bigquery, and snowflake each require an extra to install (pip install posture[gcs] / posture[s3] / posture[bigquery] / posture[snowflake] — see Installation) and authenticate the way their respective SDK always does (Application Default Credentials for gcs/bigquery; the standard boto3 credential chain for s3). snowflake has no default authenticator — every tenant states its own auth method ("SNOWFLAKE" for password, "WORKLOAD_IDENTITY" with a workload_identity_provider, key-pair via private_key_file, etc.) explicitly via config or POSTURE_SNOWFLAKE_AUTHENTICATOR; role/warehouse are optional with no tenant-specific default either — omit them to use the connecting user's own account defaults.

gcs/s3 own an opinionated object-key layout rather than taking a path prefix — <name>/<tenant>.parquet for truncate, where tenant comes from the TENANT env var (default "default"). For append:

  • gcs<name>/<tenant>/<YYYY-MM-DD>.parquet
  • s3<name>/<tenant>/YEAR=<yyyy>/MONTH=<mm>/DAY=<dd>/<name>.parquet, Hive-style partitioning so the output is directly queryable by Athena/Glue without a separate partition-projection config.

mode controls both overwrite behaviour and history. For the local file backends (csv/json/parquet), every path is rooted <path>/<tenant>/<name>... — tenant first, then table name, then date — from the TENANT env var (default "default"), so a query engine like DuckDB can glob/prune by tenant without touching other tenants' files:

  • "truncate" (the default — latest load is all posture cares about by default) — overwrites/replaces in place: output/default/hosts.csv, or output/default/hosts.parquet.
  • "append" — keeps a dated snapshot per day: output/default/hosts/2026/08/22/hosts.csv.

For the database backends (sqlite/duckdb/postgres/bigquery/snowflake), every row also carries a tenant column (from the TENANT env var), so a table can be shared by several tenants without one tenant's write clobbering another's rows — "truncate" here means tenant-scoped, not table-scoped: it deletes only the current tenant's existing rows before inserting the fresh set, leaving other tenants' rows in the same table untouched. "append" just inserts on top of whatever's already there. Either way, opt into "append" deliberately — it has real storage growth implications the default doesn't.

Every file write goes through a temp file and an atomic rename, so a failure partway through never leaves a broken file at the real path. For a paginated collection, use write_page() on a backend instance instead of write_storage() — see Paginated retrieval above.

Supported sources

See docs/index.md for the full list of collectors, each with its required environment variables, an example query, and the column schema for every table it exposes.

Development

pip install -e ".[dev]"
pytest
ruff check src tests
black src tests

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

posture-0.19.0.tar.gz (278.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

posture-0.19.0-py3-none-any.whl (201.2 kB view details)

Uploaded Python 3

File details

Details for the file posture-0.19.0.tar.gz.

File metadata

  • Download URL: posture-0.19.0.tar.gz
  • Upload date:
  • Size: 278.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for posture-0.19.0.tar.gz
Algorithm Hash digest
SHA256 82f8cf6b3db99f58a6b9d72dc672445aec0c0cf07481ce4c406080d4c21f1b2b
MD5 f057c07653ce5ca344339977c81192d9
BLAKE2b-256 528acb39122d6569fbde00367d6a1231a404c370ac5c1d1e72499c0f30867995

See more details on using hashes here.

File details

Details for the file posture-0.19.0-py3-none-any.whl.

File metadata

  • Download URL: posture-0.19.0-py3-none-any.whl
  • Upload date:
  • Size: 201.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for posture-0.19.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e1bbc1334ce25724c7a53a2157eb375b96dd0f86df0150532e5bdae8d1f03ea3
MD5 7ed3d561b03a156450668f009030198e
BLAKE2b-256 46147182089e537b88710ade94a612bb1bc7e5828ab9bc5a9a23268589dbd81a

See more details on using hashes here.

Release history Release notifications | RSS feed

1.3.0

2 files

1.2.0

2 files

1.1.0

2 files

1.0.1

2 files

1.0.0

2 files

0.23.0

2 files

0.21.0

2 files

0.20.1

2 files

0.20.0

2 files

0.19.5

2 files

0.19.3

2 files

0.19.2

2 files

0.19.1

2 files

This release

0.19.0 This release

2 files

0.18.1

2 files

0.18.0

2 files

0.17.5

2 files

0.17.4

2 files

0.17.3

2 files

0.17.2

2 files

0.17.1

2 files

0.17.0

2 files

0.16.1

2 files

0.16.0

2 files

0.15.2

2 files

0.15.1

2 files

0.15.0

2 files

0.14.0

2 files

0.13.3

2 files

0.13.2

2 files

0.13.1

2 files

0.13.0

2 files

0.12.1

2 files

0.12.0

2 files

0.11.0

2 files

0.10.1

2 files

0.10.0

2 files

0.9.5

2 files

0.9.4

2 files

0.9.3

2 files

0.9.2

2 files

0.9.1

2 files

0.9.0

2 files

0.8.4

2 files

0.8.3

2 files

0.8.2

2 files

0.8.1

2 files

0.8.0

2 files

0.7.1

2 files

0.7.0

2 files

0.6.1

2 files

0.6.0

2 files

0.5.2

2 files

0.5.0

2 files

0.4.6

2 files

0.4.4

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.3.2

2 files

0.3.1

2 files

0.2.6

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page