Skip to main content

SymageDocs Python SDK

Generate synthetic documents, identities, and tabular datasets for testing, ML training, and compliance.

Installation

pip install symagedocs

For progress bars during long jobs:

pip install symagedocs[progress]

Quick Start

from symagedocs import Client

client = Client(api_key="sk_live_...")

# List available forms
forms = client.forms.list()
for f in forms:
    print(f"{f.id}: {f.name} ({f.credit_cost} credits)")

# Generate 100 W-2 documents
# JSON ground truth and CSV are always included in the dataset zip — no need to request them.
job = client.generate.create(
    "irs_w2_single_page_2025",
    quantity=100,
    output_formats=["pdf_typed"],  # see "Output formats" for all valid tokens
    # Augmentation knobs. `degradation_profile` affects credit cost —
    # `scanned`/`faxed` add 20%, `photographed` 30%, `mixed` 25% (`clean` = no surcharge).
    # `coherence_mode` controls cross-form identity correlation in multi-form jobs.
    degradation_profile="scanned",
    coherence_mode="coherent",
)
result = client.generate.wait(job.job_id)  # polls until complete
# "dataset" = one zip with every artifact + manifest.json; "json" and "csv"
# fetch just that ground-truth slice — see "Downloading results".
client.generate.download(job.job_id, "dataset", "./w2_documents.zip")

# Per-item training data
job = client.generate.create(
    form_id="irs_w2_single_page_2025",
    quantity=10,
    output_formats=["pdf_typed", "bio"],
    idempotency_key="my-retry-safe-key",
)
client.generate.wait(job.job_id)
for example in client.generate.iter_training_examples(job.job_id, format="bio"):
    print(example.item_id, len(example.bio.tokens))

# Generate tabular data from a description
schema = client.tabular.parse("name, age, SSN, city, state, annual income")
tab_job = client.tabular.generate(columns=schema.columns, quantity=5000)
client.tabular.wait(tab_job.job_id)
client.tabular.download(tab_job.job_id, "csv", "./dataset.csv")

# Check credit balance
balance = client.account.balance()
print(f"Credits used: {balance.credits_used}")

Authentication

Get your API key at symagedocs.ai/account?tab=api.

# Pass directly
client = Client(api_key="sk_live_...")

# Or set environment variable
# export SYMAGEDOCS_API_KEY=sk_live_...
client = Client()  # reads from env

Async Support

from symagedocs import AsyncClient

async with AsyncClient(api_key="sk_live_...") as client:
    forms = await client.forms.list()
    job = await client.generate.create("irs_w2_single_page_2025", quantity=10)
    result = await client.generate.wait(job.job_id)

Configuration

client = Client(
    api_key="sk_live_...",
    base_url="https://symagedocs.ai",  # custom server
    timeout=30.0,                       # request timeout (seconds)
    max_retries=3,                      # retry on 429/5xx
)

Method Reference

Forms

Method Description
forms.list(category=None) List available forms, optionally filtered by category
forms.get(form_id) Get detailed form info including field definitions

Generation

Method Description
generate.create(form_id=None, *, form_ids=None, quantity=1, output_formats=["pdf_typed"], config=None, seed=None, webhook_url=None, ink_color=None, ink_color_distribution=None, writer_consistency=None, degradation_profile=None, coherence_mode=None, fill_scenarios=None, idempotency_key=None) Create an async generation job. Pass either form_id (single form) or form_ids (coherent multi-form generation across the same identity). output_formats values and their pairing rules are listed under output formats. ink_color must be "black", "blue", or "red"; ink_color_distribution (when set) is a weight map over those same colors that must sum to exactly 100 and overrides ink_color. writer_consistency is "per_document" (default) or "per_field". degradation_profile and coherence_mode are typed kwargs over what used to live inside config={...} — see the augmentation knobs section; config={"label_scheme": ...} selects the ML label vocabulary — see training data. fill_scenarios is a typed kwarg for the declarative partial-fill / payer policy (list of FillScenario) — see partial fill. idempotency_key attaches an Idempotency-Key header so retries within 24 hours return the original job_id and don't double-charge. The deprecated realism_level API field is intentionally not exposed; call the REST API directly if you need it.
generate.list_jobs(limit=50, cursor=None, status=None) List generation jobs (cursor-paginated)
generate.get_job(job_id) Get full job status and progress
generate.list_downloads(job_id) List per-artifact presigned download URLs for a completed job
generate.download(job_id, format="dataset", path=".") Download job output to a local file. format is exactly one of "dataset" (default), "json", "csv" — anything else raises ValueError client-side. Allowed for terminal-but-not-completed jobs (CANCELED / FAILED / EXPIRED) so partial output is recoverable. Details under downloading results.
generate.download_dataset(job_id, out_dir, parallel=8, resume=True) Download a dataset (single or sharded layout) into a directory: manifest, README, and archive(s), with parallel shard fetch and size verification. resume=True skips shards already on disk. See downloading results.
generate.wait(job_id, poll_interval=3.0) Poll until the job reaches a terminal state. Returns the final Job on completion; raises ConflictError if the job failed. Shows a progress bar when tqdm is installed (pip install symagedocs[progress]).
generate.cancel(job_id) Cancel a running job. Idempotent. Items rendered before the cancel observed remain downloadable via download(format="dataset").
generate.list_items(job_id, limit=50, cursor=None) List per-item records for a job. Cursor-paginated; each item carries its presigned download URLs.
generate.download_item(job_id, item_id) Presigned S3 URLs for one item's files.
generate.get_bio_labels(job_id, item_id) Client-side helper: fetches the item's _bio.json sidecar and returns a parsed BioDataset.
generate.get_word_annotations(job_id, item_id) Client-side helper: fetches the item's _words.json sidecar and returns parsed WordAnnotations.
generate.iter_training_examples(job_id, format="bio") Client-side helper: iterates all items, yielding training examples in the chosen format ("bio" (default), "funsd", "donut").

client.generation alias. client.generation and client.generate reference the same resource — use whichever name you prefer.

Identities

Method Description
identities.generate(quantity=1, config=None, seed=None) Generate raw synthetic identities as JSON

Tabular

Method Description
tabular.parse(prompt) Convert natural language to a column schema (LLM-powered)
tabular.generate(columns, quantity=100, output_formats=["csv"], seed=None) Create a tabular generation job
tabular.status(job_id) Get tabular job progress and ETA
tabular.download(job_id, format, path) Download tabular output to a local file. format is "csv" or "json".
tabular.wait(job_id, poll_interval=2.0) Poll until tabular job completes or fails

Account

Method Description
account.balance() Get credit balance (credits_used, credits_allocated)
account.usage(days=30) Get usage summary for the specified period

Pricing

The pricing endpoints are public/unauthenticated on the backend, but the SDK still requires an API key at construction time for consistency; the auth header is sent and ignored by these routes.

Method Description
pricing.rates() Get the current credit rate constants (CSV per-row rate, PDF base + surcharge bands, multipliers, …)
pricing.estimate(*, field_count, output_formats, record_count, degradation_profile=None) Estimate the credit cost of a hypothetical job before submitting it

Health

Method Description
client.health() Lightweight reachability probe (GET /api/v1/health). Returns the parsed JSON body. Works on both Client and AsyncClient.

Output formats

generate.create(output_formats=[...]) accepts exactly these tokens; any other value is rejected with 400 code=invalid_output_format:

Token Produces
pdf_typed Filled PDF with typed text
pdf_handwritten Filled PDF rendered in synthetic handwriting
pdf_filled Completed PDF with values in live, editable AcroForm widgets (fillable forms only)
png_typed Per-page PNG rasterizations of pdf_typed (requires pdf_typed)
png_handwritten Per-page PNG rasterizations of pdf_handwritten (requires pdf_handwritten)
bio BIO-tagged tokens with spatial positions (ML)
coco COCO object-detection annotations (ML)
yolo YOLO detection annotations (ML)
donut Donut gt_parse ground truth (ML)

Rules enforced at job creation:

  • Foundational ground truth is always included. Per-instance JSON, tabular CSV, and FUNSD per-page annotations ship in every dataset automatically; "csv", "json", and "funsd" are not requestable tokens and return 400.
  • PNG travels with its PDF. png_typed requires pdf_typed in the same request; png_handwritten requires pdf_handwritten.
  • ML formats are feature-gated. bio/coco/yolo/donut require the ml-output-formats-enabled feature flag on your account (400 code=ml_formats_disabled otherwise) and at least one render format (pdf_typed, pdf_handwritten, or png_typed) in the same request, since annotations are derived from the render pipeline.
  • pdf_filled is fillable-forms-only and not a render surface. Requesting it for a non-fillable form rejects the whole submission with 400 code=format_unsupported_for_form (error.details.unsupported_form_ids lists the offenders). It delivers live, editable AcroForm widgets rather than a flattened render, so it cannot satisfy the ML-format render dependency, is never paired with a PNG, and is unaffected by degradation_profile. Filled PDFs land under pdfs/filled/ in the bundle.

Downloading results

generate.download(job_id, format="dataset", path=".") accepts exactly three formats:

  • dataset (default) — one zip with every artifact the job produced (PDFs, PNGs, ML annotations, per-item JSON ground truth, tabular CSV) plus a manifest.json describing the contents. The response is streamed to disk in 64 KiB chunks, so multi-GB datasets download with flat memory use.
  • json — flat per-instance JSON array (no images/PDFs).
  • csv — tabular identity data.

Anything else raises ValueError client-side before a request is made. (The pre-rename token bundle is not accepted by the SDK.)

When path is a directory (the default "."), a filename is appended automatically: symagedocs_<job_id>.zip for dataset, <job_id>.json for json, <job_id>.csv for csv.

Sharded datasets. Very large jobs are stored as a sharded dataset, which has no single archive; download(format="dataset") then raises ValueError pointing you at download_dataset(). download_dataset(job_id, out_dir, parallel=8, resume=True) is the universal accessor and works for both layouts: it fetches manifest.json and README.md, then either dataset.zip (single layout) or preview.zip followed by shards/shard_NNNNN.zip downloaded in parallel with size verification against the manifest (sharded layout). With resume=True, shards already on disk whose size matches the manifest are skipped, so a partially-failed download can be retried cheaply.

Job states. Downloads are allowed for terminal-but-not-completed jobs (CANCELED / FAILED / EXPIRED) so partial output is recoverable; downloading a job that is still running returns 409.

Tabular jobs have their own surface: tabular.download(job_id, format, path) accepts "csv" or "json".

Training data

Request ML annotation formats alongside a render format, then iterate per-item training examples:

job = client.generate.create(
    "irs_w2_single_page_2025",
    quantity=50,
    output_formats=["pdf_typed", "bio"],
    config={"label_scheme": "nist3"},  # default: "semantic_concept"
)
client.generate.wait(job.job_id)
for ex in client.generate.iter_training_examples(job.job_id, format="bio"):
    print(ex.item_id, len(ex.bio.tokens))  # BIO tags + word boxes
  • iter_training_examples(job_id, format=...) yields "bio" (default), "funsd" (one example per page, with page_index set), or "donut" examples.
  • config["label_scheme"] selects the annotation vocabulary: semantic_concept (default — concept names like social_security_number), nist3 (3-class name/ssn/data), field_id (form field IDs), or field_type (e.g. ssn, currency, text). Unknown values return 400 code=invalid_label_scheme.
  • Per-item sidecars are also fetchable directly: get_bio_labels(job_id, item_id) and get_word_annotations(job_id, item_id).

See the API User Manual's Training Data section for the full annotation schemas and Donut consumer conventions.

Augmentation knobs

Two of the most-used keys in the freeform config={...} dict on generate.create are also exposed as typed kwargs:

  • degradation_profile: Literal["clean", "scanned", "faxed", "photographed", "mixed"] | None
  • coherence_mode: Literal["coherent", "shuffled", "random"] | None

Why bother? Two reasons:

  1. degradation_profile affects credit cost. Non-clean profiles need extra rendering work (rasterization, noise, paper warp), so the billing engine applies a multiplier: scanned/faxed are billed at 1.2×, mixed at 1.25×, and photographed at 1.3×. A typo on the freeform config={...} form silently falls back to the default 1.0× multiplier — meaning you don't get the degradation you asked for AND the typo isn't caught until you notice the artifacts (or don't). The typed kwarg form catches typos at type-check time.
  2. Pre-flight validation. The Literal types fence off unknown values at edit time in any IDE that supports type checking. The backend also rejects unknown values with 400 for both knobs, so even untyped callers get a fast failure — but the typed form catches the mistake before the network round-trip.

The SDK exports the canonical value tuples too:

from symagedocs import DEGRADATION_PROFILES, COHERENCE_MODES

assert "scanned" in DEGRADATION_PROFILES
assert "coherent" in COHERENCE_MODES

If you pass a value via both forms (e.g. config={"degradation_profile": "X"} AND degradation_profile="Y"), the value in config wins and a RuntimeWarning is emitted so the conflict isn't silent.

# Typed kwarg form — recommended.
job = client.generate.create(
    "irs_w2_single_page_2025",
    quantity=100,
    degradation_profile="scanned",   # billed at 1.2× — see above
    coherence_mode="coherent",
)

# Equivalent freeform form — still supported, but typos cost money.
job = client.generate.create(
    "irs_w2_single_page_2025",
    quantity=100,
    config={"degradation_profile": "scanned", "coherence_mode": "coherent"},
)

Partial fill (fill_scenarios)

fill_scenarios is a declarative partial-fill / payer policy: instead of filling every field, you describe named, weighted scenarios and the job quantity is split across them deterministically by seed. It is the typed kwarg alias for config["fill_scenarios"] (ADR-072). The canonical use case is healthcare intake: a payer or EHR prefills the patient, insurance, and diagnosis/procedure blocks, and the provider completes (or leaves blank) the rest.

Each scenario assigns every field one of three stages:

  • prefilled — value synthesized (or pinned via values); rendered as typed text on every surface, including inside handwritten output. Models machine prefill by the payer/EHR.
  • completed — value synthesized, rendered surface-native. The default stage and the pre-existing behavior.
  • blank — no value; renders empty and appears as "" in the answer key, absent from every bbox label surface.

Fields are addressed in prefilled / blank by exact field id, or by a semantic-concept glob "concept:<glob>" matched against the field's semantic_concept (fetch the addressable set from client.forms.get(form_id) — each field carries semantic_concept, structural_role, and entity_role). values pins exact field ids to scalar literals (which forces those fields to prefilled). default_stage ("completed" or "blank") covers every field not matched by a selector or pin. The same scenario definition therefore produces the intake artifact (default_stage="blank") or the completed document (default_stage="completed").

from symagedocs import Client, FillScenario

client = Client(api_key="sk_live_...")

# Weighted: 3-in-4 documents are payer-prefilled intake artifacts (everything
# the provider hasn't filled yet is genuinely blank); 1-in-4 are fully
# completed. One shared prefill set, two default stages.
scenarios: list[FillScenario] = [
    {
        "name": "payer-intake",
        "weight": 3,
        "default_stage": "blank",
        "prefilled": [
            "concept:healthcare.patient.*",
            "concept:insurance.*",
            "concept:healthcare.diagnosis.*",
            "concept:healthcare.procedure.*",
        ],
        "values": {"box11c_plan_name": "AETNA PPO"},  # pin a literal
    },
    {
        "name": "completed",
        "weight": 1,
        "default_stage": "completed",
        "prefilled": ["concept:healthcare.patient.*", "concept:insurance.*"],
    },
]

job = client.generate.create(
    "cms_1500_standard_02_12",
    quantity=1000,
    output_formats=["pdf_typed"],
    seed=42,
    fill_scenarios=scenarios,
)

Ground truth records the policy: each per-document JSON carries _metadata.fill = {"scenario": <name>, "stages": {field_id: stage}}, so blank-by-policy is distinguishable from blank-because-unset.

Two rules to know:

  • Duplicate is an error. Passing fill_scenarios via both the kwarg and config={"fill_scenarios": [...]} raises ValueError client-side (unlike the scalar augmentation knobs, which resolve to a config-wins RuntimeWarning). Raw config["fill_scenarios"] passthrough with no kwarg still works.
  • v1 restriction. fill_scenarios is mutually exclusive with coherence_mode="shuffled" (both drive the per-field override channel); the API rejects the combination with 422 code=invalid_fill_scenarios.

See the API User Manual's "Partial fill and payer scenarios" section for the full selector grammar and multi-scenario examples.

Error Handling

The SDK raises typed exceptions for API errors and retries automatically on 429 and 5xx:

from symagedocs import Client, AuthenticationError, RateLimitError, NotFoundError

try:
    forms = client.forms.list()
except AuthenticationError:
    print("Invalid API key")
except RateLimitError:
    print("Too many requests — SDK retries automatically")
except NotFoundError:
    print("Resource not found")

All error classes:

Exception HTTP Code Description
SymageDocsError Base exception for all SDK errors
AuthenticationError 401 Invalid or revoked API key
PermissionDeniedError 403 Key missing required scope
NotFoundError 404 Resource not found
ValidationError 400 Invalid request parameters
InsufficientCreditsError 402 Not enough credits for the operation
ConflictError 409 Resource in unexpected state (e.g., downloading incomplete job)
RateLimitError 429 Rate limit exceeded (SDK retries automatically)
ServerError 5xx Server-side error (SDK retries automatically)

Examples

The examples/ directory (in the repository and the source distribution; not installed with the wheel) contains complete working scripts:

  • list_forms.py — Browse available forms and credit costs
  • generate_w2s.py — Full pipeline: create job, wait, download the dataset zip
  • tabular_dataset.py — Parse NL description, generate 5k rows, download CSV
  • train_kie_model.py — Create a job with NIST3 labels and BIO output, iterate training examples, fetch word annotations

Documentation

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

symagedocs-1.0.6.tar.gz (90.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

symagedocs-1.0.6-py3-none-any.whl (45.7 kB view details)

Uploaded Python 3

File details

Details for the file symagedocs-1.0.6.tar.gz.

File metadata

  • Download URL: symagedocs-1.0.6.tar.gz
  • Upload date:
  • Size: 90.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.13

File hashes

Hashes for symagedocs-1.0.6.tar.gz
Algorithm Hash digest
SHA256 782ecf492b50b14cfc13ffc1840266bd642ce3eea5b75cd5303e808c6ce0e90e
MD5 dd03146f123d74d64f580f27aa3e2bb7
BLAKE2b-256 637ea190b1c16c329b42612e4720ae61334d1a0209ed22ad76805c5b1c3c04c7

See more details on using hashes here.

File details

Details for the file symagedocs-1.0.6-py3-none-any.whl.

File metadata

  • Download URL: symagedocs-1.0.6-py3-none-any.whl
  • Upload date:
  • Size: 45.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.13

File hashes

Hashes for symagedocs-1.0.6-py3-none-any.whl
Algorithm Hash digest
SHA256 06120cd3fe30413e1f3a336fa48dc0485222d42217e186dc09340606c5a944d7
MD5 1e4e515879a8f33a4796bcd1a2872601
BLAKE2b-256 476de1abbb4ad32640c68deba75e237d6b63024ce0577f140b6ed4c30355ac46

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

1.0.6 This release

2 files

1.0.5

2 files

1.0.4

2 files

1.0.3

2 files

1.0.2

2 files

1.0.1

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page