VCF-RDFizer is a Docker-first CLI wrapper for:
- VCF -> RDF (N-Triples) with RMLStreamer
- Optional RDF compression/decompression, into queryable HDT and COTTAS artifacts
- Semantic validation of a compressed RDF graph against its source VCF
- Data linking, which writes a provenance-tracked side-graph of external links
- Policy attachment, which applies ODRL policies to files, regions and variants and writes checked release views
The conversion targets the VCF Core vocabulary, published at
https://w3id.org/vcf-core/vocab# (prefix
vcfc:). It replaces the retired VCF-RDFizer vocabulary
(https://w3id.org/vcf-rdfizer/vocab#), which is not redirected: that namespace
serves a deprecation document linking every former term to its successor. VCF
Core is a semantic target any conversion system can adopt, and its version line
is independent of this converter's — the emitted IRIs reference the namespace,
not a pinned version.
This README covers day-to-day use: installation, the CLI flags, and worked
commands. The documentation set in docs/ covers
everything else — how the tool works internally, why it is designed that way,
and where it stops working. Start with
Architecture, and read
Limitations before you rely on the output.
Requirements
- Python 3.10+
- Docker (installed and running)
When VCF-RDFizer is connected to an interactive terminal, it shows a
lightweight Rich spinner/progress display. With redirected output or without
Rich installed, it instead prints compact status lines; CI output remains
quiet. Validation uses the same display and reports each preflight/SPARQL
query as it starts and completes. Use --quiet to suppress these terminal
progress displays (and validation's per-query/summary chatter) while keeping
the progress sidecar, command log, and metrics collection active. Use
--no-progress when the sidecar and terminal progress should both be disabled.
RMLStreamer progress reports the bytes and output parts already written;
partitioned HDT/COTTAS runs report source triples, chunks, and the currently
active merge/index stage. These updates are best-effort and do not scan RDF
content a second time or retain progress history in memory.
Install options:
pip install vcf-rdfizer
or
pipx install vcf-rdfizer
or
conda install -c conda-forge vcf-rdfizer
or pull the prebuilt Docker image directly:
docker pull ecrum19/vcf-rdfizer:latest
Release maintainers: see scripts/RELEASING.md for the PyPI,
Docker Hub, and conda-forge release procedure.
Important CLI Rule
--out is required for all modes.
This is the run output root directory. VCF-RDFizer places:
- final RDF/compression outputs
- run metrics/logs
- hidden intermediates
inside this directory.
Modes
full: VCF -> TSV -> RDF -> compression (and optional semantic validation with--validate)tsv: VCF -> TSV only (benchmarking)compress: compress an existing.ntor.nt.gzdecompress: decompress.nt.gz,.nt.br,.hdt,.cottas,.cottas.gz, or.cottas.brvalidation: compare a source VCF with its.ntor.nt.gzRDF using six semantic SPARQL queriesindex: only generate or regenerate the query index for an existing.hdtor.cottas
In full mode with multiple VCF inputs, failures are isolated per input:
- the run continues with remaining files
- failed inputs are summarized in
run_metrics/<INPUT_LABEL>__<RUN_ID>/reports/failed_inputs.csv
Main Flags (Most Used)
-m, --mode {full,compress,decompress,tsv,index,validation,link}-o, --outrequired output root directory--rdf-compressionfinal raw RDF codecs:gzip,brotli, ornone--representationsqueryable RDF outputs:hdt,cottas, ornone--artifact-compressionpackaging codecs for selected representations:gzip,brotli, ornone--hdt-strategy {auto,partitioned,single}HDT generation policy--chunk-target-bytes,--chunk-min-bytes,--chunk-max-bytesshared record-safe chunk sizing--sample-representation {expanded,condensed}genotype graph shape (expandedby default)--validate(or--run-validation) run semantic VCF/RDF validation for every input in a full run--filter-oracle {auto,bcftools,cyvcf2}FILTER oracle used by validation (autoby default)--quietsuppress terminal progress displays while retaining sidecar/log/metrics tracking--no-progressdisable terminal progress and progress sidecar creation-I, --imageDocker image repo (defaultecrum19/vcf-rdfizer)-v, --image-versionDocker tag/version-b, --buildforce Docker build-B, --no-buildfail if image not found-h, --helpshow full usage
Validation Mode
Validate one source VCF against the .nt or .nt.gz aggregate from the same
conversion. Gzip input is decompressed, parsed, and queried only inside the
Docker container; any temporary raw N-Triples are removed before the container
exits. To run
the same checks as part of a full conversion, add --validate; validation then
runs once per input after RDF/compression and accepts either the generated
.nt or .nt.gz aggregate.
vcf-rdfizer --mode validation \
--input ./cohort.vcf.gz \
--rdf ./results/cohort/cohort.nt.gz \
--sample-representation condensed \
--out ./validation-results
Use expanded for the default graph shape and condensed for the vector-based
cohort graph. Reports from standalone and full-run validation use the canonical
metrics tree:
run_metrics/<input-label>__<run-id>/reports/validation/<dataset-id>/, with a
stage summary at stages/validation/<dataset-id>.json. See Semantic VCF/RDF
validation for query definitions, preflight checks, result
statuses, and cleanup evidence.
Artifacts and engines
--rdf accepts any artifact the pipeline produces - .nt, .nt.gz, .nt.br,
.hdt, .cottas, .cottas.gz, .cottas.br. A compressed or indexed artifact
is decoded back to N-Triples inside the container and then put through the
full semantic suite, which proves it decodes to a graph that still reproduces
every VCF summary - stronger than the triple-count round-trip that runs during
compression.
In full mode, --validate-artifacts {aggregate,hdt,cottas,all} chooses which
produced artifacts to check; each is validated independently with its own
report:
vcf-rdfizer --mode full -i ./cohort.vcf.gz \
--rdf-storage-mode space-optimized --representations hdt,cottas \
--validate --validate-artifacts all -o ./results
--validation-engine selects the SPARQL backend. Comunica (default) queries the
file in memory; QLever builds an
on-disk index inside the container and serves it, which is what makes
cohort-scale graphs queryable; hdt and cottas query the compressed
artifact in place, without decoding it, through comunica-sparql-hdt and
pycottas's rdflib store. All four answer identical queries, so the choice is
never semantic, and every report records which engine ran.
vcf-rdfizer --mode validation -i ./cohort.vcf.gz --rdf ./results/cohort/cohort.hdt \
--validation-engine qlever --qlever-memory-gb 32 -o ./validation-results
Several engines can run in one pass - --validation-engine comunica,qlever or
all. Each answers the whole query set, so the run cross-checks them against
each other (engine-agreement.json) and times them against each other and
against the cyvcf2 oracle computing the same answers, in benchmark.json and a
long-format benchmark.csv.
vcf-rdfizer --mode validation -i ./cohort.vcf.gz --rdf ./results/cohort/cohort.nt.gz \
--validation-engine all -o ./validation-results
Tuning: --qlever-memory-gb, --qlever-port, --qlever-startup-timeout,
--validation-query-timeout, and repeatable --qlever-index-arg /
--qlever-server-arg escape hatches.
Validation runs three independent layers: exact aggregate comparison against
the VCF, a predicate/class census plus per-record and per-value identity
digests, and — with --shacl-shapes — SHACL conformance against the
vocabulary's published shapes.
vcf-rdfizer --mode validation -i ./cohort.vcf.gz --rdf ./results/cohort/cohort.nt.gz \
--shacl-shapes ./vocabulary/shacl/vcf-core-vocabulary.shacl.ttl \
--strict-conformance -o ./validation-results
What a PASS means. Coverage is measured, not asserted: a mutation harness corrupts a correct graph in 42 named ways and records which are detected (currently 76/78). See
docs/vcf-coverage.mdfor the element-by-element matrix and the remaining gaps, anddocs/validation-methodology.mdfor how the number is produced.
Compression Plan
Compression is configured as three independent decisions:
- RDF staging:
--rdf-storage-modecontrols how RMLStreamer output is assembled before chunking.plaincreates one.ntaggregate;space-optimizedstreams the parts into one.nt.gzaggregate and removes each source part immediately. - Raw RDF artifacts:
--rdf-compression gzip,brotlicreates compressed copies of the RDF aggregate. Use--rdf-compression nonewhen RDF is only a temporary input to HDT/COTTAS. - Queryable representations and packaging:
--representations hdt,cottascreates the selected indexed formats.--artifact-compression gzip,brotlipackages each selected representation, producing.hdt.gz,.hdt.br,.cottas.gz, or.cottas.brin addition to the queryable base artifact.
Each selector accepts comma-separated values. Use none by itself to disable
that stage; do not combine none with another value. --artifact-compression
requires at least one selected value in --representations.
The gzip used by --rdf-storage-mode space-optimized is staging storage; it is
not automatically a final raw RDF artifact. The default plan preserves the
historical outputs: --rdf-compression gzip,brotli and
--representations hdt. For the smallest final output, select
--rdf-compression none, one representation, and
--remove-rdf-storage-output.
Packaged .hdt.gz, .hdt.br, .cottas.gz, and .cottas.br files are archives,
not directly queryable indexed files. Keep the unwrapped .hdt/.cottas file
when queries must run without a decompression step.
Use --mode decompress to decode either base representation. COTTAS packages
are unpacked inside the Docker container before pycottas writes the decoded
N-Triples output, so the temporary unwrapped COTTAS file is not added to the
host filesystem.
Full Mode Flags
-i, --inputrequired VCF file or directory-r, --rulesmapping rules file (.ttl)- default:
rules/default_rules.ttl
- default:
--sample-representation {expanded,condensed}sample genotype representationexpanded(default): oneSampleCallper record/sample and oneFormatFieldValueper FORMAT keycondensed: reusable file-level samples plus one ordered value vector per record/FORMAT key
--validaterun semantic VCF/RDF validation once per input after RDF/compression; detailed results are stored beneathrun_metrics/.../reports/validation/--filter-oracle {auto,bcftools,cyvcf2}FILTER oracle for--validate--quietsuppress terminal progress and validation query chatter while retaining logs/metrics--no-progressdisable progress sidecars and terminal progress displays--rdf-storage-mode {plain,space-optimized}full-mode aggregate storage policy (default:space-optimized)plain: merge RMLStreamer parts into one uncompressed.ntspace-optimized: gzip each part into one.nt.gzaggregate and delete the source part immediately
--rdf-compression {gzip,brotli,none}raw RDF artifacts to retain--representations {hdt,cottas,none}queryable primary representations--cottas-indexespermutations ofspoorspog, comma-separated;allselects six triple orders,all-quadsselects 24 dataset orders (defaultspo)--artifact-compression {gzip,brotli,none}optional packaging applied to each selected representation--hdt-strategy {auto,partitioned,single}auto: in full mode, build smaller HDT chunks and merge them with nativehdtcpartitioned: always use chunked HDT generation for HDT-based methodssingle: always use onerdf2hdtrun per RDF input- with
space-optimized, useautoorpartitioned;singlecannot consume the gzip stream without expanding it
--chunk-target-bytestarget uncompressed bytes per HDT/COTTAS chunk--chunk-min-bytesminimum uncompressed bytes before flushing a chunk group--chunk-max-bytesmaximum uncompressed bytes in a chunk; boundaries remain on complete NT lines-P, --spark-partitionsoptional Spark partition hint (positive integer)- low-cost way to tune RMLStreamer parallelism while the wrapper still produces one aggregate RDF output
-k, --keep-tsvkeep hidden TSV intermediates-R, --keep-rmlstreamer-rdf-outputkeep the aggregate RDF output produced by RMLStreamer--remove-rdf-storage-outputexplicitly remove the aggregate.nt/.nt.gzafter successful compression-e, --estimate-sizepreflight size estimate
VCF Coverage
Full-mode conversion covers every VCF column. Three options control how much structure is emitted; all default to the richer form.
| Option | Effect |
|---|---|
--sample-representation {expanded,condensed} |
Genotype shape (see below) |
--info-representation {structured,raw} |
structured adds one vcfc:InfoFieldValue per record and key, with typed values, alongside vcfc:infoRaw |
--vcf-version {auto,4.1,4.2,4.3,4.4,4.5} |
auto (the default) reads each input's ##fileformat line, so nothing has to be supplied; an explicit version overrides it for a file whose declaration is missing or wrong |
--header-representation {structured,basic} |
structured types each ## line with its vocabulary subclass, emits the ordered vcfc:HeaderAttribute resources the SHACL profile requires, and lifts the INFO/FORMAT/FILTER/ALT/contig/META/SAMPLE/PEDIGREE attributes into their own properties |
ID, ALT, QUAL, FILTER and infoRaw are always emitted by the wrapper
rather than by the RML mapping, because each may be the VCF missing token and
the vocabulary requires that as "."^^vcfc:Null — a datatype that depends on
the row, which RML cannot express. QUAL additionally takes xsd:decimal for a
finite value and vcfc:VCFFloat for the INF/INFINITY/NAN spellings.
VCF versions
VCF Core ships a SHACL overlay per VCF version (4.1–4.5). The converter detects
each input's version from its ##fileformat line — required to be the first
line of a conforming VCF — and follows it per input, so a directory of mixed
versions converts correctly without any flag. The version selects the
vcfc:VCF4xFile class in the mapping and the behaviour that genuinely differs
between versions: whether CIPOS/CIEND carry one pair per record (4.1–4.3) or
one per ALT allele (4.4+), whether CILEN/CICN and EVENTTYPE exist at all,
and whether the local-allele and base-modification FORMAT families are defined.
VCF 4.0 has no conformance overlay, so it converts with the newest supported rules and no version class, and the run says so rather than assuming silently.
--info-representation structured also carries the allele layer (ordered
vcfc:ReferenceAllele/vcfc:AltAllele with their kind, symbolic type and
breakend components), the vcfc:FieldValueItem decomposition of
Number=A/R/G/P values, and the structural-variant carriers.
--header-representation basic is the smallest header form but does not
produce a SHACL-conformant graph.
docs/vcf-coverage.md maps every VCF element to its RDF
terms and to the validation check that covers it.
Sample Representation Modes
Full mode has exactly two explicit sample workflows. There is no automatic sample-count threshold, so the same command always produces the same graph shape and downstream consumers can select the contract they support.
See Expanded and Condensed Knowledge Representations for a detailed, worked explanation of the graph shapes, scaling behavior, and selection trade-offs.
Expanded (default)
Use --sample-representation expanded for single-sample and low-sample VCFs. It
preserves the per-sample vocabulary model:
- every record/sample pair is a
vcfc:SampleCall, linked to a reusablevcfc:VCFSamplewithvcfc:forSample; - every represented FORMAT slot is a
vcfc:FormatFieldValue, linked to its declaration withvcfc:declaredBy; GTis parsed into avcfc:Genotypewith an explicitvcfc:phasingStatus—vcfc:MixedPhasingwhen its positions disagree — and onevcfc:GenotypeAlleleCallper position carrying its ownvcfc:phaseIndicator, andFT,PS/PSL,LAA(with per-samplevcfc:LocalAlleleMembershipordinals), the copy-number and haplotype keys, and theM/DPM/ADMbase-modification families in both their numeric and aliased spellings get their own resources;- the VCF file declares
vcfc:representationProfile vcfc:ExpandedRepresentation.
With the default rules, these triples are appended directly from records.tsv;
the large materialized helper TSVs are not created. The final graph is still
expanded and grows approximately with variants × samples × FORMAT fields.
vcf-rdfizer --mode full \
--input ./small.vcf \
--sample-representation expanded \
--rdf-storage-mode plain \
--out ./results
Condensed
Use --sample-representation condensed for large multi-sample cohorts:
- sample columns are declared once as an ordered
vcfc:SampleSetof reusablevcfc:VCFSampleresources; - each genotype-bearing call has one
vcfc:CohortCallMatrix; - each FORMAT key has one
vcfc:FormatValueVector, rather than one RDF value resource per sample; - the VCF file declares
vcfc:representationProfile vcfc:CondensedRepresentation.
Condensed mode deliberately stops at the vector: it does not derive the
per-sample genotype, phase-set or base-modification resources that expanded mode
does, because that is exactly the per-sample materialization the profile exists
to avoid. Every value stays recoverable by decoding a vector against its FORMAT
definition and the matrix SampleSet. Both profiles do emit the file-level
vcfc:SampleSet, which costs one resource per sample column for the whole file.
vcfc:encodedValues uses vcfc:VCFTextVector: one tab-separated lexical item
per sample in vcfc:sampleIndex order. Commas inside a FORMAT value remain part
of that item, and absent values are emitted as . so all vectors stay aligned.
Consumers reconstruct sample i's value for a FORMAT key by selecting position
i from its vector. This changes genotype graph growth to approximately
samples + variants × FORMAT fields; the literal payload still contains all
source values, but they no longer cause per-value RDF structural triples.
vcf-rdfizer --mode full \
--input ./large-cohort.vcf.gz \
--sample-representation condensed \
--rdf-storage-mode space-optimized \
--representations hdt \
--out ./results
The workflow resolver runs only one sample emitter. Condensed mode rejects
custom mappings that consume materialized sample_calls.tsv or
sample_format_values.tsv, because running those helper-table mappings alongside
the condensed emitter would create both representations and restore the semantic
inflation this mode is designed to avoid. Remove those helper-table consumers
or select expanded mode. Custom rules with no helper-table consumers remain
compatible with condensed emission.
TSV Mode Flags
-i, --inputrequired VCF file or directory- Outputs per-run benchmark summary in
run_metrics/<INPUT_LABEL>__<RUN_ID>/tsv_metrics.csv - Writes container timing and structured TSV metrics under
timings/tsv/andstages/tsv/in that run directory
Compression Mode Flags
--rdfrequired input.ntor.nt.gzfile
Decompression Mode Flags
-C, --compressed-inputrequired.nt.gz,.nt.br,.hdt,.cottas,.cottas.gz, or.cottas.br-d, --decompress-outoptional explicit output.ntpath (must be inside--out)
Index-Only Mode Flags
-H, --hdtexisting.hdtfile; creates or regenerates its sibling sidecar--cottasexisting.cottasfile; rebuilds its embedded query index in place--cottas-indexesselects the replacement order and optional additional index copies- Exactly one of
--hdtor--cottasis required. - HDT indexing is Java-free: the image uses
hdtc1.1.0 to generate the canonical v1-1 sidecar named<file>.hdt.index.v1-1. - Each COTTAS copy stores its index inside the Parquet-based
.cottasfile; the primary is rewritten atomically and there is no index sidecar. - Existing indexes are intentionally replaced. Use this mode when an HDT sidecar is missing/stale or when a COTTAS file needs its query ordering and zone-map metadata rebuilt. No VCF conversion, RDF conversion, packaging, or decompression output is produced.
- The operation is also run automatically after each partitioned HDT merge.
Quick Start
Show help:
vcf-rdfizer --help
Full pipeline (plain aggregate RDF):
vcf-rdfizer \
--mode full \
--input ./vcf_files \
--rdf-storage-mode plain \
--rdf-compression none \
--representations none \
--out ./results
Full pipeline (plain aggregate, chunked HDT + native hdtc merge):
vcf-rdfizer \
--mode full \
--input ./vcf_files \
--rdf-storage-mode plain \
--rdf-compression none \
--representations hdt \
--hdt-strategy partitioned \
--chunk-target-bytes 536870912 \
--chunk-min-bytes 134217728 \
--chunk-max-bytes 1073741824 \
--out ./results
Full pipeline (space-optimized aggregate with shared HDT and COTTAS chunks):
vcf-rdfizer \
--mode full \
--input ./vcf_files \
--rdf-storage-mode space-optimized \
--rdf-compression none \
--representations hdt,cottas \
--chunk-target-bytes 536870912 \
--chunk-min-bytes 134217728 \
--chunk-max-bytes 1073741824 \
--out ./results
Full pipeline with a Spark partition hint:
vcf-rdfizer \
--mode full \
--input ./vcf_files \
--rdf-storage-mode space-optimized \
--spark-partitions 8 \
--rdf-compression none \
--representations hdt \
--out ./results
Full pipeline with custom rules + keep RMLStreamer RDF output:
vcf-rdfizer \
--mode full \
--input ./vcf_files \
--rules ./rules/my_rules.ttl \
--rdf-storage-mode plain \
--rdf-compression brotli \
--representations hdt \
--keep-rmlstreamer-rdf-output \
--out ./results
Ultra-small full pipeline:
vcf-rdfizer \
--mode full \
--input ./vcf_files \
--rdf-storage-mode space-optimized \
--rdf-compression none \
--representations hdt \
--hdt-strategy partitioned \
--remove-rdf-storage-output \
--out ./results
Queryable HDT and COTTAS plus gzip/Brotli packages:
vcf-rdfizer \
--mode full \
--input ./vcf_files \
--rdf-storage-mode space-optimized \
--rdf-compression none \
--representations hdt,cottas \
--artifact-compression gzip,brotli \
--out ./results
TSV-only benchmark:
vcf-rdfizer \
--mode tsv \
--input ./vcf_files \
--out ./results
Compression-only:
vcf-rdfizer \
--mode compress \
--rdf ./results/sample/sample.nt \
--rdf-compression none \
--representations hdt \
--artifact-compression gzip \
--out ./results
Compression-only from a space-optimized aggregate:
vcf-rdfizer \
--mode compress \
--rdf ./results/sample/sample.nt.gz \
--rdf-compression none \
--representations hdt,cottas \
--chunk-target-bytes 536870912 \
--out ./results
Decompression-only:
vcf-rdfizer \
--mode decompress \
--compressed-input ./results/sample/sample.hdt \
--out ./results
COTTAS decompression, including an externally packaged COTTAS file:
vcf-rdfizer \
--mode decompress \
--compressed-input ./results/sample/sample.cottas.gz \
--out ./results
Initialize an index for an existing HDT:
vcf-rdfizer \
--mode index \
--hdt ./results/sample/sample.hdt \
--out ./results
Regenerate the embedded index for an existing COTTAS file:
vcf-rdfizer \
--mode index \
--cottas ./results/sample/sample.cottas \
--out ./results
--mode index is deliberately an in-place maintenance operation. It mounts
only the directory containing the selected artifact and writes metrics under
<out>/run_metrics/<INPUT_LABEL>__<RUN_ID>/stages/index/. HDT indexing creates a
versioned sidecar beside the input. COTTAS indexing rewrites the existing
.cottas file through a bounded streaming Parquet rewrite, keeping the data
in the same artifact while rebuilding its embedded index. If the operation
fails, the original COTTAS file is left in place; HDT's previous sidecars are
restored.
HDT merging and indexing do not start a Java HDT tool. They use native hdtc:
hdtc create merges partitioned HDTs and hdtc index streams an HDT through
disk-backed external sorters. Both default to a 512 MiB soft memory budget.
Override index creation for any wrapper mode with an environment variable such
as:
HDT_INDEX_MEMORY_LIMIT=2G vcf-rdfizer \
--mode index \
--hdt ./results/sample/sample.hdt \
--out ./results
Use HDT_MERGE_MEMORY_LIMIT in the same way to tune a partitioned merge.
Accepted values use an M or G suffix. Lower values reduce in-memory sort
buffers and may increase temporary I/O; higher values can improve performance
when memory is available. Temporary files live in the container's /work area
and are removed after the attempt.
Choose COTTAS orders with --cottas-indexes spo,pso,pos or --cottas-indexes all.
Datasets (--mode compress --rdf dataset.nq or .nq.gz) use graph-aware orders,
for example --representations cottas --cottas-indexes spog,gspo, or
--cottas-indexes all-quads for all 24 permutations. Named and default graphs
are preserved; triple-only orders and HDT are rejected for dataset inputs.
Graph-aware orders also work on VCF/triple input, using the default graph.
The first order keeps sample.cottas; additional whole-graph copies are named
sample.pso.cottas, sample.pos.cottas, etc. Query one copy at a time. Each RDF
chunk is parsed once; extra orders sort its deduplicated Parquet data. Each order
gets its own streaming merge, round-trip check and requested gzip/Brotli packages.
JSON metrics list every index; existing CSV sizes refer to the primary copy.
See COTTAS representations for details.
COTTAS avoids both a global in-memory DISTINCT and a global external sort.
Each chunk is written in each requested order, so the final stage performs a
bounded k-way Parquet merge: it holds one small batch from each chunk, writes
one copy of each adjacent equal triple, and preserves that index. Its
memory use is controlled by COTTAS_MERGE_BATCH_ROWS (default 2048), not by
the total RDF graph size or a temporary DuckDB sort area. Override it only to
tune the memory/throughput tradeoff, for example:
COTTAS_MERGE_BATCH_ROWS=4096 vcf-rdfizer \
--mode compress \
--rdf ./results/cohort/cohort.nt.gz \
--rdf-compression none \
--representations cottas \
--out ./results
In full mode, an HDT sidecar-index failure is non-fatal when the HDT data
itself remains readable; the run continues and the HDT can be repaired later
with the standalone command above. If COTTAS generation/indexing cannot
produce a usable artifact, COTTAS-specific outputs are skipped while the rest
of the full pipeline continues. These warnings are printed in the run output
and written to run_metrics/<INPUT_LABEL>__<RUN_ID>/reports/index_warnings.json. The raw RDF is
retained when a representation-dependent output was unavailable so the
standalone index command or a later rerun has a recoverable source.
Output Layout
Given --out ./results:
- final outputs:
./results/<sample>/...
- per-run metrics/logs:
./results/run_metrics/<INPUT_LABEL>__<RUN_ID>/...
- hidden intermediates:
./results/.intermediate/tsv/
For an existing RDF input named test-larger.nt or test-larger.nt.gz,
--mode compress --out ./results always uses one directory:
./results/test-larger/test-larger.hdt
./results/test-larger/test-larger.hdt.index.v1-1
./results/test-larger/test-larger.cottas
./results/test-larger/test-larger.nt.gz
The same basename rule applies in full mode. VCF-RDFizer performs an output
collision check before Docker or conversion starts and never overwrites a
planned pipeline artifact. The exception is the deliberate --mode index
maintenance operation, which regenerates the selected artifact's index in
place. Choose a new --out directory, or rename/remove a conflicting pipeline
artifact before rerunning.
Intermediates are hidden by default.
Raw RDF files are removed after successful compression by default. Use
--remove-rdf-storage-output to make that cleanup explicit, or use
--keep-rmlstreamer-rdf-output to retain the aggregate RDF output instead.
The space-optimized mode retains the .nt.gz aggregate when gzip is selected
because that file is the gzip artifact itself.
Metrics
Each invocation receives a descriptive metrics directory:
run_metrics/<INPUT_LABEL>__<RUN_ID>/
<INPUT_LABEL> is the source filename without its recognized VCF/RDF or
representation suffix (for example, 1000G_phase3_chr20). A multi-file input
directory uses a batch label such as batch-vcf_data-4-inputs. This makes a
metrics directory recognizable without opening a timestamp-named folder.
Within each run directory, VCF-RDFizer writes:
run.json: source identity, resolved input paths, requested workflow configuration, and image selectionsummary.json: final status, wrapper wall time, summary table rows, and an index of every stage report and logmetrics.csv,tsv_metrics.csv, andwrapper_execution_times.csv: compact analysis-ready tables when applicablelogs/wrapper.logandlogs/progress.logtimings/<stage>/...: raw GNUtime -voutput from inside the relevant Docker containerstages/tsv/,stages/conversion/,stages/compression/,stages/compression_operations/,stages/decompression/, andstages/index/, andstages/validation/: structured stage results.compression_operations/preserves the underlying per-RDF operation and validation reports, whilecompression/provides the final output-level summary.stages/partitioned/: the full result handoff from the temporary partitioned-compression container, including every chunk build, merge, validation, workspace free-space sample, exit code, CPU time, and peak RSSreports/index_warnings.jsonandreports/failed_inputs.csvwhen applicablereports/validation/<dataset-id>/: detailed semantic-validation reports (summary.json, query results, preflight checks, and cleanup evidence) when validation is requested
input_vcf_size_bytes in stages/conversion/*.json and metrics.csv is the
uncompressed size of the source VCF, so that ratios are comparable between
plain and compressed inputs. The accompanying input_vcf_size_method records
how it was obtained:
| Method | Meaning |
|---|---|
stat |
Input was not compressed; the on-disk size is the answer |
bgzf |
Exact, summed from a bgzip/BGZF file's block headers with no decompression |
gzip-sample |
Exact; the whole single-member stream fitted in the sampling budget |
gzip-trailer |
Exact; the 32-bit ISIZE trailer resolved against the file's measured compression ratio |
inflate / inflate-shell |
Fallback full decompression pass, used when the file's structure cannot settle the answer (for example concatenated non-BGZF members) |
Only the fallback costs a full pass over the input. Because indexed .vcf.gz
files from bcftools/tabix/htslib are BGZF, the usual case is measured in
milliseconds rather than minutes. The same machinery makes --estimate-size
report a real uncompressed input size instead of an assumed expansion factor;
it says so when it had to fall back to the assumption.
Compression metrics now include per-method:
wall_seconds_*user_seconds_*sys_seconds_*max_rss_kb_*
When HDT or COTTAS is selected, compression also validates the final base
artifact before packaging or RDF cleanup. The validator reads the source
triple count, streams the artifact back through the native decoder, and
requires equal counts. HDT validation also initializes the versioned .hdt.index.*
sidecar. Validation results and source_triples/decoded_triples are stored
in the per-run compression JSON and in the HDT/COTTAS columns of metrics.csv.
Compression fails closed if the artifact cannot be decoded or the counts do
not match. In compression-only mode, the source count is obtained by a
streaming fallback when no upstream conversion metrics are available.
For full runs, a readable HDT whose sidecar index could not be created is
validated with the index check skipped, marked with index_status: "failed",
and reported in index_warnings.json; this allows packaging and later stages
to continue. COTTAS failures are reported the same way, but dependent COTTAS
artifacts are marked as not generated because the COTTAS file itself is not
usable. Explicit standalone --mode index runs remain strict and return a
failure status when regeneration fails.
For partitioned HDT/COTTAS runs, the final method metric reports one
sample-level result while stages/partitioned/<sample>.json retains the full
container-stage history. It includes chunk conversion, merge strategy and
rounds, validation, generated chunk plan, workspace free-space samples, CPU
time, peak RSS, exit codes, and bounded stderr diagnostics. This report is
preserved even when the temporary Docker volume is deleted after a failure.
Metrics may use internal stage names such as hdt_gzip and cottas_brotli.
These correspond to the public combination of --representations and
--artifact-compression; users do not need to pass those compound names.
When --validate is used in full mode, the same metrics.csv row also carries
validation_status, validation_exit_code, validation wall/CPU/RSS timings,
the detailed report path, and the RDF path that was validated. The validation
stage JSON retains the full status and temporary-RDF cleanup metadata.
Chunked Compression
Full mode always uses one of the two aggregate storage modes. Both modes create one logical N-Triples aggregate. The space-optimized mode streams each RMLStreamer part through gzip and deletes that part before processing the next one, so it avoids retaining both the part files and a full uncompressed aggregate.
When HDT or COTTAS is selected, the aggregate is read sequentially and split
into complete N-Triples records. Only one uncompressed chunk is present at a
time: it is consumed by both converters and removed before the next chunk is
read. This is especially important for space-optimized .nt.gz aggregates,
which must not be expanded into a second full raw-RDF copy. HDT chunks are
merged with the Java-free hdtc create command, which accepts existing HDT
inputs; the final HDT index is generated after merging with hdtc index.
COTTAS chunk conversion uses pycottas.rdf2cottas(..., disk=True). The final
COTTAS merge deliberately does not call pycottas.cat: version 1.1.0 runs
its global DISTINCT plus ORDER BY through an unbounded in-memory DuckDB
connection, which can be killed on large condensed graphs. VCF-RDFizer instead
uses a PyArrow k-way merge of the already index-sorted Parquet chunks. It keeps
only a configurable batch from each input, drops adjacent duplicate triples,
and writes the final COTTAS file incrementally—no graph-wide DuckDB hash table
or external-sort spill directory is created. The merge emits processed-source
and distinct-written triple counts in the terminal progress display. If the
stage fails in full mode, the warning contains the failing exit code and, when
available, a stderr_tail, maximum resident set size, and Docker-workspace
free-space samples; the raw RDF remains available for a retry.
After each final HDT/COTTAS base artifact is produced, VCF-RDFizer performs a
streaming decode/count check. This verifies both readability and that the
decoded artifact contains exactly the number of source triples. The check is
performed before .hdt.gz, .hdt.br, .cottas.gz, or .cottas.br packaging,
and before raw RDF cleanup.
Partitioned HDT/COTTAS compression runs in an ephemeral Docker-managed workspace. Temporary RDF chunks, COTTAS conversion scratch data, intermediate representations, and merge files are not written to the output directory. After a successful or failed run, the temporary workspace is removed; only the selected final artifacts and normal run metrics remain on the host. Each COTTAS conversion receives a fresh container-local DuckDB workspace, which is removed as soon as that operation completes. The final streaming merge does not need a DuckDB workspace or a full-data temporary sort file.
HDT merging and index generation use the pinned Rust hdtc 1.1.0 executable,
not hdtCat, hdtSearch.sh, or another Java HDT process. hdtc create
merges the chunk HDTs with disk-backed external sorts, and hdtc index reads
BitmapTriples as a stream to build the object/predicate orderings. This avoids
the JVM heap path that can fail with java.lang.OutOfMemoryError while
producing the same canonical HDT v1-1 sidecar,
<file>.hdt.index.v1-1, used by hdt-java and hdt-cpp.
For standalone index mode, existing versioned sidecars are moved aside while
regeneration runs and restored if indexing fails. Incomplete replacements are
removed before restoration, so a failed or interrupted attempt does not leave
a partial index. Sort runs use /work and are removed when the command exits.
The image defaults both HDT_INDEX_MEMORY_LIMIT and
HDT_MERGE_MEMORY_LIMIT to 512M. The wrapper forwards an explicitly set
host value of either variable into the relevant Docker command.
HDT_INDEX_WORK_ROOT can override the scratch root when invoking
/opt/vcf-rdfizer/ensure_hdt_index.sh directly inside the container.
COTTAS does not expose a separate index sidecar. Its index is part of the
Parquet artifact and is selected when the artifact is written. Standalone
COTTAS index mode writes a new temporary COTTAS file through the same
bounded streaming rewrite with the default spo index, then atomically
replaces the original. This is still index-only from the pipeline's point of
view: it does not rerun VCF-to-RDF conversion or create an RDF output, but it
rewrites the artifact once.
The record-safe chunk plan and per-stage timings are retained in the raw partitioned-compression metrics JSON for diagnostics. The temporary chunk files and guide are not retained as host files.
For the default mapping, multi-sample VCF columns remain compact in
records.tsv. In expanded mode, canonical SampleCall and FormatFieldValue
triples are streamed directly into the final .nt or .nt.gz aggregate rather
than first writing materialized helper rows. In condensed mode, the same input pass
emits shared samples, call matrices, and FORMAT vectors, avoiding both the
helper-table multiplier and the per-sample RDF structural multiplier.
The implementation keeps COTTAS conversion scratch state inside the Docker container and removes temporary unpacked package files when decompression finishes.
Rules
- default rules file:
rules/default_rules.ttl - rules guide:
rules/README.md
Data linking plug-ins
Three installed examples cover declarative dbSNP links (rsid-dbsnp), interval
joins against a synthetic GFF3 bundle (gene-demo), and an Ensembl API resolver
(rsid-ensembl). Add --link <ids> in full mode, or link an existing aggregate
without Docker:
vcf-rdfizer --mode link --rdf ./results/sample/sample.nt.gz \
--link rsid-dbsnp --offline -o ./linked-results
Links go into sample.links.nt; the base graph is unchanged. The
vcf-rdfizer-link companion CLI lists, scaffolds, checks and previews plug-ins.
See Data linking for all three worked examples, reference
and network safeguards, provenance, and the remaining design limitations.
Policy attachment plug-in
vcf-rdfizer-policy attaches ODRL policies to a converted graph and writes one
release view per request, without Docker. A policy can target a file, a region
or a variant, and withholding a record withholds everything it owns (its call,
alleles and genotypes). Selectors are SPARQL declared in Turtle, so adding one
needs no code. check confirms that a view withholds exactly what the policy
says, optionally against the source VCF text:
vcf-rdfizer-policy evaluate --rdf converted/P00*.nt.gz --policy policy.ttl \
--assignee https://example.org/party/alz-consortium --purpose DUO:0000007 -o views/alz
vcf-rdfizer-policy check --view views/alz --rdf converted/P00*.nt.gz \
--policy policy.ttl --vcf P00*.vcf
This is governed release, not anonymization: a released genotype still
identifies the person it came from. See Policy attachment
and the runnable cohort in examples/policy/.
Custom RML Mappings
--rules accepts any RML mapping, so you can change what RDF the pipeline
produces without touching the wrapper. A custom mapping has to honour a small
contract, and vcf-rdfizer-rules (installed alongside vcf-rdfizer) makes it
discoverable and checkable:
vcf-rdfizer-rules columns
Lists the five TSV sources the pipeline generates and every column each one provides, so you know what a mapping can reference.
vcf-rdfizer-rules init -o my_rules.ttl
Writes an annotated copy of the shipped default mapping to start from.
vcf-rdfizer-rules check my_rules.ttl
Validates the mapping before you spend hours on a run. It reports:
- logical-source paths the wrapper cannot rewrite per input,
- referenced columns no generated TSV provides (typos such as
CHROMOSOME), - which
--sample-representationvalues remain usable, - whether the mapping forces the large sample helper tables to be materialized.
Exit code is 0 when the mapping is usable and 1 when it is not; add
--json for scripted use. Then run it:
vcf-rdfizer --mode full -i ./cohort.vcf.gz --rules my_rules.ttl --rdf-storage-mode plain -o ./results
The contract
-
Keep the five
csvw:urlvalues exactly as they are. Full mode processes one VCF at a time and rewrites those literal strings to the per-input file names (/data/tsv/records.tsvbecomes/data/tsv/<sample>.records.tsv, and so on). Any other path is left untouched and will not resolve.Logical source Contents /data/tsv/records.tsvOne row per VCF data line /data/tsv/header_lines.tsvOne row per ##header line/data/tsv/file_metadata.tsvOne row summarising the source VCF /data/tsv/sample_calls.tsvHelper: one row per variant x sample /data/tsv/sample_format_values.tsvHelper: one row per variant x sample x FORMAT key -
Only reference columns the pipeline writes.
vcf-rdfizer-rules columnsis authoritative; a unit test pins those lists to whatsrc/vcf_as_tsv.shactually emits, so they cannot drift. -
Think before consuming the two helper tables. The four built-in sample maps are recognised by the wrapper, which then keeps those tables header-only and streams the genotype RDF itself. A mapping that consumes them in any other way forces them to be materialized in full - the largest intermediate the pipeline can produce - and is rejected in
--sample-representation condensed, which would otherwise emit both genotype representations at once.checkwarns about this explicitly.
The last column of records.tsv is the whitespace-joined sample ids from the
#CHROM line (or SAMPLES when the VCF declares none), so its name varies
per input and it cannot be referenced by a fixed name. Genotype RDF is emitted
by the wrapper from that column instead; see
docs/sample-representation-guide.md.
Repository Layout
VCF-RDFizer is deliberately split into a thin host-side CLI and a set of container-side stages. Nothing on the host walks the RDF itself; it plans work, launches Docker, and reads back the JSON/CSV reports each stage writes.
| Path | Role |
|---|---|
vcf_rdfizer.py |
Host CLI: argument validation, output-collision planning, Docker orchestration, metrics assembly, mode dispatch. Also emits the multi-sample genotype RDF (see below). |
vcf_rdfizer_rules.py |
vcf-rdfizer-rules CLI: scaffold, document, and validate custom RML mappings. |
vcf_rdfizer_link.py, vcf_rdfizer_linking/ |
Linker authoring CLI and shared token/interval/API runner. |
vcf_rdfizer_data/linkers/ |
Packaged examples of all three plug-in tiers. |
vcf_rdfizer_policy.py, vcf_rdfizer_policies/ |
vcf-rdfizer-policy CLI and its select → partition → decide engine. |
vcf_rdfizer_data/policy/ |
The VCF Core profile, a DUO subset and the vcfp: vocabulary. |
vcf_rdfizer_gzip.py |
Uncompressed size of a gzip/BGZF VCF without decompressing it. Used by the host preflight estimate and, inside the image, by run_conversion.sh. |
src/vcf_as_tsv.sh |
VCF -> per-input records/header_lines/file_metadata TSV, in one awk pass. |
src/run_conversion.sh |
Runs RMLStreamer, normalizes Spark part files, merges them into one .nt/.nt.gz aggregate, records conversion metrics. |
src/partitioned_compression.py |
Record-safe RDF chunking plus chunked HDT/COTTAS generation and pairwise merge, inside an ephemeral Docker volume. |
src/cottas_tool.py |
COTTAS convert / merge / reindex / decompress adapter over pycottas, with a bounded-memory streaming merge. |
src/ensure_hdt_index.sh |
Java-free canonical .hdt.index.v1-1 sidecar generation via hdtc, with restore-on-failure. |
src/validate_compression.py |
Round-trip check: decode a .hdt/.cottas artifact and compare its triple count against the source. |
src/validation/ |
Semantic VCF-vs-RDF validation: cyvcf2/bcftools oracle, SPARQL queries per representation, comparison report. |
rules/default_rules.ttl |
Default RML mapping (also shipped as package data in vcf_rdfizer_data/). |
test/ |
unittest suite; the shell/pipeline tests stub java, docker, and friends so no real external tool is needed. |
scripts/release.py |
Version bump + release metadata automation (see scripts/RELEASING.md). |
Genotype RDF is the one deliberate exception to "all data processing happens in
the container": append_expanded_sample_rdf and append_condensed_sample_rdf
in vcf_rdfizer.py append it directly to the aggregate, because the equivalent
RML maps would first have to materialize variants x samples (x FORMAT keys)
helper TSV rows. See docs/sample-representation-guide.md.
Further reading: docs/ is the in-depth documentation set -
how each part of the tool works, why, and where it stops working.
| Document | Covers |
|---|---|
| Architecture | Host/container split, failure policy, pinned toolchain |
| Conversion | VCF -> TSV -> RDF, stage by stage |
| Representations | Compression, HDT/COTTAS, chunking, round-trip checks |
| Output and metrics | Run layout, reports, progress, exit codes |
| Custom RML mappings | The --rules contract in full |
| Sample representations | Expanded vs condensed genotype shapes |
| Validation | The semantic suite, and what it does not test |
| Validation methodology | How coverage is measured, not asserted |
| VCF coverage matrix | Element by element, with the mutation that proves each row |
| CLI reference | Every flag, with constraints and interactions |
| Limitations | Everything the tool cannot do, in one place |
| Roadmap | Planned work, known defects, and rejected options |
| Data linking | Runnable examples of all three plug-in tiers, authoring, safeguards, and provenance |
| Data linking design | Broader proposal and remaining work |
| Policy attachment | Implemented v0.1.0: ODRL policies on files, regions and variants, release views, and checks |
| Privacy policy design | Proposal: ODRL-based granular disclosure control over the graph |
ACKNOWLEDGEMENTS.md- funding and attribution- Releases - release notes per version
Troubleshooting
If Docker permission issues occur, rerun with a Docker-allowed user (or configure Docker group/sudo access on your system).
If COTTAS indexing/merging fails on a very large RDF file, first inspect the
cottas-merge-stream stage in the raw partitioned-compression metrics JSON and the
run wrapper log. An exit_code=-9 means the child was killed by SIGKILL,
which is normally the kernel/Docker OOM killer; a shell wrapper may report the
same event as 137. The warning's stderr_tail, max_rss_kb, and workspace
samples distinguish COTTAS data/schema problems from Docker disk-space errors.
The streaming merge has no graph-wide DuckDB spill area: it reads at most
COTTAS_MERGE_BATCH_ROWS rows per input at a time (default 2048) and writes
the result incrementally. If the host is especially memory-constrained, lower
that value; if the merge is CPU-bound and RAM is available, raise it gradually.
The final COTTAS artifact still needs ordinary output disk space. Rebuild the
image after upgrading so the streaming merge workflow is installed.
If COTTAS is optional for the experiment, rerun with
--representations hdt; the HDT path is independent and can remain the
queryable artifact even when COTTAS cannot fit the available memory. If COTTAS
is required and the streaming merge still receives -9, lower
COTTAS_MERGE_BATCH_ROWS, verify that Docker has enough RAM for the selected
batch size, and
check the host kernel log for an external kill; this is a resource limit, not a
vocabulary or RDF-validity problem.
If HDT compression fails on very large RDF files, use
--rdf-storage-mode space-optimized or --rdf-storage-mode plain with
--hdt-strategy partitioned, then lower --chunk-target-bytes and
--chunk-max-bytes to reduce each converter's working set. Both final HDT
merge and index creation are disk-backed, so ensure the Docker data volume has
enough temporary space for their external sorts. To reduce their bounded
in-memory buffers further, set HDT_MERGE_MEMORY_LIMIT and/or
HDT_INDEX_MEMORY_LIMIT (for example, 512M); lower limits can require more
temporary I/O. Free space in the output filesystem alone does not increase the
Docker volume capacity.
Safe termination:
- Press
Ctrl+Cto interrupt a run. - The wrapper exits with code
130, writes progress torun_metrics/<INPUT_LABEL>__<RUN_ID>/logs/progress.log, and performs best-effort cleanup of tracked intermediates. - Raw RDF cleanup on interrupt follows
--keep-rmlstreamer-rdf-output:- with
--keep-rmlstreamer-rdf-output, raw RDF files are preserved - without it, tracked raw RDF files are removed during interrupt cleanup
- with
Citation
If you use VCF-RDFizer in a publication, please cite:
VCF-RDFizer maintainers. (2026). VCF-RDFizer (Version 3.2.0) [Computer software]. GitHub. https://github.com/ecrum19/VCF-RDFizer
BibTeX:
@software{vcf_rdfizer_2026,
author = {{VCF-RDFizer maintainers}},
title = {VCF-RDFizer},
year = {2026},
version = {3.2.0},
url = {https://github.com/ecrum19/VCF-RDFizer},
note = {Computer software}
}
You can also use the machine-readable citation file: CITATION.cff.
Contributing
Contributions are welcome. If you want to improve VCF-RDFizer:
- Open an issue first for bug reports, feature requests, or design changes.
- Fork the repo and create a feature branch from
main. - Keep changes focused and include/update tests for behavior changes.
- Run the unit tests locally before opening a PR:
python3 -m unittest discover -s test -p "test_*_unit.py" -q
- In your PR, include what changed, why it changed, and how you validated it.
- Use clear commit messages (for Docker publish control, include
[publish-docker]only when intended).
Licensing
- Project license:
LICENSE(MIT) - Third-party runtime notices:
THIRD_PARTY_NOTICES.md
Release files for vcf-rdfizer 3.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vcf_rdfizer-3.2.0.tar.gz | 258.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vcf_rdfizer-3.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 503.1 kB
Release files / vcf_rdfizer-3.2.0.tar.gz
| Download URL | vcf_rdfizer-3.2.0.tar.gz |
|---|---|
| Size | 258.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9ccc4cb68b0373f06750078873814e157d0c5a0196348ef468f78e94cefbeb65
|
|
BLAKE2b-256 checksum How to use checksums |
638684d82b007eb533000195fc11e86f5ddc6c6c0e72c22df1208caef092ac48
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency logRelease files / vcf_rdfizer-3.2.0-py3-none-any.whl
| Download URL | vcf_rdfizer-3.2.0-py3-none-any.whl |
|---|---|
| Size | 245.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1ec23435cb9fb0c44d4b97d568fec6e1985ddf2f6ec131793da89ef9654aabaf
|
|
BLAKE2b-256 checksum How to use checksums |
2a101b03e08f72d28b33ceb3e567d113af20bcede9f0d40e7a9b78d337f33fde
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency log