Skip to main content

filelens

Parse cXML, XML, JSON, and messy data files into clean tables — in one command.

Most tools give you an XML tree. filelens gives you rows you can actually use.

filelens is a CLI that helps you understand and clean messy data files. Most of the time, the hardest part is just understanding the file.

Ever opened a file where:

  • headers start on row 6
  • metadata is mixed with data
  • columns are inconsistent

filelens lets you:

  • inspect structure and issues
  • infer a schema
  • convert to a clean table (Parquet)

No config. No guessing. Deterministic output. Built for real-world data engineering workflows.

Quick start

macOS (recommended)

Install with Homebrew:

brew tap kraftaa/filelens
brew trust --formula kraftaa/filelens/filelens
brew install filelens

Homebrew 6 and later require explicit trust for formulae from third-party taps.

Upgrade with:

brew update
brew upgrade filelens

pipx fallback

If Homebrew is unavailable, install the PyPI package in an isolated environment:

pipx install filelens

Upgrade a pipx installation with:

pipx upgrade filelens

Use it:

filelens inspect file.csv
filelens convert file-or-folder --out-dir output/

File contracts in five minutes

Create a deterministic contract from a trusted supplier delivery:

filelens contract create examples/public/contract/trusted_supplier.csv \
  --out supplier.contract.json

The reviewable JSON records contract format version 1, the resolved parser and format, one-based header position, metadata-row count, delimiter, canonical field names, and logical types. Parser overrides are supported while creating a contract, for example --parser json or --parser cxml --cxml-mode both; the resolved parser and cXML mode are stored and reused during checks.

Check a compatible delivery:

filelens contract check examples/public/contract/compatible_supplier.csv \
  --against supplier.contract.json
Contract check: COMPATIBLE
- [INFO] numeric_widening: field "unit_price" widened compatibly from int to float
  expected: int; observed: float
- [WARNING] additional_field: additional field "notes" was observed
  expected: absent; observed: string
Summary: 0 breaking, 1 warning, 1 informational

Check a breaking delivery:

filelens contract check examples/public/contract/breaking_supplier.csv \
  --against supplier.contract.json
Contract check: VIOLATION
- [BREAKING] missing_field: expected field "quantity" is missing
  expected: int; observed: missing
- [BREAKING] missing_field: expected field "unit_price" is missing
  expected: int; observed: missing
Summary: 2 breaking, 0 warning, 0 informational

Add --json for a machine-readable report. Reports contain field/path names and expected versus observed types, but never source rows or field values.

Compatibility and exit policy

Change Severity Exit status
Missing expected field/path breaking 1
Incompatible logical type breaking 1
Header position, delimiter, or input format changed breaking 1
Additional field/path warning 0 when no breaking finding exists
Compatible int to float widening informational 0
Invalid invocation or unreadable/invalid contract operational error 2
Incoming file cannot be parsed breaking operational error 2

Parse failure uses exit 2, not 1: FileLens could not evaluate compatibility, so it must not claim that a valid observed structure violated the contract. The policy is deterministic, column order is ignored, and contract check never modifies the incoming file.

Use the status in CI or a shell script:

filelens contract check incoming.csv --against supplier.contract.json
status=$?

case "$status" in
  0) echo "compatible" ;;
  1) echo "contract violation" >&2; exit 1 ;;
  2) echo "contract could not be evaluated" >&2; exit 2 ;;
esac

Example

Before

Metadata + mixed rows + unclear structure:

Metadata: Device=LabX
Date: 2024-01-01

Sample ID,Value,Unit
S1,0.45,mg/mL
S2,0.50,mg/mL

Inspect:

filelens inspect sample.csv

Inspect messy CSV

Output:

Detected:
- header row: 4
- metadata rows: 1-3
- columns: 3

Warnings:
- none

One command

filelens convert sample.cxml --out sample.parquet

Convert cXML both

What it does:

  • detects structure
  • skips metadata
  • infers schema
  • writes sample.parquet

After

Clean table:

sample_id | value | unit
S1        | 0.45  | mg/mL
S2        | 0.50  | mg/mL

Read in pandas:

import pandas as pd
df = pd.read_parquet("sample.parquet")
df.head()

That's it

For most use cases, you only need:

filelens inspect file.csv
filelens inspect file.cxml
filelens convert file-or-folder --out-dir output/

Everything below is optional (advanced formats, dbt integration, pipelines).

Default workflow

Golden path:

filelens convert <file-or-folder> --out-dir output/

What it does:

  • auto-detects parser from extension/content
  • keeps cXML default mode as mapped (canonical columns)
  • writes parquet output files
  • appends _source_file, _source_kind, _record_id columns
  • writes sidecar report: output/_filelens_report.json (schema + warnings + status per file)

If your cXML needs extra nested fields, use:

filelens convert po.cxml --parser cxml --cxml-mode both --out-dir output/

Folder-specific command:

filelens batch examples/public/trade --out-dir output/trade

batch prints:

  • files processed / succeeded / failed
  • total rows written
  • top warnings

Supported inputs

Supports common messy data formats used in analytics and healthcare.

  • Excel / CSV (messy tabular files): .xlsx, .xlsm, .xls, .csv, .tsv, .psv, .txt
  • JSON (nested data): .json, .ndjson
  • XML (including cXML / CDA / NAACCR): .xml, .cxml, .xcml
  • HL7 (basic extraction): .hl7, .msg
  • Compressed text variants: .gz wrappers for supported text formats

Design

  • deterministic
  • no config required
  • optimized for messy real-world files

When to use filelens

  • You opened a file and do not understand its structure
  • Your Excel export has metadata rows and broken headers
  • You need to convert XML/JSON into a table quickly
  • You want clean input for dbt or a data warehouse

More examples

Inspect:

filelens inspect data/file.xlsx
filelens inspect data/order.cxml
filelens inspect data/patient-example.json
filelens inspect data/oru_r01.msg
filelens inspect data/clinical.xml
filelens inspect data/patient-example.ttl
filelens inspect data/patient-example.ttl.html

Schema:

filelens schema data/file.xlsx
filelens schema data/patient-example.json --parser fhir

Convert:

filelens convert data/file.xlsx --out data/file.parquet
filelens convert data/order.cxml --out data/order.parquet
filelens convert data/nested_lab_result.json --out data/nested_lab_result.parquet
filelens convert data/oru_r01.msg --out data/oru_r01.parquet
filelens convert data/patient-example.ttl --out data/patient-example.ttl.parquet

Optional parser override:

filelens inspect data/file.xml --parser cda
filelens inspect data/file.json --parser json
filelens inspect data/file.json --parser fhir
filelens inspect data/file.msg --parser hl7
filelens inspect data/file.ttl --parser rdf

CXML extraction mode:

cXML mode controls which columns are emitted:

  • mapped (default): canonical columns only (order_id, line_number, quantity, ...)
  • auto: extracted path-based columns only (x_*)
  • both: union of mapped + auto

If you do not pass --cxml-mode, filelens uses mapped.

cXML semantic type contract

cXML output uses a semantic contract so files from the same family do not change Parquet types based on the values present in one file:

Fields Stable type Policy
order_id, payload_id, notice_id, quote_id, line_number, supplier_part_id, supplier_part_auxiliary_id, classification, classification_domain, item_classification, address_id, ship_to_postal_code, bill_to_postal_code string Identifiers and codes retain leading zeros and alphanumeric values.
order_date, quote_date, payload_timestamp, requested_delivery_date string The complete source value is retained, including ISO-8601 time and UTC offset.
quantity, unit_price, line_total, shipping_amount, discount_amount, tax_amount float Whole and decimal values share one Parquet numeric type (Float64).
Other canonical fields and dynamic extrinsics string Values remain lossless unless a stable semantic numeric meaning is known.
All auto-extracted x_* fields string Path extraction is intentionally lexical; the XML path alone is not enough to infer a safe semantic type.

The x_* rule also applies in both mode: a canonical quantity is a float, while the corresponding auto-extracted quantity remains a string. FileLens does not apply organization-specific price, currency, or catalog validation because those rules cannot be derived from the XML document alone.

# curated canonical fields only
filelens schema data/order.cxml --parser cxml --cxml-mode mapped

# path-based auto-captured fields only (x_* columns)
filelens schema data/order.cxml --parser cxml --cxml-mode auto

# both canonical + path-based fields
filelens convert data/order.cxml --parser cxml --cxml-mode both --out data/order.parquet

If running from source, use ./target/release/filelens instead of filelens.

Optional: use with dbt

filelens outputs Parquet files that can be loaded into warehouses and modeled with dbt.

Use it in this order:

  1. Convert files to parquet.
  2. Load parquet into Postgres raw.filelens_lines.
  3. Run dbt models.
  4. Query typed marts.

Setup env vars:

export PGHOST=localhost
export PGPORT=5432
export PGUSER=...
export PGPASSWORD=...
export PGDATABASE=postgres
export DBT_PROFILES_DIR=dbt

One-command local pipeline (public examples only):

scripts/auto_load_and_run_dbt.sh --parquet-glob "$PWD/output/public/**/*.parquet" --full-refresh

What this command does:

  • loads parquet into raw.filelens_lines
  • syncs raw into raw_procurement and raw_clinical
  • runs staging models
  • runs marts (including typed marts)
  • runs tests
  • prints row counts and next query hints

Which tables to query:

  • analytics_marts.fct_procurement_lines for procurement analytics
  • analytics_marts.fct_fhir_resources for FHIR analytics
  • analytics_marts.fct_naaccr_cases for NAACCR analytics
  • analytics_marts.fct_record_attributes for generic key/value search across all extracted attributes

analytics_registry.idx_filelens_records is a cross-format registry/index table (lineage + canonical fields). It is not the primary end-user analytics table.

Why keep raw -> internal -> marts:

  • raw: ingestion/debug layer (what got loaded)
  • analytics_internal: normalization layer (map parser-specific columns into stable canonical fields)
  • marts: consumption layer (deduped and typed tables for analysts/apps)

Example consumer queries:

select * from analytics_marts.fct_procurement_lines limit 20;
select * from analytics_marts.fct_fhir_resources limit 20;
select * from analytics_marts.fct_naaccr_cases limit 20;
select * from analytics_marts.fct_record_attributes limit 20;

Trace NAACCR attributes back to original source ids:

select
  source_file,
  record_key,
  attribute_scope,
  attribute_source_id,
  attribute_name,
  attribute_value
from analytics_marts.fct_record_attributes
where source_kind = 'naaccr'
  and attribute_source_id in ('grade', 'patientidnumber', 'tumorrecordnumber')
limit 20;

Examples

See examples/ for real sample inputs:

  • procurement (cXML / xCML)
  • healthcare (FHIR, HL7, CDA, NAACCR)
  • RDF/Turtle (.ttl, .ttl.html)
  • messy CSV/TSV/PSV/TXT
  • hard cXML edge-case fixtures for parser testing: examples/hard/cxml

Workflow

What this workflow does:

  • converts only examples/public/** into Parquet under output/public/**
  • loads only output/public/**/*.parquet into Postgres raw.filelens_lines
  • runs dbt staging + marts with --full-refresh (and tests)
  • does not include non-public example paths unless you change the command

Why --full-refresh in this demo workflow:

  • it rebuilds marts from scratch so the demo is deterministic after parser/model changes
  • it avoids stale incremental state while iterating locally
  • for recurring production loads, omit --full-refresh and use incremental dbt runs
scripts/convert_inputs.sh --input-dir examples/public --output-dir output/public

export PGHOST=localhost
export PGPORT=5432
export PGUSER=...
export PGPASSWORD=...
export PGDATABASE=postgres
export DBT_PROFILES_DIR=dbt

scripts/auto_load_and_run_dbt.sh --parquet-glob "$PWD/output/public/**/*.parquet" --full-refresh

scripts/convert_inputs.sh uses ./target/release/filelens by default. Use --bin to point to another binary path.

scripts/convert_inputs.sh is non-strict by default (skips failures and continues). Add --strict to fail on first conversion error.

Advanced formats

  • RDF/Turtle (.ttl, .rdf) — experimental support
  • HTML pages containing RDF/Turtle <pre> blocks (for example *.ttl.html)

Why not pandas?

pandas can read files, but it does not:

  • detect likely header/metadata layout
  • explain quality issues up front
  • normalize mixed file families with one deterministic CLI pass

filelens is focused on that first cleanup step before your pipeline.

Metadata

Release files for filelens 0.1.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for filelens 0.1.4
File Size Uploaded
filelens-0.1.4.tar.gz 627.0 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for filelens 0.1.4
File Interpreter ABI Platform
filelens-0.1.4-py3-none-win_amd64.whl Python 3 none Windows x86-64 Details
filelens-0.1.4-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl Python 3 none Linux glibc 2.17+ x86-64 Details
filelens-0.1.4-py3-none-macosx_11_0_arm64.whl Python 3 none macOS 11.0+ ARM64 Details

Total release size: 19.2 MB

Release files / filelens-0.1.4.tar.gz

Download URL filelens-0.1.4.tar.gz
Size 627.0 kB
Tags Source
SHA-256 checksum
How to use checksums
a0f453c7af8d3ce5472ad393242d620fdbdfb82760e5d71a490a68c2dbbca17d
BLAKE2b-256 checksum
How to use checksums
d6607e9e3cdba3c94e12a6697dbb0a8a736b9d48b44ec857120f69f44618eea4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.

Transparency log

Release files / filelens-0.1.4-py3-none-win_amd64.whl

Download URL filelens-0.1.4-py3-none-win_amd64.whl
Size 6.6 MB
Tags Python 3 Windows x86-64
SHA-256 checksum
How to use checksums
c18e05a1d115c9723c2eb582806c4e626de611c0a604abe240ae06fa6f6fee1c
BLAKE2b-256 checksum
How to use checksums
adcb0b15ebd06edb1512f79a2750511bce20072405896838be90b9b218241716
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.

Transparency log

Release files / filelens-0.1.4-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL filelens-0.1.4-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 6.4 MB
Tags Linux glibc 2.17+ x86-64 Python 3
SHA-256 checksum
How to use checksums
7b8495c54ee74b0379aa946de85ee4be5f567c54aa9e433b1176a2dac6accb97
BLAKE2b-256 checksum
How to use checksums
224211fbb46573e7f2270dadcf120e4a2a1948c7bcc4ebee728ec15c73956a1a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.

Transparency log

Release files / filelens-0.1.4-py3-none-macosx_11_0_arm64.whl

Download URL filelens-0.1.4-py3-none-macosx_11_0_arm64.whl
Size 5.6 MB
Tags Python 3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
4fd0c6e97332dd92316ad35e836a06d1b965787b53736f88fa5c27c30bb35620
BLAKE2b-256 checksum
How to use checksums
dd0884bd66c3c98821c994b3fc516f187da2f5fc4f4d5cd52056438e6e8f217d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.4 This release

4 release files

0.1.3

4 release files

0.1.2

4 release files

0.1.1

4 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page