Skip to main content

Steganodf

PyPi Version PyPi Python Versions

Steganodf hides a secret message inside a dataframe.

Give it a dataframe and a few bytes to hide, steganodf gives you back a dataframe that still looks and reads like the original but carries your secret message. Anyone with the password (or with none, if you did not set one) can read the message back out. Typical use: watermarking a dataset you are about to share, so that a leaked copy can be traced back to the recipient it was issued to.

Hiding data in a table is a trade-off between three properties, and no single scheme wins on all three at once:

  • Capacity — how many bytes fit in the dataset.
  • Robustness — does the message survive what happens to data
  • Invisibility — how much the original data is distorted, and how easily an observer can tell that a watermark is there at all.

Steganodf ships four algorithms that sit at different points of that trade-off.

Usage

Try it in your browser

The web app runs the whole library in your browser, compiled to WebAssembly: nothing is uploaded, the dataframe never leaves your machine — https://dridk.github.io/steganodf/

From the command line

With uv, uvx steganodf ... runs the command line in a throwaway environment, without installing anything. Otherwise install it the usual way:

pip install steganodf
steganodf --version

# Encoding
steganodf encode -m hello host.csv stegano.csv
steganodf encode -m hello host.parquet stegano.parquet
steganodf encode -m hello -p password host.parquet stegano.parquet

# Decoding
steganodf decode stegano.csv
steganodf decode stegano.csv -p password

# Choosing an algorithm, and the carrier column for bitvote / bitghost
steganodf encode -m hello -a bitvote -c price host.csv stegano.csv
steganodf decode -a bitvote -c price stegano.csv

# Decoding without knowing the algorithm: tries all four, names the one that matched
steganodf decode -a auto stegano.csv

The CLI reads and writes .csv and .parquet, and exposes --password, --column and --algorithm. Tuning parameters such as bit_per_row or data_size are Python-only.

From Python

import steganodf
import polars as pl

df = pl.read_parquet("my_dataset.parquet")

# The payload is bytes, not str
watermarked = steganodf.encode(df, b"made by steganodf", password="secret")

# Extract your message from the watermarked dataframe
message = steganodf.decode(watermarked, password="secret")

decode returns empty bytes when no complete message could be recovered, so "no watermark here" and "the watermark says nothing" look alike; decode_details and try_decode (below) both report a success flag instead.

Pass algorithm= to encode and decode. Both sides must agree on the algorithm and on every parameter that affects the framing (password, bit_per_row, data_size, sort_columns…).

# A shuffle-resistant watermark
watermarked = steganodf.encode(df, b"made by steganodf", algorithm="bitvote")
shuffled = watermarked.sample(fraction=1.0, shuffle=True)

steganodf.decode(shuffled, algorithm="bitvote")   # b'made by steganodf'

For anything beyond the defaults, instantiate the class directly:

from steganodf.algorithms import BitPool, BitSync, BitVote, BitGhost

# 8 bits per row instead of 1: ~7x the capacity on a large dataframe
algorithm = BitPool(bit_per_row=8, password="secret")
algorithm.get_max_payload_size(df)          # conservative estimate, in bytes
watermarked = algorithm.encode(df, b"a much longer message ...")

If you receive a watermarked dataframe without being told which algorithm wrote it, pass algorithm="auto". Every algorithm validates what it reads with a CRC32, so one that was not used to encode reports a failure rather than returning garbage; steganodf tries them in turn and stops at the first success.

steganodf.decode(df, algorithm="auto", password="secret")   # b'made by steganodf'

try_decode does the same thing but also tells you which algorithm matched:

steganodf.try_decode(df, password="secret")
# {'payload': b'made by steganodf', 'success': True, 'votes': 100000,
#  'margin_min': 469, 'algorithm': 'bitvote', 'tried': ['bitvote']}

Only the default parameters are tried. bit_per_row, data_size and redundancy must match between encoding and decoding and cannot be guessed, so a dataframe watermarked with BitPool(bit_per_row=8) will not be found by auto decoding — but everything the command line produces will, since it always uses the defaults. The password is not guessed either.

The candidates are tried cheapest first (bitvote, bitghost, bitpool, bitsync), so the slow ones only run once the fast ones have failed.

Algorithms in detail

The three methods:

  • permutation — the message lives in the order of the rows. Nothing is written to the data itself, so the dataset is bit-for-bit the same multiset of rows. The price is that any operation that re-orders the table erases the message.
  • alteration — the message lives in the least significant bits of one numeric column. The values change, by one unit in the last place, and the row order becomes irrelevant.
  • synthesis — the message lives in extra rows that steganodf fabricates and inserts. Existing values are never touched, but the dataset gains records that were not in it.

DETAIL.md covers each algorithm in turn — what it is built on, what it is good at, where it breaks — plus the threat model.

Benchmark

Measured on a 100 000-row, 4-column frame (Int64, Utf8, two Float64) with a 16-byte payload.

Algorithm and settings Method Destroys original data Max capacity (100k rows) Tolerates cell edits Tolerates row deletion Survives sorting Invisibility
bitpool bit_per_row=1 permutation no — rows are only reordered 4.4 kB 2 % 2 % ❌ perfect — not one cell changed
bitpool bit_per_row=8 permutation no — rows are only reordered 29 kB 12 % 20 % ❌ perfect — not one cell changed
bitsync default max_drift (44 s decode) permutation no — rows are only reordered 2.1 kB 1.5 % 4 % ❌ perfect — not one cell changed
bitsync max_drift=256 (5.7 s decode) permutation no — rows are only reordered 2.1 kB 2 % 3 % ❌ perfect — not one cell changed
bitvote data_size=24 alteration 1 ULP on one numeric column 19 B 45 % 95 % ✅ very high — relative change of 2.2e-16
bitvote data_size=260 alteration 1 ULP on one numeric column 255 B 15 % 85 % ✅ very high — relative change of 2.2e-16
bitghost redundancy=8 synthesis +168 fabricated rows 250 B 18 % 50 % ✅ low — the fake rows are visible
bitghost redundancy=32 synthesis +672 fabricated rows 250 B 40 % 80 % ✅ low — the fake rows are visible

Development and releases

steganodf requires Python 3.11+ and is developed with uv:

uv sync --extra dev
uv run pytest --doctest-modules steganodf tests   # or: make test

Every push and pull request on main and dev runs the test suite on Python 3.11 and 3.13 (.github/workflows/tests.yml).

Releasing to PyPI is automated (.github/workflows/publish.yml, trusted publishing over OIDC — no API token). To cut a release:

  1. On dev: bump version in pyproject.toml, run uv lock, commit.
  2. Merge dev into main and push.
  3. git tag 0.3.0 && git push origin 0.3.0

The tag triggers the tests, then a build whose version must match pyproject.toml (the workflow fails otherwise), then the upload to PyPI. make publish prints this checklist with the current version.

Citation

Sacha Schutz, Meganne Souprayen. Watermark tabular datasets with rows permutations and fountain code. TechRxiv. April 28, 2025. DOI: 10.36227/techrxiv.174585796.61215338/v1

@article{schutz2025steganodf,
  title   = {Watermark tabular datasets with rows permutations and fountain code},
  author  = {Schutz, Sacha and Souprayen, Meganne},
  year    = {2025},
  month   = {4},
  journal = {TechRxiv},
  doi     = {10.36227/techrxiv.174585796.61215338/v1},
  url     = {https://www.techrxiv.org/doi/full/10.36227/techrxiv.174585796.61215338/v1}
}

Metadata

Release files for steganodf 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for steganodf 0.3.1
File Size Uploaded
steganodf-0.3.1.tar.gz 45.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for steganodf 0.3.1
File Interpreter ABI Platform
steganodf-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 90.1 kB

Release files / steganodf-0.3.1.tar.gz

Download URL steganodf-0.3.1.tar.gz
Size 45.5 kB
Tags Source
SHA-256 checksum
How to use checksums
fe50e12ce0e02c4c150afa62df678d4490de2744f989e6aeb56dd7fed35f39f8
BLAKE2b-256 checksum
How to use checksums
8943ec07d43d60396362beaac0dbd71974a50bc7015929786f722e31a2417b82
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 30, 2026.

Transparency log

Release files / steganodf-0.3.1-py3-none-any.whl

Download URL steganodf-0.3.1-py3-none-any.whl
Size 44.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8bc6ee1c66789bf05d2111f8d15d440e4b46a278d75654da3fd758dd68c264b7
BLAKE2b-256 checksum
How to use checksums
11c81482d926ba633e3f3e90cf29ce86c8c66aad949a37d5b2acc5b3e1eda735
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 release files

0.3.0

2 release files

0.2.5

1 release file

0.2.4

1 release file

0.2.3

1 release file

0.2.2

1 release file

0.2.1

2 release files

0.2.0

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page