Skip to main content

Stegobench

Stegobench measures how good a steganography detector is. You give it images whose answers are already known (this one is clean, this one hides a payload), it runs a detector over every one of them, and it reports how often the detector was right, in a versioned JSON document that names the exact bytes the number was measured on.

It will also run those detectors over images of your own. If your question is "is something hidden in these pictures", stegobench examine runs the detectors you name over the files you name and puts their answers in one table. That's the opposite direction from a measurement, and it's held apart from one deliberately: an examination writes no result document and nothing it prints is quotable, because there's no labelled answer for a detector to be right or wrong about. stegobench help scope has the difference in full, and docs/guide/examine-your-own-images.md is the guide.

Most published steganalysis results can't be checked by the people reading them. The corpora usually can't be redistributed, and the discipline that keeps a measurement honest (a clean image and its stego twin must differ in nothing but the payload, a cover and its stego twin must land on the same side of a train/test split) is described in a paper rather than enforced by the tool that produced the number. Stegobench is an attempt at fixing both.

What to install

To score detectors, you need one thing: the stegobench binary.

git clone https://github.com/elementmerc/stegobench
cd stegobench
cargo install --path crates/stegobench-cli

No configuration files and nothing to point at afterwards: the registry, the fixtures and a starter corpus are compiled into the binary as it builds, so the installed command works from any directory on a fresh machine and the clone is only needed to build it.

cargo install stegobench-cli from crates.io is not the same thing. Those files live beside the crate rather than inside it, so a binary built from the published crate alone starts with an empty registry and tells you so on the first command. It's usable, but you have to supply a registry yourself with --registry, and most people want the line above instead.

There's a second, separate program in this repository, pentimento, written in Python. You need it only if you're building a corpus of your own. Most people never will; stegobench fetch gets you a corpus to score without it. generators/README.md covers it if you do.

What it does You need it when
stegobench runs detectors over a labelled corpus and scores them always
pentimento builds, audits, packs and verifies a corpus from scratch only if you're making a corpus

They're two names on purpose. Two different programs sharing one name on one PATH is a worse problem than the one a single entry point would solve.

If all you want is the corpus and your own code, you need neither: Pentimento is published on Internet Archive, HuggingFace and Kaggle. It's a JPEG-decompressed spatial corpus and is not comparable to BOSSbase; read docs/design/pentimento.md before quoting a number from it.

Your first five minutes

Every command below is one you can run, and the output shown is output it printed.

1. What is this? Type the bare name.

$ stegobench
stegobench measures how good a steganography detector is, by running it over images whose answers are already known.
It can also run those detectors over images of your own, which is an answer rather than a measurement (`stegobench help scope`).

  stegobench list detectors    what this installation can run
  stegobench examine <image>   what the detectors say about a file
  stegobench doctor            what is installed, and what it needs
  stegobench help              the reasoning, one topic at a time
  stegobench --help            every command and flag

  stegobench describe pentimento-core    10,000 covers with their licences attached, to quote a number from

2. What can it run? The list is read from the registry, so it can't go stale the way a README can.

$ stegobench list detectors
aletheia-rich MIT               container stegobench/aletheia-rich  stegobench can't drive it
aletheia-rs   MIT               container stegobench/aletheia
aletheia-spa  MIT               container stegobench/aletheia
stegashield   proprietary       container 5iprojects/stegashield  needs STEGASHIELD_LICENCE
stegcore      AGPL-3.0-or-later local     stegcore
stegexpose    none-granted      container stegobench/stegexpose
zsteg         MIT               container stegobench/zsteg

container means pinned by image digest, sandboxed, no network. local means a program you installed, pinned by the hash of the file that ran.

3. Can this machine actually run them? doctor checks each tool's container or binary and prints the line to type for each one that's missing. Drop --no-selftest and it also asks each installed tool to flag a known planted signal and clear a known clean fixture, in both directions, because a tool that answers "stego" to everything passes a one-sided check.

stegobench doctor

4. Get a corpus. The starter corpus is compiled into the binary, so this needs no clone and no network:

stegobench fetch stegobench-starter --tier nano --out ./starter

Six covers and twelve stego images: enough to watch the machinery work, far too few to quote a number from.

5. Check what a run would cost before you commit to it. plan takes the command you'd type, so it can't describe a different run from the one that would happen. It counts the corpus rather than guessing from its size, and says the time is unknown where a tool declares no measured rate.

stegobench plan score --corpus ./starter --detector all

6. Score it. This is the job.

stegobench score --corpus ./starter --detector zsteg --out result.json

It asks the detector about every image, writes each answer as it goes, and emits a validated result-v1 document. Interrupt it and run the same command again and it picks up where it stopped. --detector all runs every registered detector over the same bytes in one pass.

Point it at a folder of your own photographs and it refuses, and explains why rather than reporting a missing file:

$ stegobench score --corpus ./holiday-photos --detector all
./holiday-photos holds 3 image(s) and no records saying which of them hides anything, so there is nothing to be right or wrong about.

`score` measures a detector, which needs the answers in advance. Asking the detectors what they make of these files is `examine`, and what comes back is an answer rather than a measurement.
`stegobench examine ./holiday-photos --detector <name>`
`stegobench help scope`      the difference between the two
`stegobench list detectors`  what is registered here
`stegobench help pairing`    what a corpus has to carry first

It exits 3, a pre-flight refusal, which a script can tell apart from an error and knows not to retry. score needs the labels; running the same detectors over those photographs without them is what examine is for:

stegobench examine ./holiday-photos --detector zsteg --detector stegexpose

One row per image, one column per detector, and no result document, because there's nothing there for a detector to be right or wrong about.

7. Turn results into a table. report renders the conditions into every row, so a figure can't be lifted out without them. Rows are ordered by arm and then detector, never by score: it isn't a ranking and no ranking can be derived from it.

This repository ships 21 real result documents under results/v1, so the command below needs a clone of it. Point report at a folder of your own results and it behaves the same way.

stegobench report results/v1 --format markdown

Every row in those 21 is flagged confounded, which is the point: they're real measurements, and the report says what's wrong with them in the same cell as the number.

The commands

schema, validate, verify, list, describe, embed, examine, doctor, plan, score, metrics, fetch, report, completions, help. Every one is built and tested. stegobench --help has the flags. check and scan, and the other words people guess, are aliases for examine.

Every subcommand accepts --json, which puts machine-readable output on stdout and leaves progress and human text on stderr, so stegobench doctor --json | jq works while you can still watch it run.

Reasoning that doesn't fit on a --help line lives behind stegobench help <topic>: scope is what this measures and what it doesn't, pairing is why a clean image and its stego twin must differ in nothing but the payload, and results is how to judge a number somebody else produced.

What it checks that you'd otherwise have to remember

score enforces the two rules rather than describing them:

  • A corpus that puts a cover and its stego twin on opposite sides of a train and test split stops the run, instead of producing an inflated number nobody could spot afterwards.
  • For every stego image that names the cover it came from, the headers of both are compared. A difference in format, size, bit depth or channel count means something other than the payload changed, and the result says so and names the images. Matching headers prove nothing on their own, so the result distinguishes "looked and found nothing" from "could not look".

Point it at a registered corpus with --corpus-id and it checks that claim rather than taking it. A run is marked named only when the entry declares the digest of its records, the directory matches it, and every image turns out to be the file its own record describes. Otherwise the run is custom, which is a perfectly good run: comparable with itself rather than with somebody else's.

One honest limit today: score reads an unpacked corpus directory, so a packed tier has to be extracted first.

Adding your own detector

A detector or embedder is a TOML file under plugins/registry/, not code wired into the harness:

name = "my-detector"
kind = "detector"
licence = "MIT"

[image]
reference = "ghcr.io/you/my-detector@sha256:..."   # a tag is refused
size_mb = 200
bundled = true            # derived from the size, not chosen: true at or
                          # below 750 MB, and an entry whose flag disagrees
                          # with its own size is refused

[emits]
output = "score"          # or "verdict" if it only says yes or no

[accepts]
formats = ["png"]

[selftest]
must_detect = "fixtures/lsb-0.4bpp.png"   # it must flag this
must_clear = "fixtures/clean.png"          # and clear this

Both fixtures are required: a tool that answers "stego" to everything would otherwise pass a one-sided check.

A tool is registered either as [image] (a container pinned by digest, run with no network, identical bytes on two machines) or [binary] (a program you installed, hashed on your machine), never both.

A detector that answers over HTTP is registered the same way, with host = true and a small adapter script. docs/guide/http-detector.md walks that case end to end. The address of your instance is never written into the entry: an entry naming a loopback or private-network address is refused, because a shipped address is scored against whatever answers on it.

See stegobench help plugins for the full reasoning, and docs/guide/ for registering a corpus, which is a different shape because a corpus is data you point at rather than code you run.

Thirteen tools are registered

Six embedders (steghide, outguess, openstego, stegosuite, hstego, and Stegcore's embed side) and seven detectors (Aletheia's SPA, RS and rich-model estimators, StegExpose, zsteg, plus Stegcore and StegaShield as subjects rather than references). Stegcore appears twice because it does both jobs, and hiding a payload and judging one are different measurements that shouldn't share an identifier.

stegoveritas has never built against current dependencies here, so it isn't registered. F5, jsteg and jphide aren't present either. They're candidates for later, not silently dropped.

Citing this work

If you're citing a number, cite the corpus. A measurement belongs to the images it was taken on, and naming the instrument doesn't tell a reader which bytes produced the figure.

@misc{pentimento,
  author       = {Daniel Iwugo},
  title        = {Pentimento: a matched-pair steganalysis corpus},
  howpublished = {\url{https://github.com/elementmerc/pentimento}},
  note         = {JPEG-decompressed spatial covers. Not comparable to BOSSbase.
                  State the tier and the arm alongside any figure}
}

Say which tier (Nano, Lite or Core) and which arm the number came from. The tiers are strict prefixes of one another, so the tier is part of what was measured.

If you're citing the tool, because you used it, extended it or are comparing methodologies, cite Stegobench. CITATION.cff carries the machine-readable version.

@misc{stegobench,
  author       = {Daniel Iwugo},
  title        = {Stegobench: a reproducible benchmark for steganalysis},
  howpublished = {\url{https://github.com/elementmerc/stegobench}}
}

Neither entry carries a version, a date or a DOI: nothing has been tagged and no archive has minted an identifier yet, so every one of those fields would be a guess. They go in when they're true.

Verifying a download

Every file attached to a release is signed, SHA256SUMS included, because a checksum on its own only tells you the bytes match a list, and whoever serves you a tarball can serve you a matching list. A signature tells you the file came out of this repository's release workflow.

Nothing is released yet, so these describe what a release will carry. Replace v1.2.3 with the tag you downloaded.

With GitHub's gh tool, which checks GitHub's own record of which workflow run built the file:

gh attestation verify stegobench-v1.2.3-x86_64-unknown-linux-musl.tar.gz \
    --repo elementmerc/stegobench

From anywhere, including a mirror or an archived copy, with cosign:

cosign verify-blob \
    --certificate SHA256SUMS.pem \
    --signature SHA256SUMS.sig \
    --certificate-identity-regexp '^https://github\.com/elementmerc/stegobench/\.github/workflows/release\.yml@refs/tags/v' \
    --certificate-oidc-issuer https://token.actions.githubusercontent.com \
    SHA256SUMS

sha256sum -c SHA256SUMS

Verifying SHA256SUMS and then running sha256sum -c covers every file in one step. There's no key to fetch first: the signature carries a short-lived certificate naming the workflow that made it.

The long --certificate-identity-regexp line is the part that matters. It says which repository, which workflow file and which kind of ref the signature has to come from. Drop it and cosign will accept a signature from anybody at all. SECURITY.md says what this does and doesn't prove.

Running the test suites

cargo test --workspace
.venv/bin/python -m unittest discover -s generators -p "test_*.py"
.venv/bin/python -m unittest discover -s tools/release -p "test_*.py"

These are the suites CI runs, so a green result here is a green result there.

Licence

AGPL-3.0-or-later. See LICENSE.

Metadata

Release files for pentimento-corpus 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pentimento-corpus 1.0.0
File Size Uploaded
pentimento_corpus-1.0.0.tar.gz 353.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pentimento-corpus 1.0.0
File Interpreter ABI Platform
pentimento_corpus-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 764.1 kB

Release files / pentimento_corpus-1.0.0.tar.gz

Download URL pentimento_corpus-1.0.0.tar.gz
Size 353.4 kB
Tags Source
SHA-256 checksum
How to use checksums
07a8666b3d58df646a0416a9d801b81657c03682db875ab38cd665f478e2ae23
BLAKE2b-256 checksum
How to use checksums
d4a71d1efd8123248e837e6f6765cdfdcb4bdbec5c7d66f23229aca412381039
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / pentimento_corpus-1.0.0-py3-none-any.whl

Download URL pentimento_corpus-1.0.0-py3-none-any.whl
Size 410.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8fe3e89cab8e480122a8a07faef9b298c5c918782bf3dbc4249ed5393a86f5fe
BLAKE2b-256 checksum
How to use checksums
f68788d01a0f17679ea6a529ad919d2dff6c18eddbe0d1ac687d3b627ff5a565
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page