Stegobench
Stegobench measures how good a steganography detector is. You give it images whose answers are already known (this one is clean, this one hides a payload), it runs a detector over every one of them, and it reports how often the detector was right, in a versioned JSON document that names the exact bytes the number was measured on.
It will also run those detectors over images of your own. If your question
is "is something hidden in these pictures", stegobench examine runs the
detectors you name over the files you name and puts their answers in one
table. That's the opposite direction from a measurement, and it's held apart
from one deliberately: an examination writes no result document and nothing it
prints is quotable, because there's no labelled answer for a detector to be
right or wrong about. stegobench help scope has the difference in full, and
docs/guide/examine-your-own-images.md is the guide.
Most published steganalysis results can't be checked by the people reading them. The corpora usually can't be redistributed, and the discipline that keeps a measurement honest (a clean image and its stego twin must differ in nothing but the payload, a cover and its stego twin must land on the same side of a train/test split) is described in a paper rather than enforced by the tool that produced the number. Stegobench is an attempt at fixing both.
What to install
To score detectors, you need one thing: the stegobench binary.
git clone https://github.com/elementmerc/stegobench
cd stegobench
cargo install --path crates/stegobench-cli
No configuration files and nothing to point at afterwards: the registry, the fixtures and a starter corpus are compiled into the binary as it builds, so the installed command works from any directory on a fresh machine and the clone is only needed to build it.
cargo install stegobench-cli from crates.io is not the same thing. Those
files live beside the crate rather than inside it, so a binary built from the
published crate alone starts with an empty registry and tells you so on the
first command. It's usable, but you have to supply a registry yourself with
--registry, and most people want the line above instead.
There's a second, separate program in this repository, pentimento, written
in Python. You need it only if you're building a corpus of your own. Most
people never will; stegobench fetch gets you a corpus to score without it.
generators/README.md covers it if you do.
| What it does | You need it when | |
|---|---|---|
stegobench |
runs detectors over a labelled corpus and scores them | always |
pentimento |
builds, audits, packs and verifies a corpus from scratch | only if you're making a corpus |
They're two names on purpose. Two different programs sharing one name on one PATH is a worse problem than the one a single entry point would solve.
If all you want is the corpus and your own code, you need neither:
Pentimento is published on
Internet Archive, HuggingFace and Kaggle. It's a JPEG-decompressed spatial
corpus and is not comparable to BOSSbase; read docs/design/pentimento.md
before quoting a number from it.
Your first five minutes
Every command below is one you can run, and the output shown is output it printed.
1. What is this? Type the bare name.
$ stegobench
stegobench measures how good a steganography detector is, by running it over images whose answers are already known.
It can also run those detectors over images of your own, which is an answer rather than a measurement (`stegobench help scope`).
stegobench list detectors what this installation can run
stegobench examine <image> what the detectors say about a file
stegobench doctor what is installed, and what it needs
stegobench help the reasoning, one topic at a time
stegobench --help every command and flag
stegobench describe pentimento-core 10,000 covers with their licences attached, to quote a number from
2. What can it run? The list is read from the registry, so it can't go stale the way a README can.
$ stegobench list detectors
aletheia-rich MIT container stegobench/aletheia-rich stegobench can't drive it
aletheia-rs MIT container stegobench/aletheia
aletheia-spa MIT container stegobench/aletheia
stegashield proprietary container 5iprojects/stegashield needs STEGASHIELD_LICENCE
stegcore AGPL-3.0-or-later local stegcore
stegexpose none-granted container stegobench/stegexpose
zsteg MIT container stegobench/zsteg
container means pinned by image digest, sandboxed, no network. local means
a program you installed, pinned by the hash of the file that ran.
3. Can this machine actually run them? doctor checks each tool's
container or binary and prints the line to type for each one that's missing.
Drop --no-selftest and it also asks each installed tool to flag a known
planted signal and clear a known clean fixture, in both directions, because a
tool that answers "stego" to everything passes a one-sided check.
stegobench doctor
4. Get a corpus. The starter corpus is compiled into the binary, so this needs no clone and no network:
stegobench fetch stegobench-starter --tier nano --out ./starter
Six covers and twelve stego images: enough to watch the machinery work, far too few to quote a number from.
5. Check what a run would cost before you commit to it. plan takes the
command you'd type, so it can't describe a different run from the one that
would happen. It counts the corpus rather than guessing from its size, and
says the time is unknown where a tool declares no measured rate.
stegobench plan score --corpus ./starter --detector all
6. Score it. This is the job.
stegobench score --corpus ./starter --detector zsteg --out result.json
It asks the detector about every image, writes each answer as it goes, and
emits a validated result-v1 document. Interrupt it and run the same command
again and it picks up where it stopped. --detector all runs every registered
detector over the same bytes in one pass.
Point it at a folder of your own photographs and it refuses, and explains why rather than reporting a missing file:
$ stegobench score --corpus ./holiday-photos --detector all
./holiday-photos holds 3 image(s) and no records saying which of them hides anything, so there is nothing to be right or wrong about.
`score` measures a detector, which needs the answers in advance. Asking the detectors what they make of these files is `examine`, and what comes back is an answer rather than a measurement.
`stegobench examine ./holiday-photos --detector <name>`
`stegobench help scope` the difference between the two
`stegobench list detectors` what is registered here
`stegobench help pairing` what a corpus has to carry first
It exits 3, a pre-flight refusal, which a script can tell apart from an error
and knows not to retry. score needs the labels; running the same detectors
over those photographs without them is what examine is for:
stegobench examine ./holiday-photos --detector zsteg --detector stegexpose
One row per image, one column per detector, and no result document, because there's nothing there for a detector to be right or wrong about.
7. Turn results into a table. report renders the conditions into every
row, so a figure can't be lifted out without them. Rows are ordered by arm and
then detector, never by score: it isn't a ranking and no ranking can be
derived from it.
This repository ships 21 real result documents under results/v1, so the
command below needs a clone of it. Point report at a folder of your own
results and it behaves the same way.
stegobench report results/v1 --format markdown
Every row in those 21 is flagged confounded, which is the point: they're
real measurements, and the report says what's wrong with them in the same
cell as the number.
The commands
schema, validate, verify, list, describe, embed, examine,
doctor, plan, score, metrics, fetch, report, completions,
help. Every one is built and tested. stegobench --help has the flags.
check and scan, and the other words people guess, are aliases for
examine.
Every subcommand accepts --json, which puts machine-readable output on
stdout and leaves progress and human text on stderr, so
stegobench doctor --json | jq works while you can still watch it run.
Reasoning that doesn't fit on a --help line lives behind
stegobench help <topic>: scope is what this measures and what it doesn't,
pairing is why a clean image and its stego twin must differ in nothing but
the payload, and results is how to judge a number somebody else produced.
What it checks that you'd otherwise have to remember
score enforces the two rules rather than describing them:
- A corpus that puts a cover and its stego twin on opposite sides of a train and test split stops the run, instead of producing an inflated number nobody could spot afterwards.
- For every stego image that names the cover it came from, the headers of both are compared. A difference in format, size, bit depth or channel count means something other than the payload changed, and the result says so and names the images. Matching headers prove nothing on their own, so the result distinguishes "looked and found nothing" from "could not look".
Point it at a registered corpus with --corpus-id and it checks that claim
rather than taking it. A run is marked named only when the entry declares
the digest of its records, the directory matches it, and every image turns out
to be the file its own record describes. Otherwise the run is custom, which
is a perfectly good run: comparable with itself rather than with somebody
else's.
One honest limit today: score reads an unpacked corpus directory, so a
packed tier has to be extracted first.
Adding your own detector
A detector or embedder is a TOML file under plugins/registry/, not code
wired into the harness:
name = "my-detector"
kind = "detector"
licence = "MIT"
[image]
reference = "ghcr.io/you/my-detector@sha256:..." # a tag is refused
size_mb = 200
bundled = true # derived from the size, not chosen: true at or
# below 750 MB, and an entry whose flag disagrees
# with its own size is refused
[emits]
output = "score" # or "verdict" if it only says yes or no
[accepts]
formats = ["png"]
[selftest]
must_detect = "fixtures/lsb-0.4bpp.png" # it must flag this
must_clear = "fixtures/clean.png" # and clear this
Both fixtures are required: a tool that answers "stego" to everything would otherwise pass a one-sided check.
A tool is registered either as [image] (a container pinned by digest, run
with no network, identical bytes on two machines) or [binary] (a program you
installed, hashed on your machine), never both.
A detector that answers over HTTP is registered the same way, with
host = true and a small adapter script. docs/guide/http-detector.md walks
that case end to end. The address of your instance is never written into the
entry: an entry naming a loopback or private-network address is refused,
because a shipped address is scored against whatever answers on it.
See stegobench help plugins for the full reasoning, and
docs/guide/ for registering a corpus, which is a different shape because
a corpus is data you point at rather than code you run.
Thirteen tools are registered
Six embedders (steghide, outguess, openstego, stegosuite, hstego, and Stegcore's embed side) and seven detectors (Aletheia's SPA, RS and rich-model estimators, StegExpose, zsteg, plus Stegcore and StegaShield as subjects rather than references). Stegcore appears twice because it does both jobs, and hiding a payload and judging one are different measurements that shouldn't share an identifier.
stegoveritas has never built against current dependencies here, so it isn't
registered. F5, jsteg and jphide aren't present either. They're candidates for
later, not silently dropped.
Citing this work
If you're citing a number, cite the corpus. A measurement belongs to the images it was taken on, and naming the instrument doesn't tell a reader which bytes produced the figure.
@misc{pentimento,
author = {Daniel Iwugo},
title = {Pentimento: a matched-pair steganalysis corpus},
howpublished = {\url{https://github.com/elementmerc/pentimento}},
note = {JPEG-decompressed spatial covers. Not comparable to BOSSbase.
State the tier and the arm alongside any figure}
}
Say which tier (Nano, Lite or Core) and which arm the number came from. The tiers are strict prefixes of one another, so the tier is part of what was measured.
If you're citing the tool, because you used it, extended it or are
comparing methodologies, cite Stegobench. CITATION.cff carries the
machine-readable version.
@misc{stegobench,
author = {Daniel Iwugo},
title = {Stegobench: a reproducible benchmark for steganalysis},
howpublished = {\url{https://github.com/elementmerc/stegobench}}
}
Neither entry carries a version, a date or a DOI: nothing has been tagged and no archive has minted an identifier yet, so every one of those fields would be a guess. They go in when they're true.
Verifying a download
Every file attached to a release is signed, SHA256SUMS included, because a
checksum on its own only tells you the bytes match a list, and whoever serves
you a tarball can serve you a matching list. A signature tells you the file
came out of this repository's release workflow.
Nothing is released yet, so these describe what a release will carry. Replace
v1.2.3 with the tag you downloaded.
With GitHub's gh tool, which checks GitHub's own record of which
workflow run built the file:
gh attestation verify stegobench-v1.2.3-x86_64-unknown-linux-musl.tar.gz \
--repo elementmerc/stegobench
From anywhere, including a mirror or an archived copy, with cosign:
cosign verify-blob \
--certificate SHA256SUMS.pem \
--signature SHA256SUMS.sig \
--certificate-identity-regexp '^https://github\.com/elementmerc/stegobench/\.github/workflows/release\.yml@refs/tags/v' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com \
SHA256SUMS
sha256sum -c SHA256SUMS
Verifying SHA256SUMS and then running sha256sum -c covers every file in
one step. There's no key to fetch first: the signature carries a short-lived
certificate naming the workflow that made it.
The long --certificate-identity-regexp line is the part that matters. It
says which repository, which workflow file and which kind of ref the signature
has to come from. Drop it and cosign will accept a signature from anybody at
all. SECURITY.md says what this does and doesn't prove.
Running the test suites
cargo test --workspace
.venv/bin/python -m unittest discover -s generators -p "test_*.py"
.venv/bin/python -m unittest discover -s tools/release -p "test_*.py"
These are the suites CI runs, so a green result here is a green result there.
Licence
AGPL-3.0-or-later. See LICENSE.
Metadata
Release files for pentimento-corpus 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pentimento_corpus-1.0.0.tar.gz | 353.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pentimento_corpus-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 764.1 kB
Release files / pentimento_corpus-1.0.0.tar.gz
| Download URL | pentimento_corpus-1.0.0.tar.gz |
|---|---|
| Size | 353.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
07a8666b3d58df646a0416a9d801b81657c03682db875ab38cd665f478e2ae23
|
|
BLAKE2b-256 checksum How to use checksums |
d4a71d1efd8123248e837e6f6765cdfdcb4bdbec5c7d66f23229aca412381039
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / pentimento_corpus-1.0.0-py3-none-any.whl
| Download URL | pentimento_corpus-1.0.0-py3-none-any.whl |
|---|---|
| Size | 410.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8fe3e89cab8e480122a8a07faef9b298c5c918782bf3dbc4249ed5393a86f5fe
|
|
BLAKE2b-256 checksum How to use checksums |
f68788d01a0f17679ea6a529ad919d2dff6c18eddbe0d1ac687d3b627ff5a565
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log