Skip to main content

imgtrail

PyPI Python CI Licence Checked with mypy Ruff

Find out where else on the web your own photos show up.

Point it at your Instagram data export. It hashes every photo, collapses the near-duplicates so you never pay to search the same picture twice, runs each unique one through reverse image search, and then downloads every candidate and compares it against your original before putting it in the report. What you get back is a list you can trust, not a pile of URLs.

imgtrail scan ~/Downloads/instagram-export.zip --dry-run
imgtrail scan ~/Downloads/instagram-export.zip
imgtrail report --open

Two engines, because one index is not the web

Reverse image search is not one thing. Cloud Vision's WEB_DETECTION and the Google Lens you get by dragging a photo into the search box are different indexes, and they disagree.

Measured on one photograph — a drone shot of a castle:

Vision Lens
Verified copies found 19, across nine Facebook pages 5
Pinterest never mentions it, in any field 3 boards, verified
Cost $3.50 / 1,000, first 1,000 free monthly subscription, 250 free monthly

Neither is a superset of the other. So both are here:

imgtrail scan EXPORT                    # vision: cheap and wide, the default
imgtrail scan EXPORT --engine lens      # lens: a different index, for the photos you care about

A photo already searched by one engine is still new to the other, so --engine lens --limit 20 spends twenty searches on the twenty you have not covered yet. Results from both land in the same report and are verified the same way.

What it finds, and what it doesn't

It searches Google's index, so it finds your photos on blogs, news sites, Pinterest, Tumblr, forums, scraper mirrors and shops that lifted your pictures — and on public Facebook posts and groups, which is where a photograph of somewhere recognisable tends to end up.

It will not find a repost on another Instagram account. Instagram blocks crawling of post images, so they aren't in anyone's index — the only way such a repost surfaces here is indirectly, via one of the many "Instagram viewer" mirror sites that are indexed. Telegram, WhatsApp, TikTok and private accounts are invisible to it too. If your question is "is someone reposting me inside Instagram", this is the wrong tool and there isn't a good one.

What it cannot prove, it says so. A candidate the search named and the site would not serve — TikTok, Facebook's lookaside — is listed apart under "found, but not verified": a place to go and look, not a claim. Pages named with no image at all are not kept: of the page-level claims that could be checked against the original, 9.6% held.

Your own Facebook page is not filtered out: from a group post there is no telling whose it is. If you cross-post everything from Instagram, --ignore-domain facebook.com.

Install

pip install imgtrail

Getting your photos

Instagram → Settings → Accounts Centre → Your information and permissions → Download your information. Ask for JSON, high quality. You'll get a ZIP; hand it straight to imgtrail scan. No scraping, nothing against the terms of service, no rate limits.

A plain folder of images works just as well.

Getting an API key

Vision, the default. Create a project at console.cloud.google.com, enable the Cloud Vision API, then Credentials → Create credentials → API key.

export IMGTRAIL_API_KEY=AIza...

The first 1,000 images each month are free, then $3.50 per 1,000. A typical profile costs nothing.

Lens, optional, through SerpApi.

export SERPAPI_KEY=...

250 searches a month on the free plan, which is enough for the way it is meant to be used: a second opinion on the photos you care about, not a second pass over everything. Beyond that it is a subscription, around $15 per 1,000 — four times Vision. Your photos are uploaded, never published to a URL.

Run --dry-run first with either one and it will tell you exactly how many searches it would make and what they would cost before spending anything.

How the verification works

Reverse image search returns a lot of near-misses. For every candidate, imgtrail downloads the image and compares perceptual hashes against your original:

Hamming distance Verdict Meaning
≤ 8 confirmed The same image, possibly recompressed
≤ 16 likely Cropped, filtered or heavily edited
> 16 rejected Not your photo

Only confirmed and likely reach the report. visuallySimilarImages is dropped entirely — it means "semantically alike", not "this is your photo", and it drowns the report in noise.

Commands

imgtrail scan SOURCE          index, dedupe, search and verify — resumable
  --engine vision|lens        which index to search (default vision)
  --dry-run                   count the searches and their cost, call nothing
  --limit N                   search at most N unique photos
  --threshold N               pHash distance for "same photo" (default 6)
  --ignore-domain DOMAIN      exclude a domain from results (repeatable)
  --again                     search everything again, paying for it again
  --no-verify                 skip the download-and-compare pass
imgtrail reparse              re-read the stored answers under today's filters
  --ignore-domain DOMAIN      exclude a domain from results (repeatable)
imgtrail trace PHOTO          everything the engine said about one photo, and its fate
imgtrail report --open        build the HTML report and open it
imgtrail status               what's in the database so far

State lives in ./imgtrail-data. Everything is idempotent: re-running scan searches only what it hasn't searched before, so an interrupted run costs nothing to resume. A scan stays inside the source you point it at — the database may hold other folders, and they are not what you asked to search.

trace answers "why is my photo not in the report" without reading the source: it prints what the engine said about that one photo and what each filter did with it. It reads the archive, so it costs nothing.

Every answer a search engine gives is kept verbatim. Filtering is a pile of judgement calls — which platforms are yours, which candidates are worth downloading — and at least one of them is wrong. reparse re-reads what you already paid for under the current rules, without calling anything, so correcting a filter costs nothing.

Privacy

Your photos are sent to Google Cloud Vision, and nowhere else. Nothing is uploaded to any server of mine — there isn't one. The database, the extracted export and the report all stay on your machine.

Architecture

Ports and adapters, sized to the problem: the rules sit in the middle and know nothing about Google, SQLite or HTTP, so swapping a search backend touches exactly one file.

domain.py     fingerprints, grouping, verdicts, what counts as "your own platform"
              — pure; no I/O, no SQL, no network
ports.py      the boundaries: PhotoSource, ImageLoader, SearchEngine, ImageFetcher,
              PhotoRepository, MatchRepository, ReportWriter
services.py   the use cases: index, plan, search, verify, report
adapters/     the details: sqlite_repository, vision, http_fetcher, local_files, html_report
cli.py        the composition root — the one module that knows every layer

Adding TinEye or Yandex means writing one SearchEngine and wiring it in cli.py. Nothing in domain.py or services.py changes.

Development

uv sync --all-groups
uv run pytest             # 85 tests, no network, no mocks
uv run ruff check .
uv run ruff format .
uv run mypy               # strict, and it passes on the tests too

The test doubles are real implementations, not mocks: an in-memory DictPhotoSource, a FakeSearchEngine that records what it was asked, and — where the wire itself is what needs testing — a real local HTTP server speaking Vision's JSON.

Licence

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

imgtrail-0.4.1.tar.gz (43.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

imgtrail-0.4.1-py3-none-any.whl (36.4 kB view details)

Uploaded Python 3

File details

Details for the file imgtrail-0.4.1.tar.gz.

File metadata

  • Download URL: imgtrail-0.4.1.tar.gz
  • Upload date:
  • Size: 43.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for imgtrail-0.4.1.tar.gz
Algorithm Hash digest
SHA256 56b535e556089fdaea0b2bb636ed083efacac59782cd1e890eec8dcff95f1ea8
MD5 f222e4c6e859704f87d7a4cdc909cf49
BLAKE2b-256 582bec411cbc39df45c372e295c6bbc70bd891232479c7d4b97792eb6865d8cc

See more details on using hashes here.

Provenance

The following attestation bundles were made for imgtrail-0.4.1.tar.gz:

Publisher: release.yml on Endika/imgtrail

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file imgtrail-0.4.1-py3-none-any.whl.

File metadata

  • Download URL: imgtrail-0.4.1-py3-none-any.whl
  • Upload date:
  • Size: 36.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for imgtrail-0.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 cdba955af401be7af77f72852f0bbb19d2eaff250d9db4d2e2c25b57a8b02d36
MD5 008f2b051993becb7e459af8058db714
BLAKE2b-256 fadd66e9ba30f07cc247a74628e91b389d23e7b3739ef1051f766c7f221cd863

See more details on using hashes here.

Provenance

The following attestation bundles were made for imgtrail-0.4.1-py3-none-any.whl:

Publisher: release.yml on Endika/imgtrail

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.5.2

2 files

0.5.1

2 files

0.5.0

2 files

This release

0.4.1 This release

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page