Skip to main content

Photo metadata from vision LLMs: verbatim transcription of what's written on scans, plus captions, keywords, and cautious date/location guesses.

Project description

photokin

Run scanned photos and documents through a vision model and gets archival metadata back: a verbatim transcription of whatever is written on the front or back, a scene caption, keywords, and deliberately cautious date/location guesses - as JSON, NDJSON streams, or changesets that ExifTool can write into the files.

Compatible with OpenAI, Anthropic, Gemini, and OpenRouter API keys.

Why should I use this?

I have inherited thousands of family photos and documents. While I would love to have the time to manually review each one, I realized that having an LLM do a first pass of them could help me manually review them later. So I started experimenting with the capabilities of LLMs and I was pleasantly surprised at the results. This library is my attempt to automate the process and get the key data I need from every photo and document to make my manual review easier.

Quick start for a single photo

pip install "photokin[openai]"          # or [anthropic] / [gemini] / [all]
export OPENAI_API_KEY=sk-...            # setx on Windows

photokin scan_042.jpg --back scan_042_back.jpg

That prints one JSON result, keyed by the image path. Abridged:

{
  "result": {
    "scan_042.jpg": {
      "keywords": ["Postcard", "1940s", "Military personnel", "..."],
      "caption": "[Back]\n27 november 44\nAlthough, I personally did not see this cathedral...",
      "ai_caption": "[AI Analysis]: A printed postcard showing... Inferred date: 1944-11-27 (confidence 0.95; evidence: handwritten date on back).",
      "category": "Postcard",
      "location_guess": {"country": "France", "city": "Le Mans", "confidence": 0.9},
      "date_guess": {"iso": "1944-11-27", "confidence": 0.95, "pattern": "Y!M!D!"}
    }
  }
}

The transcription (caption) and the interpretation (ai_caption) are kept strictly separate — the model is not allowed to "improve" what's actually written on the object. That separation is most of the reason this tool exists.

Folders and batches

Point it at a folder and it works through everything in it (non-recursive). Filename suffixes group scans of the same physical object automatically — photo-a.jpg / photo-b.jpg are variants, photo-back.jpg is the reverse side, album-page1.jpg / album-page2.jpg are pages of one document — and each group is analyzed together as one object.

SIDE NOTE ON EXPECTED Naming conventions. The full suffix grammar is name[letter][-front|-back|-negative|-pageN][-crop], case-insensitive, applied right to left:

Example Meaning
box3_025.jpg the photo itself (primary front)
box3_025-b.jpg or box3_025b.jpg another scan of the same object (variant letter, with or without dash after a digit)
box3_025-back.jpg the reverse side (-front and -negative work the same way)
album-page1.jpg, album-page2.jpg ordered pages of one document
box3_025-back-crop.jpg a cropped detail, used as a supporting view of its parent, never as its own object

The variant letter comes before the part suffix (025b-back-crop.jpg), and a file with no explicit -pageN is only treated as page 1 if its group contains other explicitly numbered pages.

Folder mode

photokin --folder ./scans/ --provider anthropic > results.json

Folder mode prints one aggregate JSON to stdout. For bigger or more repeatable jobs, manifest mode takes a JSON file instead of a folder, and adds --output-file: a .ndjson path streams one record per finished photo (you can watch progress, and a crash doesn't lose completed work), while a .json path writes a single aggregate object atomically at the end.

Manifest mode

A manifest is a JSON file listing exactly what to process — an items array where each entry needs only a path. The sample below declares one physical object, a front scan and its back, and one line of batch-wide background context. Flags like is_back are optional when the filename already says it; they exist so files that don't follow the naming conventions can still be grouped correctly.

batch.json:

{
  "items": [
    {"path": "scans/box3_017.jpg"},
    {"path": "scans/box3_017_back.jpg", "is_back": true}
  ],
  "photo_context_text": "Church family photos, mostly New Jersey, 1930s-1950s."
}
photokin --manifest batch.json --output-file results.ndjson --changeset true

Existing metadata aware enrichment

Items may have existing metadata (face tags, existing captions and comments) that can be forwarded to the model as context. Additionally, you can supply photo_context_text as free-text additional context to a single photo or a folder. Such as "these photos are all part of a wedding album." The model treats it as truth for the whole batch. Both make a real difference on hard photos.

For auditing, --changeset true emits a changeset NDJSON alongside the results: a record of proposed field writes that the ExifTool wrapper can apply to the actual files, either in the same run (--exiftool-write true --exiftool-fields EXIF:UserComment) or later and separately:

python -m photokin.exiftool --changeset batch_changeset.ndjson --enabled [--dry-run]

API keys

Keys are plain environment variables, one per provider. Photokin reads them when it builds the provider client and nowhere else. They never end up in results, changesets, or debug dumps.

Provider Variable
OpenAI OPENAI_API_KEY
Anthropic ANTHROPIC_API_KEY
Gemini GEMINI_API_KEY
OpenRouter OPENROUTER_API_KEY

For the current terminal session:

export OPENAI_API_KEY=sk-...            # macOS / Linux
$env:OPENAI_API_KEY = "sk-..."          # Windows PowerShell

To make it stick across sessions, add the export line to your shell profile (~/.bashrc, ~/.zshrc), or on Windows run setx OPENAI_API_KEY sk-... once (takes effect in new terminals, not the current one). If you keep keys in a file, keep that file out of version control.

You only need the key for the provider you're actually calling. And since a batch run makes one paid API call per photo group, it's worth using a key with a spend cap set in the provider's dashboard — a typo in a folder path is a lot cheaper that way.

Setting up ExifTool

ExifTool is optional, and can be both used for reading initial values from the photos (so you don't just overwrite everything) and also write to the photos after. Before analysis, it reads fields straight out of the files so metadata already living in an image rides along to the model as context (hydration). After analysis, it writes approved changeset fields back into the files or their sidecars (apply).

Where each result field lands when written:

Result field Tag
ai_caption (the AI analysis) EXIF:UserComment
caption (the verbatim transcription) XMP:dc:Description
keywords XMP:dc:Subject
title XMP:dc:Title
date_guess (when confident enough) EXIF:DateTimeOriginal
location_guess (when confident enough) IPTC:Country-PrimaryLocationName / Province-State / City / Sub-location

Setup. On Windows, run python -m photokin.exiftool.fetch once — it downloads the official ExifTool from exiftool.org (SHA256-verified) into ~/.photokin/bin, no system install needed. On macOS/Linux install it yourself: brew install exiftool or apt install libimage-exiftool-perl. At runtime the binary is found in this order: an explicit --exiftool-path / EXIFTOOL_PATH, then the downloaded copy in ~/.photokin/bin, then whatever exiftool is on your PATH. If none is found, analysis still runs — hydration and apply are skipped with a warning rather than failing the batch.

Writing during a run. In manifest mode with --changeset true, add --exiftool-write true --exiftool-fields EXIF:UserComment and the approved fields are written into the files right after analysis. The same settings are available as env vars (EXIFTOOL_WRITE_ENABLED, EXIFTOOL_FIELDS, EXIFTOOL_PATH), with flags winning over env over defaults.

Writing later, with an audit step. Since the changeset NDJSON is a plain record of proposed writes, you can inspect it first and apply it separately:

python -m photokin.exiftool --changeset batch_changeset.ndjson --enabled --dry-run   # counts what would be written
python -m photokin.exiftool --changeset batch_changeset.ndjson --enabled            # actually writes

The standalone applier also takes --fields to narrow which tags may be written, --write-sidecar-only to write .xmp sidecars instead of touching the originals, --no-overwrite-original to keep ExifTool's _original backup files, and --output summary.json for a machine-readable result. Date tags (EXIF:DateTimeOriginal, EXIF:CreateDate) are normalized to EXIF's YYYY:MM:DD HH:MM:SS format on the way in; unparseable dates become warnings, not writes.

All flags

Input modes

Flag What it does
image (positional) Path to the main/front image (single-photo mode)
--back PATH Back-side image for single-photo mode
--meta PATH Original metadata JSON for single-photo mode
--folder DIR Process a whole folder (non-recursive)
--manifest PATH Process a manifest JSON (items + optional metadata/context)

Provider and model

Flag What it does
--provider {openai,anthropic,gemini,openrouter} Which backend to call (default openai)
--openai-model NAME OpenAI model (default gpt-4o)
--claude-model NAME Claude model alias (sonnet or haiku); resolves to a current model id (default sonnet)
--gemini-model NAME Gemini model (default gemini-2.5-flash)
--openrouter-model SLUG Any vision-capable OpenRouter slug (default moonshotai/kimi-k3)

Image handling

Flag What it does
--max-edge N Downscale the longest edge before upload; 0 keeps original size. Smaller is cheaper, larger reads fine print better (default 1024)
--jpeg-quality N JPEG quality 1-100 for the uploaded copy (default 80)

Context

Flag What it does
--photo-context-text TEXT Inline background context, treated as authoritative
--photo-context-file PATH Same, from a UTF-8 text file

Grouping and apply behavior

Flag What it does
--update-policy {master_exact,merge_per_variant} Whether results apply to just the primary file or to every variant in a group (default merge_per_variant)
--process-all-variants Also analyze -b/-c variant scans and fold their outputs together (default off — only the primary scan of each group is analyzed)
--date-confidence-threshold X Minimum model confidence before a date guess is applied, 0-1 (default 0.7)
--location-confidence-threshold X Same, for location guesses (default 0.7)
--no-update-vocab Don't append newly proposed keywords to the vocabulary file

Output

Flag What it does
--output-file PATH .ndjson streams one record per finished photo; .json writes one aggregate object atomically (manifest mode)
--output-sidecars Also write a per-photo sidecar JSON next to each image (default off)
--batch-id ID Identifier stored in output metadata and logs
--changeset {true,false} Emit a changeset NDJSON of proposed file writes (manifest mode; default false)
--dry-run Analyze without applying anything downstream; records are marked dry_run=true (default off)

ExifTool write-back

Flag What it does
--exiftool-write {true,false} Apply changeset fields to the files after analysis (default true)
--exiftool-fields TAGS Comma-separated tags ExifTool may write (default EXIF:UserComment)
--exiftool-path PATH ExifTool binary to use (default: auto-detect)

Debug

Flag What it does
--debug-dump-llm-request {true,false} Save full request payloads to disk before each model call (default false)
--debug-dump-dir DIR Where those dumps go (default <output-dir>/debug)

Providers

OpenAI (default), Anthropic, Gemini, and OpenRouter (any vision-capable slug — Kimi, Grok, Qwen, ...). Select with --provider or LLM_PROVIDER. Only the SDK for the provider you use needs to be installed, and only that provider's key needs to be set.

Layout

Layer Where What it does
Core library photokin/ Prompts, provider dispatch, JSON parsing/repair, metadata merge, changeset emission. No ExifTool dependency.
ExifTool wrapper photokin/exiftool/ Hydration (read before analysis) and apply (write after).

Dependency direction: the wrapper imports from the core; the core never imports the wrapper. The CLI (photokin/cli.py) composes them into the full pipeline: hydrate, analyze, apply. Embedders who don't want ExifTool can call the core directly — core.process_manifest_stream takes any metadata_hydrator callable, or none.

Tests

cd python
python -m pytest

Runs photokin/tests/ and tests/. Python 3.11+.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

photokin-0.1.1.tar.gz (105.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

photokin-0.1.1-py3-none-any.whl (114.9 kB view details)

Uploaded Python 3

File details

Details for the file photokin-0.1.1.tar.gz.

File metadata

  • Download URL: photokin-0.1.1.tar.gz
  • Upload date:
  • Size: 105.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for photokin-0.1.1.tar.gz
Algorithm Hash digest
SHA256 0a305127fd5228d8773b5273939c25d599d2950be8924fc56029db99706b8aa2
MD5 2355614b4a9b5866be15738eb8afbb74
BLAKE2b-256 0aae6fa768867d979de5f4a8e08278f73d399b86c5abfd44136736c9fbd0e621

See more details on using hashes here.

Provenance

The following attestation bundles were made for photokin-0.1.1.tar.gz:

Publisher: release.yml on asielen/photokin

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file photokin-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: photokin-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 114.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for photokin-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 0458c6d8ff24f7d67cda815811c01643969fbce0ddad43e11e53c2cccb91671b
MD5 99f40863dcd280c43c9462ae4315297f
BLAKE2b-256 6928494b381d478cc362c62f13df82076c23f80f6610640eddb26d5c0db7ac34

See more details on using hashes here.

Provenance

The following attestation bundles were made for photokin-0.1.1-py3-none-any.whl:

Publisher: release.yml on asielen/photokin

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page