image-evaluator
image-evaluator is a lightweight CLI utility for generative AI practitioners and researchers. It evaluates ten core quality dimensions of synthetic images: predicted visual aesthetics, text-to-image semantic alignment, image editing directional change (Directional CLIP), facial identity preservation, pairwise image fidelity (deep perceptual distance LPIPS, structural similarity SSIM, and peak signal-to-noise ratio PSNR), dataset distribution fidelity (Fréchet Inception Distance FID and Kernel Inception Distance KID), and human preference alignment (PickScore).
Evaluation Workflow
flowchart LR
A[Task Goal] --> B{Choose --metrics}
B -->|aesthetic| C[LAION Aesthetic]
B -->|clip| D[CLIP Similarity]
B -->|directional_clip| U[Directional CLIP]
B -->|arcface| E[ArcFace Distance]
B -->|lpips| F[LPIPS Distance]
B -->|ssim| G[SSIM Similarity]
B -->|psnr| H[PSNR Ratio]
B -->|fid| O[FID Distance]
B -->|kid| P[KID Distance]
B -->|pickscore| S[PickScore Preference]
C --> I[Float, Higher Better]
D --> J[Cosine Sim, Higher Better]
U --> V[Directional Sim, Higher Better]
E --> K[Cosine Dist, Lower Better]
F --> L[Distance, Lower Better]
G --> M[Index, Higher Better]
H --> N[dB, Higher Better]
O --> Q[Wasserstein-2, Lower Better]
P --> R[MMD, Lower Better]
S --> T[Calibrated Likelihood, Higher Better]
Metric Selection
| Target Goal | Metric Name | Required Options | Output Direction | Technical Details |
|---|---|---|---|---|
| Visual appeal & quality | aesthetic |
--image |
Higher is better | docs/aesthetic-score.md |
| Prompt semantic match | clip |
--image, --prompt |
Higher is better | docs/clip-similarity.md |
| Image editing direction | directional_clip |
--image, --reference, --prompt, --prompt-src |
Higher is better | docs/directional-clip.md |
| Facial identity consistency | arcface |
--image, --reference |
Lower is better | docs/arcface-distance.md |
| Deep perceptual similarity | lpips |
--image, --reference |
Lower is better | docs/pairwise-fidelity.md |
| Structural degradation | ssim |
--image, --reference |
Higher is better | docs/pairwise-fidelity.md |
| Pixel reconstruction SNR | psnr |
--image, --reference |
Higher is better | docs/pairwise-fidelity.md |
| Population distribution distance | fid |
--image <dir>, --reference <dir> |
Lower is better | docs/fid.md |
| Unbiased kernel MMD distance | kid |
--image <dir>, --reference <dir> |
Lower is better | docs/kid.md |
| Human preference & alignment | pickscore |
--image, --prompt |
Higher is better | docs/pickscore.md |
Quick Start
Installation
image-evaluator requires Python 3.11–3.14. Install the release via pip:
python -m pip install image-evaluator==0.4.0
The supported runtime path is macOS or Linux with CPU ONNX Runtime. Linux
users who want the GPU runtime can replace onnxruntime with
onnxruntime-gpu after installation. Windows is currently unverified.
Tutorial
The CLI enforces explicit metric selection via --metrics and initializes only selected models:
--metrics(required): One or more ofaesthetic,clip,arcface,lpips,ssim,psnr,fid,kid,pickscore.--image(required): Path to an image file or directory (must be a directory whenfidorkidis selected).--prompt: Required whencliporpickscoreis selected; prohibited otherwise.--reference: Required when reference-based metrics (arcface,lpips,ssim,psnr,fid,kid) are selected; prohibited otherwise (must be a directory whenfidorkidis selected).--format: Output format, eithertext(default) orjson(for CI/CD andjqpipelines).
-
Aesthetic evaluation only:
image-evaluator --metrics aesthetic --image path/to/image.png
-
CLIP text alignment only:
image-evaluator --metrics clip --image path/to/image.png --prompt "a cat in oil painting style"
-
Pairwise fidelity triad (LPIPS, SSIM, PSNR):
image-evaluator --metrics lpips ssim psnr --image path/to/image.png --reference path/to/ref.png
-
Structured JSON output for CI/CD automation:
image-evaluator --metrics ssim psnr --image img.png --reference ref.png --format json | jq .metrics.ssim
-
Dataset distribution metrics (FID and KID):
image-evaluator --metrics fid kid \ --reference path/to/real_images/ \ --image path/to/generated_images/
-
PickScore human preference evaluation:
image-evaluator --metrics pickscore \ --image path/to/image.png \ --prompt "a cat in oil painting style"
-
Multi-metric evaluation across folders:
image-evaluator --metrics aesthetic clip arcface lpips ssim psnr fid kid pickscore \ --image path/to/images/ \ --prompt "a cat in oil painting style" \ --reference path/to/refs/
Python SDK (Zero-Disk-I/O Memory Stream)
For single-image and pairwise metrics (aesthetic, clip, arcface, lpips, ssim, psnr, pickscore), evaluate in-memory PIL.Image, torch.Tensor, or numpy.ndarray objects directly without saving to disk (dataset metrics fid and kid require directory paths):
import torch
from image_evaluator import evaluate
# Evaluate in-memory PyTorch tensors directly
t_img = torch.rand(3, 256, 256)
t_ref = torch.rand(3, 256, 256)
scores = evaluate(
metrics=["ssim", "psnr"],
image=t_img,
reference=t_ref,
device="cpu",
)
print(scores) # {'ssim': 0.0012, 'psnr': 8.142}
Batch Evaluation & Performance Best Practices
When evaluating multiple images, avoid running the CLI in shell loops (for f in *.png; do ...), which re-initializes deep models ($N$ times) and introduces massive cold-start friction (~25-35s for 10 images).
Instead, choose between two high-throughput native paradigms:
-
Native CLI Directory Mode (Single model initialization, ~2-3s for 10 images):
image-evaluator --metrics clip --image ./generated/ --prompt "a golden retriever on a sunny lawn" image-evaluator --metrics lpips ssim psnr --image ./generated/ --reference ./ground_truth/
-
Python Predictor Reuse (Zero reloading overhead, memory-resident tensors/PIL, ~1.5-2.0s for 10 images):
from PIL import Image from image_evaluator import ClipScorePredictor, SSIMPredictor # Instantiate once; weights stay resident in memory/VRAM clip_pred = ClipScorePredictor(device="cpu") ssim_pred = SSIMPredictor() # High-throughput in-memory loop samples = [("img1.png", "prompt1"), ("img2.png", "prompt2")] for path, prompt in samples: score = clip_pred.evaluate_clip_score(path, prompt)
See the full Task Selection Guide and Batch Performance Guide on the documentation site.
Interpretation & Protocol Guidelines
- Protocol Consistency: Always compare scores under identical model backbones and preprocessing pipelines.
- Relative Comparison: Avoid universal absolute thresholds; interpret scores relative to a baseline control.
- Fail-Fast Spatial Dimension Policy: Pairwise metrics (
lpips,ssim,psnr) strictly reject mismatched image dimensions withValueErrorto prevent artificial interpolation distortion. Align sizes beforehand via downsampling or super-resolution. - SSIM Minimum Size: SSIM requires both image dimensions to be at least 11 pixels because it uses the documented 11 × 11 Gaussian window. Smaller inputs fail with
ValueError. - Dataset Distribution Contract: Both
fidandkidrequire existing directories on both sides; passing single files is rejected immediately. Neither metric requires filename stem matching ($N_{\text{ref}} \ne N_{\text{gen}}$ is allowed). - Sample Size Sensitivity & Unbiasedness: FID is a biased estimator with finite-sample bias that increases significantly on smaller sample sizes (triggering a
UserWarningreminding users of finite-sample sensitivity). KID is an unbiased U-statistic estimator that can produce small negative values near zero; these reflect normal statistical fluctuations and must not be truncated. KID defaults to deterministicseed=0for bit-exact reproducibility. - Human Preference Alignment: PickScore assesses text-image alignment against trained human preference choices from the Pick-a-Pic dataset (
yuvalkirstain/PickScore_v1). Scores are uncalibrated logits where higher scores indicate stronger preference; win rate is evaluated strictly via the sigmoid difference between candidate pairs conditioned on identical prompts. PickScore strictly requires an explicit text prompt (--prompt).
Release Status
| Area | 0.3.0 status |
|---|---|
| Metrics | Aesthetic, CLIP, ArcFace, LPIPS, SSIM, PSNR, FID, KID, PickScore |
| macOS | Verified on Apple Silicon with Python 3.11 |
| Linux | Clean install and test suite verified on GitHub Actions with Python 3.11 |
| Windows | External Python 3.12 smoke test passed for CLIP, LPIPS, SSIM, and PSNR; full support remains unverified |
| Dataset metrics | FID and KID fully implemented via clean-fid 0.1.35 under clean Inception-v3 protocol with deterministic seed=0 |
| Human preference | PickScore fully implemented via PickScore_v1 (yuvalkirstain/PickScore_v1) with calibrated likelihood logits |
See CHANGELOG.md for the accepted user-facing changes and known limitations.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file image_evaluator-0.4.0.tar.gz.
File metadata
- Download URL: image_evaluator-0.4.0.tar.gz
- Upload date:
- Size: 32.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dd29a4d650710570a3fccec67a6fe92037f4a6a805e565e2b53817136125d9c3
|
|
| MD5 |
93591fc11295ae98976605ce03e46123
|
|
| BLAKE2b-256 |
8e386ee719542b2dd8e0b57a21a61a85a05852dc015f82ea2443958e6cb2858d
|
File details
Details for the file image_evaluator-0.4.0-py3-none-any.whl.
File metadata
- Download URL: image_evaluator-0.4.0-py3-none-any.whl
- Upload date:
- Size: 39.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c3ef814ebcb4f157632c6ab9bbf3647177e75983145a75959dc0cf776d68562c
|
|
| MD5 |
0c317a213749f7338de17bf9e427dc6b
|
|
| BLAKE2b-256 |
9335a897b9bc17c0b648291b6151a54696867d3291d5169283637a546eeeee7d
|