photo-s-plugin-auto-tone
PhotoS 官方 AI 自动调色插件:CLIP/SigLIP+MLP 模型预测 9 字段 Lightroom 调色参数(exposure / contrast / saturation / vibrance / wb_temp / wb_tint / clarity / texture / dehaze),附带置信度评估、RAG 检索增强, 可选 Qwen3-VL 美学评分、修图建议、风格化调色与场景偏置。
- 训练数据:1295 张 Lightroom 修图记录(XMP sidecar → 参数)
- 基础模型 v7_clean(CLIP ViT-L-14):PSNR 29.10 / SSIM 0.9775; RAG 在困难样本上 +1.93 dB
- 风格化主模型 siglip_h192_d03(SigLIP ViT-L-16-384):PSNR 32.21, 比 v7_clean 高 2.93 dB
安装
pip install photo-s-plugin-auto-tone # 轻量安装(无重依赖)
pip install 'photo-s-plugin-auto-tone[model]' # + torch / open_clip(核心推理)
pip install 'photo-s-plugin-auto-tone[qwen]' # + transformers / peft(美学评分、修图建议、Qwen 风格解析)
权重不打进 wheel:首次调用时从本仓库 GitHub Release 下载到
~/.cache/photo-s/models/ 并做 sha256 校验(每次使用前重新校验)。
核心权重在 tag auto-tone-v0.1.0(约 4.6MB);v2.1 风格化权重
auto_tone_siglip_h192_d03.pt(~850KB)在 tag auto-tone-v2.1.0。
可选 Qwen LoRA 共约 400MB,基座 Qwen3-VL-2B(约 4.3GB)需自备,
通过 PHOTOS_AUTO_TONE_QWEN_BASE 指向本地快照或 HF model id;
SigLIP 视觉塔(约 2.6GB)/ CLIP 塔(约 1.7GB)/ SigLIP tokenizer 经
modelstore 下载校验:默认 HuggingFace,国内推荐
PHOTOS_AUTO_TONE_TOWER_SOURCE=modelscope 走 ModelScope 镜像
(auto = 先 HF 失败自动回落镜像;镜像按上游 sha256 校验,不一致即报错;
HF hub 缓存已命中则零重复下载)。离线可用 PHOTOS_AUTO_TONE_TOWER_URL /
_SHA256 指向自备文件,其余权重变量见 models.py 文档字符串。
插件自身权重另有 ModelScope 镜像仓
dwphoto/photo-s-auto-tone-v2:
PHOTOS_AUTO_TONE_WEIGHT_SOURCE=auto|github|modelscope(auto=GitHub
优先失败回落;镜像仓可用 PHOTOS_AUTO_TONE_MODELSCOPE_REPO 覆盖)。
2026-09-02 起全部 8 个权重(含两个 LoRA,上传自训练机原件并回读
sha 复核)均有镜像——与 TOWER_SOURCE=modelscope 一起即 100% 国内源。
镜像差异说明:siglip 主模型为重存编码(weights_only=True 兼容,权重
逐位一致);两个 LoRA config JSON 为 ModelScope 上传时的 CRLF→LF 规范化
(语义等价,差 1 字节),均按各自来源 sha 钉死校验。
v2.4 新增
- 局部调整词汇表:
auto_tone输出可选local: [{region, params}](region ∈ subject/person/object:label)。引擎经真实管线应用全部 9 个全局字段 + 蒙版局部调整(旧接线只落 3 个字段)。checkpoint 携带 局部头即可启用(local_state_dict等键,训练侧见主仓 TRAINING.md §5.1)。 - 美学验证(verify operation):
verify_aesthetic(image, prefer)= SigLIP 回归头(毫秒级,aesthetic_head.pt由主仓tools/train_verifier.py用 LR 星级评分训练)+ Qwen VLM LoRA 终审。photo-s audit IMG --aesthetic 6即美学闸门。
用法
from photo_s_plugin_auto_tone import auto_tone, auto_tone_with_style, auto_tone_with_scene
# 普通自动调色
result = auto_tone("/path/to/photo.jpg", strength=0.8)
# {"options": {...9 字段...}, "confidence": 0.72, "warnings": [], ...}
# 风格化调色(v2.1):任意自然语言风格描述;None 时 SigLIP 自动视觉分析
styled = auto_tone_with_style("/path/to/photo.jpg", "忧郁蓝调", strength=0.8)
# {"schema_version": 2, "options": {...}, "bias": {...}, "bias_source": "preset",
# "style_desc": "忧郁蓝调", "visual_styles": [...top-3...], ...}
# 场景自适应(v2.1):552 张 LR 目录统计的 7 场景数据驱动偏置
scene = auto_tone_with_scene("/path/to/photo.jpg", "portrait", strength=0.5)
风格化组合三种能力:SigLIP 视觉分析(16 风格 top-K)、Qwen3-VL 文本解析
(自然语言 → 9 字段偏置;use_qwen=False 或 Qwen 不可用时回退 8 种手工
预设,无需任何额外下载)。analyze_visual_style(path) 单独返回视觉风格。
MCP 工具(auto_tone / aesthetic_score / verify_aesthetic(v2.4)/
tone_advisor / batch_auto_tone / auto_tone_with_style /
analyze_visual_style,batch_auto_tone 支持 style_desc 参数)通过
api/mcp_tools.register_mcp_tools(mcp) 注册;REST 路由通过
api/rest.register_routes(handler_class) 挂到 photo-s serve。
LangChain 封装见 api/langchain.py(get_style_tool() /
get_visual_style_tool())。
平台支持
- 推理设备自动选择:CUDA → Apple MPS → CPU
- Windows / macOS / Linux 均可运行;纯 CPU 可跑核心推理(较慢)
- Qwen 美学评分 / 建议 / 风格解析建议使用 CUDA(CPU 上可用但显著变慢)
许可
- 代码:MIT(与 photo_s 主仓库一致)
- 模型权重(GitHub Release 与 ModelScope 镜像
dwphoto/photo-s-auto-tone-v2上的文件):双许可 — CC-BY-NC 4.0(署名-非商用, 个人与非商业用途永久免费)+ 商业授权(职业交付/企业使用/产品集成/再分发 需购买):邮箱 1634103640@qq.com · dwphoto.top/message。权重由个人 Lightroom 修图记录训练(个人修图风格模型),不适合以 MIT 形式无限制商用。 边界判定表与 FAQ 见主仓docs/COMMERCIAL.md;完整许可文本(含商业 授权条款)见本目录LICENSE-WEIGHTS.txt。
上游依赖:OpenAI CLIP ViT-L-14(MIT)、SigLIP webli 权重(Apache-2.0)、 Qwen3-VL-2B(Apache-2.0)均允许再发布衍生权重。
Release files for photo-s-plugin-auto-tone 2.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| photo_s_plugin_auto_tone-2.3.0.tar.gz | 61.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| photo_s_plugin_auto_tone-2.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 135.6 kB
Release files / photo_s_plugin_auto_tone-2.3.0.tar.gz
| Download URL | photo_s_plugin_auto_tone-2.3.0.tar.gz |
|---|---|
| Size | 61.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5f8ce791afac868c35efc7dd16fad5b4c571995e7977b6e75b27837786d61878
|
|
BLAKE2b-256 checksum How to use checksums |
266ef907ff50bb9dcbb7676deb893dbf95daeed3de5f9b8acb88242c210549dc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 2, 2026.
Transparency logRelease files / photo_s_plugin_auto_tone-2.3.0-py3-none-any.whl
| Download URL | photo_s_plugin_auto_tone-2.3.0-py3-none-any.whl |
|---|---|
| Size | 73.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f283331989c017bea0c2bab73a0ecb2fe47810919c9d6b173dca451654b6fe90
|
|
BLAKE2b-256 checksum How to use checksums |
eac5e624ed0bdf682581b203860f62007498e2686b90955d40d1267cfddac628
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 2, 2026.
Transparency log