lightvl is a lightweight Vision-Language Model (VLM) quantization toolkit
Project description
LightVL is a lightweight Vision-Language Model (VLM) quantization toolkit supporting FP8, INT8, FP8-Block, and INT4-FP8 mixed quantization schemes. It integrates with vLLM for high-throughput inference and supports Qwen3-VL, Qwen3.5, InternVL-Chat, and Gemma-4 models.
Key capabilities:
- CPU or GPU fast quantization
- Automated side-by-side comparison of original vs quantized model latency and accuracy based vllm and transformers
- Batch inference with concurrent request execution
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
No source distribution files available for this release.See tutorial on generating distribution archives.
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
lightvl-0.0.5-py3-none-any.whl
(59.1 kB
view details)
File details
Details for the file lightvl-0.0.5-py3-none-any.whl.
File metadata
- Download URL: lightvl-0.0.5-py3-none-any.whl
- Upload date:
- Size: 59.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fe85b10cdeb6d793d8b39588852f5f11df177c93c49e3f3e76f50bb6ca113d99
|
|
| MD5 |
b1e9c8bc024c8fe5fc691a01cdfe4f14
|
|
| BLAKE2b-256 |
1d6fe6e48545ad4638f82607492c4d1790aaee7339e3763b0670a35938e6400c
|