Skip to main content

lightvl is a lightweight Vision-Language Model (VLM) quantization toolkit

Project description

LightVL is a lightweight Vision-Language Model (VLM) quantization toolkit supporting FP8, INT8, FP8-Block, and INT4-FP8 mixed quantization schemes. It integrates with vLLM for high-throughput inference and supports Qwen3-VL, Qwen3.5, InternVL-Chat, and Gemma-4 models.

Key capabilities:

  • CPU or GPU fast quantization
  • Automated side-by-side comparison of original vs quantized model latency and accuracy based vllm and transformers
  • Batch inference with concurrent request execution

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

lightvl-0.0.4-py3-none-any.whl (58.7 kB view details)

Uploaded Python 3

File details

Details for the file lightvl-0.0.4-py3-none-any.whl.

File metadata

  • Download URL: lightvl-0.0.4-py3-none-any.whl
  • Upload date:
  • Size: 58.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.12

File hashes

Hashes for lightvl-0.0.4-py3-none-any.whl
Algorithm Hash digest
SHA256 6ffe88836ad3e96856dbf258200ddfe25f1532c710d66ec6581c5e119bc4b656
MD5 254ce5882605b99733bf9e88c604a950
BLAKE2b-256 31004cf60fbd911cf5f492a79a7c054d7b32ae408bec1a80aa802af6a8d8b506

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page