Skip to main content

Quantized matmul in CUDA, with a PyTorch interface

Original code from FasterTransformer / TensorRT-LLM: https://github.com/NVIDIA/TensorRT-LLM/tree/main/cpp/tensorrt_llm/kernels

Adapted to support a different quantization scheme.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

quant_matmul-1.2.0.tar.gz (11.6 kB view details)

Uploaded Source

File details

Details for the file quant_matmul-1.2.0.tar.gz.

File metadata

  • Download URL: quant_matmul-1.2.0.tar.gz
  • Upload date:
  • Size: 11.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.10.13

File hashes

Hashes for quant_matmul-1.2.0.tar.gz
Algorithm Hash digest
SHA256 5c5d3dcf2618f97500770a0afcfafab2a13760e7e36543f8cee3b04c7b6d9456
MD5 1da26ca040b15ef5d4a5bae69877d10a
BLAKE2b-256 c94744c0eea623bc07148a7b494beaa6662334e99f74d0463c9bf8885798a69f

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

1.2.0 This release

1 file

1.1.1

1 file

1.1.0.post1

1 file

0.0.0

1 file

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page