Quantized matmul in CUDA, with a PyTorch interface
Original code from FasterTransformer / TensorRT-LLM: https://github.com/NVIDIA/TensorRT-LLM/tree/main/cpp/tensorrt_llm/kernels
Adapted to support a different quantization scheme.
Metadata
Release files for quant-matmul 1.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| quant_matmul-1.2.0.tar.gz | 11.6 kB | Details |
Release files / quant_matmul-1.2.0.tar.gz
| Download URL | quant_matmul-1.2.0.tar.gz |
|---|---|
| Size | 11.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5c5d3dcf2618f97500770a0afcfafab2a13760e7e36543f8cee3b04c7b6d9456
|
|
BLAKE2b-256 checksum How to use checksums |
c94744c0eea623bc07148a7b494beaa6662334e99f74d0463c9bf8885798a69f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.0.0 CPython/3.10.13
|