cuTENSOR is a high-performance CUDA library for tensor primitives.
Key Features
Extensive mixed-precision support:
FP64 inputs with FP32 compute.
FP32 inputs with FP16, BF16, or TF32 compute.
Complex-times-real operations.
Conjugate (without transpose) support.
Support for up to 64-dimensional tensors.
Arbitrary data layouts.
Trivially serializable data structures.
Multi-process distributed tensor contractions with cuTENSORMp (Linux only).
Main computational routines:
Direct (i.e., transpose-free) tensor contractions.
Support just-in-time compilation of dedicated kernels.
Tensor reductions (including partial reductions).
Element-wise tensor operations:
Support for various activation functions.
Support for padding of the output tensor
Arbitrary tensor permutations.
Conversion between different data types.
Documentation
Please refer to https://docs.nvidia.com/cuda/cutensor/index.html for the cuTENSOR documentation.
Installation
The cuTENSOR wheel can be installed as follows:
pip install cutensor-cuXX
where XX is the CUDA major version (currently CUDA 12 & 13 are supported). The package cutensor (without the -cuXX suffix) is deprecated. If you have cutensor installed, please remove it prior to installing cutensor-cuXX.
On Linux, install the optional mp extra to use cuTENSORMp. This installs the matching NVIDIA NCCL runtime:
pip install "cutensor-cuXX[mp]"
The mp extra installs NCCL only. The NCCL-backed cuTENSORMp library is not linked against MPI, so the wheel deliberately does not select an MPI implementation. MPI-based applications and samples require an external MPI development installation and launcher, such as Open MPI, MPICH, or NVIDIA HPC-X. Compiling the samples also requires a CUDA development toolkit with nvcc.
Metadata
Release files for cutensor-cu13 2.8.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cutensor_cu13-2.8.0-py3-none-win_amd64.whl | Python 3 | none | Windows x86-64 | Details |
| cutensor_cu13-2.8.0-py3-none-manylinux2014_x86_64.whl | Python 3 | none | Linux glibc 2.17+ x86-64 | Details |
| cutensor_cu13-2.8.0-py3-none-manylinux2014_aarch64.whl | Python 3 | none | Linux glibc 2.17+ ARM64 | Details |
Total release size: 602.4 MB
Release files / cutensor_cu13-2.8.0-py3-none-win_amd64.whl
| Download URL | cutensor_cu13-2.8.0-py3-none-win_amd64.whl |
|---|---|
| Size | 186.9 MB |
| Tags | Python 3 Windows x86-64 |
|
SHA-256 checksum How to use checksums |
07c72f22827901159490d3b26f8d1f0fe28082cc1f29b176b447f70110dd3a91
|
|
BLAKE2b-256 checksum How to use checksums |
270260281e3cca413d6fec8a8308bacc173c9c17c24b6cacb1e02d059f401014
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / cutensor_cu13-2.8.0-py3-none-manylinux2014_x86_64.whl
| Download URL | cutensor_cu13-2.8.0-py3-none-manylinux2014_x86_64.whl |
|---|---|
| Size | 208.6 MB |
| Tags | Linux glibc 2.17+ x86-64 Python 3 |
|
SHA-256 checksum How to use checksums |
60981a8b8e9f57415978a03766e7b4364cf28443635900cf3be2ac73c256a482
|
|
BLAKE2b-256 checksum How to use checksums |
15f51a3965415b1cdc21a85f5c7f4d1ce7ff580372741979f15be34c862c4472
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / cutensor_cu13-2.8.0-py3-none-manylinux2014_aarch64.whl
| Download URL | cutensor_cu13-2.8.0-py3-none-manylinux2014_aarch64.whl |
|---|---|
| Size | 206.9 MB |
| Tags | Linux glibc 2.17+ ARM64 Python 3 |
|
SHA-256 checksum How to use checksums |
6974c00b869aebeb19fe4a5e064721a27bafd400583858d419d6a630b62260ba
|
|
BLAKE2b-256 checksum How to use checksums |
0d5ce637de5d06a13e061c14e39d69e7ffd5b8a8345f590a47e9a66f2eec8cdf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|