Neural network inference engine that delivers GPU-class performance for sparsified models on CPUs

These details have not been verified by PyPI

Project links

Homepage

Project description

DeepSparse Engine

Neural network inference engine that delivers GPU-class performance for sparsified models on CPUs

Overview

The DeepSparse Engine is a CPU runtime that delivers GPU-class performance by taking advantage of sparsity within neural networks to reduce compute required as well as accelerate memory bound workloads. It is focused on model deployment and scaling machine learning pipelines, fitting seamlessly into your existing deployments as an inference backend.

This repository includes package APIs along with examples to quickly get started benchmarking and inferencing sparse models.

Sparsification

Sparsification is the process of taking a trained deep learning model and removing redundant information from the overprecise and over-parameterized network resulting in a faster and smaller model. Techniques for sparsification are all encompassing including everything from inducing sparsity using pruning and quantization to enabling naturally occurring sparsity using activation sparsity or winograd/FFT. When implemented correctly, these techniques result in significantly more performant and smaller models with limited to no effect on the baseline metrics. For example, pruning plus quantization can give noticeable improvements in performance while recovering to nearly the same baseline accuracy.

The Deep Sparse product suite builds on top of sparsification enabling you to easily apply the techniques to your datasets and models using recipe-driven approaches. Recipes encode the directions for how to sparsify a model into a simple, easily editable format.

Download a sparsification recipe and sparsified model from the SparseZoo.
Alternatively, create a recipe for your model using Sparsify.
Apply your recipe with only a few lines of code using SparseML.
Finally, for GPU-level performance on CPUs, deploy your sparse-quantized model with the DeepSparse Engine.

Full Deep Sparse product flow:

Compatibility

The DeepSparse Engine ingests models in the ONNX format, allowing for compatibility with PyTorch, TensorFlow, Keras, and many other frameworks that support it. This reduces the extra work of preparing your trained model for inference to just one step of exporting.

Quick Tour

To expedite inference and benchmarking on real models, we include the sparsezoo package. SparseZoo hosts inference-optimized models, trained on repeatable sparsification recipes using state-of-the-art techniques from SparseML.

Quickstart with SparseZoo ONNX Models

ResNet-50 Dense

Here is how to quickly perform inference with DeepSparse Engine on a pre-trained dense ResNet-50 from SparseZoo.

from deepsparse import compile_model
from sparsezoo.models import classification

batch_size = 64

# Download model and compile as optimized executable for your machine
model = classification.resnet_50()
engine = compile_model(model, batch_size=batch_size)

# Fetch sample input and predict output using engine
inputs = model.data_inputs.sample_batch(batch_size=batch_size)
outputs, inference_time = engine.timed_run(inputs)

ResNet-50 Sparsified

When exploring available optimized models, you can use the Zoo.search_optimized_models utility to find models that share a base.

Try this on the dense ResNet-50 to see what is available:

from sparsezoo import Zoo
from sparsezoo.models import classification

model = classification.resnet_50()
print(Zoo.search_sparse_models(model))

Output:

[
    Model(stub=cv/classification/resnet_v1-50/pytorch/sparseml/imagenet/pruned-conservative), 
    Model(stub=cv/classification/resnet_v1-50/pytorch/sparseml/imagenet/pruned-moderate), 
    Model(stub=cv/classification/resnet_v1-50/pytorch/sparseml/imagenet/pruned_quant-moderate), 
    Model(stub=cv/classification/resnet_v1-50/pytorch/sparseml/imagenet-augmented/pruned_quant-aggressive)
]

We can see there are two pruned versions targeting FP32 and two pruned, quantized versions targeting INT8. The conservative, moderate, and aggressive tags recover to 100%, >=99%, and <99% of baseline accuracy respectively.

For a version of ResNet-50 that recovers close to the baseline and is very performant, choose the pruned_quant-moderate model. This model will run nearly 7x faster than the baseline model on a compatible CPU (with the VNNI instruction set enabled). For hardware compatibility, see the Hardware Support section.

from deepsparse import compile_model
import numpy

batch_size = 64
sample_inputs = [numpy.random.randn(batch_size, 3, 224, 224).astype(numpy.float32)]

# run baseline benchmarking
engine_base = compile_model(
    model="zoo:cv/classification/resnet_v1-50/pytorch/sparseml/imagenet/base-none", 
    batch_size=batch_size,
)
benchmarks_base = engine_base.benchmark(sample_inputs)
print(benchmarks_base)

# run sparse benchmarking
engine_sparse = compile_model(
    model="zoo:cv/classification/resnet_v1-50/pytorch/sparseml/imagenet/pruned_quant-moderate", 
    batch_size=batch_size,
)
if not engine_sparse.cpu_vnni:
    print("WARNING: VNNI instructions not detected, quantization speedup not well supported")
benchmarks_sparse = engine_sparse.benchmark(sample_inputs)
print(benchmarks_sparse)

print(f"Speedup: {benchmarks_sparse.items_per_second / benchmarks_base.items_per_second:.2f}x")

Quickstart with Custom ONNX Models

We accept ONNX files for custom models, too. Simply plug in your model to compare performance with other solutions.

> wget https://github.com/onnx/models/raw/master/vision/classification/mobilenet/model/mobilenetv2-7.onnx
Saving to: ‘mobilenetv2-7.onnx’

from deepsparse import compile_model
from deepsparse.utils import generate_random_inputs
onnx_filepath = "mobilenetv2-7.onnx"
batch_size = 16

# Generate random sample input
inputs = generate_random_inputs(onnx_filepath, batch_size)

# Compile and run
engine = compile_model(onnx_filepath, batch_size)
outputs = engine.run(inputs)

For a more in-depth read on available APIs and workflows, check out the examples and DeepSparse Engine documentation.

Hardware Support

The DeepSparse Engine is validated to work on x86 Intel and AMD CPUs running Linux operating systems.

It is highly recommended to run on a CPU with AVX-512 instructions available for optimal algorithms to be enabled.

Here is a table detailing specific support for some algorithms over different microarchitectures:

x86 Extension	Microarchitectures	Activation Sparsity	Kernel Sparsity	Sparse Quantization
AMD AVX2	Zen 2, Zen 3	not supported	optimized	not supported
Intel AVX2	Haswell, Broadwell, and newer	not supported	optimized	not supported
Intel AVX-512	Skylake, Cannon Lake, and newer	optimized	optimized	emulated
Intel AVX-512 VNNI (DL Boost)	Cascade Lake, Ice Lake, Cooper Lake, Tiger Lake	optimized	optimized	optimized

Installation

This repository is tested on Python 3.6+, and ONNX 1.5.0+. It is recommended to install in a virtual environment to keep your system in order.

Install with pip using:

pip install deepsparse

Then if you want to explore the examples, clone the repository and any install additional dependencies found in example folders.

Notebooks

For some step-by-step examples, we have Jupyter notebooks showing how to compile models with the DeepSparse Engine, check the predictions for accuracy, and benchmark them on your hardware.

Available Models and Recipes

A number of pre-trained baseline and recalibrated models models in the SparseZoo can be used with the engine for higher performance. The types available for each model architecture are noted in its SparseZoo model repository listing.

Resources and Learning More

DeepSparse Engine Documentation, Notebooks, Examples
DeepSparse API
Debugging and Optimizing Performance
SparseML Documentation
Sparsify Documentation
SparseZoo Documentation
Neural Magic Blog, Resources, Website

Contributing

We appreciate contributions to the code, examples, and documentation as well as bug reports and feature requests! Learn how here.

Join the Community

For user help or questions about the DeepSparse Engine, sign up or log in: Deep Sparse Community Discourse Forum and/or Slack. We are growing the community member by member and happy to see you there.

You can get the latest news, webinar and event invites, research papers, and other ML Performance tidbits by subscribing to the Neural Magic community.

For more general questions about Neural Magic, please email us at learnmore@neuralmagic.com or fill out this form.

License

The project's binary containing the DeepSparse Engine is licensed under the Neural Magic Engine License.

Example files and scripts included in this repository are licensed under the Apache License Version 2.0 as noted.

Release History

Official builds are hosted on PyPI

stable: deepsparse
nightly (dev): deepsparse-nightly

Track this project via GitHub Releases.

Citation

Find this project useful in your research or other communications? Please consider citing Neural Magic's paper:

@inproceedings{pmlr-v119-kurtz20a, 
    title = {Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks}, 
    author = {Kurtz, Mark and Kopinsky, Justin and Gelashvili, Rati and Matveev, Alexander and Carr, John and Goin, Michael and Leiserson, William and Moore, Sage and Nell, Bill and Shavit, Nir and Alistarh, Dan}, 
    booktitle = {Proceedings of the 37th International Conference on Machine Learning}, 
    pages = {5533--5543}, 
    year = {2020}, 
    editor = {Hal Daumé III and Aarti Singh}, 
    volume = {119}, 
    series = {Proceedings of Machine Learning Research},
    address = {Virtual}, 
    month = {13--18 Jul}, 
    publisher = {PMLR}, 
    pdf = {http://proceedings.mlr.press/v119/kurtz20a/kurtz20a.pdf},, 
    url = {http://proceedings.mlr.press/v119/kurtz20a.html}, 
    abstract = {Optimizing convolutional neural networks for fast inference has recently become an extremely active area of research. One of the go-to solutions in this context is weight pruning, which aims to reduce computational and memory footprint by removing large subsets of the connections in a neural network. Surprisingly, much less attention has been given to exploiting sparsity in the activation maps, which tend to be naturally sparse in many settings thanks to the structure of rectified linear (ReLU) activation functions. In this paper, we present an in-depth analysis of methods for maximizing the sparsity of the activations in a trained neural network, and show that, when coupled with an efficient sparse-input convolution algorithm, we can leverage this sparsity for significant performance gains. To induce highly sparse activation maps without accuracy loss, we introduce a new regularization technique, coupled with a new threshold-based sparsification method based on a parameterized activation function called Forced-Activation-Threshold Rectified Linear Unit (FATReLU). We examine the impact of our methods on popular image classification models, showing that most architectures can adapt to significantly sparser activation maps without any accuracy loss. Our second contribution is showing that these these compression gains can be translated into inference speedups: we provide a new algorithm to enable fast convolution operations over networks with sparse activations, and show that it can enable significant speedups for end-to-end inference on a range of popular models on the large-scale ImageNet image classification task on modern Intel CPUs, with little or no retraining cost.} 
}

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

1.9.0

Jun 2, 2025

1.8.0

Jul 19, 2024

1.7.1

Mar 19, 2024

1.7.0

Mar 15, 2024

1.6.1

Dec 20, 2023

1.6.0

Dec 4, 2023

1.5.3

Aug 17, 2023

1.5.2

Jul 6, 2023

1.5.1

Jun 21, 2023

1.5.0

Jun 2, 2023

1.4.2

Mar 25, 2023

1.4.1

Mar 18, 2023

1.4.0

Feb 16, 2023

1.3.2

Jan 17, 2023

1.3.1

Jan 4, 2023

1.3.0

Dec 17, 2022

1.2.0

Oct 27, 2022

1.1.0

Aug 25, 2022

1.0.2

Jul 12, 2022

1.0.1

Jul 7, 2022

1.0.0

Jun 28, 2022

0.12.2

Jun 2, 2022

0.12.1

May 5, 2022

0.12.0

Apr 21, 2022

0.11.2

Mar 23, 2022

0.11.1

Mar 21, 2022

0.11.0

Mar 11, 2022

0.10.0

Feb 3, 2022

0.9.1

Dec 14, 2021

0.9.0

Dec 1, 2021

0.8.0

Oct 26, 2021

0.7.0

Sep 11, 2021

0.6.1

Aug 11, 2021

0.6.0

Jul 30, 2021

0.5.1

Jun 30, 2021

0.5.0

Jun 28, 2021

0.4.0

Jun 4, 2021

This version

0.3.1

May 13, 2021

0.3.0

Apr 30, 2021

0.2.0

Mar 31, 2021

0.1.1

Mar 1, 2021

0.1.0

Feb 4, 2021

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

deepsparse-0.3.1.tar.gz (30.1 MB view details)

Uploaded May 13, 2021 Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

deepsparse-0.3.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (32.7 MB view details)

Uploaded May 13, 2021 CPython 3.8manylinux: glibc 2.17+ x86-64

deepsparse-0.3.1-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (32.7 MB view details)

Uploaded May 13, 2021 CPython 3.7mmanylinux: glibc 2.17+ x86-64

deepsparse-0.3.1-cp36-cp36m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (32.7 MB view details)

Uploaded May 13, 2021 CPython 3.6mmanylinux: glibc 2.17+ x86-64

File details

Details for the file deepsparse-0.3.1.tar.gz.

File metadata

Download URL: deepsparse-0.3.1.tar.gz
Upload date: May 13, 2021
Size: 30.1 MB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.3.0 pkginfo/1.7.0 requests/2.24.0 setuptools/45.2.0 requests-toolbelt/0.9.1 tqdm/4.56.0 CPython/3.8.5

File hashes

Hashes for deepsparse-0.3.1.tar.gz
Algorithm	Hash digest
SHA256	`06a39deacf2835b80ea49b7bda5dae637edbefd6393554d65d1695fbf90994f3`
MD5	`e8da52763bb5fea495d49a728bb4cd1c`
BLAKE2b-256	`28b61ece355005043800c5f7490fdd8906994db17ea965aaefd75844d95d7738`

See more details on using hashes here.

File details

Details for the file deepsparse-0.3.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

Download URL: deepsparse-0.3.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Upload date: May 13, 2021
Size: 32.7 MB
Tags: CPython 3.8, manylinux: glibc 2.17+ x86-64
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.3.0 pkginfo/1.7.0 requests/2.24.0 setuptools/45.2.0 requests-toolbelt/0.9.1 tqdm/4.56.0 CPython/3.8.5

File hashes

Hashes for deepsparse-0.3.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm	Hash digest
SHA256	`9f12e6bc0665da61cd0839a3529e5cc531503d9a6d700bf3248096c55623b731`
MD5	`109c318a67344041792adc67855bac48`
BLAKE2b-256	`1a883edb953ebcfb03e097937ad4269e247704642568766f2bcf05756add63b4`

See more details on using hashes here.

File details

Details for the file deepsparse-0.3.1-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

Download URL: deepsparse-0.3.1-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Upload date: May 13, 2021
Size: 32.7 MB
Tags: CPython 3.7m, manylinux: glibc 2.17+ x86-64
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.3.0 pkginfo/1.7.0 requests/2.24.0 setuptools/45.2.0 requests-toolbelt/0.9.1 tqdm/4.56.0 CPython/3.8.5

File hashes

Hashes for deepsparse-0.3.1-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm	Hash digest
SHA256	`e0dbb2c0d38b363c383a5e2daba2068a9873fe386baa01687c68ac1497b2a1cb`
MD5	`33ae8ab424a4cfc7d892dc3c1492a054`
BLAKE2b-256	`f458aef8468a6140167da82661ed112268218f554c0c9bc053f189a574c118f3`

See more details on using hashes here.

File details

Details for the file deepsparse-0.3.1-cp36-cp36m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

Download URL: deepsparse-0.3.1-cp36-cp36m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Upload date: May 13, 2021
Size: 32.7 MB
Tags: CPython 3.6m, manylinux: glibc 2.17+ x86-64
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.3.0 pkginfo/1.7.0 requests/2.24.0 setuptools/45.2.0 requests-toolbelt/0.9.1 tqdm/4.56.0 CPython/3.8.5

File hashes

Hashes for deepsparse-0.3.1-cp36-cp36m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm	Hash digest
SHA256	`19abbad7ad45af917a33515d36459d794afeb3078c37ad1ac170ca22045d2b89`
MD5	`ae1dd2e98b7099b69a112a1653e5fad4`
BLAKE2b-256	`a32d4f573be7a32c4b7c3ccb910b19e9f0d11ac3142dec2bb68df9b0fd051d02`

See more details on using hashes here.

deepsparse 0.3.1

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

DeepSparse Engine

Neural network inference engine that delivers GPU-class performance for sparsified models on CPUs

Overview

Sparsification

Compatibility

Quick Tour

Quickstart with SparseZoo ONNX Models

Quickstart with Custom ONNX Models

Hardware Support

Installation

Notebooks

Available Models and Recipes

Resources and Learning More

Contributing

Join the Community

License

Release History

Citation

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distributions

File details

File metadata

File hashes

File details

File metadata

File hashes

File details

File metadata

File hashes

File details

File metadata

File hashes