VFormer
A modular PyTorch library for vision transformers models
Library Features
- Contains implementations of prominent ViT architectures broken down into modular components like encoder, attention mechanism, and decoder
- Makes it easy to develop custom models by composing components of different architectures
- Utilities for visualizing attention using techniques such as gradient rollout
Installation
From source (recommended)
git clone https://github.com/SforAiDl/vformer.git
cd vformer/
python setup.py install
From PyPI
pip install vformer
Models supported
- Vanilla ViT
- Swin Transformer
- Pyramid Vision Transformer
- CrossViT
- Compact Vision Transformer
- Compact Convolutional Transformer
- Visformer
- Vision Transformers for Dense Prediction
- CvT
- ConViT
- ViViT
- Perceiver IO
- Memory Efficient Attention
Example usage
To instantiate and use a Swin Transformer model -
import torch
from vformer.models.classification import SwinTransformer
image = torch.randn(1, 3, 224, 224) # Example data
model = SwinTransformer(
img_size=224,
patch_size=4,
in_channels=3,
n_classes=10,
embed_dim=96,
depths=[2, 2, 6, 2],
num_heads=[3, 6, 12, 24],
window_size=7,
drop_rate=0.2,
)
logits = model(image)
VFormer has a modular design and allows for easy experimentation using blocks/modules of different architectures. For example, if desired, you can use just the encoder or the windowed attention layer of the Swin Transformer model.
from vformer.attention import WindowAttention
window_attn = WindowAttention(
dim=128,
window_size=7,
num_heads=2,
**kwargs,
)
from vformer.encoder import SwinEncoder
swin_encoder = SwinEncoder(
dim=128,
input_resolution=(224, 224),
depth=2,
num_heads=2,
window_size=7,
**kwargs,
)
Please refer to our documentation to know more.
References
Metadata
Release files for vformer 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vformer-0.1.3.tar.gz | 60.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vformer-0.1.3-py2.py3-none-any.whl | Python 3, Python 2 | none | any | Details |
Total release size: 134.2 kB
Release files / vformer-0.1.3.tar.gz
| Download URL | vformer-0.1.3.tar.gz |
|---|---|
| Size | 60.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ee43cb736e9cc155ea470f58241f6d15b3c21df8410a4fa01f9f053f0733e750
|
|
BLAKE2b-256 checksum How to use checksums |
f58b511ece4f144d8bb2f84588702dd7a661446b970d8a24bb6f2fd13f0f8327
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.1 CPython/3.9.13
|
Release files / vformer-0.1.3-py2.py3-none-any.whl
| Download URL | vformer-0.1.3-py2.py3-none-any.whl |
|---|---|
| Size | 73.9 kB |
| Tags | Python 2 Python 3 |
|
SHA-256 checksum How to use checksums |
198ad9ee15c0b09ea1d68fc181c022697eaf1fa7903bcee19a1e3f1f90676fc6
|
|
BLAKE2b-256 checksum How to use checksums |
29e8b156db70933b2d166f6096e2e3d1ffc0f0ff984d9ac8567e43c8aaabe0a9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.1 CPython/3.9.13
|