Skip to main content

Multi-Modality

MobileVLM

Implementation of the LDP module block in PyTorch and Zeta from the paper: "MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices"

Install

pip3 install mobilevlm

Usage

# Import the necessary libraries
import torch
from mobilevlm import LDP

# Create an instance of the LDP model
ldp = LDP(in_channels=128, out_channels=128)

# Create an example input tensor
input_tensor = torch.randn(1, 128, 64, 64)

# Pass the input tensor through the LDP model to get the output
output = ldp(input_tensor)

# Print the shape of the output tensor
print(output.shape)

Lightweight Downsample Projection (LDP) Layer

The Lightweight Downsample Projection (LDP) Layer is a component designed for efficient feature extraction and dimensionality reduction in convolutional neural networks. The LDP layer is particularly suited for mobile and edge devices where computational resources are limited.

The LDP layer combines depthwise separable convolutions with pointwise convolutions and skip connections, allowing for a reduced number of parameters while maintaining a rich feature representation. The incorporation of Layer Normalization stabilizes the training process and allows for faster convergence.

Architecture

The LDP layer is structured as follows:

  1. Initial Pointwise Convolution: This is a 1x1 convolution that transforms the input feature map to the desired number of channels. It is computationally efficient and serves as a channel-wise feature transformation.

  2. GELU Activation: After the initial pointwise convolution, we apply a Gaussian Error Linear Unit (GELU) activation function. GELU provides non-linearity to the model, allowing it to learn more complex patterns.

  3. First Depthwise Convolution: A depthwise convolution with a stride of 1 follows, which applies a single filter per input channel. It is used for spatial feature extraction without altering the dimensionality of the feature map.

  4. First Skip Connection: The output of the first depthwise convolution is added back to the output of the initial pointwise convolution. This skip connection allows gradients to flow directly through the network, mitigating the vanishing gradient problem and enabling deeper architectures.

  5. Second Pointwise Convolution: Another 1x1 convolution is applied to further mix the channel-wise features.

  6. Layer Normalization: Normalization is applied over the channel dimension to stabilize the mean and variance of activations, leading to improved training dynamics.

  7. Second GELU Activation: A second GELU activation function is applied for additional non-linearity.

  8. Second Depthwise Convolution: This depthwise convolution has a stride of 2, halving the spatial dimensions of the feature map and effectively downsampling the input.

  9. Second Skip Connection: A pixel-wise addition combines the downsampled input to the block with the output of the second depthwise convolution. This connection helps to preserve information lost due to downsampling.

  10. Third Pointwise Convolution: A final 1x1 convolution adjusts the channel dimensions if necessary and refines the features before passing them to subsequent layers.

  11. Layer Normalization: Another layer normalization is applied to the output of the final pointwise convolution.

Why It Works

The LDP layer is designed to capture the essence of the input features while reducing the spatial resolution in a computationally efficient manner. The use of depthwise separable convolutions significantly decreases the number of parameters compared to standard convolutions, reducing both the computational cost and the risk of overfitting.

Skip connections not only help to preserve information throughout the layer but also improve gradient flow during backpropagation, allowing for deeper network architectures. Layer Normalization is known to accelerate training and make the model less sensitive to initialization and learning rate choices.

This combination of efficiency and robustness makes the LDP layer a versatile component in designing neural networks for resource-constrained environments.

Citation

@misc{chu2023mobilevlm,
    title={MobileVLM : A Fast, Reproducible and Strong Vision Language Assistant for Mobile Devices}, 
    author={Xiangxiang Chu and Limeng Qiao and Xinyang Lin and Shuang Xu and Yang Yang and Yiming Hu and Fei Wei and Xinyu Zhang and Bo Zhang and Xiaolin Wei and Chunhua Shen},
    year={2023},
    eprint={2312.16886},
    archivePrefix={arXiv},
    primaryClass={cs.CV}
}

License

MIT

Metadata

Release files for mobilevlm 0.0.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mobilevlm 0.0.3
File Size Uploaded
mobilevlm-0.0.3.tar.gz 5.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mobilevlm 0.0.3
File Interpreter ABI Platform
mobilevlm-0.0.3-py3-none-any.whl Python 3 none any Details

Total release size: 10.6 kB

Release files / mobilevlm-0.0.3.tar.gz

Download URL mobilevlm-0.0.3.tar.gz
Size 5.4 kB
Tags Source
SHA-256 checksum
How to use checksums
ccd07720419d238542fe032ae9967b504438b64be62e9325d1022ed5c5393068
BLAKE2b-256 checksum
How to use checksums
42c99e71bc94e4af424c96896b6bdc3cae5a6bd38a657bd3a3fbbda41bfb25d4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.3.2 CPython/3.11.0 Darwin/22.4.0

Release files / mobilevlm-0.0.3-py3-none-any.whl

Download URL mobilevlm-0.0.3-py3-none-any.whl
Size 5.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2bbe1c5fb12f99323541eb0f169f7988a02ad02371945ab68ea5281b43c8938c
BLAKE2b-256 checksum
How to use checksums
0c3488f9e319ff3d834ad395a23db32242207089ad5e9cfe7aee4507c1591ed5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.3.2 CPython/3.11.0 Darwin/22.4.0

Release history Release notifications | RSS feed

This release

0.0.3 This release

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page