ZeroModels
📖 Introduction
ZeroModels is a collection of models with pretrained weights, built entirely with Keras 3. It supports a range of tasks, including classification, object detection (DETR, RT-DETR, RT-DETRv2, RF-DETR, D-FINE, OWL-ViT, OWLv2, Grounding DINO), segmentation (SAM, SAM2, SAM3, SegFormer, DeepLabV3, EoMT, MaskFormer, Mask2Former, OneFormer, MobileViT-DeepLabV3, RF-DETR), monocular depth estimation (Depth Anything V1, Depth Anything V2, TIPSv2-DPT), feature extraction (DINO, DINOv2, DINOv3), vision-language modeling (CLIP, SigLIP, SigLIP2, MetaCLIP 2, TIPSv2), speech recognition (Whisper, Speech2Text, Moonshine), speech-aware language modeling (Granite Speech, Granite Speech Plus), text encoding and masked language modeling (BERT, ModernBERT, ELECTRA, RoBERTa, XLM-RoBERTa, DeBERTa, DeBERTa-v2, DeBERTa-v3), text generation with large language models (GPT, GPT-2, Qwen2, Qwen2-MoE, Qwen3, Qwen3-MoE, Qwen3-Next, Qwen3.5, GPT-OSS, Llama 2, Llama 3, Llama 4, Mistral, Mixtral, Gemma, Gemma 2, MiniMax-Text-01, MiniMax-M2, DeepSeek-V2, DeepSeek-V3, DeepSeek-V4, GLM-4, GLM-4-0414, GLM-4.5/GLM-4.6, GLM-5/GLM-5.1/GLM-5.2), text-to-text encoder-decoder modeling (T5), multimodal vision-language generation (Qwen2-VL, Qwen2.5-VL, Qwen3-VL, Qwen3-VL-MoE, Qwen3.5-MoE, InternVL3, Gemma 3, Gemma 3n, Gemma 4, Gemma 4 Unified, Mistral 3, DeepSeek-VL, Janus-Pro, MiniMax-M3-VL, GLM-4V, GLM-4.5V, Kimi K2.5, Kimi K2.6, Kimi K2.7-Code), vision-language grounding across object detection, OCR, pointing, and referring (LocateAnything), and more. It includes hybrid architectures like MaxViT alongside traditional CNNs and pure transformers. zeromodels includes custom layers and backbone support, providing flexibility and efficiency across various applications. For backbones, there are various weight variants like in1k, in21k, fb_dist_in1k, ms_in22k, fb_in22k_ft_in1k, ns_jft_in1k, aa_in1k, cvnets_in1k, augreg_in21k_ft_in1k, augreg_in21k, and many more.
⚡ Installation
From PyPI (recommended)
pip install -U zeromodels
From Source
pip install -U git+https://github.com/IMvision12/ZeroModels
📑 Documentation
📖 imvision12.github.io/ZeroModels — the rendered docs, with search.
Per-model guides - with architecture notes, usage examples, and available pretrained weights, cover one page per model across every supported task (classification, object detection, segmentation, depth estimation, feature extraction, vision-language, speech recognition, text encoding, and language modeling). Classification backbones share a single page since they all follow the same XModel / XImageClassify two-class structure; each other model has its own. Every example on those pages prints its real, measured output.
The Markdown sources live in docs/ if you would rather read them in the repo.
📑 Models
📝 Text Models
-
Text Encoders (text → embeddings, masked LM, classification)
🏷️ Model Name 📜 Reference Paper 📦 Source of Weights BERT BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding transformersModernBERT Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder transformersELECTRA ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators transformersRoBERTa RoBERTa: A Robustly Optimized BERT Pretraining Approach transformersXLM-RoBERTa Unsupervised Cross-lingual Representation Learning at Scale transformersDeBERTa DeBERTa: Decoding-enhanced BERT with Disentangled Attention transformersDeBERTa-v2 DeBERTa: Decoding-enhanced BERT with Disentangled Attention transformersDeBERTa-v3 DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing transformersT5 Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer transformers -
Text LLMs (text → text)
👁️ Vision Models
-
Backbones
-
Object Detection
🏷️ Model Name 📜 Reference Paper 📦 Source of Weights D-FINE D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution Refinement transformersDETR End-to-End Object Detection with Transformers transformersRT-DETR DETRs Beat YOLOs on Real-time Object Detection transformersRT-DETRv2 RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformers transformersRF-DETR RF-DETR: Neural Architecture Search for Real-Time Detection Transformers transformersOWL-ViT Simple Open-Vocabulary Object Detection with Vision Transformers transformersOWLv2 Scaling Open-Vocabulary Object Detection transformersGrounding DINO Marrying DINO with Grounded Pre-Training for Open-Set Object Detection transformers
-
Segmentation
-
Feature Extraction
🏷️ Model Name 📜 Reference Paper 📦 Source of Weights DINO Emerging Properties in Self-Supervised Vision Transformers torch.hubDINOv2 DINOv2: Learning Robust Visual Features without Supervision transformersDINOv3 DINOv3: Self-Supervised Visual Representation Learning at Scale transformers(gated)
-
Depth Estimation
🏷️ Model Name 📜 Reference Paper 📦 Source of Weights Depth Anything V1 Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data transformersDepth Anything V2 Depth Anything V2 transformersTIPSv2-DPT TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment transformers
🖼️ Multimodal Models
-
Vision-Language Encoders
🏷️ Model Name 📜 Reference Paper 📦 Source of Weights CLIP Learning Transferable Visual Models From Natural Language Supervision transformersMetaCLIP 2 MetaCLIP 2: A Worldwide Scaling Recipe transformersSigLIP Sigmoid Loss for Language Image Pre-Training transformersSigLIP2 SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features transformersTIPSv2 TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment transformers
-
Multimodal LLMs (image + text → text)
-
Vision-Language Grounding (object detection, OCR, pointing, referring)
🏷️ Model Name 📜 Reference Paper 📦 Source of Weights LocateAnything LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding transformers
🔊 Audio Models
-
Speech (speech → text)
🏷️ Model Name 📜 Reference Paper 📦 Source of Weights Whisper Robust Speech Recognition via Large-Scale Weak Supervision transformersSpeech2Text fairseq S2T: Fast Speech-to-Text Modeling with fairseq transformersMoonshine Moonshine: Speech Recognition for Live Transcription and Voice Commands transformers
-
Speech LLMs (audio + text → text)
🏷️ Model Name 📜 Reference Paper 📦 Source of Weights Granite Speech Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities transformersGranite Speech Plus Granite Speech Plus Technical Report transformers
📜 License
This project leverages timm and transformers for converting pretrained weights from PyTorch to Keras. For licensing details, please refer to the respective repositories.
- 🔖 zeromodels Code: This repository is licensed under the Apache 2.0 License.
🌟 Credits
- The Keras team for their powerful and user-friendly deep learning framework
- The Transformers library for its robust tools for loading and adapting pretrained models
- The pytorch-image-models (timm) project for pioneering many computer vision model implementations
- All contributors to the original papers and architectures implemented in this library
Citing
BibTeX
@misc{gc2025zeromodels,
author = {Gitesh Chawda},
title = {ZeroModels},
year = {2025},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/IMvision12/ZeroModels}}
Metadata
Release files for zeromodels 1.2.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| zeromodels-1.2.6.tar.gz | 1.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| zeromodels-1.2.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.6 MB
Release files / zeromodels-1.2.6.tar.gz
| Download URL | zeromodels-1.2.6.tar.gz |
|---|---|
| Size | 1.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1ce040dea888148face7657046d6fdd2ad8d1f55f68ab1f2a331169f52296ad3
|
|
BLAKE2b-256 checksum How to use checksums |
b60cd4272209952bce65db7ef4528e1f3e3024dd5382ba9188dea00555d7b99f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.
Transparency logRelease files / zeromodels-1.2.6-py3-none-any.whl
| Download URL | zeromodels-1.2.6-py3-none-any.whl |
|---|---|
| Size | 2.0 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
dbbae0b86e927685a4f9c86e9a8289217b6d836069ccfe8cac90b39e63bbc7bd
|
|
BLAKE2b-256 checksum How to use checksums |
f96f078734e9a0a9c6ac039be23c8b5919209ccadb9a248f65b38e564fe3548a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.
Transparency log