Last released Jun 16, 2026
SHADA: a self-supervised hierarchical hybrid (CNN + Transformer) multi-modal model for vision and text
Supported by