Trial runs for a number of OpenCLIP NaFlex ViT encoder based multimodal (CLIP, SigLIP, MaMMUT, CoCa) image-text models.
-
rwightman/naflex_ViT-B-32.cc12m-pj24-32-40-siglip
Zero-Shot Image Classification • Updated • 24 -
rwightman/naflex_ViT-B-32.cc12m-pj24-32-40
Zero-Shot Image Classification • Updated • 31 • 1 -
rwightman/naflex_ViT-B-32.cc12m
Zero-Shot Image Classification • Updated • 22 -
rwightman/mammut2-naflex_ViT-B-32.cc12m
Zero-Shot Image Classification • Updated • 24