Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Google announces PaliGemma, a new family of open vision language models combining SigLIP image encoder and Gemma text decoder. The release includes three types of checkpoints (pretrained, mix, fine-tuned) in various resolutions and precisions, available on Hugging Face Hub.
From the source
PaliGemma is a new family of vision language models from Google. PaliGemma can take in an image and a text and output text. The team at Google has released three types of models: the pretrained (pt) models, the mix models, and the fine-tuned (ft) models, each with different resolutions and available in multiple precisions for convenience.
huggingface.co