# Hugging Face — PaliGemma – Google's Cutting-Edge Open Vision Language Model

- Company: Hugging Face (huggingface.co)
- Announced: 2024-05-14T00:00:00+00:00
- Category: new-model
- Subject: Platform
- Open weights: yes
- Models affected: PaliGemma
- Source: https://huggingface.co/blog/paligemma
- Record: https://forck.live/items/1890-paligemma-google-s-cutting-edge-open-vision-language-model

Google announces PaliGemma, a new family of open vision language models combining SigLIP image encoder and Gemma text decoder. The release includes three types of checkpoints (pretrained, mix, fine-tuned) in various resolutions and precisions, available on Hugging Face Hub.

## Evidence

Verbatim from https://huggingface.co/blog/paligemma:

> PaliGemma is a new family of vision language models from Google. PaliGemma can take in an image and a text and output text. The team at Google has released three types of models: the pretrained (pt) models, the mix models, and the fine-tuned (ft) models, each with different resolutions and available in multiple precisions for convenience.

---

Record: https://forck.live/items/1890-paligemma-google-s-cutting-edge-open-vision-language-model
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
