# Hugging Face — Zero-shot image-to-text generation with BLIP-2

- Company: Hugging Face (huggingface.co)
- Announced: 2023-02-15
- Category: new-model
- Coverage: not counted
- Announcement: yes
- Group: models
- Source: https://huggingface.co/blog/blip-2
- Record: https://forck.live/items/2096-zero-shot-image-to-text-generation-with-blip-2
- Subject: Platform
- Models affected: BLIP-2

Hugging Face introduces BLIP-2, a visual-language model from Salesforce Research that uses a Q-Former to bridge frozen image encoders and LLMs, enabling zero-shot image captioning, visual question answering, and more. The model is now available in the Transformers library.

## Evidence

Verbatim from https://huggingface.co/blog/blip-2:

> This guide introduces BLIP-2 from Salesforce Research that enables a suite of state-of-the-art visual-language models that are now available in 🤗 Transformers.

---

Record: https://forck.live/items/2096-zero-shot-image-to-text-generation-with-blip-2
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
