Zero-shot image-to-text generation with BLIP-2
Hugging Face introduces BLIP-2, a visual-language model from Salesforce Research that uses a Q-Former to bridge frozen image encoders and LLMs, enabling zero-shot image captioning, visual question answering, and more. The model is now available in the Transformers library.
