Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
smolagents now supports vision language models (VLMs) by allowing images to be passed to the agent at start or dynamically via callbacks, enabling use cases like autonomous web browsing.
From the source
We have added vision support to smolagents, which unlocks the use of vision language models in agentic pipelines natively.
huggingface.co