Hugging Face releases ScreenSuite, a comprehensive evaluation suite for GUI agents that includes 13 benchmarks covering perception, grounding, single-step actions, and multi-step agents, and…
This blog post explains the implementation of KV caching from scratch in the nanoVLM repository, a small pure PyTorch codebase for training Vision Language Models.
A software engineer and music producer built a personal sound generation app that runs entirely on an Arm-based CPU using the open-source Stable Audio Open model, enabling real-time audio generation…
H Company announces the release of Holo1, a new family of open-source Action Vision Language Models designed for GUI automation, including Holo1-3B and Holo1-7B.
TRL now supports co-located vLLM, allowing training and inference to share the same GPUs, reducing idle time and improving throughput without extra hardware.
SmolVLA is a compact (450M) open-source Vision-Language-Action model for robotics, trained on community data, that outperforms larger models on simulation and real-world tasks.
OpenAI is opposing a court order sought by The New York Times and other plaintiffs that would require indefinite retention of user data from ChatGPT and the API, citing privacy and data protection…
Replicate announces the community's exploration of FLUX.1 Kontext, a new image editing model from Black Forest Labs, highlighting its capabilities for text-based edits, community examples, and the…