This is a tutorial blog post about profiling attention mechanisms in PyTorch, part of a series. It does not announce any AI product or model.
The transformers vLLM backend now achieves native or better throughput compared to custom vLLM implementations for many LLM architectures, using torch.fx and AST to apply runtime fusions.
Hugging Face and Amazon SageMaker AI announced a deep-link integration that allows developers to go from model discovery on Hugging Face to experimentation in SageMaker Studio with a single click,…
Hugging Face announced a curated catalog of open-weight models from its ecosystem, refreshed weekly, that can be deployed in one click onto Microsoft Foundry Managed Compute, a new managed GPU…
This new release of LeRobot v0.6.0 adds new world model policies, VLAs, reward models, depth sensing, VLM-based dataset annotation, custom video encoding, cloud training on HF Jobs, and a simplified…
Hugging Face and SkyPilot announce integration allowing users to mount Hugging Face storage (buckets, repos) into SkyPilot jobs with zero egress fees, enabling cross-cloud GPU usage without data…
This is a technical blog post (Part 4 of the PRX series) describing the data strategy used to assemble training data for the PRX model, including data sources, captioning philosophy, and data…
Hugging Face announces major updates to the Kernels project, including a new kernel repository type on the Hub, improved security with trusted publishers and code signing, revamped CLIs, and extended…