# Hugging Face — Run a vLLM Server on HF Jobs in One Command

- Company: Hugging Face (huggingface.co)
- Announced: 2026-06-26T00:00:00+00:00
- Category: capability-change
- Subject: Platform
- Models affected: Qwen/Qwen3-4B, Qwen/Qwen3.5-122B-A10B
- Pricing: pay-per-second; a10g-large runs at $1.50/hour
- Source: https://huggingface.co/blog/vllm-jobs
- Record: https://forck.live/items/1472-run-a-vllm-server-on-hf-jobs-in-one-command

Hugging Face announces that users can now spin up a private, OpenAI-compatible LLM endpoint on HF Jobs with a single command, using vLLM, with pay-per-second billing and no server provisioning.

## Evidence

Verbatim from https://huggingface.co/blog/vllm-jobs:

> You can spin up a private, OpenAI-compatible LLM endpoint on Hugging Face infrastructure with a single command — no servers to provision, no Kubernetes, pay-per-second.

---

Record: https://forck.live/items/1472-run-a-vllm-server-on-hf-jobs-in-one-command
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
