# Hugging Face — tokenizers v1: encode, decode and scaling, measured

- Company: Hugging Face (huggingface.co)
- Announced: 2026-09-21
- Category: developer-tool-release
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/tokenizers-v1
- Record: https://forck.live/items/12710-tokenizers-v1-encode-decode-and-scaling-measured
- Subject: Platform

Hugging Face released tokenizers v1, a major performance update to its tokenization library that aims to be tens of times faster than v0.23 while preserving identical token IDs, API, vocabulary and merge ranks. The update introduces optimizations including SIMD-based bitstream splitting instead of regex, thread-local word caching for repeated pre-tokens, and improvements across normalization, pre-tokenization, model, and post-processing stages to prevent tokenization from becoming a bottleneck in machine learning workflows.

## Evidence

Verbatim from https://huggingface.co/blog/tokenizers-v1:

> This is why we have chosen to heavily focus on performance for the upcoming version 1 of tokenizers. Tokenization should be light and should scale with your workflow. Your GPUs should never sit idle waiting for the CPU to complete its tokenization.

---

Record: https://forck.live/items/12710-tokenizers-v1-encode-decode-and-scaling-measured
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
