From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Tokenizers v1 achieves major performance gains through algorithmic optimization
Hugging Face released tokenizers v1, a major performance update to its tokenization library that aims to be tens of times faster than v0.23 while preserving identical token IDs, API, vocabulary and merge ranks.
The update introduces optimizations including SIMD-based bitstream splitting instead of regex, thread-local word caching for repeated pre-tokens, and improvements across normalization, pre-tokenization, model, and post-processing stages to prevent tokenization from becoming a bottleneck in machine learning workflows.
From the source
This is why we have chosen to heavily focus on performance for the upcoming version 1 of tokenizers. Tokenization should be light and should scale with your workflow. Your GPUs should never sit idle waiting for the CPU to complete its tokenization.
huggingface.co