# Databricks — Scaling document classification to 100k+ labels

- Company: Databricks (databricks.com)
- Announced: 2026-07-20T18:15:00+00:00
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://www.databricks.com/blog/scaling-document-classification-100k-labels
- Record: https://forck.live/items/17503-scaling-document-classification-to-100k-labels
- Subject: Mosaic AI

Mapping text to large taxonomies of 100,000+ labels, whether that's biomedical entity linking, vendor normalization, or company deduplication, is a common production problem where regex, trained classifiers, and direct LLM calls all struggle on cost, maintenance, and context limits. Our solution pairs vector search with the Databricks AI Classify function: retrieve a shortlist of candidate labels per document, then let the AI Classify function pick from that shortlist instead of the full taxonomy. Across three benchmarks spanning these use cases, SQL-native vector search plus AI Classify beat the best cost-efficient frontier model by five points of accuracy at roughly a hundredth of the token cost. Across Databricks, thousands of customers build production workloads that map freeform text to normalized taxonomies of 100k+ labels. A few common use cases include: Biomedical entity linking. Clinical notes and research papers mention diseases, drugs, and procedures that must be matched to a concept in the Unified Medical Language System , a vocabulary with thousands of biomedical concept identifiers. Vendor normalization. …

---

Record: https://forck.live/items/17503-scaling-document-classification-to-100k-labels
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
