From the source
Most SQL is machine-written, but humans alone likely write billions of custom SQL queries each month Based on our internal estimates and publicly available data, such as from Snowflake filings . prompted by business questions.
They are quite good at it — humans score 92.96% on BIRD , a realistic benchmark for translating natural-language questions into SQL.
However, AI performance on text-to-SQL has lagged behind.
LLM scores on the BIRD leaderboard improved from just below 70% in 2024 to 82% today.
Frontier models like GPT-5.6 Sol Ultra and Claude Fable 5 can score in the mid-80s, albeit at a cost that is prohibitive for high-volume applications.
This isn’t for lack of training data: SQL is widely represented in the internet content used in LLM pretraining.
The challenge for AI is in navigating the ambiguous questions and highly-contextual schema that characterize real-world examples.
A common approach for improving AI performance on tasks people understand well is building agentic scaffolding.
…






