# Thinking Machines Lab — Putting Task Expertise into RL Achieves State-of-the-Art Performance on Text-to-SQL

- Company: Thinking Machines Lab (thinkingmachines.ai)
- Announced: 2026-08-27
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://thinkingmachines.ai/news/putting-task-expertise-into-rl
- Record: https://forck.live/items/18553-putting-task-expertise-into-rl-achieves-state-of-the-art-performance-on-text
- Subject: Thinking Machines / Inkling

Many industries rely on relational databases that are queried with SQL. Most SQL is machine-written, but humans alone likely write billions of custom SQL queries each month Based on our internal estimates and publicly available data, such as from Snowflake filings . prompted by business questions. They are quite good at it — humans score 92.96% on BIRD , a realistic benchmark for translating natural-language questions into SQL. However, AI performance on text-to-SQL has lagged behind. LLM scores on the BIRD leaderboard improved from just below 70% in 2024 to 82% today. Frontier models like GPT-5.6 Sol Ultra and Claude Fable 5 can score in the mid-80s, albeit at a cost that is prohibitive for high-volume applications. This isn’t for lack of training data: SQL is widely represented in the internet content used in LLM pretraining. The challenge for AI is in navigating the ambiguous questions and highly-contextual schema that characterize real-world examples. A common approach for improving AI performance on tasks people understand well is building agentic scaffolding. …

---

Record: https://forck.live/items/18553-putting-task-expertise-into-rl-achieves-state-of-the-art-performance-on-text
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
