9 repos
This cluster focuses on natural language processing applied to SQL generation and database querying, centered around synthetic and curated datasets for training models to convert natural language queries into SQL. The core repos include multiple versions of OmniSQL (a text-to-SQL model series) and SynSQL (a synthetic dataset), alongside supporting work on dataset creation and NLP pipelines. Visitors here will find resources for understanding semantic parsing, few-shot learning for SQL generation, and the datasets needed to benchmark these systems.