Text-to-SQL and NLP Datasets

9 repos

This cluster focuses on natural language processing applied to SQL generation and database querying, centered around synthetic and curated datasets for training models to convert natural language queries into SQL. The core repos include multiple versions of OmniSQL (a text-to-SQL model series) and SynSQL (a synthetic dataset), alongside supporting work on dataset creation and NLP pipelines. Visitors here will find resources for understanding semantic parsing, few-shot learning for SQL generation, and the datasets needed to benchmark these systems.

Python · 2
HTML · 1
machine-learning ·1,802
natural-language-interface ·1,802
natural-language ·1,802
dataset ·1,802
database ·1,802
natural-language-processing ·1,802
SQL ·1,264
code ·1,201
text-to-sql ·961
datadesigner ·695