datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multidomain-complex-text-pool
Complex Text Pool Dataset
Overview
A curated collection of complex, long-form English texts sampled from 9 diverse domains. Each document has been truncated to a maximum of 4,000 characters, preserving clean sentence boundaries. The dataset is designed to provide challenging, real-world text samples across multiple subject areas.
Categories and Sample Counts
Category
Samples
news
9,999
encyclopedic
10,000
conversational
10,000… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/multidomain-complex-text-pool.synthetic-complex-Text-to-SQL
Synthetic Complex Text-to-SQL
Synthetic Text-to-SQL using multiple joins, WHERE statements, window and aggregate functions on filtered SQL from bigcode/the-stack.
