datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zhtw-roleplay-space-grimoire
Space Grimoire RP Corpus (Traditional Chinese)
Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0.
中文說明在下方
Dataset Summary
Source text
283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.sql-grimoire
Image generated by DALL-E.
Grimoire of SQL
Overview
Grimoire of SQLite is a comprehensive dataset tailored for training and evaluating text-to-SQL models. It consolidates and enhances multiple existing datasets, including Spider, BirdBench, and Gretel, by correcting errors, refining natural language queries, and validating SQL queries for runnability. The dataset is specifically designed to support high-quality fine-tuning of models like GPT-4 and its variants… See the full description on the dataset page: https://huggingface.co/datasets/data-maki/sql-grimoire.Grimoire
