datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-llm-synthetic-qa
Tiny-LLM: Synthetic Question-Answering Dataset
Dataset Description
This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch.
It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.TinyLLMPretrainingCore
Synthetic Simple-English Subject Explanations Dataset
Dataset Summary
This dataset contains synthetic, GPT-generated texts that explain a wide range of subjects using simple English.Each subject is expanded into multiple long-form explanations that repeat key ideas across different styles, perspectives, and framing strategies.
The dataset is designed to emphasize clarity, redundancy, and consistency, making it useful for educational NLP, simplification tasks, and… See the full description on the dataset page: https://huggingface.co/datasets/MaxHastings/TinyLLMPretrainingCore.
