datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-llm-synthetic-qa
Tiny-LLM: Synthetic Question-Answering Dataset
Dataset Description
This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch.
It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.cleand_openthought312_dif9_tiny元データ: https://huggingface.co/datasets/LLMTeamAkiyama/clean_openthought312_difficulty_9_filterd
データ件数: 1,456
平均トークン数: 5,894
最大トークン数: 8,186
合計トークン数: 8,581,562
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 33.2 MB
加工内容:
元データに対して、token数を8912以下に制限したテスト用tiny版
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/openthoughts3/clean_openthoughts3_tiny_pickup.ipynb
