datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-llm-synthetic-qa
Tiny-LLM: Synthetic Question-Answering Dataset
Dataset Description
This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch.
It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.TinyLLMPretrainingCore
Synthetic Simple-English Subject Explanations Dataset
Dataset Summary
This dataset contains synthetic, GPT-generated texts that explain a wide range of subjects using simple English.Each subject is expanded into multiple long-form explanations that repeat key ideas across different styles, perspectives, and framing strategies.
The dataset is designed to emphasize clarity, redundancy, and consistency, making it useful for educational NLP, simplification tasks, and… See the full description on the dataset page: https://huggingface.co/datasets/MaxHastings/TinyLLMPretrainingCore.cleand_openthought312_dif9_tiny元データ: https://huggingface.co/datasets/LLMTeamAkiyama/clean_openthought312_difficulty_9_filterd
データ件数: 1,456
平均トークン数: 5,894
最大トークン数: 8,186
合計トークン数: 8,581,562
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 33.2 MB
加工内容:
元データに対して、token数を8912以下に制限したテスト用tiny版
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/openthoughts3/clean_openthoughts3_tiny_pickup.ipynb
bond005__meno-tiny-0.1-details
Dataset Card for Evaluation run of bond005/meno-tiny-0.1
Dataset automatically created during the evaluation run of model bond005/meno-tiny-0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bond005__meno-tiny-0.1-details.prithivMLmods__Bellatrix-Tiny-1.5B-R1-details
Dataset Card for Evaluation run of prithivMLmods/Bellatrix-Tiny-1.5B-R1
Dataset automatically created during the evaluation run of model prithivMLmods/Bellatrix-Tiny-1.5B-R1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Bellatrix-Tiny-1.5B-R1-details.prithivMLmods__Bellatrix-Tiny-1B-v2-details
Dataset Card for Evaluation run of prithivMLmods/Bellatrix-Tiny-1B-v2
Dataset automatically created during the evaluation run of model prithivMLmods/Bellatrix-Tiny-1B-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Bellatrix-Tiny-1B-v2-details.prithivMLmods__FastThink-0.5B-Tiny-details
Dataset Card for Evaluation run of prithivMLmods/FastThink-0.5B-Tiny
Dataset automatically created during the evaluation run of model prithivMLmods/FastThink-0.5B-Tiny
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__FastThink-0.5B-Tiny-details.V3N0M__Jenna-Tiny-2.0-details
Dataset Card for Evaluation run of V3N0M/Jenna-Tiny-2.0
Dataset automatically created during the evaluation run of model V3N0M/Jenna-Tiny-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/V3N0M__Jenna-Tiny-2.0-details.
