CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gabriel8 /tiny-llm-synthetic-qa Tiny-LLM: Synthetic Question-Answering Dataset Dataset Description This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch. It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.textquestion-answering100K<n<1M2 likes113 downloads11mo agoHugging Face02MaxHastings /TinyLLMPretrainingCore Synthetic Simple-English Subject Explanations Dataset Dataset Summary This dataset contains synthetic, GPT-generated texts that explain a wide range of subjects using simple English.Each subject is expanded into multiple long-form explanations that repeat key ideas across different styles, perspectives, and framing strategies. The dataset is designed to emphasize clarity, redundancy, and consistency, making it useful for educational NLP, simplification tasks, and… See the full description on the dataset page: https://huggingface.co/datasets/MaxHastings/TinyLLMPretrainingCore.texttext-generation10K<n<100K0 likes25 downloads9mo agoHugging Face03LLMTeamAkiyama /cleand_openthought312_dif9_tiny元データ: https://huggingface.co/datasets/LLMTeamAkiyama/clean_openthought312_difficulty_9_filterd データ件数: 1,456 平均トークン数: 5,894 最大トークン数: 8,186 合計トークン数: 8,581,562 ファイル形式: JSONL ファイル分割数: 1 合計ファイルサイズ: 33.2 MB 加工内容: 元データに対して、token数を8912以下に制限したテスト用tiny版 使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/openthoughts3/clean_openthoughts3_tiny_pickup.ipynb tabularquestion-answering1K<n<10K0 likes21 downloads1y agoHugging Face04open-llm-leaderboard /bond005__meno-tiny-0.1-detailsgated Dataset Card for Evaluation run of bond005/meno-tiny-0.1 Dataset automatically created during the evaluation run of model bond005/meno-tiny-0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bond005__meno-tiny-0.1-details.tabular10K<n<100K0 likes17 downloads2y agoHugging Face05open-llm-leaderboard /prithivMLmods__Bellatrix-Tiny-1.5B-R1-detailsgated Dataset Card for Evaluation run of prithivMLmods/Bellatrix-Tiny-1.5B-R1 Dataset automatically created during the evaluation run of model prithivMLmods/Bellatrix-Tiny-1.5B-R1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Bellatrix-Tiny-1.5B-R1-details.tabular10K<n<100K0 likes15 downloads2y agoHugging Face06open-llm-leaderboard /prithivMLmods__Bellatrix-Tiny-1B-v2-detailsgated Dataset Card for Evaluation run of prithivMLmods/Bellatrix-Tiny-1B-v2 Dataset automatically created during the evaluation run of model prithivMLmods/Bellatrix-Tiny-1B-v2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Bellatrix-Tiny-1B-v2-details.tabular10K<n<100K0 likes14 downloads2y agoHugging Face07open-llm-leaderboard /prithivMLmods__FastThink-0.5B-Tiny-detailsgated Dataset Card for Evaluation run of prithivMLmods/FastThink-0.5B-Tiny Dataset automatically created during the evaluation run of model prithivMLmods/FastThink-0.5B-Tiny The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__FastThink-0.5B-Tiny-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face08open-llm-leaderboard /V3N0M__Jenna-Tiny-2.0-detailsgated Dataset Card for Evaluation run of V3N0M/Jenna-Tiny-2.0 Dataset automatically created during the evaluation run of model V3N0M/Jenna-Tiny-2.0 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/V3N0M__Jenna-Tiny-2.0-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.