CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-llm-leaderboard-old /details_BramVanroy__llama2-13b-ft-mc4_nl_cleaned_tiny Dataset Card for Evaluation run of BramVanroy/llama2-13b-ft-mc4_nl_cleaned_tiny Dataset Summary Dataset automatically created during the evaluation run of model BramVanroy/llama2-13b-ft-mc4_nl_cleaned_tiny on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_BramVanroy__llama2-13b-ft-mc4_nl_cleaned_tiny.0 likes260 downloads3y agoHugging Face02open-llm-leaderboard-old /details_freecs__Tiny-Llama-3-7b Dataset Card for Evaluation run of freecs/Tiny-Llama-3-7b Dataset automatically created during the evaluation run of model freecs/Tiny-Llama-3-7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_freecs__Tiny-Llama-3-7b.1 likes200 downloads3y agoHugging Face03open-llm-leaderboard-old /details_dball__zephyr-tiny-sft-qlora-quantized-2 Dataset Card for Evaluation run of dball/zephyr-tiny-sft-qlora-quantized-2 Dataset automatically created during the evaluation run of model dball/zephyr-tiny-sft-qlora-quantized-2 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dball__zephyr-tiny-sft-qlora-quantized-2.1 likes140 downloads3y agoHugging Face04Gabriel8 /tiny-llm-synthetic-qa Tiny-LLM: Synthetic Question-Answering Dataset Dataset Description This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch. It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.textquestion-answering100K<n<1M2 likes113 downloads11mo agoHugging Face05open-llm-leaderboard-old /details_dball__zephyr-tiny-dpo-qlora Dataset Card for Evaluation run of dball/zephyr-tiny-dpo-qlora Dataset automatically created during the evaluation run of model dball/zephyr-tiny-dpo-qlora on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dball__zephyr-tiny-dpo-qlora.0 likes102 downloads3y agoHugging Face06open-llm-leaderboard-old /details_phanerozoic__Tiny-Pirate-1.1b-v0.1 Dataset Card for Evaluation run of phanerozoic/Tiny-Pirate-1.1b-v0.1 Dataset automatically created during the evaluation run of model phanerozoic/Tiny-Pirate-1.1b-v0.1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_phanerozoic__Tiny-Pirate-1.1b-v0.1.0 likes87 downloads2y agoHugging Face07open-llm-leaderboard-old /details_Danielbrdz__Barcenas-Tiny-1.1b-DPO Dataset Card for Evaluation run of Danielbrdz/Barcenas-Tiny-1.1b-DPO Dataset automatically created during the evaluation run of model Danielbrdz/Barcenas-Tiny-1.1b-DPO on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Danielbrdz__Barcenas-Tiny-1.1b-DPO.0 likes57 downloads3y agoHugging Face08open-llm-leaderboard-old /details_bigcode__tiny_starcoder_py Dataset Card for Evaluation run of bigcode/tiny_starcoder_py Dataset Summary Dataset automatically created during the evaluation run of model bigcode/tiny_starcoder_py on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigcode__tiny_starcoder_py.1 likes56 downloads3y agoHugging Face09tinyllms /aime-1983-2023-trajectoriestext1K<n<10K0 likes55 downloads6mo agoHugging Face10open-llm-leaderboard-old /details_uukuguy__GDC-Tiny-L1-1.8B Dataset Card for Evaluation run of uukuguy/GDC-Tiny-L1-1.8B Dataset automatically created during the evaluation run of model uukuguy/GDC-Tiny-L1-1.8B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_uukuguy__GDC-Tiny-L1-1.8B.0 likes52 downloads2y agoHugging Face11tinyllms /gpqa-extended-trajectoriestext1K<n<10K0 likes40 downloads6mo agoHugging Face12open-llm-leaderboard-old /details_phanerozoic__Tiny-Cowboy-1.1b-v0.1 Dataset Card for Evaluation run of phanerozoic/Tiny-Cowboy-1.1b-v0.1 Dataset automatically created during the evaluation run of model phanerozoic/Tiny-Cowboy-1.1b-v0.1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_phanerozoic__Tiny-Cowboy-1.1b-v0.1.0 likes35 downloads2y agoHugging Face13open-llm-leaderboard-old /details_anton-l__gpt-j-tiny-random Dataset Card for Evaluation run of anton-l/gpt-j-tiny-random Dataset Summary Dataset automatically created during the evaluation run of model anton-l/gpt-j-tiny-random on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_anton-l__gpt-j-tiny-random.0 likes33 downloads3y agoHugging Face14zesen-kth /tiny-llm tiny-llm retained training evidence 1782 seed/horizon records across 594 configurations: cosine-to-zero and WSD schedules, synchronous and four- and eight-worker decentralized training, at 20 to 160 global tokens per parameter, for 20.4M-parameter models on C4. Snapshot 2026-09-17. This mirrors doc/data/current-training/ in WangZesen/tiny-llm. The website built from it is at https://wangzesen.github.io/tiny-llm/. Layout Retained artifacts are grouped into gzipped… See the full description on the dataset page: https://huggingface.co/datasets/zesen-kth/tiny-llm.1K<n<10K0 likes32 downloads4d agoHugging Face15shekhar1536 /tinyllm-data0 likes31 downloads24d agoHugging Face16open-llm-leaderboard-old /details_dpv__finetuned-gpt2-tiny Dataset Card for Evaluation run of dpv/finetuned-gpt2-tiny Dataset Summary Dataset automatically created during the evaluation run of model dpv/finetuned-gpt2-tiny on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dpv__finetuned-gpt2-tiny.0 likes30 downloads3y agoHugging Face17open-llm-leaderboard-old /details_raidhon__coven_tiny_1.1b_32k_orpo_alpha0 likes30 downloads2y agoHugging Face18open-llm-leaderboard-old /details_Josephgflowers__Qllama-tiny-.5B-test-10 likes27 downloads2y agoHugging Face19tinyllms /gpqa-main-trajectoriestext1K<n<10K0 likes27 downloads6mo agoHugging Face20open-llm-leaderboard-old /details_Josephgflowers__TinyLlama-Cinder-Tiny-Agent0 likes26 downloads2y agoHugging Face21MaxHastings /TinyLLMPretrainingCore Synthetic Simple-English Subject Explanations Dataset Dataset Summary This dataset contains synthetic, GPT-generated texts that explain a wide range of subjects using simple English.Each subject is expanded into multiple long-form explanations that repeat key ideas across different styles, perspectives, and framing strategies. The dataset is designed to emphasize clarity, redundancy, and consistency, making it useful for educational NLP, simplification tasks, and… See the full description on the dataset page: https://huggingface.co/datasets/MaxHastings/TinyLLMPretrainingCore.texttext-generation10K<n<100K0 likes25 downloads9mo agoHugging Face22LLMTeamAkiyama /cleand_openthought312_dif9_tiny元データ: https://huggingface.co/datasets/LLMTeamAkiyama/clean_openthought312_difficulty_9_filterd データ件数: 1,456 平均トークン数: 5,894 最大トークン数: 8,186 合計トークン数: 8,581,562 ファイル形式: JSONL ファイル分割数: 1 合計ファイルサイズ: 33.2 MB 加工内容: 元データに対して、token数を8912以下に制限したテスト用tiny版 使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/openthoughts3/clean_openthoughts3_tiny_pickup.ipynb tabularquestion-answering1K<n<10K0 likes21 downloads1y agoHugging Face23open-llm-leaderboard /bond005__meno-tiny-0.1-detailsgated Dataset Card for Evaluation run of bond005/meno-tiny-0.1 Dataset automatically created during the evaluation run of model bond005/meno-tiny-0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bond005__meno-tiny-0.1-details.tabular10K<n<100K0 likes17 downloads2y agoHugging Face24open-llm-leaderboard-old /details_phanerozoic__Tiny-Knight-1.1b-v0.10 likes15 downloads2y agoHugging Face25open-llm-leaderboard /prithivMLmods__Bellatrix-Tiny-1.5B-R1-detailsgated Dataset Card for Evaluation run of prithivMLmods/Bellatrix-Tiny-1.5B-R1 Dataset automatically created during the evaluation run of model prithivMLmods/Bellatrix-Tiny-1.5B-R1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Bellatrix-Tiny-1.5B-R1-details.tabular10K<n<100K0 likes15 downloads2y agoHugging Face26open-llm-leaderboard /prithivMLmods__Bellatrix-Tiny-1B-v2-detailsgated Dataset Card for Evaluation run of prithivMLmods/Bellatrix-Tiny-1B-v2 Dataset automatically created during the evaluation run of model prithivMLmods/Bellatrix-Tiny-1B-v2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Bellatrix-Tiny-1B-v2-details.tabular10K<n<100K0 likes14 downloads2y agoHugging Face27RonanMcGovern /eval-llm-lingo-tiny-llm-lingo-20251226-0016textn<1K0 likes14 downloads9mo agoHugging Face28open-llm-leaderboard-old /details_Jiayi-Pan__Tiny-Vicuna-1B Dataset Card for Evaluation run of Jiayi-Pan/Tiny-Vicuna-1B Dataset Summary Dataset automatically created during the evaluation run of model Jiayi-Pan/Tiny-Vicuna-1B on the Open LLM Leaderboard. The dataset is composed of 1 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Jiayi-Pan__Tiny-Vicuna-1B.0 likes13 downloads3y agoHugging Face29RonanMcGovern /eval-llm-lingo-tiny-llm-lingo-20251226-0008textn<1K0 likes12 downloads9mo agoHugging Face30TinyLLM /hand_gesture_dataThe study is conducted on a total of 7 participants. The participants were instructed to perform three hand gestures (Hold, Single Tap and Double Tap) under different light conditions (low(100-200 lux, medium (600-750 lux) and high (1500-1600 lux)) and at different distances from the light sensor (low(2-4 cm) and high(8-10 cm)) tabular10K<n<100K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.