CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wannaphong /KhanomTanLLM-pretrained-dataset KhanomTanLLM pretrained dataset This daataset collect all raw text for pretraining LLM. Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Tokens 53,376,211,711 Tokens English: 31,629,984,243 Tokens Thai: 12,785,565,497 Tokens Code: 8,913,084,300 Toekns Parallel data: 190,310,686 Tokens Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer All subset Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.texttext-generation10M<n<100M1 likes235 downloads2y agoHugging Face02wannaphong /KhanomTanLLM-pretrained-dataset-thai-subset KhanomTanLLM pretrained dataset (Thai subset) This daataset collect all raw text for pretraining LLM. (Thai subset) Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0 pythainlp/thai-tnhc2-books pythainlp/thai-constitution-corpus pythainlp/thai-it-books pythainlp/prd_news_3011202 pythainlp/thailand-policy-statements pythainlp/thai-cc-license pythainlp/blognone_news pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.texttext-generation10M<n<100M0 likes64 downloads2y agoHugging Face03open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1000k_fineweb-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face04open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face05liechticonsulting /filtered_apertus_pretrainedUsed to create this dataset: import json import os from datasets import load_dataset from tqdm import tqdm # --- Configuration --- DATASET_NAME = "swiss-ai/apertus-pretrain-swiss" SUBSTRING_TO_FILTER = "entscheidsuche_html" COLUMN_TO_CHECK = "id" OUTPUT_FILENAME = "filtered_apertus_pretrain_swiss.jsonl" def filter_function(example): """ Returns True to keep the example, False to discard it. We keep the row only if the substring is NOT in the 'id' column. """ return… See the full description on the dataset page: https://huggingface.co/datasets/liechticonsulting/filtered_apertus_pretrained.text100K<n<1M0 likes22 downloads1y agoHugging Face06open-llm-leaderboard /athirdpath__Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit-detailsgated Dataset Card for Evaluation run of athirdpath/Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit Dataset automatically created during the evaluation run of model athirdpath/Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/athirdpath__Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face07open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1400k_fineweb_uncovai_selected-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1400k_fineweb_uncovai_selected Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1400k_fineweb_uncovai_selected The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1400k_fineweb_uncovai_selected-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face08open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_600k_fineweb-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_600k_fineweb Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_600k_fineweb The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_600k_fineweb-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face09open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1000k_fineweb_uncovai_human_removed-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb_uncovai_human_removed Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb_uncovai_human_removed The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb_uncovai_human_removed-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face10open-llm-leaderboard /FlofloB__smollm2_pretrained_200k_fineweb-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2_pretrained_200k_fineweb Dataset automatically created during the evaluation run of model FlofloB/smollm2_pretrained_200k_fineweb The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2_pretrained_200k_fineweb-details.tabular10K<n<100K0 likes9 downloads2y agoHugging Face11open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_400k_fineweb-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_400k_fineweb Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_400k_fineweb The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_400k_fineweb-details.tabular10K<n<100K0 likes9 downloads2y agoHugging Face12open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1200k_fineweb_uncovai_human_removed-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1200k_fineweb_uncovai_human_removed Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1200k_fineweb_uncovai_human_removed The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1200k_fineweb_uncovai_human_removed-details.tabular10K<n<100K0 likes9 downloads2y agoHugging Face13open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_800k_fineweb-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face14open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_200k_fineweb_uncovai_human_removed-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_200k_fineweb_uncovai_human_removed Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_200k_fineweb_uncovai_human_removed The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_200k_fineweb_uncovai_human_removed-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face15open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_400k_fineweb_uncovai_human_removed-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_400k_fineweb_uncovai_human_removed Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_400k_fineweb_uncovai_human_removed The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_400k_fineweb_uncovai_human_removed-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face16open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1400k_fineweb_uncovai_human_removed-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1400k_fineweb_uncovai_human_removed Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1400k_fineweb_uncovai_human_removed The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1400k_fineweb_uncovai_human_removed-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face17open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_200k_fineweb_uncovai_selected-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_200k_fineweb_uncovai_selected Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_200k_fineweb_uncovai_selected The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_200k_fineweb_uncovai_selected-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face18open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1200k_fineweb_uncovai_selected-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1200k_fineweb_uncovai_selected Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1200k_fineweb_uncovai_selected The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1200k_fineweb_uncovai_selected-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face19open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_human_removed-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_human_removed Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_human_removed The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_human_removed-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face20open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_600k_fineweb_uncovai_human_removed-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_600k_fineweb_uncovai_human_removed Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_600k_fineweb_uncovai_human_removed The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_600k_fineweb_uncovai_human_removed-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face21AiAF /LOGS-Pretrained-Codename-75567-V1textn<1K1 likes7 downloads2y agoHugging Face22vilm /viet-pretrained-001gatedtext10K<n<100K0 likes6 downloads3y agoHugging Face23open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_400k_fineweb_uncovai_selected-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_400k_fineweb_uncovai_selected Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_400k_fineweb_uncovai_selected The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_400k_fineweb_uncovai_selected-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face24open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_600k_fineweb_uncovai_selected-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_600k_fineweb_uncovai_selected Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_600k_fineweb_uncovai_selected The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_600k_fineweb_uncovai_selected-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face25open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1000k_fineweb_uncovai_selected-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb_uncovai_selected Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb_uncovai_selected The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb_uncovai_selected-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face26open-llm-leaderboard /ontocord__wide_3b-stage1_shuf_sample1_jsonl-pretrained-detailsgated Dataset Card for Evaluation run of ontocord/wide_3b-stage1_shuf_sample1_jsonl-pretrained Dataset automatically created during the evaluation run of model ontocord/wide_3b-stage1_shuf_sample1_jsonl-pretrained The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b-stage1_shuf_sample1_jsonl-pretrained-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face27Tohirju /PretrainedDatagated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. 100K<n<1M0 likes6 downloads1mo agoHugging Face28open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1200k_fineweb-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1200k_fineweb Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1200k_fineweb The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1200k_fineweb-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face29mingye94 /generic_brand_count_pretrained_dataCount of generic name and brand name of cancer drugs in various pretrained corpora. textn<1K0 likes4 downloads2y agoHugging Face30open-llm-leaderboard /SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-detailsgated Dataset Card for Evaluation run of SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo Dataset automatically created during the evaluation run of model SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details.tabular10K<n<100K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.