datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KhanomTanLLM-pretrained-dataset
KhanomTanLLM pretrained dataset
This daataset collect all raw text for pretraining LLM.
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Tokens
53,376,211,711 Tokens
English: 31,629,984,243 Tokens
Thai: 12,785,565,497 Tokens
Code: 8,913,084,300 Toekns
Parallel data: 190,310,686 Tokens
Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer
All subset
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.KhanomTanLLM-pretrained-dataset-thai-subset
KhanomTanLLM pretrained dataset (Thai subset)
This daataset collect all raw text for pretraining LLM. (Thai subset)
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0
pythainlp/thai-tnhc2-books
pythainlp/thai-constitution-corpus
pythainlp/thai-it-books
pythainlp/prd_news_3011202
pythainlp/thailand-policy-statements
pythainlp/thai-cc-license
pythainlp/blognone_news
pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.FlofloB__smollm2-135M_pretrained_1000k_fineweb-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb-details.FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-details.filtered_apertus_pretrainedUsed to create this dataset:
import json
import os
from datasets import load_dataset
from tqdm import tqdm
# --- Configuration ---
DATASET_NAME = "swiss-ai/apertus-pretrain-swiss"
SUBSTRING_TO_FILTER = "entscheidsuche_html"
COLUMN_TO_CHECK = "id"
OUTPUT_FILENAME = "filtered_apertus_pretrain_swiss.jsonl"
def filter_function(example):
"""
Returns True to keep the example, False to discard it.
We keep the row only if the substring is NOT in the 'id' column.
"""
return… See the full description on the dataset page: https://huggingface.co/datasets/liechticonsulting/filtered_apertus_pretrained.athirdpath__Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit-details
Dataset Card for Evaluation run of athirdpath/Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit
Dataset automatically created during the evaluation run of model athirdpath/Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/athirdpath__Llama-3.1-Instruct_NSFW-pretrained_e1-plus_reddit-details.FlofloB__smollm2-135M_pretrained_1400k_fineweb_uncovai_selected-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1400k_fineweb_uncovai_selected
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1400k_fineweb_uncovai_selected
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1400k_fineweb_uncovai_selected-details.FlofloB__smollm2-135M_pretrained_600k_fineweb-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_600k_fineweb
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_600k_fineweb
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_600k_fineweb-details.FlofloB__smollm2-135M_pretrained_1000k_fineweb_uncovai_human_removed-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb_uncovai_human_removed
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb_uncovai_human_removed
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb_uncovai_human_removed-details.FlofloB__smollm2_pretrained_200k_fineweb-details
Dataset Card for Evaluation run of FlofloB/smollm2_pretrained_200k_fineweb
Dataset automatically created during the evaluation run of model FlofloB/smollm2_pretrained_200k_fineweb
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2_pretrained_200k_fineweb-details.FlofloB__smollm2-135M_pretrained_400k_fineweb-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_400k_fineweb
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_400k_fineweb
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_400k_fineweb-details.FlofloB__smollm2-135M_pretrained_1200k_fineweb_uncovai_human_removed-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1200k_fineweb_uncovai_human_removed
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1200k_fineweb_uncovai_human_removed
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1200k_fineweb_uncovai_human_removed-details.FlofloB__smollm2-135M_pretrained_800k_fineweb-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb-details.FlofloB__smollm2-135M_pretrained_200k_fineweb_uncovai_human_removed-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_200k_fineweb_uncovai_human_removed
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_200k_fineweb_uncovai_human_removed
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_200k_fineweb_uncovai_human_removed-details.FlofloB__smollm2-135M_pretrained_400k_fineweb_uncovai_human_removed-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_400k_fineweb_uncovai_human_removed
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_400k_fineweb_uncovai_human_removed
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_400k_fineweb_uncovai_human_removed-details.FlofloB__smollm2-135M_pretrained_1400k_fineweb_uncovai_human_removed-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1400k_fineweb_uncovai_human_removed
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1400k_fineweb_uncovai_human_removed
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1400k_fineweb_uncovai_human_removed-details.FlofloB__smollm2-135M_pretrained_200k_fineweb_uncovai_selected-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_200k_fineweb_uncovai_selected
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_200k_fineweb_uncovai_selected
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_200k_fineweb_uncovai_selected-details.FlofloB__smollm2-135M_pretrained_1200k_fineweb_uncovai_selected-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1200k_fineweb_uncovai_selected
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1200k_fineweb_uncovai_selected
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1200k_fineweb_uncovai_selected-details.FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_human_removed-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_human_removed
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_human_removed
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_human_removed-details.FlofloB__smollm2-135M_pretrained_600k_fineweb_uncovai_human_removed-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_600k_fineweb_uncovai_human_removed
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_600k_fineweb_uncovai_human_removed
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_600k_fineweb_uncovai_human_removed-details.LOGS-Pretrained-Codename-75567-V1viet-pretrained-001FlofloB__smollm2-135M_pretrained_400k_fineweb_uncovai_selected-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_400k_fineweb_uncovai_selected
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_400k_fineweb_uncovai_selected
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_400k_fineweb_uncovai_selected-details.FlofloB__smollm2-135M_pretrained_600k_fineweb_uncovai_selected-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_600k_fineweb_uncovai_selected
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_600k_fineweb_uncovai_selected
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_600k_fineweb_uncovai_selected-details.FlofloB__smollm2-135M_pretrained_1000k_fineweb_uncovai_selected-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb_uncovai_selected
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb_uncovai_selected
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb_uncovai_selected-details.ontocord__wide_3b-stage1_shuf_sample1_jsonl-pretrained-details
Dataset Card for Evaluation run of ontocord/wide_3b-stage1_shuf_sample1_jsonl-pretrained
Dataset automatically created during the evaluation run of model ontocord/wide_3b-stage1_shuf_sample1_jsonl-pretrained
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b-stage1_shuf_sample1_jsonl-pretrained-details.PretrainedData
Saidzoda Lab — Gated Research Dataset
Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz).
Dataset contents, provenance, and statistics are not publicly disclosed.
Access is granted manually on request.
FlofloB__smollm2-135M_pretrained_1200k_fineweb-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1200k_fineweb
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1200k_fineweb
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1200k_fineweb-details.generic_brand_count_pretrained_dataCount of generic name and brand name of cancer drugs in various pretrained corpora.
SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details
Dataset Card for Evaluation run of SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo
Dataset automatically created during the evaluation run of model SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details.
