datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki40b
Dataset Card for "wiki40b"
Dataset Summary
Clean-up text for 40+ Wikipedia languages editions of pages
correspond to entities. The datasets have train/dev/test splits per language.
The dataset is cleaned up by page filtering to remove disambiguation pages,
redirect pages, deleted pages, and non-entity pages. Each example contains the
wikidata id of the entity, and the full Wikipedia article after page processing
that removes non-content sections and structured objects.… See the full description on the dataset page: https://huggingface.co/datasets/google/wiki40b.InfiMM-WebMath-40B
InfiMM-WebMath-40B Dataset
ArXiv| PDF
This dataset is also discussed in the survey paper A Survey of Deep Learning for Geometry Problem Solving.
The accompanying reading list/code for the survey can be found at: https://github.com/majianz/gps-survey
InfiMM-WebMath-40B is a large-scale, open-source multimodal dataset specifically designed for mathematical reasoning tasks. It incorporates both text and images, extracted from web documents, to advance the pre-training of Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Infi-MM/InfiMM-WebMath-40B.details_tiiuae__falcon-40b
Dataset Card for Evaluation run of tiiuae/falcon-40b
Dataset Summary
Dataset automatically created during the evaluation run of model tiiuae/falcon-40b on the Open LLM Leaderboard.
The dataset is composed of 124 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tiiuae__falcon-40b.climbmix-40b-az
ClimbMix 40B — Azerbaijani
A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate.
Dataset Summary
This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language.
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.details_AI-Sweden-Models__gpt-sw3-40b
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-40b
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-40b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-40b.wiki40b_jaThis dataset is a reformatted version of the Japanese portion of wiki40b dataset.
When you use this dataset, please cite the original paper:
@inproceedings{guo-etal-2020-wiki,
title = "{W}iki-40{B}: Multilingual Language Model Dataset",
author = "Guo, Mandy and
Dai, Zihang and
Vrande{\v{c}}i{\'c}, Denny and
Al-Rfou, Rami",
booktitle = "Proceedings of the Twelfth Language Resources and Evaluation Conference",
month = may,
year = "2020",
address =… See the full description on the dataset page: https://huggingface.co/datasets/fn-aka-mur/wiki40b_ja.wiki40b-ja
range3/wiki40b-ja
This dataset consists of three parquet files from the wiki40b dataset with only Japanese data extracted. It is generated by the following python code.
このデータセットは、wiki40bデータセットの日本語データのみを抽出した3つのparquetファイルで構成されます。以下のpythonコードによって生成しています。
import datasets
dss = datasets.load_dataset(
"wiki40b",
"ja",
beam_runner="DirectRunner",
)
for split,ds in dss.items():
ds.to_parquet(f"wikipedia-ja-20230101/{split}.parquet")
details_OpenBuddy__openbuddy-falcon-40b-v16.1-4k
Dataset Card for Evaluation run of OpenBuddy/openbuddy-falcon-40b-v16.1-4k
Dataset automatically created during the evaluation run of model OpenBuddy/openbuddy-falcon-40b-v16.1-4k on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_OpenBuddy__openbuddy-falcon-40b-v16.1-4k.wiki40b_lines_en40b5095340B_dataset_W_fixeddetails_tiiuae__falcon-40b
Dataset Card for Evaluation run of tiiuae/falcon-40b
Dataset automatically created during the evaluation run of model tiiuae/falcon-40b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_tiiuae__falcon-40b.wiki40b_lines_eswiki40b_qa_ja_train_smalltiiuae__falcon-40b-details
Dataset Card for Evaluation run of tiiuae/falcon-40b
Dataset automatically created during the evaluation run of model tiiuae/falcon-40b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-40b-details.wiki40b-da
Dataset Card for "wiki40b-da"
Dataset Summary
This dataset is an upload of the Danish part of the Wiki40b dataset, being a cleaned version of a dump of Wikipedia.
The dataset is identical in content to this dataset on the Hugging Face Hub, but that one requires both apache_beam, tensorflow and mwparserfromhell, which can lead to dependency issues since these are not compatible with several newer packages.
The training, validation and test splits are the original ones.… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/wiki40b-da.details_dfurman__falcon-40b-openassistant-peft
Dataset Card for Evaluation run of dfurman/falcon-40b-openassistant-peft
Dataset Summary
Dataset automatically created during the evaluation run of model dfurman/falcon-40b-openassistant-peft on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dfurman__falcon-40b-openassistant-peft.wiki40b_en_100_0_split
Dataset Card for "wiki40b_en_100_0_split"
More Information needed
Knowledge_distilled_dataset_by_Fuka2025Q2-40b_qsearch将棋AI用の知識蒸留済みのデータセットを公開します。およそ80億局面あります。 nodchip氏が公開しているtanuki-.nnue-pytorch-2024-07-30.1をhaoでqsearchシャッフルしたのちふかうら王 NewsLetterで配布されたFuka2025Q2-40bで評価値を書き換えました。Eval_Coef=600でDLモデルのvalueと評価値を変換しています。 データにバグがあるかもしれませんが、品質保証はしません。
https://huggingface.co/datasets/nodchip/tanuki-.nnue-pytorch-2024-07-30.1https://yaneurao.fanbox.cc/
yulanmini_phase26_40bmoe-40b-a18b-dataset
moe-40b-a18b-dataset
Training data for Leonther/moe-40b-a18b-lora — an experimental 40B-A18B MoE
student distilled from GLM-5.3-Flash.
Composition
file
records
source
description
gold/professor.jsonl
2
GLM-5.3-Flash (local, ~2 t/s)
gold question/answer pairs
gold/professor_raw.jsonl
2
GLM-5.3-Flash
raw batch output (ids, usage)
gold/seeds.jsonl
22
qwen38
professor prompt seeds (domain + length bucket)
fake/synthetic.jsonl
100
qwen38 (Qwen3.8-27B… See the full description on the dataset page: https://huggingface.co/datasets/Leonther/moe-40b-a18b-dataset.wiki40b-lm-en
wiki40b-lm-en
wiki40b-lm-tensorflow1-en-v1
Source: https://www.kaggle.com/models/google/wiki40b-lm/tensorFlow1/en
Qwen2.5-distill-v3-instruct-dataset-base-chat-template-40B-20260322tiiuae__falcon-40b-instruct-details
Dataset Card for Evaluation run of tiiuae/falcon-40b-instruct
Dataset automatically created during the evaluation run of model tiiuae/falcon-40b-instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-40b-instruct-details.dclm_baseline_1.0_40btwiki40b-da-clean
Dataset Card for "wiki40b-da-clean"
Dataset Summary
This dataset is an slightly modified and filtered version of Wiki40b-da daset which is a fork of this dataset on the Hugging Face Hub.
The dataset contains two sub-sets, for which the original columns "wikidata_id" and "version_id" are removed from both:
"text": Contains the filtered text of the Wikipedia paragraphs, with formatting removed (START_ARTICLE, START_PARAGRAPH and \n removed)
"sentences" Contains the… See the full description on the dataset page: https://huggingface.co/datasets/jealk/wiki40b-da-clean.AI-Sweden-Models__gpt-sw3-40b-details
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-40b
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-40b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-Sweden-Models__gpt-sw3-40b-details.infimm-webmath-40b_urls
Dataset Card for infimm-webmath-40b_urls
This dataset provides the URLs and top-level domains associated with training records in Infi-MM/InfiMM-WebMath-40B. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/infimm-webmath-40b_urls.wiki40b_binidxLaMini-LM-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML-eval-llama2
Dataset Card for "LaMini-LM-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML"
More Information needed
