CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /wiki40b Dataset Card for "wiki40b" Dataset Summary Clean-up text for 40+ Wikipedia languages editions of pages correspond to entities. The datasets have train/dev/test splits per language. The dataset is cleaned up by page filtering to remove disambiguation pages, redirect pages, deleted pages, and non-entity pages. Each example contains the wikidata id of the entity, and the full Wikipedia article after page processing that removes non-content sections and structured objects.… See the full description on the dataset page: https://huggingface.co/datasets/google/wiki40b.text10M<n<100M37 likes11k downloads3y agoHugging Face02Infi-MM /InfiMM-WebMath-40B InfiMM-WebMath-40B Dataset ArXiv| PDF This dataset is also discussed in the survey paper A Survey of Deep Learning for Geometry Problem Solving. The accompanying reading list/code for the survey can be found at: https://github.com/majianz/gps-survey InfiMM-WebMath-40B is a large-scale, open-source multimodal dataset specifically designed for mathematical reasoning tasks. It incorporates both text and images, extracted from web documents, to advance the pre-training of Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Infi-MM/InfiMM-WebMath-40B.textimage-text-to-text10M<n<100M69 likes2.7k downloads1y agoHugging Face03open-llm-leaderboard-old /details_tiiuae__falcon-40b Dataset Card for Evaluation run of tiiuae/falcon-40b Dataset Summary Dataset automatically created during the evaluation run of model tiiuae/falcon-40b on the Open LLM Leaderboard. The dataset is composed of 124 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tiiuae__falcon-40b.0 likes674 downloads3y agoHugging Face04LocalDoc /climbmix-40b-az ClimbMix 40B — Azerbaijani A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate. Dataset Summary This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language. Property Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.texttext-generation10M<n<100M2 likes511 downloads6mo agoHugging Face05open-llm-leaderboard-old /details_AI-Sweden-Models__gpt-sw3-40b Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-40b Dataset Summary Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-40b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-40b.0 likes311 downloads3y agoHugging Face06fn-aka-mur /wiki40b_jaThis dataset is a reformatted version of the Japanese portion of wiki40b dataset. When you use this dataset, please cite the original paper: @inproceedings{guo-etal-2020-wiki, title = "{W}iki-40{B}: Multilingual Language Model Dataset", author = "Guo, Mandy and Dai, Zihang and Vrande{\v{c}}i{\'c}, Denny and Al-Rfou, Rami", booktitle = "Proceedings of the Twelfth Language Resources and Evaluation Conference", month = may, year = "2020", address =… See the full description on the dataset page: https://huggingface.co/datasets/fn-aka-mur/wiki40b_ja.text100K<n<1M5 likes195 downloads3y agoHugging Face07range3 /wiki40b-ja range3/wiki40b-ja This dataset consists of three parquet files from the wiki40b dataset with only Japanese data extracted. It is generated by the following python code. このデータセットは、wiki40bデータセットの日本語データのみを抽出した3つのparquetファイルで構成されます。以下のpythonコードによって生成しています。 import datasets dss = datasets.load_dataset( "wiki40b", "ja", beam_runner="DirectRunner", ) for split,ds in dss.items(): ds.to_parquet(f"wikipedia-ja-20230101/{split}.parquet") texttext-generation100K<n<1M11 likes156 downloads4y agoHugging Face08open-llm-leaderboard-old /details_OpenBuddy__openbuddy-falcon-40b-v16.1-4k Dataset Card for Evaluation run of OpenBuddy/openbuddy-falcon-40b-v16.1-4k Dataset automatically created during the evaluation run of model OpenBuddy/openbuddy-falcon-40b-v16.1-4k on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_OpenBuddy__openbuddy-falcon-40b-v16.1-4k.0 likes155 downloads3y agoHugging Face09kotoba-speech /wiki40b_lines_entext10M<n<100M0 likes139 downloads10mo agoHugging Face10harryvar /40b509530 likes135 downloads1y agoHugging Face11Respair /40B_dataset_W_fixedtext1M<n<10M0 likes115 downloads2y agoHugging Face12OALL /details_tiiuae__falcon-40b Dataset Card for Evaluation run of tiiuae/falcon-40b Dataset automatically created during the evaluation run of model tiiuae/falcon-40b. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_tiiuae__falcon-40b.tabular100K<n<1M0 likes114 downloads2y agoHugging Face13kotoba-speech /wiki40b_lines_estext1M<n<10M0 likes106 downloads10mo agoHugging Face14Exveria /wiki40b_qa_ja_train_small0 likes64 downloads2y agoHugging Face15open-llm-leaderboard /tiiuae__falcon-40b-detailsgated Dataset Card for Evaluation run of tiiuae/falcon-40b Dataset automatically created during the evaluation run of model tiiuae/falcon-40b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-40b-details.tabular10K<n<100K0 likes62 downloads2y agoHugging Face16alexandrainst /wiki40b-da Dataset Card for "wiki40b-da" Dataset Summary This dataset is an upload of the Danish part of the Wiki40b dataset, being a cleaned version of a dump of Wikipedia. The dataset is identical in content to this dataset on the Hugging Face Hub, but that one requires both apache_beam, tensorflow and mwparserfromhell, which can lead to dependency issues since these are not compatible with several newer packages. The training, validation and test splits are the original ones.… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/wiki40b-da.texttext-generation100K<n<1M1 likes58 downloads3y agoHugging Face17open-llm-leaderboard-old /details_dfurman__falcon-40b-openassistant-peft Dataset Card for Evaluation run of dfurman/falcon-40b-openassistant-peft Dataset Summary Dataset automatically created during the evaluation run of model dfurman/falcon-40b-openassistant-peft on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dfurman__falcon-40b-openassistant-peft.0 likes55 downloads3y agoHugging Face18yunyu /wiki40b_en_100_0_split Dataset Card for "wiki40b_en_100_0_split" More Information needed tabular10M<n<100M0 likes51 downloads3y agoHugging Face19penguinkumimanu /Knowledge_distilled_dataset_by_Fuka2025Q2-40b_qsearch将棋AI用の知識蒸留済みのデータセットを公開します。およそ80億局面あります。 nodchip氏が公開しているtanuki-.nnue-pytorch-2024-07-30.1をhaoでqsearchシャッフルしたのちふかうら王 NewsLetterで配布されたFuka2025Q2-40bで評価値を書き換えました。Eval_Coef=600でDLモデルのvalueと評価値を変換しています。 データにバグがあるかもしれませんが、品質保証はしません。 https://huggingface.co/datasets/nodchip/tanuki-.nnue-pytorch-2024-07-30.1https://yaneurao.fanbox.cc/ 0 likes49 downloads10mo agoHugging Face20O2iginal /yulanmini_phase26_40b0 likes49 downloads6mo agoHugging Face21Leonther /moe-40b-a18b-dataset moe-40b-a18b-dataset Training data for Leonther/moe-40b-a18b-lora — an experimental 40B-A18B MoE student distilled from GLM-5.3-Flash. Composition file records source description gold/professor.jsonl 2 GLM-5.3-Flash (local, ~2 t/s) gold question/answer pairs gold/professor_raw.jsonl 2 GLM-5.3-Flash raw batch output (ids, usage) gold/seeds.jsonl 22 qwen38 professor prompt seeds (domain + length bucket) fake/synthetic.jsonl 100 qwen38 (Qwen3.8-27B… See the full description on the dataset page: https://huggingface.co/datasets/Leonther/moe-40b-a18b-dataset.0 likes49 downloads8d agoHugging Face22aitetic /wiki40b-lm-en wiki40b-lm-en wiki40b-lm-tensorflow1-en-v1 Source: https://www.kaggle.com/models/google/wiki40b-lm/tensorFlow1/en 0 likes47 downloads2mo agoHugging Face23O2iginal /Qwen2.5-distill-v3-instruct-dataset-base-chat-template-40B-202603220 likes44 downloads6mo agoHugging Face24open-llm-leaderboard /tiiuae__falcon-40b-instruct-detailsgated Dataset Card for Evaluation run of tiiuae/falcon-40b-instruct Dataset automatically created during the evaluation run of model tiiuae/falcon-40b-instruct The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-40b-instruct-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face25toklens /dclm_baseline_1.0_40bttabular10M<n<100M0 likes39 downloads4mo agoHugging Face26jealk /wiki40b-da-clean Dataset Card for "wiki40b-da-clean" Dataset Summary This dataset is an slightly modified and filtered version of Wiki40b-da daset which is a fork of this dataset on the Hugging Face Hub. The dataset contains two sub-sets, for which the original columns "wikidata_id" and "version_id" are removed from both: "text": Contains the filtered text of the Wikipedia paragraphs, with formatting removed (START_ARTICLE, START_PARAGRAPH and \n removed) "sentences" Contains the… See the full description on the dataset page: https://huggingface.co/datasets/jealk/wiki40b-da-clean.text1M<n<10M0 likes37 downloads2y agoHugging Face27open-llm-leaderboard /AI-Sweden-Models__gpt-sw3-40b-detailsgated Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-40b Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-40b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-Sweden-Models__gpt-sw3-40b-details.tabular10K<n<100K0 likes30 downloads2y agoHugging Face28nhagar /infimm-webmath-40b_urls Dataset Card for infimm-webmath-40b_urls This dataset provides the URLs and top-level domains associated with training records in Infi-MM/InfiMM-WebMath-40B. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/infimm-webmath-40b_urls.texttext-generation10M<n<100M0 likes27 downloads1y agoHugging Face29TaiyouIllusion /wiki40b_binidxtext100K<n<1M0 likes26 downloads3y agoHugging Face30sachithgunasekara /LaMini-LM-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML-eval-llama2 Dataset Card for "LaMini-LM-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML" More Information needed text1K<n<10K0 likes23 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.