CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ReliableAI /irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B. tabular100K<n<1M1 likes6.7k downloads2y agoHugging Face02common-pile /stackv2_edu_filtered Stack V2 Edu Description We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.tabulartext-generation10M<n<100M6 likes6.5k downloads1y agoHugging Face03JQL-AI /hplt2_edu_scores HPLT2-Edu-scores Dataset summary HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.tabulartext-ranking1B<n<10B1 likes5.7k downloads1y agoHugging Face04JQL-AI /fw2_edu_scores Fineweb2-Edu-scores Dataset summary FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.tabulartext-ranking1B<n<10B6 likes3.9k downloads1y agoHugging Face05jiviteshjn /fineweb-edu-zh-chengyu-cpt Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.tabulartext-generation1M<n<10M1 likes1.9k downloads2mo agoHugging Face06eduzrh /STER STER: Zero-shot 3D Geometric Entity Resolution Benchmark Multi-city, cross-LoD 3D building matching benchmark for the NS-D2S paper (AAAI 2026). Strictly follows the 3dSAGER (SIGMOD 2026) methodology and data format. Dataset Overview Dataset City Country Buildings LOD Source Urban Typology amsterdam Amsterdam NL 123,259 3DBAG LOD1.2/1.3/2.2 Historic canal city rotterdam Rotterdam NL 152,694 3DBAG LOD1.2/1.3/2.2 Post-war modern hague Den Haag NL 181… See the full description on the dataset page: https://huggingface.co/datasets/eduzrh/STER.tabularothern<1K0 likes1.5k downloads9d agoHugging Face07LumiOpen /hpltv2-llama33-edu-annotation HPLT version 2.0 educational annotations This dataset contains annotations derived from HPLT v2 cleaned samples. There are 500,000 annotations for each language if the source contains at least 500,000 samples. We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier. Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.tabular10M<n<100M3 likes1.2k downloads1y agoHugging Face08JQL-AI /JQL-LLM-Edu-Annotations 📚 JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.tabular10M<n<100M2 likes1.1k downloads1y agoHugging Face09RioYokotaLab /fineweb-edutabular100M<n<1B0 likes617 downloads1y agoHugging Face10JQL-AI /curated_edu_scorestabularn<1K0 likes495 downloads1y agoHugging Face11craffel /common-pile-stack-edutabular10M<n<100M0 likes488 downloads1y agoHugging Face12Eurolingua /hplt3_edu_scores HPLT3-Edu-scores Dataset summary HPLT3-JQL-Education is a model-annotated language subset of HPLT3, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. HPLT3-Edu-scores was created based on scores assigned by a deep learning classifier trained to identify educational samples using Snowflake's Arctic-embed-m-v2.0 embeddings. For all training ablations, we used… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/hplt3_edu_scores.tabulartext-ranking1B<n<10B0 likes441 downloads6mo agoHugging Face13ZhuofengLi /fineweb-edu-pretokenized-llama3-100b FineWeb-Edu Pretokenized with Llama 3.1 (100B) This repository contains the sample/100BT subset of FineWeb-Edu, pretokenized for Megatron-LM-style training with meta-llama/Meta-Llama-3.1-8B. It is a derived, pretokenized version of the upstream data, not an official Hugging Face FineWeb release. Dataset summary 140 indexed shards 97,270,686 non-empty documents 97,458,793,013 tokens English web text from FineWeb-Edu sample/100BT Source dataset revision:… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-llama3-100b.tabularn<1K0 likes305 downloads2mo agoHugging Face14eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes223 downloads1y agoHugging Face15secmlr /fineweb-edutabular10M<n<100M0 likes218 downloads2mo agoHugging Face16thepowerfuldeez /massive-yt-edu-queue Massive YouTube Educational Video Queue Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours. Description This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.tabularautomatic-speech-recognition1M<n<10M1 likes143 downloads7mo agoHugging Face17semran1 /synthetic_fw_edutabular100K<n<1M0 likes69 downloads9mo agoHugging Face18Lambent /1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157 Average tokens per entry: 5575.16 tabular1K<n<10K0 likes39 downloads2y agoHugging Face19ModalitiesTeam /FW_EDU_SUBSET_500k_docs FineWeb-Edu Subset This dataset contains 483,606 documents sampled from the FineWeb-Edu dataset. The dataset is used throughout various tutorials on modalities. For licensing, see their conditions. tabular100K<n<1M0 likes38 downloads2y agoHugging Face20JoTeqtheFirstAI /finepdfs-edu-ml300ktabular100K<n<1M0 likes38 downloads8mo agoHugging Face21NancyAbdullah11 /Educational-Lecture-Datasettabular10K<n<100K0 likes37 downloads4mo agoHugging Face22Lambent /creative-writing-2048-fineweb-edu-sampleCreative Writing: keywords: - "creative writing" - "storytelling" - "roleplaying" - "narrative structure" - "character development" - "worldbuilding" - "plot devices" - "genre fiction" - "writing techniques" - "literary elements" - "RPG storytelling" - "interactive narrative" max_entries: 2048 min_tokens: 512 max_tokens: 2048 min_int_score: 4 Total tokens in matching entries: 2218544 tabular1K<n<10K4 likes36 downloads2y agoHugging Face23Tgram3D /education-forum-jfk-dataset Education Forum JFK Assassination Debate Dataset Overview This dataset contains a structured archive of public discussion threads centered on the JFK Assassination Debate section of the Education Forum. Beginning with Version 2.0, the dataset also includes selected discussion forums from other sections of the Education Forum while retaining the original dataset name for continuity and discoverability. The Education Forum spans more than two decades of discussion… See the full description on the dataset page: https://huggingface.co/datasets/Tgram3D/education-forum-jfk-dataset.tabular100K<n<1M0 likes36 downloads3mo agoHugging Face24afrilang-edu /predictedtabular10K<n<100K0 likes30 downloads7mo agoHugging Face25NanoMatriX /finepdfs-edu-ml300ktabular100K<n<1M0 likes29 downloads8mo agoHugging Face26adcks1999 /Ai_education_datasettabular1K<n<10K0 likes24 downloads6mo agoHugging Face27semran1 /dclm-edu-3-plustabular1M<n<10M1 likes21 downloads1y agoHugging Face28secmlr /finepdfs-edutabular1M<n<10M0 likes19 downloads2mo agoHugging Face29rulins /FineWeb-Edu-1BTA subset of FineWeb-Edu randomly sampled from the whole dataset of around 1B gpt2 tokens. This dataset is created for illustration purpose in retrieval-scaling. Please do not distribute. tabular100K<n<1M1 likes17 downloads2y agoHugging Face30rulins /FineWeb-Edu-1MTA subset of FineWeb-Edu randomly sampled from the whole dataset of around 1M gpt2 tokens. This dataset is created for illustration purpose in retrieval-scaling. Please do not distribute. tabular1K<n<10K0 likes15 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.