CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01autogluon /fev_datasets Forecast evaluation datasets This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models. The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities. The datasets follow a format that is compatible with the fev package. Data format and usage Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.tabulartime-series-forecasting100K<n<1M13 likes45k downloads8mo agoHugging Face02autogluon /chronos_datasets Chronos datasets Time series datasets used for training and evaluation of the Chronos forecasting models. Note that some Chronos datasets (ETTh, ETTm, brazilian_cities_temperature and spanish_energy_and_weather) that rely on a custom builder script are available in the companion repo autogluon/chronos_datasets_extra. See the paper for more information. Data format and usage The recommended way to use these datasets is via https://github.com/autogluon/fev. All datasets… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/chronos_datasets.tabulartime-series-forecasting10M<n<100M76 likes33k downloads2y agoHugging Face03google-research-datasets /go_emotions Dataset Card for GoEmotions Dataset Summary The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral. The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test splits. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.tabulartext-classification100K<n<1M267 likes14k downloads3y agoHugging Face04M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.2k downloads3y agoHugging Face05KEDevO /crypto-market-datasets Crypto Market Data Lake Created and maintained by Eimantas Kulbe ⭐ If you use this dataset in research, a product, or any publication, please cite the author (see Citation below). It took significant infrastructure and months of collection effort — a citation is the simplest way to give credit. A continuously growing data lake of crypto market microstructure data sourced from Binance and alternative data providers. Full coverage across all 10 symbols — complete tick-level… See the full description on the dataset page: https://huggingface.co/datasets/KEDevO/crypto-market-datasets.tabulartime-series-forecasting10B<n<100B9 likes1.8k downloads3mo agoHugging Face06mkd-chanwoo /keural-datasets Keural Pretraining Datasets (Stage 2) Stage 2 final production corpus for training the Keural Korean LLM. Quality-filtered, deduplicated, and domain-balanced across 4 domains. Summary Metric Value Total processed documents (post-filter) 757,710,609 Dedup removed (Stage 2) 93,919,634 Final documents 663,790,975 Total tokens ~522B Domains English, Korean, Code, Science Source datasets 43 Format Parquet (snappy compressed, sharded) Upload… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-datasets.tabular100M<n<1B0 likes1.7k downloads4mo agoHugging Face07community-datasets /wiki_snippets Dataset Card for "wiki_snippets" Dataset Summary Wikipedia version split into plain text snippets for dense semantic indexing. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure We show detailed information for 2 configurations of the dataset (with 100 snippet passage length and 0 overlap) in English: wiki40b_en_100_0: Wiki-40B wikipedia_en_100_0: Wikipedia Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/wiki_snippets.tabulartext-generation10M<n<100M6 likes1.5k downloads2y agoHugging Face08community-datasets /glucose Dataset Card for [Dataset Name] Dataset Summary GLUCOSE: GeneraLized and COntextualized Story Explanations, is a novel conceptual framework and dataset for commonsense reasoning. Given a short story and a sentence X in the story, GLUCOSE captures ten dimensions of causal explanation related to X. These dimensions, inspired by human cognitive psychology, cover often-implicit causes and effects of X, including events, location, possession, and other attributes.… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/glucose.tabularfill-mask10K<n<100K3 likes1.4k downloads2y agoHugging Face09google-research-datasets /discofuse Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.tabular10M<n<100M6 likes1.3k downloads3y agoHugging Face10inference-optimization /speculators-ci-datasets speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.tabulartext-generation1K<n<10K0 likes1.2k downloads2mo agoHugging Face11Nacryos /ancient-scripts-datasets Ancient Scripts Decipherment Datasets Collated datasets for the paper: Deciphering Undersegmented Ancient Scripts Using Phonetic Prior Jiaming Luo, Frederik Hartmann, Enrico Santus, Regina Barzilay, Yuan Cao Transactions of the Association for Computational Linguistics, 2021 arXiv:2010.11054 This repository gathers the training datasets used in the paper — both those hosted in the authors' GitHub repos and the external cited sources. Repository Structure data/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Nacryos/ancient-scripts-datasets.tabulartext-classification10M<n<100M1 likes602 downloads6mo agoHugging Face12puyang2025 /seven-phishing-email-datasets Dataset Card for Seven Phishing/Spam Email Datasets Dataset Summary This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks. Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label). Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.tabulartext-classification100K<n<1M1 likes570 downloads8mo agoHugging Face13open-source-metrics /datasets-dependents datasets metrics This dataset contains metrics about the huggingface/datasets package. Number of repositories in the dataset: 4997 Number of packages in the dataset: 215 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 22 packages that have more than 1000 stars. There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.tabular10K<n<100K0 likes555 downloads2y agoHugging Face14AgentPublic /data-gouv-datasets-catalog 📢 Sondage 2026 : Utilisation des datasets publiques de MediaTech Vous utilisez ce dataset ou d’autres datasets de notre collection MediaTech ? Votre avis compte ! Aidez-nous à améliorer nos datasets publiques en répondant à ce sondage rapide (5 min) : 👉 https://grist.numerique.gouv.fr/o/albert/forms/gF4hLaq9VvUog6c5aVDuMw/11 Merci pour votre contribution ! 🙌 🇫🇷 Data.gouv.fr Datasets Catalog This dataset contains a processed and embedded version of the… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/data-gouv-datasets-catalog.tabular100K<n<1M4 likes533 downloads6h agoHugging Face15nesteo-datasets /nesteo-prototype NestEO: Modular and Hierarchical EO Dataset Framework NestEO is a hierarchical, resolution-aligned, UTM-based nested grid dataset framework supporting general-purpose, multi-scale multimodal Earth Observation workflows. Built from diverse EO sources and enriched with metadata for landcover, climate zones, and population, it enables scalable, representative and progressive sampling for AI4EO. Grid Levels: 120000m, 12000m, 2400m, 1200m, 600m, 300m, 150mGrid Metadata: ESA WorldCover… See the full description on the dataset page: https://huggingface.co/datasets/nesteo-datasets/nesteo-prototype.tabularimage-segmentation10K<n<100K1 likes530 downloads1y agoHugging Face16hulk10 /data_gouv_datasets_catalog-full-documents 🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data. Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes : titre et description, organisation productrice, licence, couverture spatiale et temporelle, fréquence de mise à jour, formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.tabular100K<n<1M0 likes509 downloads23h agoHugging Face17holistic-ai /holisticai-datasetstabular1M<n<10M1 likes485 downloads2y agoHugging Face18community-datasets /qanta Dataset Card for "qanta" Dataset Summary The Qanta dataset is a question answering dataset based on the academic trivia game Quizbowl. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances mode=first,char_skip=25 Size of downloaded dataset files: 170.75 MB Size of the generated dataset: 147.18 MB Total amount of disk used: 317.93 MB An example of… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/qanta.tabularquestion-answering1M<n<10M6 likes464 downloads2y agoHugging Face19community-datasets /re_dial Dataset Card for ReDial (Recommendation Dialogues) Dataset Summary ReDial (Recommendation Dialogues) is an annotated dataset of dialogues, where users recommend movies to each other. The dataset was collected by a team of researchers working at Polytechnique Montréal, MILA – Quebec AI Institute, Microsoft Research Montréal, HEC Montreal, and Element AI. The dataset allows research at the intersection of goal-directed dialogue systems (such as restaurant recommendation)… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/re_dial.tabularother10K<n<100K2 likes427 downloads2y agoHugging Face20jordiclive /scored_summarization_datasets Dataset Card for "Scored-Summarization-datasets" A collection of Text summarization datasets geared towards training a multi-purpose text summarizer. Each dataset is a parquet file with the following features. default text: a string feature. The source document summary: a string feature. The summary of the document provenance: a string feature. Information about the sub dataset. t5_text_token_count: a int64 feature. The number of tokens the text is encoded in.… See the full description on the dataset page: https://huggingface.co/datasets/jordiclive/scored_summarization_datasets.tabular1M<n<10M8 likes425 downloads4y agoHugging Face21community-datasets /gutenberg_time Dataset Card for the Gutenberg Time dataset Dataset Summary A clean data resource containing all explicit time references in a dataset of 52,183 novels whose full text is available via Project Gutenberg. Supported Tasks and Leaderboards [More Information Needed] Languages Time-of-the-day classification from excerpts. Dataset Structure Data Instances { "guten_id": 28999, "hour_reference": 12, "time_phrase": "midday"… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/gutenberg_time.tabulartext-classification100K<n<1M10 likes417 downloads2y agoHugging Face22nkasmanoff /huggingface-datasets Dataset Card for "huggingface-datasets" This dataset is a snapshot of all public datasets in HuggingFace as of 04/24/2023. It is based on the dataset metadata that can be found at the following endpoint: https://huggingface.co/api/datasets/{dataset_id} Which contains information like the dataset name, its tags, description, and more. Please note that description is different from dataset card, which is what you are reading now :-). I would love to replace this dataset with one which… See the full description on the dataset page: https://huggingface.co/datasets/nkasmanoff/huggingface-datasets.tabular10K<n<100K1 likes399 downloads3y agoHugging Face23InstaDeepAI /winnow-ms-datasets Dataset Card: winnow-ms-datasets Summary Training and benchmark datasets for demonstrating Winnow’s confidence calibration and label-free FDR control on de novo peptide sequencing outputs (via InstaNovo). Contents The repository is organised into the following directories. Spectrum splits use the following terminology throughout: full (also raw): the full search space — all spectra from the experiment labelled: spectra with database search labels… See the full description on the dataset page: https://huggingface.co/datasets/InstaDeepAI/winnow-ms-datasets.tabular1M<n<10M0 likes392 downloads16d agoHugging Face24pwc-archive /datasets [!CAUTION] This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 28th, 2025. image10K<n<100K5 likes347 downloads1y agoHugging Face25yrrhall /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B0 likes340 downloads4mo agoHugging Face26sci-datasets /fineweb-edu-scitabular10M<n<100M2 likes339 downloads1y agoHugging Face27community-datasets /disaster_response_messages Dataset Card for Disaster Response Messages Dataset Summary This dataset contains 30,000 messages drawn from events including an earthquake in Haiti in 2010, an earthquake in Chile in 2010, floods in Pakistan in 2010, super-storm Sandy in the U.S.A. in 2012, and news articles spanning a large number of years and 100s of different disasters. The data has been encoded with 36 different categories related to disaster response and has been stripped of messages with sensitive… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/disaster_response_messages.tabulartext-classification10K<n<100K10 likes305 downloads2y agoHugging Face28viktor-shcherb /combined-chat-datasets Combined Chat Datasets A standardized, unified collection of 30 conversational AI datasets -- spanning organic in-the-wild chats, voluntary sharing, side-by-side preferences, conversation trees, RLHF pairs, and crowdsourced instruction tuning data -- normalized to a single schema for easy joint use. This dataset is a re-distribution. It does not relicense the underlying data. See the Legal & Licensing section -- you must comply with each source dataset's original license.… See the full description on the dataset page: https://huggingface.co/datasets/viktor-shcherb/combined-chat-datasets.tabulartext-generation1M<n<10M1 likes262 downloads6mo agoHugging Face29HSJUSER /0529_DATASETSThis dataset was created using Physical AI Tools and LeRobot. Dataset Structure meta/info.json: { "total_episodes": 121, "total_frames": 90808, "total_videos": 363, "codebase_version": "v2.1", "robot_type": "ffw_sg2_rev1", "total_tasks": 1, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:120" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HSJUSER/0529_DATASETS.tabularrobotics10K<n<100K0 likes250 downloads4mo agoHugging Face30community-datasets /ohsumed Dataset Card for ohsumed Dataset Summary The OHSUMED test collection is a set of 348,566 references from MEDLINE, the on-line medical information database, consisting of titles and/or abstracts from 270 medical journals over a five-year period (1987-1991). The available fields are title, abstract, MeSH indexing terms, author, source, and publication type. The National Library of Medicine has agreed to make the MEDLINE references in the test database available for… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/ohsumed.tabulartext-classification100K<n<1M7 likes246 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.