CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DigitalLearningGmbH /MATH-lighteval Dataset Card for Mathematics Aptitude Test of Heuristics (MATH) dataset in lighteval format Dataset Summary The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate answer derivations and explanations. This version of the dataset contains appropriate builder configs s.t. it… See the full description on the dataset page: https://huggingface.co/datasets/DigitalLearningGmbH/MATH-lighteval.text10K<n<100K66 likes39k downloads2y agoHugging Face02Digital-Divide-Data /Luhya-ASR-Data-subset-642H Luhya ASR Data Subset 642H Luhya speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M1 likes8.4k downloads1mo agoHugging Face03ziggylott /tlott-digital-products T. Lott Digital Products Digital product files for T. Lott's online store. Products Audiobooks (MP3) eBooks (PDF) Software (ZIP) Cover images (PNG) Download URLs Files can be downloaded directly: https://huggingface.co/datasets/ziggylott/tlott-digital-products/resolve/main/{filepath} audion<1K0 likes5k downloads19d agoHugging Face04Digital-Divide-Data /Somali-ASR-Subset-68H Somali ASR Subset 68H Somali speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M3 likes3.9k downloads1mo agoHugging Face05tgsc /c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines" More Information needed text10M<n<100M1 likes3.9k downloads3y agoHugging Face06maxerbox /temperature_digital_twin Temperature Digital Twin PVVX BLE sensor readings (temperature, humidity, battery) collected via TheengsGateway → MQTT → dlt pipeline. tabular100K<n<1M0 likes3.5k downloads14d agoHugging Face07Digital-Divide-Data /khmer-speech-dataset Khmer ASR Cultural Dataset 727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.audioautomatic-speech-recognition100K<n<1M26 likes3.5k downloads3mo agoHugging Face08Digital-Divide-Data /Kamba-ASR-Data-Subset-484H Kamba ASR Data Subset 484H Kamba speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M0 likes2.8k downloads1mo agoHugging Face09LLM-Digital-Twin /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.imagetext-classification1K<n<10K33 likes2.7k downloads6mo agoHugging Face10berkeley-hci /digital-coach DigitalCoach Dataset DigitalCoach is a multimodal expert-novice computer-use coaching dataset for studying how humans teach software skills through grounded dialogue.It contains 72 coaching sessions, 22,752 dialogue turns, and 28.1 hours of screen recordings, collected across 5 software applications in creativity, engineering, and productivity-oriented workflows. Each session pairs one expert coach with one novice learner, and captures timestamped data: dialogue transcripts… See the full description on the dataset page: https://huggingface.co/datasets/berkeley-hci/digital-coach.textvideo-text-to-text10K<n<100K2 likes2.4k downloads14d agoHugging Face11Digital-Divide-Data /Gusii-ASR-Data-Subset-470H Gusii ASR Data Subset 470H Gusii speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M0 likes2k downloads1mo agoHugging Face12Digital-Divide-Data /khm-asr-cultural Khmer ASR Cultural Dataset 134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.audioautomatic-speech-recognition10K<n<100K9 likes1.1k downloads5mo agoHugging Face13torchgeo /digital_typhoonDigitial Typhoon Dataset: KITAMOTO, A., HWANG, J., VUILLOD, B., GAUTIER, L., TIAN, Y., & CLANUWAT, T. (2023, December). Digital Typhoon: Long-term Satellite Image Dataset for the Spatio-Temporal Modeling of Tropical Cyclones. NeurIPS 2023 Datasets and Benchmarks (Spotlight). This dataset was created by the Digital Typhoon project. image100K<n<1M2 likes940 downloads3y agoHugging Face14edithatogo /digitalnz DigitalNZ and RNZ Source Archive Registry status Registry ID: edithatogo/digitalnz Family: nz-cultural-heritage Repository role: mixed_source_archive Canonical dataset: edithatogo/digitalnz Operational status: active_mixed_bundle Rights status: component_specific_review_required Authoritative catalog: edithatogo/dataset-estate-registry Origin and provenance Origin repository: https://github.com/edithatogo/dnz Upstream source: DigitalNZ API… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/digitalnz.tabulartext-retrievaln<1K0 likes841 downloads23h agoHugging Face15digitalhen /us-airport-wait-times US Airport Security & Immigration Wait Times Minute-resolution TSA security checkpoint wait times for 30 US airports, plus hourly CBP immigration hall wait times for arriving international passengers. Airports publish their current wait time and then overwrite it. Nobody keeps the history. This dataset is that history: a continuous archive collected by polling each airport's public feed roughly once a minute. Collection began 1 April 2026 with the New York, Philadelphia and… See the full description on the dataset page: https://huggingface.co/datasets/digitalhen/us-airport-wait-times.tabular10M<n<100M0 likes821 downloads22h agoHugging Face16mehuldamani /big-math-digitsThis dataset is obtained from filtering Big-Math, a large-scale, high-quality math dataset for RL in LLMs. Specifically, we retain only answers that are floats to allow for near-perfect verification. We also filter to keep questions for which the Llama solve rate is between 0 and 70%. To cite Big-Math: @article{albalak2025big, title={Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models}, author={Albalak, Alon and Phung, Duy and Lile, Nathan and… See the full description on the dataset page: https://huggingface.co/datasets/mehuldamani/big-math-digits.text10K<n<100K3 likes551 downloads1y agoHugging Face17Kokoslocke /NACA_4_Digit_for_ML NACA 4-Digit Airfoil CFD Dataset Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions. Dataset Summary ~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles AoA range: −5° to +5° Reynolds number range: 100,000 – 500,000 129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.tabularothern<1K0 likes487 downloads3mo agoHugging Face18DigitalLearningGmbH /tatoeba_mt_parquet Dataset Card for DigitalLearningGmbH/tatoeba_mt_parquet This is a mirror of Helsinki-NLP/tatoeba_mt, converted to parquet for compatibility with newer huggingface requirements. Original dataset card follows. Dataset Summary The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed translations collected by Tatoeba.org and provided as parallel corpus from OPUS. This dataset includes test and development… See the full description on the dataset page: https://huggingface.co/datasets/DigitalLearningGmbH/tatoeba_mt_parquet.texttext-generation1M<n<10M1 likes472 downloads5mo agoHugging Face19LLM-Digital-Twin /Twin-2K-500-Mega-Study Twin-2K-500-Mega-Study Dataset GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study To see more details for how to process these data, please refer to this GitHub repository. This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants). Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.texttext-generation10K<n<100K2 likes393 downloads8mo agoHugging Face20zentardev /handwritten-digit-dataset Handwritten Digit Dataset This dataset contains a collection of handwritten digits (0-9) contributed by users through an interactive web-based drawing application. The dataset is continuously updated, reflecting real-world human handwriting variability. Dataset Details The images are pre-processed to match the standard machine learning format for digit recognition: Dimensions: 28x28 pixels. Format: Grayscale (single channel). Processing: Each digit is cropped to… See the full description on the dataset page: https://huggingface.co/datasets/zentardev/handwritten-digit-dataset.tabularimage-classificationn<1K0 likes393 downloads2h agoHugging Face21Digital-Divide-Data /Luhya-ASR-Data-subset-50haudio10K<n<100K0 likes355 downloads11mo agoHugging Face22gokhankocmarli /inline-digital-holography-v3 Dataset Card for Synthetic Inline Holographical Images v3 (224px Highly Diverse) This dataset provides synthetic image triplets representing inline holographical imaging in a simulated environment. This version (v3) uses a native 224x224 resolution optimized for modern Vision Transformers (ViT, Swin) and contains 25,000 samples across 8 noise configurations. Each data sample consists of: An object-domain field (ground truth), Its corresponding forward-propagated hologram (the… See the full description on the dataset page: https://huggingface.co/datasets/gokhankocmarli/inline-digital-holography-v3.tabularimage-to-image1B<n<10B0 likes350 downloads7mo agoHugging Face23DigitalUmuganda /monolingual_machine_translation_datatext100K<n<1M0 likes337 downloads3y agoHugging Face24hoangbang /speak-the-digit Speak the Digit: Spoken Digit Recognition Dataset Summary A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded. Splits Split Examples Description train 2,400 Labeled training data test 600 Public inputs with withheld target labels or annotations Data Fields Field Type audio Audio id string label string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/speak-the-digit.audioaudio-classification1K<n<10K0 likes325 downloads2mo agoHugging Face25Congo-digital-service /audios-lingala-annotatees Annotated Lingala Dataset – Full Version Description This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models. It includes: the original audio files (viewable directly in the Hugging Face viewer) text transcriptions Mel spectrograms tokenized labels Overall statistics Metric Value Total volume 5 h 0 min 18 s Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.audioautomatic-speech-recognition10K<n<100K0 likes319 downloads15d agoHugging Face26yatin-superintelligence /digital-hospital-environment Digital Hospital Environment Digital Hospital is an open-source clinical AI benchmark environment for evaluating agents that must operate inside a structured hospital workflow. It combines role-specific medical knowledge checks, patient-facing clinical operations, cross-role communication, deterministic grading, dense process rewards, and rollout capture in one downloadable runtime. The benchmark is designed for model evaluation, process-supervision datasets, offline… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/digital-hospital-environment.texttext-generationn<1K16 likes274 downloads3mo agoHugging Face27electricsheepafrica /Digital-Development-Indicators-For-African-Countries Digital Development Indicators For African Countries | Africa (World Health Organization) Size category: 1K<n<10K - Formats: csv - Sector: technology_digital - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Digital-Development-Indicators-For-African-Countries.tabulartabular-classification1K<n<10K0 likes263 downloads1mo agoHugging Face28Digital-nimbus /llama-2-oai-function-callingtext1K<n<10K3 likes245 downloads3y agoHugging Face29Digital-Dermatology /CleanPatrick CleanPatrick: A Benchmark for Data Cleaning Welcome to CleanPatrick, the first large-scale benchmark designed for data cleaning in the image domain. Built on the Fitzpatrick17k dermatology dataset, CleanPatrick is a dataset for measuring the performance in detecting three major data quality issues: off-topic samples, near-duplicates, and label errors. Overview CleanPatrick consists of dermatological images annotated with over 500,000 binary labels across three data… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Dermatology/CleanPatrick.tabularimage-classification100K<n<1M2 likes221 downloads4mo agoHugging Face30AdhyanshVerma /un-digital-library United Nations Digital Library (UNDL) Comprehensive Master Dataset 1. Executive Summary Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.tabulartext-classification10K<n<100K0 likes221 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.