datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MATH-lighteval
Dataset Card for Mathematics Aptitude Test of Heuristics (MATH) dataset in lighteval format
Dataset Summary
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems
from mathematics competitions, including the AMC 10, AMC 12, AIME, and more.
Each problem in MATH has a full step-by-step solution, which can be used to teach
models to generate answer derivations and explanations. This version of the dataset
contains appropriate builder configs s.t. it… See the full description on the dataset page: https://huggingface.co/datasets/DigitalLearningGmbH/MATH-lighteval.Luhya-ASR-Data-subset-642H
Luhya ASR Data Subset 642H
Luhya speech dataset for automatic speech recognition.
tlott-digital-products
T. Lott Digital Products
Digital product files for T. Lott's online store.
Products
Audiobooks (MP3)
eBooks (PDF)
Software (ZIP)
Cover images (PNG)
Download URLs
Files can be downloaded directly:
https://huggingface.co/datasets/ziggylott/tlott-digital-products/resolve/main/{filepath}
Somali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines
Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines"
More Information needed
temperature_digital_twin
Temperature Digital Twin
PVVX BLE sensor readings (temperature, humidity, battery) collected via
TheengsGateway → MQTT → dlt pipeline.
khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.Kamba-ASR-Data-Subset-484H
Kamba ASR Data Subset 484H
Kamba speech dataset for automatic speech recognition.
Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.digital-coach
DigitalCoach Dataset
DigitalCoach is a multimodal expert-novice computer-use coaching dataset for studying how humans teach software skills through grounded dialogue.It contains 72 coaching sessions, 22,752 dialogue turns, and 28.1 hours of screen recordings, collected across 5 software applications in creativity, engineering, and productivity-oriented workflows.
Each session pairs one expert coach with one novice learner, and captures timestamped data:
dialogue transcripts… See the full description on the dataset page: https://huggingface.co/datasets/berkeley-hci/digital-coach.Gusii-ASR-Data-Subset-470H
Gusii ASR Data Subset 470H
Gusii speech dataset for automatic speech recognition.
khm-asr-cultural
Khmer ASR Cultural Dataset
134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.digital_typhoonDigitial Typhoon Dataset:
KITAMOTO, A., HWANG, J., VUILLOD, B., GAUTIER, L., TIAN, Y., & CLANUWAT, T. (2023, December). Digital Typhoon: Long-term Satellite Image Dataset for the Spatio-Temporal Modeling of Tropical Cyclones. NeurIPS 2023 Datasets and Benchmarks (Spotlight).
This dataset was created by the Digital Typhoon project.
digitalnz
DigitalNZ and RNZ Source Archive
Registry status
Registry ID: edithatogo/digitalnz
Family: nz-cultural-heritage
Repository role: mixed_source_archive
Canonical dataset: edithatogo/digitalnz
Operational status: active_mixed_bundle
Rights status: component_specific_review_required
Authoritative catalog: edithatogo/dataset-estate-registry
Origin and provenance
Origin repository: https://github.com/edithatogo/dnz
Upstream source: DigitalNZ API… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/digitalnz.us-airport-wait-times
US Airport Security & Immigration Wait Times
Minute-resolution TSA security checkpoint wait times for 30 US airports, plus
hourly CBP immigration hall wait times for arriving international passengers.
Airports publish their current wait time and then overwrite it. Nobody keeps the
history. This dataset is that history: a continuous archive collected by polling
each airport's public feed roughly once a minute.
Collection began 1 April 2026 with the New York, Philadelphia and… See the full description on the dataset page: https://huggingface.co/datasets/digitalhen/us-airport-wait-times.big-math-digitsThis dataset is obtained from filtering Big-Math, a large-scale, high-quality math dataset for RL in LLMs. Specifically, we retain only answers that are floats to allow for near-perfect verification. We also filter to keep questions for which the Llama solve rate is between 0 and 70%.
To cite Big-Math:
@article{albalak2025big,
title={Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models},
author={Albalak, Alon and Phung, Duy and Lile, Nathan and… See the full description on the dataset page: https://huggingface.co/datasets/mehuldamani/big-math-digits.NACA_4_Digit_for_ML
NACA 4-Digit Airfoil CFD Dataset
Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions.
Dataset Summary
~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles
AoA range: −5° to +5°
Reynolds number range: 100,000 – 500,000
129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.tatoeba_mt_parquet
Dataset Card for DigitalLearningGmbH/tatoeba_mt_parquet
This is a mirror of Helsinki-NLP/tatoeba_mt, converted to parquet for compatibility with newer huggingface requirements.
Original dataset card follows.
Dataset Summary
The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed translations collected by Tatoeba.org and provided as parallel corpus from OPUS. This dataset includes test and development… See the full description on the dataset page: https://huggingface.co/datasets/DigitalLearningGmbH/tatoeba_mt_parquet.Twin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.handwritten-digit-dataset
Handwritten Digit Dataset
This dataset contains a collection of handwritten digits (0-9) contributed by users through an interactive web-based drawing application. The dataset is continuously updated, reflecting real-world human handwriting variability.
Dataset Details
The images are pre-processed to match the standard machine learning format for digit recognition:
Dimensions: 28x28 pixels.
Format: Grayscale (single channel).
Processing: Each digit is cropped to… See the full description on the dataset page: https://huggingface.co/datasets/zentardev/handwritten-digit-dataset.Luhya-ASR-Data-subset-50hinline-digital-holography-v3
Dataset Card for Synthetic Inline Holographical Images v3 (224px Highly Diverse)
This dataset provides synthetic image triplets representing inline holographical imaging in a simulated environment. This version (v3) uses a native 224x224 resolution optimized for modern Vision Transformers (ViT, Swin) and contains 25,000 samples across 8 noise configurations.
Each data sample consists of:
An object-domain field (ground truth),
Its corresponding forward-propagated hologram (the… See the full description on the dataset page: https://huggingface.co/datasets/gokhankocmarli/inline-digital-holography-v3.monolingual_machine_translation_dataspeak-the-digit
Speak the Digit: Spoken Digit Recognition
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
2,400
Labeled training data
test
600
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
audio
Audio
id
string
label
string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/speak-the-digit.audios-lingala-annotatees
Annotated Lingala Dataset – Full Version
Description
This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models.
It includes:
the original audio files (viewable directly in the Hugging Face viewer)
text transcriptions
Mel spectrograms
tokenized labels
Overall statistics
Metric
Value
Total volume
5 h 0 min 18 s
Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.digital-hospital-environment
Digital Hospital Environment
Digital Hospital is an open-source clinical AI benchmark environment for evaluating agents that must operate inside a structured hospital workflow. It combines role-specific medical knowledge checks, patient-facing clinical operations, cross-role communication, deterministic grading, dense process rewards, and rollout capture in one downloadable runtime. The benchmark is designed for model evaluation, process-supervision datasets, offline… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/digital-hospital-environment.Digital-Development-Indicators-For-African-Countries
Digital Development Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Digital-Development-Indicators-For-African-Countries.llama-2-oai-function-callingCleanPatrick
CleanPatrick: A Benchmark for Data Cleaning
Welcome to CleanPatrick, the first large-scale benchmark designed for data cleaning in the image domain.
Built on the Fitzpatrick17k dermatology dataset, CleanPatrick is a dataset for measuring the performance in detecting three major data quality issues:
off-topic samples, near-duplicates, and label errors.
Overview
CleanPatrick consists of dermatological images annotated with over 500,000 binary labels across three data… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Dermatology/CleanPatrick.un-digital-library
United Nations Digital Library (UNDL) Comprehensive Master Dataset
1. Executive Summary
Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.
