datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mint-1t-html-images-gte6-sample
Size: 6769158 images sampled from Mint-1t-html
Criteria: Data entries with greater than or equal to 6 images (gte6)
GaussianDWM-sampledsample-images-TADNEenglish-casual-speech-sample-south-african-accent
English Casual Speech Sample (South African Accent)
South African crowd-sourced participants respond to questions about their daily lives and activities.
This dataset is a sample of a larger collection from the same data collection campaign.
Changelog
FEB 2026: initial share. ASR (Chirp3) transcripts. WER: 12%
Specs
Speakers: ~550 unique South African speakers
Total duration: ~60 hours
Files sample rate: 48kHz
Actual sample rate: TBD
Language: English (SA… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/english-casual-speech-sample-south-african-accent.laion-samplesTADNE-sample-images
TADNE sample images
Images generated by the TADNE model.
Note
prediction_results/anime-face-detector
https://github.com/hysts/anime-face-detector
YOLOv3 + HRNetV2
prediction_results/deepdanbooru
https://github.com/KichangKim/DeepDanbooru
model-resnet_custom_v3.h5
prediction_results/deepdanbooru/intermediate_features
Output by the following model
4096-dim
def create_model() -> tf.keras.Model:
path = huggingface_hub.hf_hub_download('hysts/DeepDanbooru'… See the full description on the dataset page: https://huggingface.co/datasets/hysts/TADNE-sample-images.bot_sampleNNIRP-dataset-sample
NNIRP Dataset: How You Split Is What You Get
A dataset and evaluation protocol for predicting inference runtime of neural network models from their ONNX computational graphs. Contains 103,070 profiling samples from 190 source configurations spanning 6 architecture families, organized into 156 clusters across 28 sub-families.
Dataset Summary
Each sample includes three data layers:
Layer
Format
Size
Description
Profiling
.json
~130 MB
Runtime, VRAM, and RAM… See the full description on the dataset page: https://huggingface.co/datasets/nnirp/NNIRP-dataset-sample.SDO_AIA_2024_94_193_131_2000_samples
Frame sequences of patches of the Sun's surface preceding flares with an assigned GOES class. SDO AIA @ 94, 193, and 131 Å
Cosmos-Transfer1-7B-Sample-AV-Data-Example
Cosmos-Transfer1-7B-Sample-AV-Data-Example
Cosmos | Code | Paper | Paper Website
Dataset Description:
This dataset contains 10 sample data points intended to help users better utilize our Cosmos-Transfer1-7B-Sample-AV model. It includes HD Map annotations and LiDAR data, with no personally identifiable information such as faces or license plates. This dataset is intended for research and development only.
Dataset Owner(s):
NVIDIA
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos-Transfer1-7B-Sample-AV-Data-Example.vlm3r_sample_10ksample_data_imagessamples-for-demo-20250408vb_samplesample_data_tiresampled-videos-part2emo_speech_samplem3exam_augmented_samplesamples-for-demo-20250312otoSpeech-HQ-full-duplex-samples
Dataset Card for otoSpeech-HQ-full-duplex-samples: Full-Duplex Conversational Speech Dataset Samples
Dataset Summary
otoSpeech-HQ-full-duplex-samples is a curated collection of high-quality full-duplex conversational speech samples designed for commercial and production-oriented use.
This repository is derived from a private subset of otoSpeech and features carefully selected English two-speaker conversations with enhanced audio quality. The samples are intended for… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-HQ-full-duplex-samples.StableCascade_Lora_Training_sampleSTORM2_samplesamples-for-demo-20250228zello-public-channels-voice-sample
Zello Public Channels Voice Dataset Sample
Dataset summary
Total audio: 48.23 hours
Total messages: 13207
Breakdown by language
Language
Hours
Messages
Speakers
Channels
ms
10.65
2056
112
7
en
9.54
2990
362
26
id
6.98
1597
157
11
es
6.86
2144
224
21
tl
4.28
1404
78
6
pt
3.92
863
66
11
sw
2.43
182
21
1
ru
1.30
282
42
12
th
0.95
55090
3
zh
0.87
828
312
10
it
0.15
25
19
7
is
0.09
76
25
10
ko
0.05
75
43
13
vi
0.04
42
34
8
fr… See the full description on the dataset page: https://huggingface.co/datasets/zello/zello-public-channels-voice-sample.PyVision-RL-Eval-Samplessmall_model_samples_2
