datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mint-1t-html-images-gte6-sample
Size: 6769158 images sampled from Mint-1t-html
Criteria: Data entries with greater than or equal to 6 images (gte6)
GaussianDWM-sampledspeech_400ksam2act-datasets
SAM2Act
SAM2Act is a multi-view robotics transformer policy for robotic manipulation. Built on RVT-2, it combines multi-resolution upsampling with visual embeddings from the SAM2 foundation model to improve 3D action prediction, multitask learning, and generalization. SAM2Act+ extends this policy with a memory bank, memory encoder, and memory attention so the agent can condition on prior observations and actions for spatial memory-dependent tasks.
For full project details, code… See the full description on the dataset page: https://huggingface.co/datasets/hqfang/sam2act-datasets.speech_dataSAM-LLaVA-Captions10Msample-images-TADNEcc12m-sam2-parse-treeSAM-SGT
SGT: Semantic Generative Tuning for Unified Multimodal Models
This repository hosts checkpoints fine-tuned with Semantic Generative Tuning (SGT) — a training
paradigm that couples visual understanding and generation in Unified Multimodal Models (UMMs)
by using image segmentation as a generative proxy.
Unified multimodal models typically optimize understanding and generation with misaligned
objectives (sparse text tokens vs. dense pixel targets), which isolates the two capabilities.… See the full description on the dataset page: https://huggingface.co/datasets/Two-hot/SAM-SGT.english-casual-speech-sample-south-african-accent
English Casual Speech Sample (South African Accent)
South African crowd-sourced participants respond to questions about their daily lives and activities.
This dataset is a sample of a larger collection from the same data collection campaign.
Changelog
FEB 2026: initial share. ASR (Chirp3) transcripts. WER: 12%
Specs
Speakers: ~550 unique South African speakers
Total duration: ~60 hours
Files sample rate: 48kHz
Actual sample rate: TBD
Language: English (SA… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/english-casual-speech-sample-south-african-accent.laion-samplessamromur_asr
Dataset Card for samromur_asr
Dataset Summary
This is a modfied copy of the dataset from The Language and Voice Laboratory in RU.
This is the first release of the Samrómur Icelandic Speech corpus that contains 100.000 validated utterances.
The corpus is a result of the crowd-sourcing effort run by the Language and Voice Lab at the Reykjavik University, in cooperation with Almannarómur, Center for Language Technology.
Languages
The audio is in Icelandic.
The… See the full description on the dataset page: https://huggingface.co/datasets/DavidErikMollberg/samromur_asr.bot_sampleTADNE-sample-images
TADNE sample images
Images generated by the TADNE model.
Note
prediction_results/anime-face-detector
https://github.com/hysts/anime-face-detector
YOLOv3 + HRNetV2
prediction_results/deepdanbooru
https://github.com/KichangKim/DeepDanbooru
model-resnet_custom_v3.h5
prediction_results/deepdanbooru/intermediate_features
Output by the following model
4096-dim
def create_model() -> tf.keras.Model:
path = huggingface_hub.hf_hub_download('hysts/DeepDanbooru'… See the full description on the dataset page: https://huggingface.co/datasets/hysts/TADNE-sample-images.NNIRP-dataset-sample
NNIRP Dataset: How You Split Is What You Get
A dataset and evaluation protocol for predicting inference runtime of neural network models from their ONNX computational graphs. Contains 103,070 profiling samples from 190 source configurations spanning 6 architecture families, organized into 156 clusters across 28 sub-families.
Dataset Summary
Each sample includes three data layers:
Layer
Format
Size
Description
Profiling
.json
~130 MB
Runtime, VRAM, and RAM… See the full description on the dataset page: https://huggingface.co/datasets/nnirp/NNIRP-dataset-sample.speech_40kFAST_3D_SAMCosmos-Transfer1-7B-Sample-AV-Data-Example
Cosmos-Transfer1-7B-Sample-AV-Data-Example
Cosmos | Code | Paper | Paper Website
Dataset Description:
This dataset contains 10 sample data points intended to help users better utilize our Cosmos-Transfer1-7B-Sample-AV model. It includes HD Map annotations and LiDAR data, with no personally identifiable information such as faces or license plates. This dataset is intended for research and development only.
Dataset Owner(s):
NVIDIA
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos-Transfer1-7B-Sample-AV-Data-Example.MMRS-1MA huggingface version of MMRS-1M dataset.
Files are downloaded from https://github.com/wivizhang/EarthGPT
vlm3r_sample_10karabic_speech_data_8.tarimgeditEuroSpeech-WebDatasetpascal-parts-sam3durg-university-noticessample_data_imagesSAM_8KRUOK_ver_samvb_samplesam3d_latents
