datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vlsp2020_vinai_100h
unofficial mirror of VLSP 2020 - VinAI - ASR challenge dataset
official announcement:
tiếng việt: https://institute.vinbigdata.org/events/vinbigdata-chia-se-100-gio-du-lieu-tieng-noi-cho-cong-dong/
in eglish: https://institute.vinbigdata.org/en/events/vinbigdata-shares-100-hour-data-for-the-community/
VLSP 2020 workshop: https://vlsp.org.vn/vlsp2020
official download: https://drive.google.com/file/d/1vUSxdORDxk-ePUt-bUVDahpoXiqKchMx/view?usp=sharing
contact: info@vinbigdata.org… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vlsp2020_vinai_100h.vlm-voice-audio
🎙️ VLM Robotics Voice Commands (Audio)
Natural speech commands for Vision-Language-Model robot control.
This dataset contains 9,999 audio recordings of human voice commands
for controlling robots — covering pick & place, navigation, manipulation, observation,
multi-step tasks, spatial commands, safety, household chores, and conversational feedback.
🎯 Purpose
Training omni-modal VLMs that understand spoken robot commands. The audio captures
natural speech patterns… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/vlm-voice-audio.ukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples… See the full description on the dataset page: https://huggingface.co/datasets/vladsfa/ukr-dialects-audio-dataset.
