datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Complete_Data_Source_100K_HOURS
Multi-Language Audio Collection (100K Hours)
This repository is physically reorganized for Absolute 100% Data Visibility.
🏗️ Global Consolidator
Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here.
mixed_multilingual_commonvoice_all_languages_100kBuild from mozilla commonvoice 13 using the script commited in this repo.
Used to teach a model to ignore languages that are not french
FineLAP-100kvls-100k
VLS 100K
74,936 MS-COCO images paired with a long written description, a one-sentence
spoken summary of that description, the spoken audio, and that audio pre-encoded
with a neural codec. Images and audio are embedded in the parquet, so the viewer
renders them and one call opens the set:
from datasets import load_dataset
ds = load_dataset("seonglae/vls-100k", split="train")
ds[0]["image"] # PIL image
ds[0]["audio"] # decoded waveform
ds[0]["sst"] # the sentence that was… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/vls-100k.VITS_DATASET_60k_100KUAE_100Kperuvian_speech_more_100kEgyptian-dialect-100kqwen3-asr-hebrew-100kIndicVoices-ML-100ksaudi-tts-synthetic-100kmj_emo_embed_100k
Dataset Card for "mj_emo_embed_100k"
More Information needed
FineLAP-100kComplete_100k_DataSOVA-audiobooks-100ken_asr_mls_sub0_al_50-100k
