datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Malaysian-Emilia-annotated
Malaysian Emilia Annotated
Annotate Malaysian-Emilia using Data-Speech pipeline.
Malaysian Youtube
Originally from malaysia-ai/crawl-youtube
Total 3168.8 hours.
Gender prediction, filtered-24k_processed_24k_gender.zip
Language prediction, filtered-24k_processed_language.zip
Force alignment.
Post cleaned to 24k and 44k sampling rates,
24k, filtered-24k_processed_24k.zip
44k, filtered-24k_processed_44k.zip
Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.YouTube-Cantonese-Emilia
YouTube Cantonese — Emilia
2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by
running alvanlii/cantonese-youtube
through the Emilia
speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering).
Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn
label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in
both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.Malaysian-Emilia
Malaysian Emilia
Gather Malaysian Emilia from,
https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2
https://huggingface.co/datasets/Scicom-intl/Malaysian-Chinese-Emilia
https://huggingface.co/datasets/mesolitica/Malaysian-Emilia#malaysian-dialect
And do,
Trim silent.
Permutation for Voice Conversion include post-filtering during permutation.
Convert to Neucodec speech tokens.
Malaysian-Chinese-Emilia
Malaysian-Chinese-Emilia
Use https://github.com/mesolitica/Emilia to pseudo-label Malaysian Chinese audio.
Total rows: 605169
Total hours: 1857.611445057867 hours
Permutation for Voice Conversion
Also we already calculated speaker permutation to prepare for voice conversion.
Emilia-YODAS-Voice-Conversion
Emilia-YODAS-Voice-Conversion
We sample https://huggingface.co/datasets/amphion/Emilia-Dataset YODAS set for voice conversion.
Filter transcriptions based on character repetitiveness and word ngrams.
Filter speaker similarity using https://huggingface.co/nvidia/speakerverification_en_titanet_large during speaker permutation.
Convert audio to speech tokens using https://huggingface.co/neuphonic/neucodec
We also upload the full permutation as zip files.
Speech Tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Emilia-YODAS-Voice-Conversion.emilia-yodas-en-mimiMalaysian-Emilia-v2
Malaysian Emilia v2
This version 2 should fixed https://github.com/open-mmlab/Amphion/issues/436, an Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Malaysian and Singaporean Speech Generation. Replicating Emilia on,
Dataset
Clone and Extract
We upload as split zip files so you can clone and extract distributedly,
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2.emilia-yodas-english-neucodec
Dataset Card for NeuCodec Emilia-YODAS
Dataset Summary
The NeuCodec Emilia-YODAS dataset is an English-language dataset containing >30M audio samples (>78k hours), taken from the English-language subset of Emilia-YODAS and compressed with NeuCodec.
Usage
import torch
from datasets import load_dataset
from neucodec import NeuCodec
# load dataset and model
dataset = load_dataset("neuphonic/emilia-yodas-english-neucodec", split="train"… See the full description on the dataset page: https://huggingface.co/datasets/neuphonic/emilia-yodas-english-neucodec.emilia_full_filtered_3fbc6c4fJA_Emilia_Yodas_ScribeEvents
JA Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/JA_Emilia_Yodas_266h containing only samples with ElevenLabs Scribe v1 audio events.
Changes from source
Filtered to rows where events_scribe is non-empty (4433 rows kept)
Bracket format unified: (event) in text_scribe replaced with [event]
Event types include
Vocal bursts: <laughs>, <sighs>, <clears throat>, etc.
Background: <background noise>, <music>, etc.
Other: <pause>, <unintelligible>… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/JA_Emilia_Yodas_ScribeEvents.emilia_lhotse_manifestEN_Emilia_Yodas_ScribeEvents
EN Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/EN_Emilia_Yodas_616h containing only samples with ElevenLabs Scribe v1 audio events (vocal bursts, background sounds, etc.).
Changes from source
Filtered to only include rows where events_scribe is non-empty (16017 rows out of 228,265 original)
Bracket format unified: Round brackets (laughs) in text_scribe replaced with square brackets [laughs] for consistency with vocal burst annotation format… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/EN_Emilia_Yodas_ScribeEvents.emilia-subset-annotatedDE_Emilia_Yodas_ScribeEvents
DE Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/DE_Emilia_Yodas_680h containing only samples with ElevenLabs Scribe v1 audio events.
Changes from source
Filtered to rows where events_scribe is non-empty (12173 rows kept)
Bracket format unified: (event) in text_scribe replaced with [event]
Event types include
Vocal bursts: <laughs>, <sighs>, <clears throat>, etc.
Background: <background noise>, <music>, etc.
Other: <pause>, <unintelligible>… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/DE_Emilia_Yodas_ScribeEvents.Emilia-wip-2emilia-subset-tagsEmilia-fr-tts-full-descriptions-v1emilia-yodas-fr-filteredECG-Cardiac-Insuffisance-Dataset
About this Dataset
Context
Cardiac Insufficiency Classification Dataset
Abstract
This dataset comprises heartbeat signals structured similarly to well-known ECG heartbeat datasets, designed to classify the degree of cardiac insufficiency. The signals represent electrocardiogram (ECG) waveforms corresponding to various levels of cardiac insufficiency, ranging from no insufficiency to severe stages, including a possible insufficiency class. Each heartbeat signal… See the full description on the dataset page: https://huggingface.co/datasets/Pasko-Emiliano/ECG-Cardiac-Insuffisance-Dataset.synthesized-s2s-emiliaMalaysian-Emilia-Sidon
Malaysian-Emilia-Sidon
Apply sarulab-speech/sidon-v0.1 on,
https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2
https://huggingface.co/datasets/Scicom-intl/Malaysian-Chinese-Emilia
https://huggingface.co/datasets/Scicom-intl/Malaysian-Tamil-Emilia
Emilia-ManifestsEmilia-DE-TextEmilia-fr-tts-tags-full-v1emilia-subset-tags-textdataset_v3_synth_top50custom_emilia_clarke_left_2eval_act_your_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1788,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/emiliano-ng/eval_act_your_dataset.small_ko_emilia_emotionMalaysian-Tamil-Emilia
Malaysian-Tamil-Emilia
Use https://github.com/mesolitica/Emilia to pseudo-label Malaysian Tamil audio.
emilia-yodas-english-neucodec-VJKL-250k-prep
