datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese-male-voice-A-datasetafrispeak_kinyarwanda_male_tts_datasetbpsd-unirqvae3-unidac4-ytsv
BPSD score-image, audio and notation tokens (U-MusT)
Tokenized Beethoven Piano Sonata Dataset v2 for
U-MusT — the test-only split, and the only corpus in the
collection carrying all four modalities: score-image tokens, audio tokens, and LMX notation.
Because it is held out for evaluation, the image tokens here are not shift-augmented: they have
shape (1, 1, H, W, 4), a single tokenization. The audio tokens retain the 9-variant stack.
BPSD ships no system-level image alignment… See the full description on the dataset page: https://huggingface.co/datasets/malerlab/bpsd-unirqvae3-unidac4-ytsv.maestro-unidac4-ytsv
MAESTRO + ASAP audio and MIDI tokens (U-MusT)
Tokenized MAESTRO v3.0.0 for
U-MusT: DAC audio tokens and MT3-style MIDI event arrays,
covering roughly 199 hours of Disklavier-captured piano performance with precisely aligned MIDI.
This repository also contains ASAP-derived data. lmx/ and asap_note_events/ come from the
ASAP dataset, whose audio is itself MAESTRO. Both
carry the same license, so nothing conflicts, but the repository name mentions only one of the two
corpora it… See the full description on the dataset page: https://huggingface.co/datasets/malerlab/maestro-unidac4-ytsv.public_youtube1120radio_2public_youtube700fluent_speech_commands_malepublic_youtube1120_hqlibrispeech_malevoxceleb_malemaleo-short-1.5H
Dataset Card for Maleo Short 1.5H
Dataset Description
Dataset Summary
Maleo Short 1.5H is a manually curated, rigorously annotated speaker diarization dataset designed to benchmark State-of-the-Art (SOTA) models against complex, "in-the-wild" media domains. While modern diarization pipelines excel in controlled acoustic environments (like telephony or reading corpora), they heavily struggle with the overlapping speech, sound effects, and rapid speaker shifts… See the full description on the dataset page: https://huggingface.co/datasets/maleo-ai/maleo-short-1.5H.turkish_malearabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.jalak
Jalak — Indonesian Multi-Speaker TTS Dataset
A Coqui-TTS-ready multi-speaker speech dataset for Indonesian, Javanese, and Sundanese,
built to accompany the maiaid/jalak-model VITS
checkpoint.
The layout matches jalak-model/config.json exactly: root dataset/ path, Coqui coqui
formatter, pipe-separated metadata audio_file|text|speaker_name.
Dataset Summary
Split / metadata file
Speakers
Clips
Source
License
metadata-javanese.csv
39 × JV-xxxxx
5,822… See the full description on the dataset page: https://huggingface.co/datasets/maleo-ai/jalak.GV_Train_100h_MaleAll_Hindi_ASR_Male_v1.1slakh-unidac4-ytsv
SLakh2100 audio and MIDI tokens (U-MusT)
Tokenized SLakh2100 for
U-MusT: DAC audio tokens and MIDI event arrays over roughly
145 hours of synthesized multi-track audio rendered from the
Lakh MIDI Dataset. Only the mixed audio was tokenized;
stems were discarded.
SLakh is pop rather than classical, and the paper trains on it but excludes it from reported
results, since the work targets Western classical music. No audio is redistributed.
Token files contain shift… See the full description on the dataset page: https://huggingface.co/datasets/malerlab/slakh-unidac4-ytsv.opensinger_maletts_lingala_malefiltered_nepali_male_dataset1bangla-emotion-maleemirates-dialect-speech-male
🌍 Emirates Dialectal Arabic Audio Dataset
This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning.
📌 Source Data & Provenance
Source Repository: https://github.com/MahaAlBlooki/alsanaa-emirati-dataset
Domain & Content: Spoken Emirati dialectal Arabic speech recordings.
Dialect Focus: Emirates / Gulf Dialectal Arabic.
Standardized Format: 22,050 Hz… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/emirates-dialect-speech-male.IndicVoices_Hindi_audio_44100_18_30_malesyspin_merged_male_ttssaudi-dialect-speech-malepersian_dataset_maleSYSPIN_Hindi_Male_TTSopenslr42-khmer-male
OpenSLR SLR42 Khmer Male Speech
This dataset is a processed version of the OpenSLR SLR42 Khmer speech dataset.
Dataset Description
This dataset contains approximately 2,906 Khmer speech recordings with corresponding Khmer transcriptions.
Each example contains:
audio: Khmer speech recording
text: Khmer transcription
Dataset Structure
Column
Type
Description
audio
Audio
Khmer speech recording
text
String
Khmer transcription… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/openslr42-khmer-male.MALE_FEMALE_VOICE_BAND
Male/Female Hindi Voice Dataset
Whisper-verified recordings with the original script retained as text.
Choose the male or female subset in the Dataset Viewer. Audio is embedded in Parquet for reliable playback and pagination.
