datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-Talker-SD
Dataset Card for Multi-Talker-SD
Dataset Description
Multi-Talker-SD is a large-scale bilingual (English–Mandarin) multi-speaker meeting dataset designed to support research on speaker diarization and meeting transcription.
Size: 1,000 simulated meetings
Participants per meeting: 10–30 speakers
Average duration: ~20 minutes per meeting, up to one hour
Languages: English, Mandarin (code-switching possible)
Audio characteristics: realistic speaker overlap… See the full description on the dataset page: https://huggingface.co/datasets/yihao005/Multi-Talker-SD.synthetic-parallel-external
Synthetic Parallel EN↔LG — external
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-external.crowd-recital-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Recital - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~78h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.CAMEO
CAMEO: Collection of Multilingual Emotional Speech Corpora
Dataset Description
CAMEO is a curated collection of multilingual emotional speech datasets.
It includes 13 distinct datasets with transcriptions, encompassing a total of 41,265 audio samples.
The collection features audio in eight languages: Bengali, English, French, German, Italian, Polish, Russian, and Spanish.
Example Usage
The dataset can be loaded and processed using the datasets library:
from… See the full description on the dataset page: https://huggingface.co/datasets/yiyandeng/CAMEO.crowd-recital-yi
About
This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project.
Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read.
Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.hung-yi_lee
Hung-yi Lee Lecture ASR/TTS
This dataset contains speech segments aligned with subtitles from Hung-yi Lee lecture videos.
It is intended for Mandarin ASR and TTS experiments.
Dataset Details
Speaker: Hung-yi Lee
Language: Traditional Chinese / Taiwan Mandarin (zh-TW)
Sampling rate: 16 kHz
Audio format in the published dataset: embedded FLAC audio in Parquet shards
Source videos with audio: 15
Segments: 29043
Total duration: 16.92 hours
Courses: 機器學習2026… See the full description on the dataset page: https://huggingface.co/datasets/voidful/hung-yi_lee.sixuxar_yijiri_mak7
Dataset Info
This dataset consists of paired audio and text data sourced from the following book:
Title: Къэрмокъуэ М. Щихухэр иджыри мэкI. Япэ тхылъ.
Publication: Нальчик: Эльбрус, 1999
Audio Specifications
Sample Rate: 16,000 Hz
Total Length: 10:36:40
Source: adigabook.ru
Processing Information
Audio-text pairs for this dataset were extracted and aligned using META AI's forced alignment algorithm.
crowd-whatsapp-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~19h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.synthetic-parallel-salt
Synthetic Parallel EN↔LG — salt
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-salt.crowd-whatsapp-yi
About
This dataset was created by crowd-sourced Whatsapp voice recordings in Yiddish as part of the ivrit.ai project.
Volunteers read a message sent to them from a predefined set of messages, recording themselves using Whasapp voice message sent to the collecting bot.
Later this data is normalized by aligning the captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi.kinyarwanda-speech-trimmed
Kinyarwanda Speech Dataset (Trimmed)
This dataset contains processed Kinyarwanda speech data with trimmed audio segments.
Dataset Structure
The dataset contains two splits:
dev_test: 9,263 samples
test: 9,265 samples
Features
Each sample contains:
id: Unique identifier for the sample
audio: Audio data
audio_language: Language of the audio (Kinyarwanda)
text: Transcription of the audio
prompt: Associated prompt or context
duration: Duration of the audio… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/kinyarwanda-speech-trimmed.synth-qa-taste-codec-chat
Synthetic QA Taste-S Codec Chat
18571 single-turn Traditional Chinese QA utterances with synthesized speech, 21.6 hours of
audio before codec extraction.
Assistant speech is represented as:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
Each text token is followed by its 16 Taste-S FSQ codes (codebooks a..p).
Configurations
default — messages (user question + assistant <SAY> speech), audio, and answer text.
Statistics
Utterances: 18571… See the full description on the dataset page: https://huggingface.co/datasets/yilele/synth-qa-taste-codec-chat.synthetic-parallel-sunbird
Synthetic Parallel EN↔LG — sunbird
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-sunbird.lug-eng-synthetic-parallel-v1
Synthetic English-Luganda Parallel Speech (149,486 pairs)
Synthetic parallel audio for English-Luganda speech-to-speech translation
research, generated with Orpheus 3B TTS and stored as parquet shards with
embedded audio.
Schema
field
type
notes
id
string
pair id, e.g. pair_000123
audio_eng
Audio(22 050 Hz mono)
English clip
audio_lug
Audio(22 050 Hz mono)
Luganda clip
text_eng
string
English transcript
text_lug
string
Luganda transcript… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/lug-eng-synthetic-parallel-v1.
