datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-Talker-SD
Dataset Card for Multi-Talker-SD
Dataset Description
Multi-Talker-SD is a large-scale bilingual (English–Mandarin) multi-speaker meeting dataset designed to support research on speaker diarization and meeting transcription.
Size: 1,000 simulated meetings
Participants per meeting: 10–30 speakers
Average duration: ~20 minutes per meeting, up to one hour
Languages: English, Mandarin (code-switching possible)
Audio characteristics: realistic speaker overlap… See the full description on the dataset page: https://huggingface.co/datasets/yihao005/Multi-Talker-SD.GitTaskBench
Dataset Card for GitTaskBench
The dataset was presented in the paper GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging.
Dataset Details
Dataset Description
GitTaskBench is a benchmark dataset designed to evaluate the capabilities of code-based intelligent agents in solving real-world tasks by leveraging GitHub repositories.It contains 54 representative tasks across 7 domains, carefully curated to reflect… See the full description on the dataset page: https://huggingface.co/datasets/Nicole-Yi/GitTaskBench.synthetic-parallel-external
Synthetic Parallel EN↔LG — external
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-external.CoDeTT
CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation
🌐 Dataset Summary
CoDeTT is a benchmark dataset for turn-taking decision evaluation in full-duplex spoken dialogue systems.It evaluates not only what action a model should take at the current moment, but also whether the underlying semantic intent is aligned.
Core action space (4 classes):
Maintain
Stop & Listen
Takeover
Dismiss
Fine-grained intent space:
14 scenario labels across two system states:… See the full description on the dataset page: https://huggingface.co/datasets/YingaoWang-casia/CoDeTT.CAMEO
CAMEO: Collection of Multilingual Emotional Speech Corpora
Dataset Description
CAMEO is a curated collection of multilingual emotional speech datasets.
It includes 13 distinct datasets with transcriptions, encompassing a total of 41,265 audio samples.
The collection features audio in eight languages: Bengali, English, French, German, Italian, Polish, Russian, and Spanish.
Example Usage
The dataset can be loaded and processed using the datasets library:
from… See the full description on the dataset page: https://huggingface.co/datasets/yiyandeng/CAMEO.crowd-recital-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Recital - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~78h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.crowd-recital-yi
About
This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project.
Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read.
Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.hung-yi_lee
Hung-yi Lee Lecture ASR/TTS
This dataset contains speech segments aligned with subtitles from Hung-yi Lee lecture videos.
It is intended for Mandarin ASR and TTS experiments.
Dataset Details
Speaker: Hung-yi Lee
Language: Traditional Chinese / Taiwan Mandarin (zh-TW)
Sampling rate: 16 kHz
Audio format in the published dataset: embedded FLAC audio in Parquet shards
Source videos with audio: 15
Segments: 29043
Total duration: 16.92 hours
Courses: 機器學習2026… See the full description on the dataset page: https://huggingface.co/datasets/voidful/hung-yi_lee.sixuxar_yijiri_mak7
Dataset Info
This dataset consists of paired audio and text data sourced from the following book:
Title: Къэрмокъуэ М. Щихухэр иджыри мэкI. Япэ тхылъ.
Publication: Нальчик: Эльбрус, 1999
Audio Specifications
Sample Rate: 16,000 Hz
Total Length: 10:36:40
Source: adigabook.ru
Processing Information
Audio-text pairs for this dataset were extracted and aligned using META AI's forced alignment algorithm.
synthetic-parallel-salt
Synthetic Parallel EN↔LG — salt
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-salt.crowd-whatsapp-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~19h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.crowd-whatsapp-yi
About
This dataset was created by crowd-sourced Whatsapp voice recordings in Yiddish as part of the ivrit.ai project.
Volunteers read a message sent to them from a predefined set of messages, recording themselves using Whasapp voice message sent to the collecting bot.
Later this data is normalized by aligning the captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi.kinyarwanda-speech-trimmed
Kinyarwanda Speech Dataset (Trimmed)
This dataset contains processed Kinyarwanda speech data with trimmed audio segments.
Dataset Structure
The dataset contains two splits:
dev_test: 9,263 samples
test: 9,265 samples
Features
Each sample contains:
id: Unique identifier for the sample
audio: Audio data
audio_language: Language of the audio (Kinyarwanda)
text: Transcription of the audio
prompt: Associated prompt or context
duration: Duration of the audio… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/kinyarwanda-speech-trimmed.synth-qa-taste-codec-chat
Synthetic QA Taste-S Codec Chat
18571 single-turn Traditional Chinese QA utterances with synthesized speech, 21.6 hours of
audio before codec extraction.
Assistant speech is represented as:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
Each text token is followed by its 16 Taste-S FSQ codes (codebooks a..p).
Configurations
default — messages (user question + assistant <SAY> speech), audio, and answer text.
Statistics
Utterances: 18571… See the full description on the dataset page: https://huggingface.co/datasets/yilele/synth-qa-taste-codec-chat.synthetic-parallel-sunbird
Synthetic Parallel EN↔LG — sunbird
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-sunbird.lj_speechThis is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading
passages from 7 non-fiction books in English. A transcription is provided for each clip. Clips vary in length
from 1 to 10 seconds and have a total length of approximately 24 hours.
Note that in order to limit the required storage for preparing this dataset, the audio
is stored in the .wav format and is not converted to a float32 array. To convert the audio
file to a float32 array, please make use of the `.map()` function as follows:
```python
import soundfile as sf
def map_to_array(batch):
speech_array, _ = sf.read(batch["file"])
batch["speech"] = speech_array
return batch
dataset = dataset.map(map_to_array, remove_columns=["file"])
```lug-eng-synthetic-parallel-v1
Synthetic English-Luganda Parallel Speech (149,486 pairs)
Synthetic parallel audio for English-Luganda speech-to-speech translation
research, generated with Orpheus 3B TTS and stored as parquet shards with
embedded audio.
Schema
field
type
notes
id
string
pair id, e.g. pair_000123
audio_eng
Audio(22 050 Hz mono)
English clip
audio_lug
Audio(22 050 Hz mono)
Luganda clip
text_eng
string
English transcript
text_lug
string
Luganda transcript… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/lug-eng-synthetic-parallel-v1.
