datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
symile-m3
Dataset Card for Symile-M3
Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice.
Paper: https://arxiv.org/abs/2411.01053
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.arsivkana-sounds
Kana Sounds
147 short spoken clips, one for every hiragana and katakana character used by
Kana Trainer: the 46 seion, 20 dakuon,
5 handakuon, 33 yoon and 43 tokushon. They come from a single reader on
FUN Japanese Learning.
Dataset structure
audio/
seion/ 46 clips a.mp3, i.mp3, ka.mp3, ... n.mp3
dakuon/ 20 clips ga.mp3, za.mp3, ji.mp3, ... bo.mp3
handakuon/ 5 clips pa.mp3, pi.mp3, pu.mp3, pe.mp3, po.mp3
yoon/ 33 clips kya.mp3… See the full description on the dataset page: https://huggingface.co/datasets/arsalan-anwari/kana-sounds.StoryTTS
StoryTTS
STORYTTS: A HIGHLY EXPRESSIVE TEXT-TO-SPEECH DATASET WITH RICH TEXTUAL EXPRESSIVENESS ANNOTATIONS
StoryTTS is a highly expressive text-to-speech dataset that contains rich expressiveness both in acoustic and textual perspective, from the recording of a Mandarin storytelling show (评书), which is delivered by a female artist, Lian Liru(连丽如). It contains 61 hours of consecutive and highly prosodic speech equipped with accurate text transcriptions and rich textual… See the full description on the dataset page: https://huggingface.co/datasets/Arsenal/StoryTTS.tts-crh-arslan
Open Source Crimean Tatar Text-to-Speech datasets
This is subset of Arslan voice with train/test splits.
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Statistics
Quality: high
Duration: 1h20m
Frequency: 48 kHz
Cite this work
@misc {smoliakov_2025,
author = { {Smoliakov} },
title = { qirimtatar-tts (Revision c2ceec6) }… See the full description on the dataset page: https://huggingface.co/datasets/speech-uk/tts-crh-arslan.Instance_pvparsa_voice_datasetarsi-mizo-asr
Mizo ASR Dataset (Low-Resource Speech Recognition)
📌 Overview
This repository presents a curated dataset designed for fine-tuning Automatic Speech Recognition (ASR) models for the Mizo language, a low-resource language with limited publicly available speech data.
The goal of this project is to support the development of robust speech recognition systems for Mizo by focusing on high-quality data curation, preprocessing, and model fine-tuning workflows.
⚠️ Note: The full… See the full description on the dataset page: https://huggingface.co/datasets/arsi-consultancy/arsi-mizo-asr.ar-SA-xVectorereuse-enhancedadffasdfadaudio-fake-real-datasetpodcast-filter-samplesNarendraA
