datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dsb_audio_corpus
Acknowledgements
Thanks to all speakers that contributed to this dataset!
Thanks to "Ludowe Nakładnistwo Domowina" and "Rěčny Centrum WITAJ" for donation of their recordings!
kiraat
KIRAAT — A Turkish Read-Speech Corpus
A sentence-aligned read-speech corpus built from publicly available
recordings on Turkish audiobook YouTube channels. The channel credits are
in the table at the end of this card; every clip carries the channel it came
from in the channel column.
clips
1,840,404
duration
3,105.7 hours
recommended subset
1,547,494 clips / 2,575.2 hours
channels
27
speakers (clustered)
90
source recordings
2,680
words (ASR)
21,695,774… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/kiraat.Galgame_Speech_SER_16kHz
Dataset Card for Galgame_Speech_SER_16kHz
[!IMPORTANT]The following rules (in the original repository) must be followed:
必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关!
训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。
English:
You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_SER_16kHz.turkish-audiobook-raw
Turkish Audiobook Speech Corpus (Raw)
Türkçe konuşma araştırmaları için derlenmiş, işlenmemiş uzun-form ses kayıtlarından
oluşan bir koleksiyon. Kayıtlar çeşitli kaynaklardan bir araya getirilmiştir ve
konuşmacı, kayıt ortamı, süre ve ses kalitesi bakımından geniş bir çeşitlilik gösterir.
İçerik
Uzun-form Türkçe konuşma kayıtları (m4a / mp3)
Kaynağa göre klasörlenmiş düz dizin yapısı
Transkript, hizalama veya segmentasyon içermez — ham hâldedir… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-audiobook-raw.turkish-tts-audiobooks
Turkish TTS Audiobooks
Turkish read-speech corpus for text-to-speech training, built from Turkish
audiobook and spoken-article recordings by an automatic pipeline: VAD
segmentation → technical QC → acoustic event tagging → DNSMOS → speaker
embedding/consistency → double-pass Whisper ASR → text policy → leakage-free
splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards.
The pipeline that produced it — every stage, every threshold, the export and
audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.audios-lingala-annotatees
Annotated Lingala Dataset – Full Version
Description
This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models.
It includes:
the original audio files (viewable directly in the Hugging Face viewer)
text transcriptions
Mel spectrograms
tokenized labels
Overall statistics
Metric
Value
Total volume
5 h 0 min 18 s
Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.Kubu-hai
Dataset Card for kubu-hai.model 🙅♂️🤖
^|D Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
kubu-hai is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of… See the full description on the dataset page: https://huggingface.co/datasets/Seriki/Kubu-hai.audios-lingala-annotatees-v2
Annotated Lingala Audio — canonical corpus
Annotated Lingala speech for open automatic speech recognition research and for
fine-tuning speech models.
This release is a full reconstruction of the corpus from its source
recordings and annotations. It supersedes
Congo-digital-service/audios-lingala-annotatees,
which is deprecated — see Relationship to the previous release below.
What this dataset contains
Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.wolof-french-asr
Wolof-French ASR Dataset
Description
Dataset unifié pour l'entraînement de modèles de reconnaissance automatique de la parole (ASR) en wolof et français. Le wolof est une langue d'Afrique de l'Ouest parlée principalement au Sénégal par plus de 10 millions de locuteurs. Les locuteurs wolof pratiquent fréquemment le code-switching (alternance wolof/français), ce qui rend indispensable un modèle ASR capable de transcrire les deux langues.
Composition du dataset… See the full description on the dataset page: https://huggingface.co/datasets/serge-wilson/wolof-french-asr.ivan_shamyakin_sertsa_na_daloni_all
Сэрца на далоні
Аўтар / Author: Іван ШамякінМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
7,083
Працягласць
22 гадз 58 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ivan_shamyakin_sertsa_na_daloni_all.wolof_speech_transcription
Wolof Speech Transcription
Description
Dataset de reconnaissance automatique de la parole (ASR) en wolof, une langue d'Afrique de l'Ouest parlée par plus de 10 millions de locuteurs, principalement au Sénégal.
Ce dataset est un miroir HuggingFace du corpus wolof du projet ALFFA hébergé à l'origine sur GitHub par le laboratoire GETALP (Grenoble).
Source originale
Ce dataset provient du projet ALFFA :
Repository : getalp/ALFFA_PUBLIC
Laboratoire : GETALP… See the full description on the dataset page: https://huggingface.co/datasets/serge-wilson/wolof_speech_transcription.uk-pods
uk-pods - speech datasets of Ukrainian podcasts.
Preparation
Clone the dataset repository and extract the content of clips.tar.gz archive.
git clone https://huggingface.co/datasets/taras-sereda/uk-pods
cd uk-pods && tar -zxvf clips.tar.gz
To use these manifests for training/inference with NeMo [1] modify audio_filepath to absolute locations of audio files extracted in previous step.
# data_root=<clonned_repo_dir> # /home/ubuntu/uk-pods
data_root=$(realpath .)
sed -i… See the full description on the dataset page: https://huggingface.co/datasets/taras-sereda/uk-pods.astryd_lindgren_braty_lvinae_sertsa_all
Браты Ільвінае сэрца
Аўтар / Author: Астрыд ЛіндгрэнМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
2,116
Працягласць
6 гадз 31 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона… See the full description on the dataset page: https://huggingface.co/datasets/fosters/astryd_lindgren_braty_lvinae_sertsa_all.serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
This is the first pushed Ghanaian Speech Lab ASR pipeline artifact. It is a
review artifact for the v0.1 Akan ASR pass, not a trained model checkpoint.
Expected future model repo:
teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
What This Artifact Contains
data/manifest.jsonl: harmonized Waxal + GhanaNLP manifest references.
reports/sanitize.json: sanitization report and… See the full description on the dataset page: https://huggingface.co/datasets/teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1.costumer-service-persian-1astryd_lindgren_braty_lvinae_sertsa_output_original
Браты Ільвінае сэрца — арыгінальнае аўдыё
Аўтар / Author: Астрыд ЛіндгрэнМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
astryd_lindgren_braty_lvinae_sertsa_output
Доўгасць аўдыё
7h06m
Радкоў у датасеце
1,835
Структура
Кожны радок змяшчае:
audio — арыгінальны… See the full description on the dataset page: https://huggingface.co/datasets/fosters/astryd_lindgren_braty_lvinae_sertsa_output_original.ivan_shamyakin_sertsa_na_daloni_output_original
Сэрца на далоні — арыгінальнае аўдыё
Аўтар / Author: Іван ШамякінМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
ivan_shamyakin_sertsa_na_daloni_output
Доўгасць аўдыё
24h58m
Радкоў у датасеце
5,909
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ivan_shamyakin_sertsa_na_daloni_output_original.dv-presidential-speech
Dhivehi Presidential Speech Dataset (Parquet Format)
This is a parquet-formatted version of the dash8x/dv-presidential-speech dataset, converted from the old loading script format to modern parquet format for better compatibility.
Dataset Summary
Dhivehi Presidential Speech is a Dhivehi speech dataset containing around 2.5 hours (1 GB) of speech collected from Maldives President's Office consisting of 7 speeches given by President Yaameen Abdhul Gayyoom.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Serialtechlab/dv-presidential-speech.spg_series01
SPG Series - Moore aligned speech (EP1)
Sentence-level audio <-> Moore text pairs, force-aligned from the bilingual TV show
Sid Pa Gilmde (episode 1, "Falez DK"). Built for Moore ASR / TTS.
Audio: 16 kHz mono wav clips, one per sentence.
Text: Moore (mos), in spoken order; source_fr is the French reference.
Alignment: MMS forced alignment (ctc-forced-aligner, uroman), then pause-trimming
(densest_run, gap > 1.5 s) to drop mis-anchored words across silences / laughter.… See the full description on the dataset page: https://huggingface.co/datasets/burkimbia/spg_series01.17-minute-world-languages_serbe
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/serbe/
Site à scrapper
