datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GeoSound
GeoSound
GeoSound is a geo-referenced soundscape dataset that pairs satellite/aerial imagery with
environmental audio recordings. It aggregates recordings from four crowdsourcing platforms —
Freesound, Aporee,
iNaturalist, and Flickr
(via the YFCC100M collection) — and covers a wide geographic footprint.
Research use only. See LICENSE.md for full license and attribution details.
Splits
Split
Rows
train293,718
val
4,999
test
9,931
Train/val/test… See the full description on the dataset page: https://huggingface.co/datasets/MVRL/GeoSound.asante-twi-ttsakuapem-twi-ttscode_switch_yodas_zh
Dataset Card for code-switching yodas
This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas
This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon.
Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.Common-Voice-Geo-Cleaned
Common Voice Georgian — Cleaned for TTS/STT
A high-quality subset of Mozilla Common Voice Georgian cleaned and filtered specifically for text-to-speech fine-tuning.
Dataset Summary
Total samples
21,421
Total duration
35.0 hours
Speakers
12
Sample rate
24 kHz mono WAV
Language
Georgian (kat)
Source
Mozilla Common Voice 19.0
License
CC-0 (public domain)
Splits
Split
Samples
Description
train
20,300
Training data
eval… See the full description on the dataset page: https://huggingface.co/datasets/NMikka/Common-Voice-Geo-Cleaned.Geoff_Bremner_Multimodal_Music_Corpus_SAMPLE
Geoff Bremner Multimodal Music Corpus — Sample Release
This is a single-track sample from the Geoff Bremner Multimodal Music Corpus,
a growing, research-grade, commercially licensable dataset of 100% original
music — written, recorded, and produced entirely by one artist .
If this sample meets your needs - please contact me directly for more
Geoff Bremner
https://linktr.ee/gbaudio
License
This dataset is released under CC BY-NC 4.0… See the full description on the dataset page: https://huggingface.co/datasets/geoffbremneraudio/Geoff_Bremner_Multimodal_Music_Corpus_SAMPLE.Common-Voice-Geo-Cleaned
Common Voice Georgian — Cleaned for TTS/STT
A high-quality subset of Mozilla Common Voice Georgian cleaned and filtered specifically for text-to-speech fine-tuning.
Dataset Summary
Total samples
21,421
Total duration
35.0 hours
Speakers
12
Sample rate
24 kHz mono WAV
Language
Georgian (kat)
Source
Mozilla Common Voice 19.0
License
CC-0 (public domain)
Splits
Split
Samples
Description
train
20,300
Training… See the full description on the dataset page: https://huggingface.co/datasets/CH0BRK/Common-Voice-Geo-Cleaned.ASCEND_CLEAN
Dataset Card for Dataset Name
This dataset is derived from CAiRE/ASCEND. More information is available at https://huggingface.co/datasets/CAiRE/ASCEND.
Removed 嗯 呃 um uh
Resolved [UNK]'s using whisper-medium
Usage
Default utterances with cleaned transcripts
from datasets import load_dataset
data = load_dataset("georgechang8/ASCEND_CLEAN") # add split="train" for train set, etc.
Concatenated 30s utterances with cleaned transcripts… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/ASCEND_CLEAN.cv16_30sGTZAN
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Georgejohn/GTZAN.GEOdatasetGeorgian-Speech-Dataset
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Georgian (ka)
🏷️ Tags
Audio, Speech, Speech Recognition, Georgian, ML, Machine, Machine Learning
📦 Size Category
n < 1K
georgian-speech-datasetka-geo-voice-male-v1
Dataset Card for Georgian Male Voice Dataset v1
Intended Use
Primary Use: Training and fine-tuning TTS models for Georgian language synthesis, including microsoft/speecht5_tts.
Secondary Use: Research in speech synthesis, voice conversion, or linguistic analysis.
SpeechT5 Compatibility
This dataset is specifically formatted to be compatible with microsoft/speecht5_tts fine-tuning. The dataset includes:
Audio: 22,050 Hz mono WAV files (matching SpeechT5… See the full description on the dataset page: https://huggingface.co/datasets/akalandia/ka-geo-voice-male-v1.dataset-20250728_102101-sw
GeoPoll Swahili Speech Dataset
This dataset contains speech recognition data for Swahili (sw) collected and processed by GeoPoll.
Dataset Summary
This dataset is designed for fine-tuning speech recognition models on Swahili audio data. It includes high-quality audio segments with corresponding transcriptions.
Dataset Statistics
Total samples: 11814
Total duration: 20.45 hours
Average duration: 6.23 seconds per sample
Number of speakers: 6
Language: Swahili… See the full description on the dataset page: https://huggingface.co/datasets/GeoPoll/dataset-20250728_102101-sw.TestStyle_TTS_UPSC_GEOArtieAbamsvozviniboyTTS_10s_clean_documentry_style_national_geographyGeorge-Harrison-Brainwashed-TheOsloChildgeorgii-marchuk-davyd-garadotskiia-kanony
Давыд-Гарадоцкія каноны
Metadata
Author: Георгій Марчук
Title: Давыд-Гарадоцкія каноны
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/georgii-marchuk-davyd-garadotskiia-kanony.georgii-marchuk-kryk-na-khutary-margaryta-zakharyia
Крык на хутары
Metadata
Author: Георгій Марчук
Title: Крык на хутары
Narrator: Маргарыта Захарыя
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size:… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/georgii-marchuk-kryk-na-khutary-margaryta-zakharyia.GeorgeHarrisonTalkingMixed_Training_Geo
