CoolFace
Datasetpublic

Peacockery/georgian-asr-corpus-v0

georgian-asr-corpus-v0 145.3 hours of Georgian ASR training data: 92,185 clips across FLEURS ka_ge and Common Voice Georgian (scripted 25.0 and spontaneous 3.0, via the Mozilla Data Collective). Splits: train 64,633 / dev 13,456 / test 14,096. Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=kat_Geor/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and audio_size (sample count).… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/georgian-asr-corpus-v0.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes24downloads
3 commits on main
8b1202b4mo ago

Add files using upload-large-folder tool

chikingsley
d1376a54mo ago

Upload README.md with huggingface_hub

chikingsley
a510a244mo ago

initial commit

chikingsley