CoolFace
Datasetpublic

Peacockery/georgian-asr-corpus-v0

georgian-asr-corpus-v0 145.3 hours of Georgian ASR training data: 92,185 clips across FLEURS ka_ge and Common Voice Georgian (scripted 25.0 and spontaneous 3.0, via the Mozilla Data Collective). Splits: train 64,633 / dev 13,456 / test 14,096. Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=kat_Geor/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and audio_size (sample count).… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/georgian-asr-corpus-v0.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes24downloads
settings

This repository belongs to Peacockery on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegeorgian-asr-corpus-v0
visibilitypublic
licencecc-by-4.0
gatedno
ownerPeacockery
Account settings
Peacockery/georgian-asr-corpus-v0 · CoolFace