CoolFace
Datasetpublic

espnet/wikitongues

The WikiTongues speech corpus is a collection of conversational audio across 700+ languages. It can be used for spoken language modelling or speech representation learning. This dataset includes the raw unsegmented audio in a 16kHz single channel format. Each clip is usually 2-10 minutes long, and contains one or more speakers conversing in their language(s). Sometimes, a speaker may switch languages within a single clip. The total dataset size is around 70 hours. The current version of the… See the full description on the dataset page: https://huggingface.co/datasets/espnet/wikitongues.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
4likes234downloads

espnet/wikitongues · main · files are served by the source, never re-hosted here