espnet/wikitongues
The WikiTongues speech corpus is a collection of conversational audio across 700+ languages. It can be used for spoken language modelling or speech representation learning. This dataset includes the raw unsegmented audio in a 16kHz single channel format. Each clip is usually 2-10 minutes long, and contains one or more speakers conversing in their language(s). Sometimes, a speaker may switch languages within a single clip. The total dataset size is around 70 hours. The current version of the… See the full description on the dataset page: https://huggingface.co/datasets/espnet/wikitongues.
4234
