datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
podcast-tokenized-bg2.5-enj4.5podcast-tokenized-bg3.5-enj5tokenized-voxpopuli
Tokenized Vox-Populi Dataset
This repository serves to store the tokenized audio dataset extracted from Facebook's (Meta) open source VoxPopuli dataset.
We tokenize the videos and obtain their discrete indices using WavTokenizer.
Many thanks to Meta (formally Facebook) for releasing it under the CC0 license.
Current Languages Uploaded
Lithuanian (Lt)
Estonian (Et)
Spanish (Es)
Slovak (Sk)
Croatian (Hr)
Finnish (Fi)
Dutch (Nl)
Hungarian (Hu)
Romanian (Ro)
Italian… See the full description on the dataset page: https://huggingface.co/datasets/GiftedNova/tokenized-voxpopuli.sokoban-10k-vjepa2-tokenized-shardscv17-xcodec-2.0-tokenizedlmd_1000_bass_tokenizednano4m-audio-tokenized
