datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data
podcast-tokenized-bg3.5-enj5podcast-tokenized-bg2.5-enj4.5tokenized-voxpopuli
Tokenized Vox-Populi Dataset
This repository serves to store the tokenized audio dataset extracted from Facebook's (Meta) open source VoxPopuli dataset.
We tokenize the videos and obtain their discrete indices using WavTokenizer.
Many thanks to Meta (formally Facebook) for releasing it under the CC0 license.
Current Languages Uploaded
Lithuanian (Lt)
Estonian (Et)
Spanish (Es)
Slovak (Sk)
Croatian (Hr)
Finnish (Fi)
Dutch (Nl)
Hungarian (Hu)
Romanian (Ro)
Italian… See the full description on the dataset page: https://huggingface.co/datasets/GiftedNova/tokenized-voxpopuli.sokoban-10k-vjepa2-tokenized-shardszangei_tokenizer_1M_imgs
zangei_tokenizer_1M_imgs
Packed image dataset for Zangei tokenizer training. Each record has a stable core_id
that can be used to join this image dataset with DINOv3 feature shards and future
text/metadata datasets.
Views
original_webp: original image dimensions, encoded as WebP.
center_sq_256_webp: center-square crop, resized to 256x256, encoded as WebP.
resized_256_webp: direct resize to 256x256 without cropping, encoded as WebP.
Shard target: about 10GB per shard… See the full description on the dataset page: https://huggingface.co/datasets/kingsidharth/zangei_tokenizer_1M_imgs.TokenIT_democv17-xcodec-2.0-tokenizedimage_token_training_Lumina
Image Token Training Sets
This repository contains public training token files for internal research / experimentation.
Contents
COCO_Janus_tokens_for_train/
laion_Lumina7B_tokens_for_train/
midjourney_Lumina7B_tokens_for_train/
File format
Each subset contains .pt files.
Notes
The files are provided as precomputed training artifacts.
Consumers should load them with PyTorch.
Directory structure is preserved as-is.
lmd_1000_bass_tokenizedmixed-fluac-tokenizernano4m-audio-tokenized
