datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zangei-dit-stage-1-250k-256px-imgking3zangei_tokenizer_1M_imgs
zangei_tokenizer_1M_imgs
Packed image dataset for Zangei tokenizer training. Each record has a stable core_id
that can be used to join this image dataset with DINOv3 feature shards and future
text/metadata datasets.
Views
original_webp: original image dimensions, encoded as WebP.
center_sq_256_webp: center-square crop, resized to 256x256, encoded as WebP.
resized_256_webp: direct resize to 256x256 without cropping, encoded as WebP.
Shard target: about 10GB per shard… See the full description on the dataset page: https://huggingface.co/datasets/kingsidharth/zangei_tokenizer_1M_imgs.kinetic-400_450sampleshq-kineticsAfrivoice_Kinyarwanda_old_version
Dataset summary
[need more information]
Supported tasks
[need more information]
How to use
[need more information]
Dataset structure
Data fields
[need more information]
Data splits
[need more information]
Data preprocessing
[need more information]
Licensing Information
All datasets are licensed under the Creative Commons license (CC-BY-4).
kink-catalog-slim
