datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wan22-animate-3k-opensource-data
Wan2.2 Animate Open Dataset Pack
This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment.
The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards.
Restore:
cat datasets.tar.part-* | tar -xf -
sha256sum -c SHA256SUMS
After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.color-classification-opensourceissuesopensource_augment_processedopensource_augmentgmai_vl_5m_opensource_cleaned
gmai_vl_5m_opensource_cleaned
The gmai_vl_5m_opensource family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
272,632
QA turns
1,024,908
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
77
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/gmai_vl_5m_opensource_cleaned.2018-opensourcebioreactoropensource-collect-tessellation-augmentmodal-kaggle-opensource-stack-artifactsopensourcenewopen-source-galleriesopensource-datasets
