datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMMU-LLM-R1-format
MMMU-LLM-R1 Reformatted Dataset
walton-multimodal-cold-start-r1-format-30k
walton-multimodal-cold-start-r1-format-30k
WaltonFuture/Multimodal-Cold-Start converted to multimodal-open-r1-8k-verified format with filtering
Dataset Description
This dataset was processed using the data-preproc package for vision-language model training.
Processing Configuration
Base Model: Qwen/Qwen2.5-7B-Instruct
Tokenizer: Qwen/Qwen2.5-7B-Instruct
Sequence Length: 16384
Processing Type: Vision Language (VL)
Dataset Features
input_ids:… See the full description on the dataset page: https://huggingface.co/datasets/penfever/walton-multimodal-cold-start-r1-format-30k.walton-multimodal-cold-start-r1-format
walton-multimodal-cold-start-r1-format
WaltonFuture/Multimodal-Cold-Start converted to multimodal-open-r1-8k-verified format with filtering
Dataset Description
This dataset was processed using the data-preproc package for vision-language model training.
Processing Configuration
Base Model: Qwen/Qwen2.5-7B-Instruct
Tokenizer: Qwen/Qwen2.5-7B-Instruct
Sequence Length: 16384
Processing Type: Vision Language (VL)
Dataset Features
input_ids: Tokenized… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/walton-multimodal-cold-start-r1-format.MMMU-multimodal-open-r1-formatMMMU-满血版R1蒸馏多模态Reasoning验证集参照lmms-lab/multimodal-open-r1-8k-verified格式改版
vlaa-thinking-sft-to-r1-format
vlaa-thinking-sft-to-r1-format
vlaa-thinking-sft dataset transformed to match multimodal-open-r1-8k-verified format with filtering
Dataset Description
This dataset was processed using the data-preproc package for vision-language model training.
Processing Configuration
Base Model: allenai/Molmo-7B-O-0924
Tokenizer: allenai/Molmo-7B-O-0924
Sequence Length: 8192
Processing Type: Vision Language (VL)
Dataset Features
input_ids: Tokenized input… See the full description on the dataset page: https://huggingface.co/datasets/penfever/vlaa-thinking-sft-to-r1-format.
