datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MolmoAct-Pretraining-Mixture
MolmoAct - Pretraining Mixture
Data Mixture used for MolmoAct Pretraining. Contains a subset of OXE formulated as Action Reasoning Data along with auxiliary robot data and link to Multimodal Web data.
MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Pretraining-Mixture.ABC-Pretraining-Data
ABC Pretraining Data
This dataset contains the pretraining data for ABC, an open-source multimodal embedding model that uses a vision-language model backbone to deeply integrate image features with natural language instructions, advancing the state of visual embeddings with natural language control.
This dataset is derived from Google's Conceptual Captions dataset.
Each item in the dataset contains a URL where the corresponding image can be downloaded and mined negatives for… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ABC-Pretraining-Data.Vn-OCR-PretrainingJp-OCR-Pretrainingpropagator-multimodal-pretraining-data
Propagator Multimodal Pretraining Data
This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format.
This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout.
Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.anime_pretraining_2RoboX-VQA-Pretrainingwc-vqa-continual-pretraining
garrykuwanto/wc-vqa-continual-pretraining
English subset of worldcuisines/vqa-v1.1
Task 1 (dish name prediction), packaged for continual pretraining / SFT of
small vision-language models (specifically SmolVLM2-256M).
Splits
split
rows
source
train
27000
task1 / train / lang=en
validation
300
task1 / test_small / lang=en
test
1500
task1 / test_large / lang=en
Schema
field
type
notes
image
Image
bytes from upstream images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/garrykuwanto/wc-vqa-continual-pretraining.noise-pretrainingmo436-ssl-pretraining-cropsmo436-ssl-pretrainingpre_training_test
