datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
robomme_preprocessed_data
RoboMME Training Data (Pickle Format)
Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code
This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments.
.
├── data # zipped pickle files
├── features # zipped precompute siglip embeddings
├── meta # statistics for robomme
├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.ChartQA_small_preprocessedSeeU45_PreProcessednabirds_custom_split_preprocessedllamaindex-vdr-en-train-preprocessed
llamaindex-vdr-en-train-preprocessed
This dataset is a preprocessed English subset of llamaindex/vdr-multilingual-train, prepared for training multimodal Sentence Transformer embedding models on document screenshot retrieval.
Changes from the original dataset
The original llamaindex/vdr-multilingual-train dataset stores hard negatives as a list of ID strings that reference other rows. This dataset makes two key changes:
English only: Only the English subset (53,512… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/llamaindex-vdr-en-train-preprocessed.DuMPlinGS-preprocessed-datadata_remove_v0_preprocessedzebra-cot-mistral-small-3.2-24b-preprocessed
Zebra-CoT Preprocessed — Mistral Hackathon 2026
Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct.
Format
text: formatted as [INST] question [/INST] <think> reasoning </think> answer
image: PIL JPEG image for the corresponding visual task
Usage
Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning.
Hackathon
Created for Mistral Hackaton 2026 — Fine-tuning track with W&B.
patchlet-embed-preprocessedDL3DV-benchmark-preprocessedDiabetic_Retinopathy_Preprocessed_Dataset_256x256This is dataset comes from this Kaggle Dataset
from the user Sachin Kumar.
The goal of the dataset is for the Varun AIM Projects to easily start running and download the dataset on their local computer in the HF libraries as the directory I strongly recommedn to use.
MeshAI-Preprocessed-4Kdata_remove_v0_preprocessed_1000Diabetic_Retinopathy_Detection_preprocessedPreprocessed_Neu3Dmmsr_1MP_benchmark_preprocessedDiabetic_Retinopathy_Detection_preprocessed2mammosightr-preprocessed
MammosighTR — Preprocessed Mammography Dataset (BI-RADS)
Preprocessed PNG mammograms with image-level BI-RADS labels, derived from
the nationwide Turkish breast-cancer screening dataset (MammosighTR)
released for the TEKNOFEST 2023 Artificial Intelligence in Health Competition
by the Republic of Turkey Ministry of Health. Original DICOMs are cropped to
the breast region with a YOLOX detector and exported as PNG; we add an
image-level metadata mapping built from the official… See the full description on the dataset page: https://huggingface.co/datasets/gulluk/mammosightr-preprocessed.DocLayNet-Instruct-v1-preprocessedwaifu-preprocessed-datasetsvlm-preprocessed-datasets-v2seraiki-handwritten-preprocessedMathVision_preprocessedpreprocessedChartQA_small_preprocessed_chunked_boxPearl-vdr-ar-train-preprocessed
Pearl-vdr-ar-train-preprocessed
Arabic culturally-aligned, VDR-style (query, image, hard-negatives) triplets for training multimodal embedding models with Sentence Transformers.
Dataset structure
Each row contains:
Column
Type
Description
query
string
Arabic text question about the image
category
string
High-level Arab-culture topic (Music, Landmarks, Cuisine, ...)
country
string
Country the sample is anchored to (Algeria, Saudi Arabia, ...)
image
image… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Pearl-vdr-ar-train-preprocessed.bus_cot_preprocessedG17-preprocessed-datasetpreprocessed_amazon_productssunday-intelligence-imojis-preprocessed
