datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Imagenet-1k_validationCOCO_captions_validation
Dataset Card for "COCO_captions_validation"
More Information needed
VQAv2_validation
Dataset Card for "VQAv2_validation"
More Information needed
VQAv2_sample_validation
Dataset Card for "VQAv2_sample_validation"
More Information needed
Places365-ValidationImagenet1k_sample_validation
Dataset Card for "Imagenet1k_sample_validation"
More Information needed
TextVQA_validation
Dataset Card for "TextVQA_validation"
More Information needed
imagenet-1k-validation-subsetsVPPO_MMK12_validation
Dataset Card for VPPO_MMK12_validation
Dataset Details
Dataset Description
This dataset is the official validation split used to fine-tune the VPPO-7B and VPPO-32B models presented in our paper, "Spotlight on Token Perception for Multimodal Reinforcement Learning".
This is a direct copy of the test split of FanqingM/MMK12 dataset. We have isolated it here to ensure the exact version used in our experiments is publicly available, guaranteeing reproducibility for… See the full description on the dataset page: https://huggingface.co/datasets/chamber111/VPPO_MMK12_validation.VQAv2_minival_validation_vprevious
Dataset Card for "VQA_minival_validation"
More Information needed
VizWiz_validation
Dataset Card for "VizWiz_validation"
More Information needed
alt-text-validationThis dataset contains images and alt text from various sources.
It is used to control the quality of https://huggingface.co/Mozilla/distilvit using the https://github.com/mozilla/checkvite application
This application let users try out the model on the images and classify them. The dataset is then updated.
When an image is marked as need_training it will be use to fine-tune the model to fix some of its inaccuracies
VQAv2_minival_validation
Dataset Card for "VQAv2_minival_validation_v2"
More Information needed
medical_records_parsing_validation_set
Medical Records Parsing Validation Set
Dataset Composition and Clinical Relevance
The Eka Medical Records Parsing Dataset empowers evaluation of AI systems designed to extract structured information from unstructured medical documents, enabling true digitisation of healthcare data while maintaining clinical accuracy.
The dataset comprise 288 carefully selected images of laboratory reports and prescriptions representing diverse formats and templates encountered in Indian… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/medical_records_parsing_validation_set.Chest_Xray_validation_N_Hotacouslic_ai_validationImagenette_validation
Dataset Card for "Imagenette_validation"
More Information needed
VQAv2_sample_validation
Dataset Card for "VQAv2_sample_validation"
More Information needed
smolvla_calvin_task_D_D_validationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.image": {
"dtype": "image",
"shape": [
200,
200,
3
],
"names": [
"height",
"width",
"channels"
]
},
"observation.images.image2": {… See the full description on the dataset page: https://huggingface.co/datasets/dohyun1411/smolvla_calvin_task_D_D_validation.MMMU-Reasoning-Distill-Validationvalidation_setbangla-ocr-validation_data_printed
Bangla OCR Validation Dataset (Printed + Scanned)
📌 Description
This dataset is a Bangla OCR validation dataset containing a mix of printed document images and their corresponding text annotations. It is designed to evaluate OCR and vision-language models on both clean digital text and scanned document images.
📊 Dataset Composition
1507 line-level images with text annotations
50 full-page document images with text
Data includes:
Printed/typed Bangla text… See the full description on the dataset page: https://huggingface.co/datasets/arobin79/bangla-ocr-validation_data_printed.clothesdataset_validationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 6,
"total_frames": 850,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/spokkali/clothesdataset_validation.BGData_Validationroco2-question-dataset-validationrendered_wikipedia_validationsatellite_Roofs_MASK_train_validation_v1TinyImagenet_validation
Dataset Card for "TinyImagenet_validation"
More Information needed
validation-datasetARES-hard-validation🌟 ARES — Adaptive Multimodal Reasoning FrameworkTwo-stage adaptive reasoning: cold-start + entropy-shaped RL.
🔑 HighlightsBalanced reasoning across easy & hard tasks via token-level entropy shaping.SOTA efficiency–accuracy tradeoffs on diverse multimodal and textual benchmarks.
📚 Training Pipeline
Adaptive Cold-Start — curate difficulty-aware reasoning traces
Entropy-Shaped RL (AEPO) — trigger exploration via high-window entropy, hierarchical rewards
📂 Resources
Paper: ARES:… See the full description on the dataset page: https://huggingface.co/datasets/ares0728/ARES-hard-validation.
