datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ZoomBench
ZoomBench: A Fine-Grained Multimodal Perception Benchmark
📃 Paper | 🏠 Project | 🤗 Models
Overview
ZoomBench is a challenging benchmark designed to evaluate the fine-grained multimodal perception capabilities of Multimodal Large Language Models (MLLMs). It specifically targets scenarios where decisive visual evidence is small, subtle, or easily overwhelmed by global context — situations that demand "zooming-level" perception from a single full image.
It is… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZoomBench.corruption-zoom_blur
Corruption Dataset: Zoom_Blur
Dataset Description
This dataset contains corrupted versions of ImageNet-1K images using zoom_blur corruption. It is part of the ImageNet-C benchmark for evaluating model robustness to common image corruptions.
Dataset Structure
Train: 1,281,167 corrupted images
Validation: 50,000 corrupted images
Classes: 1000 ImageNet-1K classes
Format: Arrow (Hugging Face Datasets)
Corruption Type: Zoom_Blur
Applies zoom blur… See the full description on the dataset page: https://huggingface.co/datasets/MarMaster/corruption-zoom_blur.Q-Zoom-Training
Q-Zoom Training Data
Curated training data for the Q-Zoom gated Region-of-Interest mechanism for Vision-Language Models. Companion dataset to the Q-Zoom release repository.
What this repo contains
This repo holds the question JSONLs and the ROI training pickles used to train the three components of Q-Zoom (SD-RPN, Post-SFT, Dynamic Gate). It is the companion to:
YuhengSSS/RoITraining — image archives (*.tar / *.zip) for COCO / GQA / OCR-VQA / DocVQA / TextVQA / ChartQA /… See the full description on the dataset page: https://huggingface.co/datasets/YuhengSSS/Q-Zoom-Training.ZoomBenchmldr-zoomed-100char-total-queries-documentsmldr-zoomed-1000char-total-queries-documentsmldr-zoomed-1000char-2000-queries-documentsvstar_direct_attributes_seal_zoomSriLankaLawmldr-zoomed-1000char-total1m2r-rope-zoom-fifteenThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 30,
"total_frames": 9824,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-zoom-fifteen.message-decoding-abc-zoom-inreal-world-noise-through-zoom
Real-World Noise Through Zoom (RWNTZ)
5 different real world noise settings (bedroom, crowded room, background music, rain, road with cars)
2 different speakers
various microphone distances (6 inches, 24 inches)
32 total samples with different phrases
recorded through Zoom to simulate real-world linguistic fieldwork scenarios
manually verified word level transcriptions
g2p phoneme trancriptions
audio to phoneme trancriptions with a variety of Wav2Vec2 based models
message-decoding-words-and-sequences-target-zoom-in-r11m2r-rope-zoomThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 50,
"total_frames": 16500,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-zoom.1m2r-rope-zoom-15demos-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5-panda",
"total_episodes": 30,
"total_frames": 9824,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-zoom-15demos-v30.message-decoding-words-and-sequences-target-zoom-in1m2r-rope-zoom-15demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 30,
"total_frames": 9824,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-zoom-15demos.1m2r-rope-zoom-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5-panda",
"total_episodes": 50,
"total_frames": 16500,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-zoom-v30.mldr-zoomed-100char-totalvstar_relative_position_seal_zoommessage-decoding-abc-zoom-in-r11m2r-cot-rope-zoomThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"total_videos": 0,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 10,
"splits": {},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-cot-rope-zoom.message-decoding-words-and-sequences-zoom-in-r11m2r-rope-zoom-5demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 10,
"total_frames": 3228,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-zoom-5demos.1m2r-rope-zoom-fiveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 10,
"total_frames": 3228,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-zoom-five.1m2r-rope-zoom-5demos-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5-panda",
"total_episodes": 10,
"total_frames": 3228,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-zoom-5demos-v30.zoo_med_Q_test1zoo_med_test1Drug-1
