datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gta-data-files-universalGTA5easyr1-103k-4MP-jedi-ui-vision-gta1-data
easyr1-103k-4MP-jedi-ui-vision-gta1-data
Merged dataset composed of the following sources:
datasets/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP (63031 samples in split train)
datasets/easyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b (39943 samples in split train)
Summary
Generated on: 2025-09-18 06:29:16 UTC
Split: train
Column strategy: intersection
Samples after merge: 102974
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-103k-4MP-jedi-ui-vision-gta1-data.easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter-tokenizedeasyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b
easyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b
This dataset was generated from filtered GTA shards with images streamed from ZIP archives.
Generated on: 2025-09-18 05:15:30 UTC
Script: push_easyr1_zip_shards_to_hf.py
Filters directory: /p/project1/synthlaion/awadalla1/gta-grounding-data-filters
JSONL glob: gta_shard_*zip.jsonl
Resize max: 4.0 MP
Prompt format: gta1 (output: coordinates)
Random seed: 42
Deduplicate: False
Debug images: True
System Prompt
You are… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b.easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter-coord-grid
easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter-coord-grid
Augmented version of easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter with a fixed 100px coordinate grid overlay.
Each image is overlaid with vertical and horizontal grid lines every 100
pixels at native resolution. Major ticks (every 1 steps)
are emphasized and axis labels show pixel values to help models localize
precise coordinates.
Summary
Generated on:… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter-coord-grid.easyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-not-all-correct-stage-one-temp-1_1-RL-a-keasyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-not-all-correct-stage-one-temp-1_1-RLCalliBench
🧠 CalliReader: Contextualizing Chinese Calligraphy via an Embedding-aligned Vision Language Model
📂 Code
📄 Paper
CalliBench is aimed to comprehensively evaluate VLMs' performance on the recognition and understanding of Chinese calligraphy.
📦 Dataset Summary
Samples: 3,192 image–annotation pairs
Tasks: Full-page recognition and Contextual VQA (choice of author/layout/style, bilingual interpretation, and intent analysis).
Annotations:
Metadata of author… See the full description on the dataset page: https://huggingface.co/datasets/gtang666/CalliBench.GTA5subset
GTA5 Subset for Zero-Shot Domain Adaptive Semantic Segmentation
This repository contains a curated subset of the GTA5 dataset, specifically designed for experiments in the paper Zero Shot Domain Adaptive Semantic Segmentation by Synthetic Data Generation and Progressive Adaptation. This subset includes images and labels necessary for training and evaluating models in zero-shot domain adaptive semantic segmentation scenarios.
The original GTA5 dataset was extracted from… See the full description on the dataset page: https://huggingface.co/datasets/roujin/GTA5subset.easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter
easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter
Augmented version of easyr1-57k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced with coordinate jitter.
For each original example, 1 additional copies were created. Each copy
randomly jitters the target coordinate by ±1 pixel in both X and Y. The
assistant coordinate in messages is updated, and bbox/normalized_bbox
are shifted when present.
Summary
Generated on: 2025-09-07 17:36:27 UTC
Source… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter.easyr1-21k-jedi-grounding-4MP-gta1-nores-fixed
easyr1-21k-jedi-grounding-4MP-gta1-nores-fixed
This dataset is a fixed version of /Users/anasawadalla/Desktop/cua/easyr1-21k-jedi-grounding-4MP. It applies two changes:
Rebuilds prompts/messages to GTA1 format without resolution in the system prompt
Removes samples where a 100x100 patch around the bbox midpoint is a solid color
Summary
Generated on: 2025-09-05 23:26:46 UTC
Source dataset: /Users/anasawadalla/Desktop/cua/easyr1-21k-jedi-grounding-4MP
Split: train… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-21k-jedi-grounding-4MP-gta1-nores-fixed.easyr1-60k-hard-qwen7b-easy-gta1-4MP-no-resolution-in-prompt-no-os-atlas
easyr1-60k-hard-qwen7b-easy-gta1-4MP-no-resolution-in-prompt-no-os-atlas
This dataset was generated using the EasyR1 grounding dataset pipeline.
Generation Details
Generated on: 2025-08-29 09:04:32 UTC
Script: push_easyr1_to_hf.py
Data directory: /lustre/fsw/portfolios/nvr/users/aawadalla/LLaMA-Factory/data
Parameters Used
Maximum samples: 60000
Image resize (max megapixels): 4.0 MP
Minimum native image resolution: 0.0 MP
Prompt format: gta1
Output format:… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-60k-hard-qwen7b-easy-gta1-4MP-no-resolution-in-prompt-no-os-atlas.easyr1-49k-hard-qwen7b-easy-gta1-stacked-pro-apps-no-resolution-in-prompt-ui-vision-5k-jedi-4MP
easyr1-49k-hard-qwen7b-easy-gta1-4MP-stacked-pro-apps-no-resolution-in-prompt-ui-vision-grounding-4MP-add-5k-jedi
Merged dataset composed of the following sources:
datasets/easyr1-44k-hard-qwen7b-easy-gta1-4MP-stacked-pro-apps-no-resolution-in-prompt-ui-vision-grounding-4MP (44769 samples in split train)
datasets/easyr1-21k-jedi-grounding-4MP-gta1-nores-fixed (18032 samples in split train)
Summary
Generated on: 2025-09-14 03:24:59 UTC
Split: train
Column strategy:… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-49k-hard-qwen7b-easy-gta1-stacked-pro-apps-no-resolution-in-prompt-ui-vision-5k-jedi-4MP.easyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-stage-three-temp-1_7-RL-zero-correct-to-0.2GTA5_dataseteasyr1-57k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced
easyr1-60k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced
Filtered version of /Users/anasawadalla/Desktop/cua/easyr1-60k-hard-qwen7b-easy-gta1-4MP-no-resolution-in-prompt-no-os-atlas aligned with fixed JEDI dataset /Users/anasawadalla/Desktop/cua/easyr1-21k-jedi-grounding-4MP-gta1-nores-fixed.
This dataset removes samples originating from JEDI that are not present in the fixed JEDI set.
JEDI samples are identified via exact (image_path, prompt) match with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-57k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced.GTArena-UI-Defectseasyr1-44k-hard-qwen7b-easy-gta1-stacked-pro-apps-no-resolution-in-prompt-ui-vision-4MP
easyr1-44k-hard-qwen7b-easy-gta1-4MP-stacked-pro-apps-no-resolution-in-prompt-ui-vision-grounding-4MP
Merged dataset composed of the following sources:
datasets/easyr1-38k-hard-qwen7b-easy-gta1-4MP-stacked-pro-apps-no-resolution-in-prompt (38979 samples in split train)
datasets/ui-vision-grounding-4MP (5790 samples in split train)
Summary
Generated on: 2025-09-13 22:23:37 UTC
Split: train
Column strategy: intersection
Samples after merge: 44769
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-44k-hard-qwen7b-easy-gta1-stacked-pro-apps-no-resolution-in-prompt-ui-vision-4MP.GTAV-Driving-Dataseteasyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-not-all-correct-stage-two-temp-1_1-RLeasyr1-10k-hard-qwen7b-easy-gta1-4MP-add-os-atlas-and-aria-gta1-filtered-with-qwen3b-and-gta1-7b
easyr1-10k-hard-qwen7b-easy-gta1-4MP-add-os-atlas-and-aria-gta1-filtered-with-qwen3b-and-gta1-7b
This dataset was generated using the EasyR1 grounding dataset pipeline.
Generation Details
Generated on: 2025-08-26 22:24:41 UTC
Script: push_easyr1_to_hf.py
Data directory: /lustre/fsw/portfolios/nvr/users/aawadalla/LLaMA-Factory/data
Parameters Used
Maximum samples: 10000
Image resize (max megapixels): 4.0 MP
Minimum native image resolution: 0.0 MP
Prompt… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-10k-hard-qwen7b-easy-gta1-4MP-add-os-atlas-and-aria-gta1-filtered-with-qwen3b-and-gta1-7b.easyr1-10k-hard-qwen7b-easy-gta1-4MP-add-jedi-grounding-with-no-filtering
easyr1-10k-hard-qwen7b-easy-gta1-4MP-add-jedi-grounding-with-no-filtering
This dataset was generated using the EasyR1 grounding dataset pipeline.
Generation Details
Generated on: 2025-08-27 00:53:16 UTC
Script: push_easyr1_to_hf.py
Data directory: /lustre/fsw/portfolios/nvr/users/aawadalla/LLaMA-Factory/data
Parameters Used
Maximum samples: 10000
Image resize (max megapixels): 4.0 MP
Minimum native image resolution: 0.0 MP
Prompt format: gta1_with_resolution… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-10k-hard-qwen7b-easy-gta1-4MP-add-jedi-grounding-with-no-filtering.gta5-kaggle-full
GTA5 Kaggle Full Mini Set
Full mini GTA5 dataset mirrored from Kaggle for coursework experiments.
Structure
images/
labels/
splits/train_full.txt
easyr1-10k-hard-qwen7b-easy-gta1-4MP-professional-apps-grounding-only-no-resolution-in-prompt
easyr1-10k-hard-qwen7b-easy-gta1-4MP-professional-apps-grounding-only-no-resolution-in-prompt
This dataset was generated using the EasyR1 grounding dataset pipeline.
Generation Details
Generated on: 2025-08-26 12:16:32 UTC
Script: push_easyr1_to_hf.py
Data directory: /lustre/fsw/portfolios/nvr/users/aawadalla/LLaMA-Factory/data
Parameters Used
Maximum samples: 10000
Image resize (max megapixels): 4.0 MP
Minimum native image resolution: 0.0 MP
Prompt format:… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-10k-hard-qwen7b-easy-gta1-4MP-professional-apps-grounding-only-no-resolution-in-prompt.CalliTrain
🧠 CalliReader: Contextualizing Chinese Calligraphy via an Embedding-aligned Vision Language Model
📂 Code
📄 Paper
CalliBench is aimed to comprehensively evaluate VLMs' performance on the recognition and understanding of Chinese calligraphy.
📦 Dataset Summary
Samples: 3,192 image–annotation pairs
Tasks: Full-page recognition and Contextual VQA (choice of author/layout/style, bilingual interpretation, and intent analysis).
Annotations:
Metadata of author… See the full description on the dataset page: https://huggingface.co/datasets/gtang666/CalliTrain.sim2real_gta5_to_cityscapeseasyr1-20k-hard-qwen7b-easy-gta1-4MPeasyr1-38k-hard-qwen7b-easy-gta1-4MP-stacked-pro-apps-no-resolution-in-prompt
easyr1-38k-hard-qwen7b-easy-gta1-4MP-stacked-pro-apps-no-resolution-in-prompt
This dataset was generated using the EasyR1 grounding dataset pipeline.
Generation Details
Generated on: 2025-08-24 18:17:24 UTC
Script: push_easyr1_to_hf.py
Data directory: /lustre/fsw/portfolios/nvr/users/aawadalla/LLaMA-Factory/data
Parameters Used
Maximum samples: 50000
Image resize (max megapixels): 4.0 MP
Minimum native image resolution: 0.0 MP
Prompt format: gta1
Output… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-38k-hard-qwen7b-easy-gta1-4MP-stacked-pro-apps-no-resolution-in-prompt.easyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-stage-two-temp-1_1-RL-zero-correct-to-0.3
