CoolFace
Datasetpublic

mlfoundations-cua-dev/easyr1-10k-hard-qwen7b-easy-gta1-composed-20-dual-5-montage-4MP

easyr1-10k-hard-qwen7b-easy-gta1-composed-20-dual-5-montage-4MP This dataset was generated using the enhanced EasyR1 grounding dataset pipeline with composition capabilities. Generation Details Generated on: 2025-08-24 11:02:10 UTC Script: push_easyr1_composed_to_hf.py Data directory: /lustre/fs12/portfolios/nvr/projects/nvr_lacr_llm/users/aawadalla/LLaMA-Factory/data Parameters Used Maximum samples: 10000 Image resize (max megapixels): 4.0 MP… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-10k-hard-qwen7b-easy-gta1-composed-20-dual-5-montage-4MP.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes24downloads
Dataset Card

easyr1-10k-hard-qwen7b-easy-gta1-composed-20-dual-5-montage-4MP

This dataset was generated using the enhanced EasyR1 grounding dataset pipeline with composition capabilities.

Generation Details

  • —Generated on: 2025-08-24 11:02:10 UTC
  • —Script: push_easyr1_composed_to_hf.py
  • —Data directory: /lustre/fs12/portfolios/nvr/projects/nvr_lacr_llm/users/aawadalla/LLaMA-Factory/data

Parameters Used

  • —Maximum samples: 10000
  • —Image resize (max megapixels): 4.0 MP
  • —Minimum native image resolution: 0.0 MP
  • —Prompt format: gta1_with_resolution
  • —Output format: coordinates
  • —Random seed: 42

Composition Settings

  • —Dual-screen ratio: 0.2 (20% of samples)
  • —Montage ratio: 0.05 (5% of samples)
  • —Single-screen ratio: 0.75 (75% of samples)

Dataset Groups

The following JSON/JSONL files were used to create this dataset:

Dataset Group 1

Files (intersection of kept samples across all files):

  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/pixmo-points-qwentoolcall-notgrounded-QwenQwen2.5-VL-7B-Instruct-qwentoolcall.jsonl
  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/pixmo-points-gta1-correctlygrounded-HelloKKMeGTA1-7B-gta1.jsonl

Dataset Group 2

Files (intersection of kept samples across all files):

  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/autogui-grounding-only-qwentoolcall-notgrounded-QwenQwen2.5-VL-7B-Instruct-qwentoolcall.jsonl
  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/autogui-grounding-only-gta1-correctlygrounded-HelloKKMeGTA1-7B-gta1.jsonl

Dataset Group 3

Files (intersection of kept samples across all files):

  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/seeclick-qwentoolcall-notgrounded-QwenQwen2.5-VL-7B-Instruct-qwentoolcall.jsonl
  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/seeclick-gta1-correctlygrounded-HelloKKMeGTA1-7B-gta1.jsonl

Dataset Group 4

Files (intersection of kept samples across all files):

  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/pc-e-grounding-only-claude-instructions-qwentoolcall-notgrounded-QwenQwen2.5-VL-7B-Instruct-qwentoolcall.jsonl
  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/pc-e-grounding-only-claude-instructions-gta1-correctlygrounded-HelloKKMeGTA1-7B-gta1.jsonl

Dataset Group 5

Files (intersection of kept samples across all files):

  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/omniact-grounding-only-qwentoolcall-notgrounded-QwenQwen2.5-VL-7B-Instruct-qwentoolcall.jsonl
  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/omniact-grounding-only-gta1-correctlygrounded-HelloKKMeGTA1-7B-gta1.jsonl

Dataset Group 6

Files (intersection of kept samples across all files):

  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/showui-desktop-qwentoolcall-notgrounded-QwenQwen2.5-VL-7B-Instruct-qwentoolcall.jsonl
  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/showui-desktop-gta1-correctlygrounded-HelloKKMeGTA1-7B-gta1.jsonl

Dataset Group 7

Files (intersection of kept samples across all files):

  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/showui-web-qwentoolcall-notgrounded-QwenQwen2.5-VL-7B-Instruct-qwentoolcall.jsonl
  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/showui-web-gta1-correctlygrounded-HelloKKMeGTA1-7B-gta1.jsonl

Dataset Group 8

Files (intersection of kept samples across all files):

  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/uground-qwentoolcall-notgrounded-QwenQwen2.5-VL-7B-Instruct-qwentoolcall.jsonl
  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/uground-gta1-correctlygrounded-HelloKKMeGTA1-7B-gta1.jsonl

Dataset Group 9

Files (intersection of kept samples across all files):

  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/waveui-qwentoolcall-notgrounded-QwenQwen2.5-VL-7B-Instruct-qwentoolcall.jsonl
  • —/lustre/fs12/portfolios/nvr/projects/nvrlacrllm/users/aawadalla/grounding-data-filters/waveui-gta1-correctlygrounded-HelloKKMeGTA1-7B-gta1.jsonl

Dataset Statistics

  • —Total training samples: 10000
  • —Composition breakdown:
  • —Single-screen: 7500
  • —Dual-screen: 2000
  • —Montage: 500
  • —Image dimensions: Variable (based on composition type)
  • —Columns: images, prompt, easyr1prompt, bbox, imagepath, messages, imagewidth, imageheight, compositiontype

System Prompt

The following system prompt is used for this dataset:

You are an expert UI element locator. Given a GUI image and a user's element description, provide the coordinates of the specified element as a single (x,y) point. The image resolution is height 2048 and width 2048. For elements with area, return the center point.

Output the coordinate pair exactly:
(x,y)

Sample Entry

  • —User prompt: <image> Find and click Sport
  • —Assistant response: (891,172)
  • —Bounding box: [865, 159, 917, 185]
  • —Image path: showui-web-images/showuiweb021834.jpg
  • —Composition type: single

Usage

python
from datasets import load_dataset

dataset = load_dataset("mlfoundations-cua-dev/easyr1-10k-hard-qwen7b-easy-gta1-composed-20-dual-5-montage-4MP")

# Access the training data
train_data = dataset['train']

# Example: Get samples by composition type
single_screen = [s for s in train_data if s.get('_composition_type') == 'single']
dual_screen = [s for s in train_data if s.get('_composition_type') == 'dual_screen']
montage = [s for s in train_data if s.get('_composition_type') == 'montage']

Composition Types

Single-Screen

Standard single image samples with UI element grounding.

Dual-Screen

Two images concatenated horizontally, simulating dual-monitor setups. Coordinates are adjusted to the correct screen half.

Montage

Multiple application windows overlaid on desktop backgrounds, simulating real multi-window desktop environments.

License

Please refer to the original dataset licenses for usage restrictions.