G-Eval
Datasets
All datasets matching “G-Eval”POINTS-GUI-G-Evaluation
Evaluation Datasets for POINTS-GUI-G
Since various evaluation benchmarks are currently quite heterogeneous, for ease of use, we have organized ScreenSpotv2, MMBench-L2, ScreenSpot-Pro, UIVision, and OSWorld-G into the same format, allowing you to perform one-click evaluation using VLMEvalKit. For detailed evaluation procedures, please refer to POINTS-GUI.
When using this dataset, you must comply with the original licenses of each respective dataset.
easyr1-osworld-g-eval-4MPeasyr1-osworld-g-eval-2MPeasyr1-osworld-g-refined-eval
Grounding dataset for vLLM eval (original-resolution)
This dataset was generated from a unified JSON + raw images and formatted
for your vLLM evaluator.
- **Images**: original resolution (no resizing)
- **Prompt**: GTA1-style header + `<image>` + instruction → `easyr1_prompt`
- **Images column**: `images` is a list with a single `Image` (so your evaluator can do `imgs[0]`)
- **Ground truth**:
- If a `polygon_xy` exists in source, we export `normalized_polygon` (and set… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-osworld-g-refined-eval.easyr1-osworld-g-refined-eval-4MP
easyr1-osworld-g-refined-eval-4MP
This dataset was generated from eval-grounding annotations using convert_eval_grounding_to_hf.py.
Parameters
JSON file: /p/project1/synthlaion/awadalla1/eval-grounding-data/osworld-g-eval-refined.json
Images dir: /p/project1/synthlaion/awadalla1/eval-grounding-data
Max samples: None
Resize max MP: 4.0
Prompt format: gta1
Output format: coordinates
Debug images: False
Dataset Statistics
Train samples: 510
Columns: id… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-osworld-g-refined-eval-4MP.easyr1-osworld-g-refined-eval-4MP-refusal
