datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PM4Bench-QGO-Train
PM4Bench QGO training data
Synthetic multilingual OCR data for QGO reinforcement learning
Overview
The paper Benchmarking and Boosting Multilingual Capabilities of LVLMs via
OCR-Centric Reinforcement Learning
uses PM4Bench to show that OCR is a key source of cross-lingual performance
gaps when text is rendered visually. QGO addresses that finding with GRPO on
synthetic OCR data, without task-specific VQA or GUI supervision. This
repository contains the… See the full description on the dataset page: https://huggingface.co/datasets/DatasetMan/PM4Bench-QGO-Train.qgeo-smolvla-depth-sogwikipedia_with_imagesimageqguideprj_gia_dataset_metaworld_assembly_v2_1111An imitation learning environment for the assembly-v2 environment, sample for the policy assembly-v2
This environment was created as part of the Generally Intelligent Agents project gia: https://github.com/huggingface/gia
Load dataset
First, clone it with
git clone https://huggingface.co/datasets/qgallouedec/prj_gia_dataset_metaworld_assembly_v2_1111
Then, load it with
import numpy as np
dataset = np.load("prj_gia_dataset_metaworld_assembly_v2_1111/dataset.npy"… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/prj_gia_dataset_metaworld_assembly_v2_1111.svgamberfloor-plans-datasettwitter-dogpaw88-2026.01.12-2010643391567343999-qgZ7fLp6ytjjbPqd-part1twitter-DwaynemillLom-2025.07.23-1947923787368038725-QG_NNBG1OcvE4S88-part1twitter-nyanchan2k3-2026.01.05-2008181276046770501-qG66J22mImRvJApu-part1twitter-TaoHuaBang-2026.02.14-2022528613112054224-8u2-hTh-qGVbBABl-part1qgrp-jiegou-exif-images
