datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies.
Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process).
UI-Genie-Agent-16kThis repository contains the Trajectory dataset from the paper UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based
Mobile GUI Agents.
Github: https://github.com/Euphoria16/UI-Genie
swiss-caselaw-web-ui
Swiss Case Law Open Dataset
962,724 published decisions from Swiss federal, cantonal, and regulatory bodies.
Full text, structured metadata, and daily updates. The March 20, 2026 snapshot contains German, French, and Italian decisions; the export schema also reserves rm for Romansh.
What this is
A structured, searchable archive of Swiss court decisions — from the Federal Supreme Court (BGer) down to cantonal courts in all 26 cantons. Every decision includes the full… See the full description on the dataset page: https://huggingface.co/datasets/ArneH/swiss-caselaw-web-ui.UIBenchKitUni-GUI-Desktop-1
Uni-GUI-Desktop-1
A large-scale desktop GUI agent trajectory dataset, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
2,685
Total Steps
~36K
Platform
Desktop (1920x1080)
Applications
10 categories
Coordinate System
Normalized to [0, 999]
Applications
App
Description
chrome
Web browsing tasks
gimp… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-Desktop-1.UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper
The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files.
Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/LiamLian0727/UIIS.Uni-GUI-OpenMobile
Uni-GUI-OpenMobile
A mobile GUI agent trajectory dataset collected on open-source Android applications via AndroidWorld, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
2,640
Total Steps
~25.9K
Platform
Android Mobile (1080x2400)
Applications
19 open-source apps
Coordinate System
Normalized to [0, 1000]… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-OpenMobile.tripitaka-mbu
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกและอรรถกถาไทยฉบับมหามกุฏราชวิทยาลัย จำนวน ๙๑ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑ (863 หน้า): พระวินัยปิฎก มหาวิภังค์ เล่ม ๑ ภาค ๑
เล่ม ๒ (664 หน้า): พระวินัยปิฎก มหาวิภังค์ เล่ม ๑ ภาค ๒
เล่ม ๓: พระวินัยปิฎก มหาวิภังค์ เล่ม ๑ ภาค ๓
เล่ม ๔: พระวินัยปิฎก มหาวิภังค์ เล่ม ๒
เล่ม ๕: พระวินัยปิฎก… See the full description on the dataset page: https://huggingface.co/datasets/uisp/tripitaka-mbu.uiui-vision
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Introduction
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/ui-vision.qb-audio
qb-audio — read-aloud audio for quizbowl questions
TTS read-aloud recordings of quizbowl questions from
qbreader, one Opus file per tossup / bonus.
Built so any quizbowl app can add listening practice over plain HTTP — no
install, no API key. Generation is in progress (July 2026); the dataset
grows as the run proceeds.
Layout
tossups/{qid[-2:]}/{qid}.opus, bonuses/{qid[-2:]}/{qid}.opus — audio,
Opus 24 kHz mono. qid is the question's qbreader _id; files shard by… See the full description on the dataset page: https://huggingface.co/datasets/uild42/qb-audio.uioUni-GUI-Mobile
Uni-GUI-Mobile
A mobile GUI agent trajectory dataset, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
871
Total Steps
~14K
Platform
Mobile (1080x2400)
Applications
10 categories
Coordinate System
Normalized to [0, 999]
Applications
App
Description
calendar
Calendar event management
chrome
Mobile… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-Mobile.cart_dataset_part1Uni-GUI-OpenCUA
Uni-GUI-OpenCUA
A post-processed desktop GUI agent trajectory dataset derived from OpenCUA, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Overview
Uni-GUI-OpenCUA contains 832 trajectories with ~14K interaction steps across 11 desktop applications and task categories. The raw OpenCUA trajectories have been cleaned, filtered, and post-processed through our Unified Cross-Platform Data… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-OpenCUA.BLEnD
BLEnD
This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track).
24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.uipieertruyjuiuoipiuytrUIIS-semanticuiupouyjy567industrial_cart_2UI-Elements-Detection-Dataset
Web UI Elements Dataset
Overview
A comprehensive dataset of web user interface elements collected from the world's most visited websites. This dataset is specifically curated for training AI models to detect and classify UI components, enabling automated UI testing, accessibility analysis, and interface design studies.
Key Features
300+ popular websites sampled
15 essential UI element classes
High-resolution screenshots (1920x1080)
Rich accessibility metadata… See the full description on the dataset page: https://huggingface.co/datasets/YashJain/UI-Elements-Detection-Dataset.uioipiu54676uiuopojhtrg345Aria-UI_Data
🖼️ Try Aria-UI! · 📖 Project Page · 📌 Paper
· ⭐ Code · 📚 Aria-UI Checkpoints
Overview of the data
Web
Mobile
Desktop
Element Caption Field
"element caption"
"long_element_caption", "short_element_caption"
"element caption"
Instruction Field
"instructions"
"instructions"
"instructions"
Collection Source
Aria-UI Common Crawl
AMEX Original Dataset
Aria-UI Ubuntu
Number of Instructions
2.9M
1.1M
150K
Number of Images
173K
104K
7.8K
Our dataset… See the full description on the dataset page: https://huggingface.co/datasets/Aria-UI/Aria-UI_Data.dhamma-scholar-book
Multi-File CSV Dataset
คำอธิบาย
หนังสือนักธรรม ตรี โท เอก จำนวน ๕๒ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑ (82 หน้า): นักธรรมตรี - นวโกวาท
เล่ม ๒ (82 หน้า): นักธรรมตรี - พุทธศาสนาสุภาษิต เล่ม ๑
เล่ม ๓ (106 หน้า): นักธรรมตรี - พุทธประวัติเล่ม ๑
เล่ม ๔: นักธรรมตรี - พุทธประวัติเล่ม ๒
เล่ม ๕: นักธรรมตรี - พุทธประวัติเล่ม ๓
เล่ม ๖:… See the full description on the dataset page: https://huggingface.co/datasets/uisp/dhamma-scholar-book.uiuiopouiauykiuopouMobileWorld-Eval-Results
MobileWorld-Eval-Results
Evaluation results of the UI-MOPD trained model (Qwen3-VL-8B-Thinking) on a mobile agent benchmark. Contains full execution trajectories including screenshots, marked action visualizations, and task outcomes across 117 diverse mobile tasks.
Evaluation Summary
Metric
Value
Model
Qwen3-VL-8B-Thinking
Total Tasks
117
Successful
12
Success Rate
10.3%
Action Space
mobile_use
Max Steps
50
Avg Steps
32.5
Total Steps
~3.8K… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/MobileWorld-Eval-Results.
