datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
computer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/computer-use-large.vidore_v3_computer_scienceViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.atari-vla-stage1-5hz
TESS-Atari Stage 1 (5Hz)
Human gameplay demonstrations from Atari games, formatted for Vision-Language-Action (VLA) model training.
Overview
Metric
Value
Source
Atari-HEAD
Games
11 (overlapping with DIAMOND benchmark)
Samples
~4M
Action Rate
5 Hz (1 action per observation)
Format
Lumine-style action tokens
Games Included
Alien, Asterix, BankHeist, Breakout, DemonAttack, Freeway, Frostbite, Hero, MsPacman, RoadRunner, Seaquest… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/atari-vla-stage1-5hz.AIRBOT_MMK2_close_the_computer
AIRBOT_MMK2_close_the_computer
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_close_the_computer.AIRBOT_MMK2_storage_computer_box
AIRBOT_MMK2_storage_computer_box
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_computer_box.atari-vla-stage1-15hz
TESS-Atari Stage 1 (15Hz)
Human gameplay demonstrations from Atari games with action chunking, formatted for Vision-Language-Action (VLA) model training.
Overview
Metric
Value
Source
Atari-HEAD
Games
11 (overlapping with DIAMOND benchmark)
Samples
~1.3M
Observation Rate
5 Hz
Action Rate
15 Hz (3 actions per observation)
Format
Lumine-style action tokens
Why Action Chunking?
VLA models run at ~5 Hz inference speed, but Atari runs at… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/atari-vla-stage1-15hz.computer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/computer-use-large.csgo-vla-stage1-5hz
CS:GO VLA Stage 1 Dataset (5Hz Chunked)
Vision-Language-Action dataset for Counter-Strike: Global Offensive with action chunking, converted from the TeaPearce CS:GO dataset.
Overview
Frame rate: 5Hz (every 3rd frame)
Action chunking: 3 actions per sample (~200ms coverage)
Total samples: ~1.8M chunks
Split: train / test following Diamond split
Map: Dust2 deathmatch
Action Format
<|action_start|> m1_x m1_y [keys1] ; m2_x m2_y [keys2] ; m3_x m3_y [keys3]… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/csgo-vla-stage1-5hz.tess-atari-5hz-384ovos-wake-word-bench-community-computer
OVOS wake_word bench — community-computer
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-computer.IndustryCorpus2_computer_programming_code
IndustryCorpus2: Programming
This repository contains the IndustryCorpus2: Programming domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year = {2024}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_computer_programming_code.ovos-wake-word-bench-picovoice-computer
OVOS wake_word bench — picovoice-computer
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-computer.pi-computer-use-sessions
Coding agent session traces for thomasmustier/pi-computer-use-sessions
This dataset contains redacted coding agent session traces collected while working on https://github.com/tmustier/pi-computer-use. The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction, secret scanning, visual review where applicable, and LLM review.
Source git repo: https://github.com/tmustier/pi-computer-use
Data… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-computer-use-sessions.tess-atari-15hz-384
TESS-Atari Stage 1 - Preprocessed (15Hz, 384x384)
Training-ready version of the 15Hz dataset with images pre-resized to 384x384 (SmolVLM native resolution).
Overview
Metric
Value
Source
TESS-Computer/atari-vla-stage1-15hz
Samples
1,340,293
Image Size
384x384 (pre-resized)
Action Rate
15 Hz (3 actions per observation)
Format
Lumine-style action tokens
Why Preprocessed?
Training VLMs requires resizing images to the model's native… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/tess-atari-15hz-384.vn-provinces-household-computer-rate
Vietnam household computer ownership rate
Vietnam household computer ownership rate. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (378 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
national (6 rows)
data/national.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-household-computer-rate.ovos-wake-word-bench-synthetic-wakewords-hey_computer
OVOS wake_word bench — synthetic-wakewords-hey_computer
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_computer.computer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS… See the full description on the dataset page: https://huggingface.co/datasets/Whiteglove44/computer-use-large.pine-computer-saasbench-results
Pine Computer — SaaSBench Evaluation Results
Pine Computer · Pine AI
Task-level results and evidence hashes for Pine Computer on SaaS-Bench. Raw traces and screenshots remain private. Separate releases retain their own task versions, selection rules and benchmark suite; they are not pooled.
Latest: SaaS-Bench 1.1 · September 18–19, 2026
78.33% checks passed · 29/106 tasks fully resolved (27.36%) · 106/106 tasks with valid results.
The checkpoint numerator is 1,077… See the full description on the dataset page: https://huggingface.co/datasets/pine-ai/pine-computer-saasbench-results.csgo-vla-stage1-16hz
CS:GO VLA Stage 1 Dataset (16Hz)
Vision-Language-Action dataset for Counter-Strike: Global Offensive, converted from the TeaPearce CS:GO dataset.
Overview
Frame rate: 16Hz (native, 1 action per frame)
Total samples: ~5.5M frames
Split: train (5M) / test (500K) following Diamond split
Map: Dust2 deathmatch
Action Format
<|action_start|> mouse_x mouse_y [keys] <|action_end|>
Examples:
<|action_start|> 0 0 <|action_end|> # idle… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/csgo-vla-stage1-16hz.vidore_v3_computer_science_embeddingNOTE
ViDoRe V3: Computer Science dataset ColQwen2 Embeddings
This dataset contains pre-computed embeddings for the ViDoRe V3 : Computer Science dataset using the ColQwen2 model.
ViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/vidore_v3_computer_science_embedding.computer-use-large-actions
computer-use-large-actions
9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video).
Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0.
Split by software
category
examples
vscode
2,500
autocad
2,500
blender
1,000
excel
1,000
photoshop
1,000
salesforce
1,000
VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.je-veux-un-pack-de-computer-vision-pour-detecter-les-objets
Je veux un pack de computer vision pour detecter les objets,…
Je veux un pack de computer vision pour detecter les objets, dans des salles de bains pour le segment home avec des depth map
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/je-veux-un-pack-de-computer-vision-pour-detecter-les-objets.hello-je-veux-un-pack-de-computer-vision-pour-detecter-les
Hello, Je veux un pack de computer vision pour detecter les …
Hello, Je veux un pack de computer vision pour detecter les objets, dans des salles de bains pour le segment home avec des depth map, 1 camera, 1 render
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/hello-je-veux-un-pack-de-computer-vision-pour-detecter-les.LeroyDyer__LCARS_AI_StarTrek_Computer-details
Dataset Card for Evaluation run of LeroyDyer/LCARS_AI_StarTrek_Computer
Dataset automatically created during the evaluation run of model LeroyDyer/LCARS_AI_StarTrek_Computer
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__LCARS_AI_StarTrek_Computer-details.epfl-computer-science-mcqacomputer-use-agent-traces-250k
Computer-Use Traces Dataset
250,000 real-world computer-use traces with screen states, browser sessions, UI actions, task instructions, and completion outcomes for training AI agents.
This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get a real sample package.
Overview
The… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/computer-use-agent-traces-250k.carla-simlingo-raw
SimLingo CARLA Dataset (Raw, 4Hz)
Raw driving data from CARLA simulator. No transformations or derived fields - all original measurements preserved as-is.
Dataset Summary
Source: SimLingo (CVPR 2025)
Scale: 228,757 frames (23 shards)
Frame Rate: 4 FPS
Resolution: 1024x512 RGB
Routes: Complete driving episodes (routes never split across shards)
Column Schema
Core Fields
Column
Type
Description
route_id
string
Route identifier
frame_idx… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/carla-simlingo-raw.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 322,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ComputerNoob/so100_test.africa-rwanda-table-4-10-computer-literacy-rate-of-the-population-by-age-7b7860dc
Table 4.10: Computer literacy rate of the population by age groups according to area of residence, province, sex and consumption quintile | Africa (Rwanda Data Sharing Platform - NISR)
17 rows - 1 Africa country/area - 2023-10-16-2024-10-15 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 17 rows from Rwanda Data Sharing Platform - NISR, covering Table 4.10: Computer literacy rate of the population by age groups according to… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-rwanda-table-4-10-computer-literacy-rate-of-the-population-by-age-7b7860dc.tess-atari-asterix-15hz-384
TESS-Atari: Asterix (15Hz, 384x384)
Single-game preprocessed dataset for VLA training.
Overview
Metric
Value
Game
Asterix
Samples
41,646
Image Size
384x384
Action Rate
15 Hz (3 actions per observation)
Format
Lumine-style action tokens
Filters Applied
score > 0 - Active gameplay only (no menus/idle)
No pure NOOP - Player actually taking actions
Action Format
<|action_start|> RIGHT ; UP ; UPRIGHT <|action_end|>
Usage… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/tess-atari-asterix-15hz-384.
