datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
computer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/computer-use-large.vidore_v3_computer_scienceViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.Computer-Science
מאגר הנתונים CS26 HIT — מדעי המחשב
מאגר זה משמש לאחסון מרכזי של נתוני לימוד ומשאבים אקדמיים עבור סטודנטים למדעי המחשב במכון הטכנולוגי חולון. המידע המצוי כאן מונגש בצורה נוחה באמצעות פורטל גישה נפרד המאפשר ניווט ויזואלי וחיפוש יעיל בתוך התיקיות השונות.
קישורים וגישה
ניתן להשתמש בפורטל בכתובת https://cs26-cs26-portal.hf.space/
תנאי שימוש וזכויות יוצרים
כל חומרי הלימוד והתכנים המופיעים במאגר זה פתוחים וחופשיים לשימוש לצורכי למידה בלבד. ניתן לקחת את… See the full description on the dataset page: https://huggingface.co/datasets/CS26/Computer-Science.computer-use-data-psai
Computer Use Dataset - PSAI
A large-scale, multimodal dataset of human-computer interactions for training and evaluating AI agents.
🔗 Access Dataset: https://huggingface.co/datasets/anaisleila/computer-use-data-psai
📊 Dataset Overview
This dataset contains 3,167 completed tasks of human-computer interactions captured with video, screenshots, DOM snapshots, and detailed interaction events. Created by Paradigm Shift AI for advancing computer use AI agent research.… See the full description on the dataset page: https://huggingface.co/datasets/anaisleila/computer-use-data-psai.minecraft-vla-stage1
Minecraft VLA Stage 1: Action Pretraining Data
Vision-Language-Action training data for Minecraft, processed from OpenAI's VPT contractor dataset.
Dataset Description
This dataset contains frame-action pairs from Minecraft gameplay, designed for training VLA models following the Lumine methodology.
Source
Original: OpenAI VPT Contractor Data (7.x subset)
Videos: 17,886 videos (330 hours of early-game gameplay)
Task: "Play Minecraft" with focus on first 30… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/minecraft-vla-stage1.Spatial_Four_Bar_Mechanism_Closed
Dataset Overview
This dataset is generated for the path synthesis of 1-DOF(degree-of-freedom) closed-loop spatial four-bar linkage mechanisms. Specifically, it contains closed paths only with their corresponding mechanisms. All mechanisms are actuated by revolute (R) joints, and include other various joints, including prismatic (P), cylindrical (C), universal (U), and spherical (S) joints. The dataset covers all possible 1-DOF(degree-of-freedom) closed-loop spatial four-bar… See the full description on the dataset page: https://huggingface.co/datasets/ComputerAidedDesignInnovation/Spatial_Four_Bar_Mechanism_Closed.synth-computer-use-cachecomputer-use
Computer Use Trajectories
Successful computer-use agent trajectories collected on OSWorld tasks.
Dataset Details
Rows: 160 (one per task trajectory)
Steps: 1,378 total across all trajectories (avg ~8.6 steps/task)
Agent: Gemini 3 Flash Preview with linearized accessibility-tree grounding
Score filter: Only trajectories with score = 1.0 (fully successful)
Domains
Domain
Tasks
Description
chrome
21
Web browsing tasks in Google Chrome
gimp
15
Image… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/computer-use.minecraft-vla-stage2
Minecraft VLA Stage 2: Instruction-Following Data
Stage 2 of the TESS-Minecraft Vision-Language-Action training pipeline.
Overview
This dataset adds task instructions to the Stage 1 visuomotor data, enabling instruction-following training.
Data Format
Field
Type
Description
id
string
Unique sample ID
video_id
string
Source video name
frame_idx
int
Frame index within video
instruction
string
Task instruction (empty for continuation frames)… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/minecraft-vla-stage2.vidore_v3_computer_science_mteb_format
Vidore3ComputerScienceRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_computer_science
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.Spatial_Four_Bar_Mechanism
Dataset Overview
This dataset is generated for the path synthesis of 1-DOF(degree-of-freedom) closed-loop spatial four-bar linkage mechanisms. Specifically, it contains both open and closed paths together with their corresponding mechanisms. All mechanisms are actuated by revolute (R) joints, and include other various joints, including prismatic (P), cylindrical (C), universal (U), and spherical (S) joints. The dataset covers all possible 1-DOF(degree-of-freedom) closed-loop spatial… See the full description on the dataset page: https://huggingface.co/datasets/ComputerAidedDesignInnovation/Spatial_Four_Bar_Mechanism.Computer-Science-Conversational-Dataset-IndicComputer-Science-Parallel-Dataset-Indiccomputer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS Code
304… See the full description on the dataset page: https://huggingface.co/datasets/adrianmele/computer-use-large.atari-vla-stage1-5hz
TESS-Atari Stage 1 (5Hz)
Human gameplay demonstrations from Atari games, formatted for Vision-Language-Action (VLA) model training.
Overview
Metric
Value
Source
Atari-HEAD
Games
11 (overlapping with DIAMOND benchmark)
Samples
~4M
Action Rate
5 Hz (1 action per observation)
Format
Lumine-style action tokens
Games Included
Alien, Asterix, BankHeist, Breakout, DemonAttack, Freeway, Frostbite, Hero, MsPacman, RoadRunner, Seaquest… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/atari-vla-stage1-5hz.tess-agentnet
TESS AgentNet Dataset
Computer use trajectories for training Vision-Language-Action models.
Features
image: Screenshot (PIL Image)
instruction: Task description
action_type: 0=MOUSE, 1=KEYBOARD
mouse_x, mouse_y: Normalized coordinates [0,1]
click_type: 0-8 (NO_CLICK, LEFT_CLICK, etc.)
keyboard_text: Text with special tokens
os_type: ubuntu, windows_macos
episode_id, step_idx: Episode structure
Click Types
Index
Type
Description
0
NO_CLICK… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/tess-agentnet.Handwritten-Computer-Science-Notes-Dataset
English Handwritten Computer Science Notes Dataset
This dataset contains high-resolution images of handwritten computer science notes written in English. It includes algorithm explanations, code snippets, flowcharts, theoretical content, and annotations. The dataset is designed to support AI research in handwriting recognition, OCR, and document understanding specifically for computer science education.
Contact
For queries or collaborations related to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Computer-Science-Notes-Dataset.synthetic-wakeword-hey_computer
synthetic-wakeword-hey_computer
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey computer".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.AIRBOT_MMK2_close_the_computer
AIRBOT_MMK2_close_the_computer
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_close_the_computer.hey-computer-speech-commands
Hey Computer: Speech Command Recognition
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
13,192
Labeled training data
test
3,295
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
audio
Audio
id
string
label
string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/hey-computer-speech-commands.AIRBOT_MMK2_storage_computer_box
AIRBOT_MMK2_storage_computer_box
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_computer_box.IndustryCorpus2_computer_communication
IndustryCorpus2: Computing & Telecommunications
This repository contains the IndustryCorpus2: Computing & Telecommunications domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_computer_communication.computer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS Code
304… See the full description on the dataset page: https://huggingface.co/datasets/nawed/computer-use-large.atari-vla-stage1-15hz
TESS-Atari Stage 1 (15Hz)
Human gameplay demonstrations from Atari games with action chunking, formatted for Vision-Language-Action (VLA) model training.
Overview
Metric
Value
Source
Atari-HEAD
Games
11 (overlapping with DIAMOND benchmark)
Samples
~1.3M
Observation Rate
5 Hz
Action Rate
15 Hz (3 actions per observation)
Format
Lumine-style action tokens
Why Action Chunking?
VLA models run at ~5 Hz inference speed, but Atari runs at… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/atari-vla-stage1-15hz.synthetic-computers-at-scale
Synthetic Computers
Paper: Synthetic Computers at Scale for Long-Horizon Productivity Simulation (arXiv:2604.28181)
A dataset of 98 synthetic computer environments designed for research on
computer-use agents, long-horizon planning, and persona-grounded reasoning.
Each row describes a single fictional user's computer — including the user's
persona, professional context, monthly objectives, collaborators, project
portfolio, filesystem policy, full file listing, and a graph of file… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/synthetic-computers-at-scale.computer-agent-arena
Computer Agent Arena: Evaluating Computer-Use Agents via Crowdsourcing from Real Users
Dataset Description
Computer Agent Arena is a comprehensive evaluation platform for multi-modal AI agents, particularly focusing on computer use and GUI interaction tasks. This dataset contains real interaction trajectories from various state-of-the-art AI agents performing complex computer tasks in controlled environments.
The dataset includes:
4,641 agent trajectories across diverse… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/computer-agent-arena.GroceryInContextcomputerAssignmentUniversity-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus
Title
University-level Mathematics, Physics, Chemistry, Computer Science Reasoning Corpus
Size
200,000+ text+ multimodal university-level problems, each with step-by-step solutions and final answers
Format
Natural language explanations with multimodal samples include images (graphs, diagrams, etc.)
Subject
Mathematics, Physics, Chemistry, Computer Science
Labeling Details
Question ID/Question Stem (Full text/content) /Subject/Question Type… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/University-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus.computer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/computer-use-large.
