datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vidore_v3_computer_scienceViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.computer-use-data-psai
Computer Use Dataset - PSAI
A large-scale, multimodal dataset of human-computer interactions for training and evaluating AI agents.
🔗 Access Dataset: https://huggingface.co/datasets/anaisleila/computer-use-data-psai
📊 Dataset Overview
This dataset contains 3,167 completed tasks of human-computer interactions captured with video, screenshots, DOM snapshots, and detailed interaction events. Created by Paradigm Shift AI for advancing computer use AI agent research.… See the full description on the dataset page: https://huggingface.co/datasets/anaisleila/computer-use-data-psai.computer-use
Computer Use Trajectories
Successful computer-use agent trajectories collected on OSWorld tasks.
Dataset Details
Rows: 160 (one per task trajectory)
Steps: 1,378 total across all trajectories (avg ~8.6 steps/task)
Agent: Gemini 3 Flash Preview with linearized accessibility-tree grounding
Score filter: Only trajectories with score = 1.0 (fully successful)
Domains
Domain
Tasks
Description
chrome
21
Web browsing tasks in Google Chrome
gimp
15
Image… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/computer-use.vidore_v3_computer_science_mteb_format
Vidore3ComputerScienceRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_computer_science
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.tess-agentnet
TESS AgentNet Dataset
Computer use trajectories for training Vision-Language-Action models.
Features
image: Screenshot (PIL Image)
instruction: Task description
action_type: 0=MOUSE, 1=KEYBOARD
mouse_x, mouse_y: Normalized coordinates [0,1]
click_type: 0-8 (NO_CLICK, LEFT_CLICK, etc.)
keyboard_text: Text with special tokens
os_type: ubuntu, windows_macos
episode_id, step_idx: Episode structure
Click Types
Index
Type
Description
0
NO_CLICK… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/tess-agentnet.Handwritten-Computer-Science-Notes-Dataset
English Handwritten Computer Science Notes Dataset
This dataset contains high-resolution images of handwritten computer science notes written in English. It includes algorithm explanations, code snippets, flowcharts, theoretical content, and annotations. The dataset is designed to support AI research in handwriting recognition, OCR, and document understanding specifically for computer science education.
Contact
For queries or collaborations related to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Computer-Science-Notes-Dataset.computer-agent-arena
Computer Agent Arena: Evaluating Computer-Use Agents via Crowdsourcing from Real Users
Dataset Description
Computer Agent Arena is a comprehensive evaluation platform for multi-modal AI agents, particularly focusing on computer use and GUI interaction tasks. This dataset contains real interaction trajectories from various state-of-the-art AI agents performing complex computer tasks in controlled environments.
The dataset includes:
4,641 agent trajectories across diverse… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/computer-agent-arena.GroceryInContextcomputerAssignmenthtml-samplecomputer-agent-trajectoriescomputer-vision
Hibou Computer Vision Dataset
The Hibou Project is a drone recognition and localization system.
It is designed to detect and localize drones in real-time, using a combination of audio and video.
Official code repo: Hibou Project
This dataset is designed to train YOLO-based models for drone detection.
Object Classes
ID
Class Name
Ratio
0
Drone
74.17%
1
Other
25.83%
Dataset Description
Column
Description
image
Image from the… See the full description on the dataset page: https://huggingface.co/datasets/Hibou-Foundation/computer-vision.bmw-computer-vision-datasetcomputer-use
Computer Use Trajectories
Successful computer-use agent trajectories collected on OSWorld tasks.
Dataset Details
Rows: 160 (one per task trajectory)
Steps: 1,378 total across all trajectories (avg ~8.6 steps/task)
Agent: Gemini 3 Flash Preview with linearized accessibility-tree grounding
Score filter: Only trajectories with score = 1.0 (fully successful)
Domains
Domain
Tasks
Description
chrome
21
Web browsing tasks in Google Chrome
gimp
15
Image… See the full description on the dataset page: https://huggingface.co/datasets/REXX-NEW/computer-use.open-computer-using-agent
Dataset Description
This dataset is associated with the ongoing 'Nous Project' - creating a computer using agent based on open source models. The data was collected using Anthropic Claude 3.5 Sonnet Latest to record conversational state along with computer state data including:
Cursor position
Active windows
Computer display dimensions
System state information
User interactions
While the original interview problem covered only button clicks, this dataset is more comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/palkarpratik84/open-computer-using-agent.Datasets_computer_visionvidore_v3_computer_science_embeddingNOTE
ViDoRe V3: Computer Science dataset ColQwen2 Embeddings
This dataset contains pre-computed embeddings for the ViDoRe V3 : Computer Science dataset using the ColQwen2 model.
ViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/vidore_v3_computer_science_embedding.je-veux-un-pack-de-computer-vision-pour-detecter-les-objets
Je veux un pack de computer vision pour detecter les objets,…
Je veux un pack de computer vision pour detecter les objets, dans des salles de bains pour le segment home avec des depth map
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/je-veux-un-pack-de-computer-vision-pour-detecter-les-objets.hello-je-veux-un-pack-de-computer-vision-pour-detecter-les
Hello, Je veux un pack de computer vision pour detecter les …
Hello, Je veux un pack de computer vision pour detecter les objets, dans des salles de bains pour le segment home avec des depth map, 1 camera, 1 render
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/hello-je-veux-un-pack-de-computer-vision-pour-detecter-les.quickdraw-circles
Quick, Draw! Circles - Trajectory Dataset
Dataset for training trajectory prediction models, specifically designed for the Qwen-DiT-Draw project.
Dataset Description
This dataset contains chunked trajectory data from the Quick, Draw! circle category, formatted for training diffusion-based trajectory prediction models.
Key Features
Variable-length trajectories with stop signals (GR00T-style)
16-point chunks with (x, y, state) format
Loss masking for handling… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/quickdraw-circles.quickdraw-circles-delta
Quick, Draw! Circles - Trajectory Dataset
Dataset for training trajectory prediction models, specifically designed for the Qwen-DiT-Draw project.
Dataset Description
This dataset contains chunked trajectory data from the Quick, Draw! circle category, formatted for training diffusion-based trajectory prediction models.
Key Features
Variable-length trajectories with stop signals (GR00T-style)
16-point chunks with (x, y, state) format
Loss masking for handling… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/quickdraw-circles-delta.Viet-ComputerScience-VQA
Dataset Overview
This dataset is was created from 6899 Vietnamese 🇻🇳 Computer Science books. Each image has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 40,000 detailed descriptions, and query-based questions and answers generated by the Gemini 1.5 Flash model, currently Google's leading model on the WildVision Arena Leaderboard. This results in a richly annotated dataset, ideal for… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-ComputerScience-VQA.vidore_v3_computer_science_english_open-ended_Chartcomputer-thoughtshello1-je-veux-un-pack-de-computer-vision-pour-detecter-les
Hello1, Je veux un pack de computer vision pour detecter les…
Hello1, Je veux un pack de computer vision pour detecter les objets, dans des salles de bains pour le segment home avec des depth map, 1 camera, 1 render
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/hello1-je-veux-un-pack-de-computer-vision-pour-detecter-les.computer_mouse_labeledvidore_v3_computer_science_english_instructionvidore_v3_computer_science_english_boolean_Othercomputer_mousevidore_v3_computer_science_english_compare-contrast_Other
