datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WideSearch
WideSearch: Benchmarking Agentic Broad Info-Seeking
Dataset Summary
WideSearch is a benchmark designed to evaluate the capabilities of Large Language Model (LLM) driven agents in broad information-seeking tasks. Unlike existing benchmarks that focus on finding a single, hard-to-find fact, WideSearch assesses an agent's ability to handle tasks that require gathering a large amount of scattered, yet easy-to-find, information.
The challenge in these tasks lies not in… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/WideSearch.WideDepth
WideDepth: Millimeter-Accurate Benchmark for Fisheye Depth Estimation
WideDepth is the first indoor dataset for fisheye depth estimation, featuring 101 scenes containing 5K high-resolution stereo pairs labeled with millimeter-level ground truth depth and disparity.
Paper: WideDepth: Millimeter-Accurate Benchmark for Fisheye Depth Estimation
Project Page: https://ilyaind.github.io/WideDepth
The dataset includes paired pinhole and fisheye samples across varying fields of view… See the full description on the dataset page: https://huggingface.co/datasets/IlyaInd/WideDepth.wider_faceWIDER FACE dataset is a face detection benchmark dataset, of which images are
selected from the publicly available WIDER dataset. We choose 32,203 images and
label 393,703 faces with a high degree of variability in scale, pose and
occlusion as depicted in the sample images. WIDER FACE dataset is organized
based on 61 event classes. For each event class, we randomly select 40%/10%/50%
data as training, validation and testing sets. We adopt the same evaluation
metric employed in the PASCAL VOC dataset. Similar to MALF and Caltech datasets,
we do not release bounding box ground truth for the test images. Users are
required to submit final prediction files, which we shall proceed to evaluate.chocopan-t3-reverse-oracle-hdf5-wide-v1
chocopan-t3-reverse-oracle-hdf5-wide-v1
Raw HDF5 output of a scripted oracle for reverse manipulation tasks in simulation -- take an object out of a container or off a plate and put it back on the table -- for the batch generated with widened object initial placements: 7,200 attempts over 45 tasks, failures included.
This is the raw, unfiltered output of the generator, in LIBERO's create_dataset.py HDF5
layout. It is published because it is bulky to regenerate, not because it is… See the full description on the dataset page: https://huggingface.co/datasets/chocopan/chocopan-t3-reverse-oracle-hdf5-wide-v1.Wider_FaceSegLitewidedepth
Dataset Card for WideDepth in FiftyOne
FiftyOne dataset for WideDepth — an indoor fisheye depth-estimation benchmark (ICRA 2026) with millimeter-accurate ground-truth depth and disparity rendered from high-resolution LiDAR scans.
We use one fixed camera configuration from the full WideDepth benchmark — 195° FOV, 300 mm focal length, CENTER stereo position — across all 101 indoor scenes.
The full Hub release has many combinations (4 FOVs × 5 focal lengths × 3 positions, plus… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/widedepth.human-proteome-wide-msa
Human Proteome ColabFold MSAs
This dataset contains ColabFold/MMseqs2 multiple sequence alignments and AlphaFold 3 JSON inputs for the human proteome query set.
Contents
a3m/shard-*/: one A3M file per protein accession, sharded to stay below repository directory file limits.
af3_json/shard-*/: one AlphaFold 3 JSON file per protein accession, sharded to stay below repository directory file limits.
manifest.tsv: tab-separated index with accession, metadata… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/human-proteome-wide-msa.WideDepth-train
WideDepth train
Dataset Summary
WideDepth train is an outdoor multi-view fisheye training dataset for depth/disparity learning and domain adaptation.
The dataset accompanies the paper WideDepth: Millimeter-Accurate Benchmark for Fisheye Depth Estimation, accepted to ICRA 2026.
It provides synchronized frames across:
left fisheye RGB
right fisheye RGB
upper equirect RGB
lower virtual equirect RGB
equirect disparity
fisheye depth
This subset was captured outdoors… See the full description on the dataset page: https://huggingface.co/datasets/IlyaInd/WideDepth-train.WideSeek-R1-train-data
Training Dataset
🌐 Project Page | 📄 Paper | 📖 Doc | 💻 Code | 📦 Dataset | 🤗 Models
We provide three datasets:
width_20k.jsonl
depth_20k.jsonl
hybrid_20k.jsonl
Dataset Sources and Relationships
width_20k.jsonl is constructed by us and is tailored for WideSearch-style tasks.
depth_20k.jsonl is sourced from ASearcher's training data.
hybrid_20k.jsonl is a balanced mixture of the two and serves as the core training set for our main training experience.
All three… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/WideSeek-R1-train-data.WideSeek-R1-test-data
Testing Dataset
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
mmu_hsc_pdr3_wide_21
mmu_hsc_pdr3_wide_21 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_hsc_pdr3_wide_21.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_hsc_pdr3_wide_21.pick_place_lego_wider_range_richardThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 20918,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seeingrain/pick_place_lego_wider_range_richard.pick_place_lego_wider_range_dongThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 57,
"total_frames": 22565,
"total_tasks": 1,
"total_videos": 114,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:57"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seeingrain/pick_place_lego_wider_range_dong.football_matcheswider_face_yolo
Wider face yolo
wider_face.zip
- train
- images
- ....jpg
- labels
- ....txt
- valid
- images
- ....jpg
- labels
- ....txt
lewam_eval_pnpt_diffusion_wideThis dataset was created using LeRobot.
Dataset Description
Real-robot evaluation rollouts recorded on a Rebot B601 7-DoF arm, from the LeWAM
project. Every episode here is a policy rollout on the physical robot — not a
teleoperated demonstration — scored by a human operator immediately after it ran.
Task: pick-and-place (b601_pnpt): the arm picks a roll of tape and places it on a target. Trained on ehalicki/b601_pusht_pick_and_place.
Policy: a LeRobot Diffusion Policy… See the full description on the dataset page: https://huggingface.co/datasets/ehalicki/lewam_eval_pnpt_diffusion_wide.widerperson
FineWiderPerson — WiderPerson in the unified detection format
Source: official WiderPerson.zip via Google Drive.
Converted by the finedet project into a unified, AutoTrain-compatible layout:
image / width / height / objects{bbox, category} with COCO-format
[x, y, w, h] boxes in absolute pixels. Boxes are clipped to the image and
empty boxes dropped; category ids are densified per the category tables
below.
Box format
objects.bbox follows the COCO convention: [x, y… See the full description on the dataset page: https://huggingface.co/datasets/finedet/widerperson.lewam_eval_pusht_diffusion_wideThis dataset was created using LeRobot.
Dataset Description
Real-robot evaluation rollouts recorded on a Rebot B601 7-DoF arm, from the LeWAM
project. Every episode here is a policy rollout on the physical robot — not a
teleoperated demonstration — scored by a human operator immediately after it ran.
Task: non-prehensile PushT (b601_pusht): the arm pushes a T-shaped block into a target pose without grasping it. Shares only its name with the simulated PushT. Trained on… See the full description on the dataset page: https://huggingface.co/datasets/ehalicki/lewam_eval_pusht_diffusion_wide.NUS-WIDE
NUS-WIDE
Dataset Overview
The NUS-WIDE dataset is a large-scale multi-label image dataset that can be widely used for image classification and multi-label learning tasks. It contains 269,648 images collected from Flickr, annotated with 81 concept labels .
Training set:161,789 images
Test set:107,859 images
Dataset Directory Structure
NUS-WIDE
├── Groundtruth
│ ├── TrainTestLabels
│ │ ├── Labels_zebra_Train.txt
│ │ ├── Labels_zebra_Test.txt
│ │… See the full description on the dataset page: https://huggingface.co/datasets/Lxyhaha/NUS-WIDE.modified_wider_facewider_face_backgroundKo-widesearch
Ko-WideSearch
A Korean breadth-search benchmark: each task asks a web agent to exhaustively
enumerate a closed set and fill every attribute cell of a table (e.g. "list every
award category at the 59th Grand Bell Awards and give each winner"). 228 tasks across
three difficulty tiers.
[!IMPORTANT]
The question and answer fields are encrypted. To keep this a fair,
leakage-aware test of web agents, the gold is not published as plain text — it is
canary-XOR obfuscated (same scheme… See the full description on the dataset page: https://huggingface.co/datasets/Minbyul/Ko-widesearch.Mistral_Trivia-QA_Dataset
Mistral Trivia QA Dataset
The Mistral Trivia QA Dataset is a collection of trivia questions and answers designed to evaluate and train question-answering models. It covers a wide range of topics and is particularly useful for assessing a model's ability to handle general knowledge and reasoning tasks.The documents are derived from WikiText-2, providing diverse and well-structured textual content suitable for extractive QA generation.
Model outputs for this dataset were generated… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Mistral_Trivia-QA_Dataset.exp2_sim_wide_blueThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 26,
"total_frames": 5017,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:26"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/ameimei/exp2_sim_wide_blue.pick_place_lego_wider_range_dangThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 10759,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seeingrain/pick_place_lego_wider_range_dang.WideSeek-R1-test-data
Testing Dataset
🌐 Project Page | 📄 Paper | 📖 Doc | 💻 Code | 📦 Dataset | 🤗 Models
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearchdataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
Acknowledgement
Thanks to WideSearch for providing a… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/WideSeek-R1-test-data.WiderPerson-Camera
WiderPerson-Camera
Per-image camera parameter annotations for the WiderPerson dataset
(a pedestrian-detection benchmark in the wild; 13,382 images), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/WiderPerson-Camera.nus_widewide-robot-object-inside-container-v0
wide-robot — object_inside_container real-camera capture (v0)
Raw multi-view video for the Phase 3A real-camera ingestion pilot of
wide-robot: a marker-based, source-independent
verifier judges object_inside_container episodes (PASS / FAIL / UNCERTAIN) from real camera
evidence. This repo holds the raw capture media that is gitignored in the code repo; the derived
JSON (calibrations, tracks, rollouts, verdicts) and the full results live in the git repo under… See the full description on the dataset page: https://huggingface.co/datasets/Alexreysa/wide-robot-object-inside-container-v0.WideSeek-R1-SFT-data
WideSeek-R1 SFT Data
This dataset contains agent-level, multi-turn supervised fine-tuning trajectories for both width-only and depth-only tasks in WideSeek-R1.
Construction
The trajectories were generated by Qwen3-235B-A22B using the WideSeek-R1 multi-agent workflow with offline retrieval tools. Width and depth trajectories are balanced at the question level.
For each question-level trajectory, we retain one main-agent session and up to three subagent sessions… See the full description on the dataset page: https://huggingface.co/datasets/WideSeek-R1/WideSeek-R1-SFT-data.
