datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
interactive-stage2-30k-v1InteractiveRadiologyTutor_datssetsqualcomm-interactive-cooking-dataset-ego-mistake-corrections
Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark
Description
This dataset contains cooking videos with timestamped instruction and feedback for task guidance.
Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp.
Dataset Details
Release files:
annotations/annotations.json
videos/*.MP4
Release statistics:
Total videos: 40
Total released annotations: 1,597
Text type counts in… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.Chinese_interactive_novels_3k
中文互动小说结构化语料
This dataset contains uncleaned (!) 3534 structured Chinese interactive novels (中文互动小说), accounting for around 0.25B (gpt-3.5) tokens in total.
All contents are parsed from certain online sources.
Usage
This dataset can be potentially used for LLM training. But be aware that you'd better clean the data yourself to remove undesired low-quality contents.
Each novel is a dict structured as follows:
class Novel:
book_title: str
book_author: str… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/Chinese_interactive_novels_3k.SWE-Bench-Pro-interactive-issue
SWE-Bench Pro / Interactive / Issue
Companion data release for the anonymous paper "Opinion: Coding-Agent Benchmarks Should Match Their Users' Task Flows" (SWE-TaskFlow). SWE-TaskFlow transforms an issue-derived benchmark into replayable multi-turn trajectories while preserving the original tasks and tests. This dataset contains the issue-solving prompt sequences (no QA turns) over the 701 SWE-Bench Pro tasks that admit a three-part decomposition (out of the 731 public tasks).… See the full description on the dataset page: https://huggingface.co/datasets/Anonym01048/SWE-Bench-Pro-interactive-issue.SWE-Bench-Pro-interactive-issue-qa
SWE-Bench Pro / Interactive / Issue+QA
Companion data release for the anonymous paper "Opinion: Coding-Agent Benchmarks Should Match Their Users' Task Flows" (SWE-TaskFlow). This dataset contains the QA-augmented trajectories: verifiable questions about repository behavior inserted before, between, or after the split issue turns. Every question ships with its hidden reference answer, an executable golden proof script, and the creation-time proof-execution report.
The plain… See the full description on the dataset page: https://huggingface.co/datasets/Anonym01048/SWE-Bench-Pro-interactive-issue-qa.interactive-sweDataset Summary
Interactive SWE-bench is a dataset developed by CMU Language Technologies Institute (LTI) that contains 500 verified samples from the SWE-bench test set. This dataset is an enhanced version of the original SWE-bench dataset, featuring both the original detailed GitHub issues and their simplified, focused versions.
The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Each entry includes both the original detailed issue description and a… See the full description on the dataset page: https://huggingface.co/datasets/cmu-lti/interactive-swe.qualcomm-interactive-cooking-dataset
Qualcomm Interactive Cooking Dataset
Description
The Qualcomm Interactive Cooking Dataset is designed to evaluate the ability of multi-modal LLMs to provide step-by-step instructions, focusing on the cooking domain.
Dataset Details
The Qualcomm Interactive Cooking Dataset includes step-by-step instructions and feedback pairs. The videos are from the CaptainCook4D dataset - licensed under Apache 2.0.
Dataset Collection Process
The text annotations and… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset.repo_universe_interactive
Repo Universe Interactive Graph
Small static-viewer payload for the Research Library repository universe.
This dataset is intentionally separate from the full repo graph export. It contains browser-friendly graph assets for the static WebGL/WASM app. Parquet is the preferred payload format; JSON/JSONL is retained as a compatibility fallback:
parquet/interactive/lod_10000.parquet
parquet/interactive/lod_50000.parquet
parquet/interactive/lod_200000.parquet… See the full description on the dataset page: https://huggingface.co/datasets/PeytonT/repo_universe_interactive.paper_universe_interactive
Paper Universe Interactive Graph
Small static-viewer payload for the Research Library paper universe.
This dataset is intentionally separate from PeytonT/paper_graph. It contains browser-friendly interactive assets needed by the static WebGL/WASM app. Parquet is the preferred payload format; JSON is retained as a compatibility fallback:
parquet/interactive/papers_50000.parquet
parquet/interactive/papers_200000.parquet
parquet/interactive/papers_all.parquet… See the full description on the dataset page: https://huggingface.co/datasets/PeytonT/paper_universe_interactive.Xianxia-Cultivation-System-Interactive-Sandbox-System-Exampleunreal-interactive-devinteractive-sports-nhl
interactive-sports: NHL research database
The database the agents in interactive_sports query. One SQLite file,
2.46 GB, covering 2010-10-07 to
2026-06-14.
Agents never see this file directly. The harness builds cutoff-scoped, tokenised
VIEWS over it — every view is filtered to game_date <= as_of_date, and every
player and team is replaced by an opaque P#### / T#### token that is minted
fresh per run. The raw tables below carry real identities; the agent surface does
not.… See the full description on the dataset page: https://huggingface.co/datasets/gilberty005/interactive-sports-nhl.qualcomm-interactive-cooking-dataset-counterfactual-mistakes
Qualcomm Interactive Cooking Dataset: Ego Counterfactual Mistakes
Description
This synthetic dataset contains mistake-intervention annotations for interactive cooking guidance. Each row contains video segment with instruction/feedback text pairs and their timestamps.
Dataset Details
Files:
annotations.json
Release statistics:
Total rows: 25,087
Unique videos (dataset + video_id): 1,110
Rows by source dataset:
CaptainCook4D: 4,969
Ego4D: 13,847
Ego-Exo4D: 6… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-counterfactual-mistakes.Interactive-PEDES-v1
📁 Interactive-PEDES Dataset
This folder contains data for the Interactive-PEDES benchmark used in the paper: LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification" (ICML 2025)
📦 Structure
.
├── CUHK-PEDES
│ ├── caption_all.json
│ ├── imgs/
│ │ ├── cam_a/
│ │ ├── cam_b/
│ │ ├── CUHK01/
│ │ ├── CUHK03/
│ │ ├── Market/
│ │ ├── test_query/
│ │ └── train_query/
├── ICFG-PEDES
│ ├── ICFG-PEDES.json
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/XLearning-SCU/Interactive-PEDES-v1.interactive-fiction-knowledgeLLMBind-GPT-Interactive-Datahumanoid-interactive-session-logs
Humanoid Interactive Session Logs
High-quality session-level interaction logs for humanoid AI agents.
humanoid-interactive-control-panel-logs
Interactive Control Panel Logs
High-fidelity logs representing professional humanoid control interfaces.
Reasoning-While-Asking-SFT-Datasetstrl-main-ec-playwright_interactive-gc-claude_client_strl_dplm-mc-claude_agent_sonnet-rb-r0DeepSeek-R1-Distill-Data-5k3d-interactive-assestsInteractive_Benchmarks
Interactive Benchmarks (IB)
Interactive Benchmarks for Evaluating Interactive Reasoning and Agent Capabilities
Usage
from datasets import load_dataset
dataset = load_dataset("interactivebench/Interactive_Benchmarks")
Citation
If you use the Interactive_Benchmarks dataset in your research, please consider citing it as follows:
@misc{interactivebench,
title={Interactive Benchmarks for Evaluating Interactive Reasoning and Agent… See the full description on the dataset page: https://huggingface.co/datasets/interactivebench/Interactive_Benchmarks.strl-main-ec-playwright_interactive-gc-claude_client_strl_summary-mc-claude_agent_sonnet-r0collabllm-20q-interactiveChatMed_Consult_Dataset_Interactivehumanoid-interactive-dialogue-states
Humanoid Interactive Dialogue States
A state-based interactive dialogue dataset for humanoid robots.
Interactive_Benchmarksinteractive-map
