datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Amazon-Reviews-2023Amazon Review 2023 is an updated version of the Amazon Review 2018 dataset.
This dataset mainly includes reviews (ratings, text) and item metadata (desc-
riptions, category information, price, brand, and images). Compared to the pre-
vious versions, the 2023 version features larger size, newer reviews (up to Sep
2023), richer and cleaner meta data, and finer-grained timestamps (from day to
milli-second).weblinx-browsergym
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the video tag.
This dataset was specifically created to allow WebLINX to be used inside the BrowserGym and Agentlab ecosystem. Please see the browsergym repository for more information.
[!NOTE]
The version associated with this library is WebLINX… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/weblinx-browsergym.WebLINX-full
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead:
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models
💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.VideoChat3-LV116k
VideoChat3-LV116K
VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments.
The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.agent-reward-bench
AgentRewardBench
💾Code
📄Paper
🌐Website
🤗Dataset
💻Demo
🏆Leaderboard
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor
Loading dataset
You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.MCP-AtlasMCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Leaderboard | MCP Atlas Paper | Github
Dataset Summary
This public release is a subset of 500 sample tasks from the MCP Atlas Benchmark dataset.
MCP Atlas is a large-scale benchmark for evaluating tool-use competency, comprising 36 real MCP servers and 220 tools.
Tasks are designed to assess tool-use competency in realistic, multi-step workflows.
Tasks use natural language prompts that avoid… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/MCP-Atlas.test-mcp-logs(Put queries first as heuristics don't detect when there are no logs)
Motion-o-MCoT-PLM-motion-keyframes
Motion-o-MCoT (PLM + motion keyframes)
Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/.
Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats).
Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.discover-toolsaria_synthetic_envs_mcmc_3dgs_newLicense Notice:This dataset is derived from the Aria Dataset.It follows the Aria Synthetic Environments Dataset License Agreement.See Aria License for details.
sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.TimeLens2-93K
TimeLens2-93K
TimeLens2-93K is a large-scale, long-video temporal grounding dataset. This release contains 23,793 videos and 93,232 text–temporal interval pairs, including 12,091 multi-span pairs. The videos range from short clips to nearly 100 minutes and cover broad web domains such as entertainment, education, sports, news, science and technology, gaming, travel, vehicles, music, and daily life.
TimeLens2-93K offers a rare combination of scale, long-context coverage… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/TimeLens2-93K.mcpmark-trajectory-log
MCPMark Trajectory Logs (mcpmark-v1-0905)
These logs are publicly available for research use. They capture end-to-end trajectories (prompts, tool calls, and outcomes) for MCPMark tasks across multiple MCP services and models.
Contents
Each task run produces a trajectory folder containing three files:
meta.json: Metadata (task id, model, timestamps, status, etc.)
messages.json: Turn-by-turn exchanges including model thoughts/tool calls
execution.log: Key execution-time… See the full description on the dataset page: https://huggingface.co/datasets/Jakumetsu/mcpmark-trajectory-log.stereoset
Dataset Card for StereoSet
Dataset Summary
StereoSet is a dataset that measures stereotype bias in language models. StereoSet consists of 17,000 sentences that measures model preferences across gender, race, religion, and profession.
Supported Tasks and Leaderboards
multiple-choice question answering
Languages
English (en)
Dataset Structure
Data Instances
#intersentence
{'bias_type': 'race', 'context': 'Many people live in… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/stereoset.airbot-fold-cloth-mcapbio-mcp-data
Bio-MCP-Data
A repository containing biological datasets that will be used by BIO-MCP MCP (Model Context Protocol) standard.
About
This repository hosts biological data assets formatted to be compatible with the Model Context Protocol, enabling AI models to efficiently access and process biological information. The data is managed using Git Large File Storage (LFS) to handle large biological datasets.
Purpose
Provide standardized biological datasets for AI… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/bio-mcp-data.security_instruct_mcq_2481McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval.
mC4-Hindi-Cleaned-3.0
Dataset Card for "mC4-Hindi-Cleaned-3.0"
More Information needed
mc4-ja
Dataset Card for "mc4-ja"
More Information needed
guertin-mcro-forensic-corpus
Guertin MCRO Forensic Corpus
The public court record of State of Minnesota v. Matthew David Guertin (Hennepin County, 27-CR-23-1886) and 2,902
other Minnesota dockets, collected from Minnesota Court Records Online (MCRO) — the Minnesota Judicial Branch's public
access to its own case records — together with the database built from them, screen recordings of the collection,
four federal cases, the email record, and a forensic analysis of all of it. Every filing is the PDF the… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus.matterport3d_region_mcmc_3dgsLicense Notice:This dataset is derived from Matterport3D.It follows the Matterport End User License Agreement for Academic Use of Model Data.See Matterport3D License for details.
A3-Synth
A3-Synth
💾 Code
📄 Paper
🌐 Website
🤗 Dataset
🤖 Models
📦 PyPI
Structured Distillation of Web Agent Capabilities Enables Generalization
Xing Han Lù, Siva Reddy
A3-Synth is a synthetic training dataset for web agents, generated using the Agent-as-Annotators (A3) framework. It contains ~16k SFT training examples produced by Gemini 3 Pro acting as the Annotator across 3,000 tasks on 6 WebArena environments.
Dataset Structure
A3-Synth/
training/… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/A3-Synth.FaithDialFaithDial is a new benchmark for hallucination-free dialogues, created by manually editing hallucinated and uncooperative responses in Wizard of Wikipedia.mctsQAEgo4D-MC-testThis benchmark was collected by QAEgo4D and updated by GroundVQA.
We conducted some processing for the experiments presented in our paper ReKV.
gpqa_diamond_mcscannet_mcmc_3dgs
Data Statistics
Scenes
Mean PSNR ↑
Mean SSIM ↑
Mean LPIPS ↓
Mean Depth L1 ↓
Mean #3DGS
Total #3DGS
1,613
30.17 dB
0.875
0.221
0.0151 m
1.000M
1.613B
License Notice:This dataset is derived from ScanNet and follows the ScanNet Terms of Use.See ScanNet Terms for details.
VideoChat3-OL617k
VideoChat3-OL617K
VideoChat3-OL617K is the online video instruction data used by VideoChat3. It is designed to train proactive streaming video assistants that continuously observe incoming video, accumulate visual evidence, and respond at the appropriate moment.
The dataset converts video-question-answer triples into causal streaming supervision. Visual clue intervals are first localized and verified, then transformed into streaming sequences with explicit response-state tokens:… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-OL617k.japan-travel-mcp-data
Japan Travel MCP — Data
The runtime data for the japan-travel-mcp
Model Context Protocol server. Comprehensive Japanese travel data for AI agents,
built from public official sources, covering all 47 prefectures and 1,938 local
government entities.
Code lives on GitHub: github.com/ookami0210/japan-travel-mcp
Data lives here. The npm package downloads this dataset on first run.
Why this dataset exists
Japan's tourism information — created to reach the world — is… See the full description on the dataset page: https://huggingface.co/datasets/open-travel/japan-travel-mcp-data.
