CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01McAuley-Lab /Amazon-Reviews-2023Amazon Review 2023 is an updated version of the Amazon Review 2018 dataset. This dataset mainly includes reviews (ratings, text) and item metadata (desc- riptions, category information, price, brand, and images). Compared to the pre- vious versions, the 2023 version features larger size, newer reviews (up to Sep 2023), richer and cleaner meta data, and finer-grained timestamps (from day to milli-second).10B<n<100B356 likes65k downloads2y agoHugging Face02McGill-NLP /weblinx-browsergym WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the video tag. This dataset was specifically created to allow WebLINX to be used inside the BrowserGym and Agentlab ecosystem. Please see the browsergym repository for more information. [!NOTE] The version associated with this library is WebLINX… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/weblinx-browsergym.image-to-text4 likes35k downloads2y agoHugging Face03McGill-NLP /WebLINX-full WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead: WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models 💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.text10K<n<100K8 likes29k downloads1y agoHugging Face04MCG-NJU /VideoChat3-LV116k VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments. The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.textvideo-text-to-text1K<n<10K15 likes22k downloads2mo agoHugging Face05McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes19k downloads1y agoHugging Face06ScaleAI /MCP-AtlasMCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers Leaderboard | MCP Atlas Paper | Github Dataset Summary This public release is a subset of 500 sample tasks from the MCP Atlas Benchmark dataset. MCP Atlas is a large-scale benchmark for evaluating tool-use competency, comprising 36 real MCP servers and 220 tools. Tasks are designed to assess tool-use competency in realistic, multi-step workflows. Tasks use natural language prompts that avoid… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/MCP-Atlas.textn<1K19 likes9.6k downloads2mo agoHugging Face07hf-mcp-server /test-mcp-logs(Put queries first as heuristics don't detect when there are no logs) textn<1K0 likes9.1k downloads9d agoHugging Face08bishoygaloaa /Motion-o-MCoT-PLM-motion-keyframes Motion-o-MCoT (PLM + motion keyframes) Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/. Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats). Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.image1K<n<10K1 likes7.4k downloads6mo agoHugging Face09mcp-tools /discover-toolstextn<1K5 likes6.9k downloads1mo agoHugging Face10ShapeSplats /aria_synthetic_envs_mcmc_3dgs_newgatedLicense Notice:This dataset is derived from the Aria Dataset.It follows the Aria Synthetic Environments Dataset License Agreement.See Aria License for details. 10K<n<100K1 likes6.6k downloads24d agoHugging Face11b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes6.3k downloads3y agoHugging Face12MCG-NJU /TimeLens2-93K TimeLens2-93K TimeLens2-93K is a large-scale, long-video temporal grounding dataset. This release contains 23,793 videos and 93,232 text–temporal interval pairs, including 12,091 multi-span pairs. The videos range from short clips to nearly 100 minutes and cover broad web domains such as entertainment, education, sports, news, science and technology, gaming, travel, vehicles, music, and daily life. TimeLens2-93K offers a rare combination of scale, long-context coverage… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/TimeLens2-93K.textvideo-text-to-text10K<n<100K13 likes6k downloads2mo agoHugging Face13Jakumetsu /mcpmark-trajectory-log MCPMark Trajectory Logs (mcpmark-v1-0905) These logs are publicly available for research use. They capture end-to-end trajectories (prompts, tool calls, and outcomes) for MCPMark tasks across multiple MCP services and models. Contents Each task run produces a trajectory folder containing three files: meta.json: Metadata (task id, model, timestamps, status, etc.) messages.json: Turn-by-turn exchanges including model thoughts/tool calls execution.log: Key execution-time… See the full description on the dataset page: https://huggingface.co/datasets/Jakumetsu/mcpmark-trajectory-log.10K<n<100K0 likes5.6k downloads1y agoHugging Face14McGill-NLP /stereoset Dataset Card for StereoSet Dataset Summary StereoSet is a dataset that measures stereotype bias in language models. StereoSet consists of 17,000 sentences that measures model preferences across gender, race, religion, and profession. Supported Tasks and Leaderboards multiple-choice question answering Languages English (en) Dataset Structure Data Instances #intersentence {'bias_type': 'race', 'context': 'Many people live in… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/stereoset.texttext-classification1K<n<10K31 likes5.5k downloads3y agoHugging Face15xiaoleezuishuai /airbot-fold-cloth-mcapimage1K<n<10K0 likes5k downloads2mo agoHugging Face16longevity-genie /bio-mcp-data Bio-MCP-Data A repository containing biological datasets that will be used by BIO-MCP MCP (Model Context Protocol) standard. About This repository hosts biological data assets formatted to be compatible with the Model Context Protocol, enabling AI models to efficiently access and process biological information. The data is managed using Git Large File Storage (LFS) to handle large biological datasets. Purpose Provide standardized biological datasets for AI… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/bio-mcp-data.text0 likes4.7k downloads1y agoHugging Face17gussieIsASuccessfulWarlock /security_instruct_mcq_2481textn<1K0 likes4.7k downloads2y agoHugging Face18Multilingual-Multimodal-NLP /McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval. texttext-generation10K<n<100K21 likes4.1k downloads2y agoHugging Face19zicsx /mC4-Hindi-Cleaned-3.0 Dataset Card for "mC4-Hindi-Cleaned-3.0" More Information needed text1M<n<10M2 likes4.1k downloads3y agoHugging Face20izumi-lab /mc4-ja Dataset Card for "mc4-ja" More Information needed text10M<n<100M6 likes4k downloads3y agoHugging Face21Matt1up /guertin-mcro-forensic-corpus Guertin MCRO Forensic Corpus The public court record of State of Minnesota v. Matthew David Guertin (Hennepin County, 27-CR-23-1886) and 2,902 other Minnesota dockets, collected from Minnesota Court Records Online (MCRO) — the Minnesota Judicial Branch's public access to its own case records — together with the database built from them, screen recordings of the collection, four federal cases, the email record, and a forensic analysis of all of it. Every filing is the PDF the… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus.tabulartext-classification1M<n<10M0 likes3.9k downloads9d agoHugging Face22GaussianWorld /matterport3d_region_mcmc_3dgsgatedLicense Notice:This dataset is derived from Matterport3D.It follows the Matterport End User License Agreement for Academic Use of Model Data.See Matterport3D License for details. 3dother1K<n<10K0 likes3.8k downloads1y agoHugging Face23McGill-NLP /A3-Synth A3-Synth 💾 Code 📄 Paper 🌐 Website 🤗 Dataset 🤖 Models 📦 PyPI Structured Distillation of Web Agent Capabilities Enables Generalization Xing Han Lù, Siva Reddy A3-Synth is a synthetic training dataset for web agents, generated using the Agent-as-Annotators (A3) framework. It contains ~16k SFT training examples produced by Gemini 3 Pro acting as the Annotator across 3,000 tasks on 6 WebArena environments. Dataset Structure A3-Synth/ training/… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/A3-Synth.text-generation10K<n<100K1 likes3.5k downloads6mo agoHugging Face24McGill-NLP /FaithDialFaithDial is a new benchmark for hallucination-free dialogues, created by manually editing hallucinated and uncooperative responses in Wizard of Wikipedia.texttext-generation10K<n<100K18 likes3.4k downloads4y agoHugging Face25yingyingzhang /mcts0 likes3.3k downloads32m agoHugging Face26Becomebright /QAEgo4D-MC-testThis benchmark was collected by QAEgo4D and updated by GroundVQA. We conducted some processing for the experiments presented in our paper ReKV. textn<1K2 likes3.1k downloads2y agoHugging Face27hendrydong /gpqa_diamond_mctextn<1K2 likes3.1k downloads2y agoHugging Face28GaussianWorld /scannet_mcmc_3dgsgated Data Statistics Scenes Mean PSNR ↑ Mean SSIM ↑ Mean LPIPS ↓ Mean Depth L1 ↓ Mean #3DGS Total #3DGS 1,613 30.17 dB 0.875 0.221 0.0151 m 1.000M 1.613B License Notice:This dataset is derived from ScanNet and follows the ScanNet Terms of Use.See ScanNet Terms for details. 3dother1K<n<10K3 likes2.9k downloads2mo agoHugging Face29MCG-NJU /VideoChat3-OL617k VideoChat3-OL617K VideoChat3-OL617K is the online video instruction data used by VideoChat3. It is designed to train proactive streaming video assistants that continuously observe incoming video, accumulate visual evidence, and respond at the appropriate moment. The dataset converts video-question-answer triples into causal streaming supervision. Visual clue intervals are first localized and verified, then transformed into streaming sequences with explicit response-state tokens:… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-OL617k.videovideo-text-to-text1K<n<10K13 likes2.8k downloads21d agoHugging Face30open-travel /japan-travel-mcp-data Japan Travel MCP — Data The runtime data for the japan-travel-mcp Model Context Protocol server. Comprehensive Japanese travel data for AI agents, built from public official sources, covering all 47 prefectures and 1,938 local government entities. Code lives on GitHub: github.com/ookami0210/japan-travel-mcp Data lives here. The npm package downloads this dataset on first run. Why this dataset exists Japan's tourism information — created to reach the world — is… See the full description on the dataset page: https://huggingface.co/datasets/open-travel/japan-travel-mcp-data.text-retrieval100K<n<1M0 likes2.8k downloads7h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.