CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01McGill-NLP /WebLINX-full WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead: WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models 💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.text10K<n<100K8 likes28k downloads1y agoHugging Face02Yelp /yelp_review_full Dataset Card for YelpReviewFull Dataset Summary The Yelp reviews dataset consists of reviews from Yelp. It is extracted from the Yelp Dataset Challenge 2015 data. Supported Tasks and Leaderboards text-classification, sentiment-classification: The dataset is mainly used for text classification: given the text, predict the sentiment. Languages The reviews were mainly written in english. Dataset Structure Data Instances A… See the full description on the dataset page: https://huggingface.co/datasets/Yelp/yelp_review_full.texttext-classification100K<n<1M149 likes20k downloads3y agoHugging Face03AlgorithmicResearchGroup /s2orc_full S2ORC Full — Semantic Scholar Open Research Corpus A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information. Dataset Description S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_full.texttext-generation10M<n<100M2 likes18k downloads5mo agoHugging Face04Spawning /pd12m-fullThis dataset is the downloaded variant of Spawning/PD12M. More specifically, this dataset is compatible with webdataset. It was made public after obtaining permission from the original authors of the dataset. You can use the following to explore the dataset with webdataset: import webdataset as wds dataset_path = "pipe:curl -s -f -L https://huggingface.co/datasets/sayakpaul/pd12m-full/resolve/main/{00155..02480}.tar" dataset = ( wds.WebDataset(dataset_path… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd12m-full.image10M<n<100M21 likes10k downloads2y agoHugging Face05weikaih /vsi-bench-qa-v3-hm3d-fullimage10K<n<100K0 likes8.5k downloads1y agoHugging Face06PromptEval /PromptEval_MMLU_full MMLU Multi-Prompt Evaluation Data Overview This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt evaluation of LLMs."… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_full.tabularquestion-answering10M<n<100M3 likes8.2k downloads2y agoHugging Face07KMK040412 /aitw-processed-labeled-full AiTW Processed Full with App Labels This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset. Why This Exists AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.imageimage-text-to-text1M<n<10M0 likes7.6k downloads4mo agoHugging Face08vkehfdl1 /banana-vidorev3-fullpipe Banana ViDoRe v3 Fullpipe Nano Banana Pro full-pipeline synthetic training data for ViDoRe v3 finance and industrial domains. This repository contains 4670 training records and 40190 unique referenced images across domain-separated ColFlor/ColQwen training splits. Images are included in the repository and paths in each JSONL are relative to that domain directory. Generated at: 2026-06-29T09:39:01.644860+00:00 Layout finance/train.jsonl finance/metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/vkehfdl1/banana-vidorev3-fullpipe.imagevisual-document-retrieval1K<n<10K0 likes7.4k downloads3mo agoHugging Face09yifishbossman /financial-analyst-data-full financial-analyst-data-full A-share historical price + valuation data packaged for financial-analyst — the 14-agent single-stock deep-dive research workstation. Published: 2026-05-24 Preset: full — 全 A 股完整包 (含历史退市股). 量化研究员 / 重度用户. lite 全 + TDX 历年财报原始 zip (用户跑 import_tdx_financial.py 解) + F10 原始文本 (公司大事/龙虎榜/主力追踪/最新提示 .txt). Size: ~14.1 GB What's included 5450 stocks daily OHLCV + 7 valuation fields (PE/PB/PS/DV/MV/CIRC_MV/turnover_rate) Date range (daily): 1990-12-19… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-full.texttime-series-forecasting1K<n<10K7 likes6.7k downloads4mo agoHugging Face10MRSHREY197 /icrm-hitek-full-db-mixed ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet). Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent). Contents The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/MRSHREY197/icrm-hitek-full-db-mixed.text1B<n<10B0 likes5.8k downloads17d agoHugging Face11OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes5.4k downloads8mo agoHugging Face12NuTonic /sat-image-boundingbox-sft-full NU-TONIC raw SFT Full Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters. Provenance Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2) Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8. Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.imageimage-text-to-text100K<n<1M14 likes5.2k downloads5mo agoHugging Face13AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes4.4k downloads1y agoHugging Face14devil-69 /ICMR-HITEK-FULL-MIXED-DBgated ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet). Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent). Contents The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/devil-69/ICMR-HITEK-FULL-MIXED-DB.text1B<n<10B0 likes4.3k downloads23d agoHugging Face15benjamin-paine /free-music-archive-full FMA: A Dataset for Music Analysis Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson. International Society for Music Information Retrieval Conference (ISMIR), 2017. We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.audioaudio-to-audio100K<n<1M20 likes3.5k downloads2y agoHugging Face16567-labs /wikipedia-bge-small-en-v1.5-fulltext1M<n<10M4 likes3.4k downloads3y agoHugging Face17wannabeyour /icrm-hitek-fulldbtext1B<n<10B0 likes3.3k downloads1mo agoHugging Face18acmc /beamit-full-texts-dataset Dataset Card for "beamit-full-texts-dataset" More Information needed text10K<n<100K0 likes3k downloads3y agoHugging Face19tonyALTR /3D_full_poly_gentextn<1K0 likes2.5k downloads1y agoHugging Face20SwayStar123 /preprocessed_DCAE-f64_1024_pd12m-fulltext1M<n<10M0 likes2.5k downloads1y agoHugging Face21tonyALTR /3D_full_poly_rot_gentextn<1K0 likes2.4k downloads1y agoHugging Face22TfqDeadlox636 /icrm-hitek-fulldbtext1B<n<10B3 likes2.1k downloads29d agoHugging Face23jm-rt /arvo-vulnsmith-full ARVO CyberGym-format smoke dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It contains 10 ARVO tasks that are outside the original CyberGym set. textn<1K0 likes2k downloads2mo agoHugging Face24jedibear /s2orc_full S2ORC Full — Semantic Scholar Open Research Corpus A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information. Dataset Description S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/jedibear/s2orc_full.texttext-generation10M<n<100M0 likes2k downloads4mo agoHugging Face25KlingTeam /FullBenchimage1K<n<10K7 likes2k downloads1y agoHugging Face26benjamin-paine /freesound-laion-640k-commercial-16khz-full About this Repository This repository is the training split of the complete FreeSound LAION 640k dataset, limited only to licenses that permit commercial works, resampled to 16khz using torchaudio.transforms.Resample. This is ideal for use cases where a variety of audio is desired but fidelity and labels are unnecessary, such as background audio for augmenting other datasets. Dataset Versions You are looking at the full dataset which contains 403,146 unique sounds… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/freesound-laion-640k-commercial-16khz-full.audioaudio-to-audio100K<n<1M2 likes1.9k downloads2y agoHugging Face27armanbabayan /full_checkbox_dropdown_radiobuttonimage1K<n<10K0 likes1.8k downloads3y agoHugging Face28DatologyAI /DatBench-Full DatBench: Discriminative, Faithful, and Efficient VLM Evaluations DatBench is a curated evaluation suite for vision–language models (VLMs) designed to be faithful, discriminative, and efficient. 📄 DatBench: Discriminative, Faithful, and Efficient VLM Evaluationshttps://arxiv.org/abs/2601.02316 Modern VLM benchmarks often overestimate model capability due to multiple-choice inflation, language-only shortcuts, annotation noise, and redundant low-signal samples. DatBench reframes… See the full description on the dataset page: https://huggingface.co/datasets/DatologyAI/DatBench-Full.image100K<n<1M77 likes1.8k downloads5mo agoHugging Face29Felldude /Birds_of_North_America_Fullimage10K<n<100K0 likes1.8k downloads8mo agoHugging Face30sqy201x /full-o46text10K<n<100K0 likes1.7k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.