CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MuteApo /RealCam-Vid RealCam-Vid Dataset News 25/04/08: We provide torch dataset demo code for example usage of our RealCam-Vid. 25/03/26: Release our dataset RealCam-Vid v1 for metric-scale camera-controlled video generation, containing ~100K video clips with dedicated short/long captions and metric-scale camera annotations. 25/02/18: Initial commit of the project, we plan to release the full dataset and data processing code in several… See the full description on the dataset page: https://huggingface.co/datasets/MuteApo/RealCam-Vid.tabularimage-to-video100K<n<1M8 likes121k downloads1y agoHugging Face02rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K276 likes20k downloads3y agoHugging Face03McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes19k downloads1y agoHugging Face04Perle-ai /multimodal-ct-radiology-reports Perle AI Multi-phase CECT and CT with Radiology Reports Summary A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering. The release has three configurations: Config Modality Subjects Pairing cect_3phase 3-phase contrast-enhanced abdominal CT (DICOM) 5 per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.tabularimage-classificationn<1K4 likes10k downloads5mo agoHugging Face05NRVBench /nrvbench-review NR Video Editing Benchmark This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions. The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.imagevideo-to-videon<1K1 likes9.5k downloads5mo agoHugging Face06opensporks /resumes Dataset Card for Resume Dataset Dataset Summary Context A collection of Resume Examples taken from livecareer.com for categorizing a given resume into any of the labels defined in the dataset. Content Contains 2400+ Resumes in string as well as PDF format. PDF stored in the data folder differentiated into their respective labels as folders with each resume residing inside the folder in pdf form with filename as the id defined in the csv. Inside the… See the full description on the dataset page: https://huggingface.co/datasets/opensporks/resumes.text1K<n<10K14 likes8.8k downloads2y agoHugging Face07vsevolodpl /REPID REPID: Rendering Evaluation of Photographic Image Dataset REPID (officially introduced as the Rendering Evaluation of Photographic Image Dataset) is a large-scale benchmark designed for Image Rendering Quality Assessment (IRQA) in paper Beyond distortions: a benchmark for subjective evaluation of image rendering quality. Unlike traditional Image Quality Assessment (IQA) which focuses on technical degradations like noise or blur, REPID aims to model subjective human aesthetic… See the full description on the dataset page: https://huggingface.co/datasets/vsevolodpl/REPID.imageimage-classification100K<n<1M4 likes7.6k downloads2mo agoHugging Face08SimplexAI /quantum-representations Epsilon-Transformers Belief Analysis Dataset This dataset contains trained neural network models and their corresponding belief state regression analysis from the Epsilon-Transformers project. The models were trained on four different stochastic processes and analyzed for their ability to learn and represent belief states. See https://github.com/adamimos/epsilon-transformers/tree/quantum-public for codebase which generated this data. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SimplexAI/quantum-representations.tabularother100K<n<1M0 likes6.2k downloads1y agoHugging Face09garak-llm /rubygems-20241031text100K<n<1M0 likes6k downloads2y agoHugging Face10hf-audio /open-asr-leaderboard-resultstabularn<1K0 likes5.5k downloads20h agoHugging Face11liamdugan /raid 🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨 🌐 Website, 🖥️ Github, 📝 Paper RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors. It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks. It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors. Load… See the full description on the dataset page: https://huggingface.co/datasets/liamdugan/raid.texttext-classification1M<n<10M28 likes5.3k downloads2y agoHugging Face12riotu-lab /Synthetic-UAV-Flight-Trajectories UAV Trajectory Dataset Summary This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training. Data Description The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Synthetic-UAV-Flight-Trajectories.tabular100K<n<1M17 likes5.1k downloads2y agoHugging Face13victor /real-or-fake-fake-jobposting-predictiontabular10K<n<100K5 likes4.8k downloads4y agoHugging Face14Longitude-Labs /spreadsheet-arena-release Spreadsheet Arena A dataset of 555 pairwise human preference votes over LLM-generated spreadsheets, spanning 124 distinct user-submitted prompts and 17 models. This is the public release accompanying the Spreadsheet Arena paper. Contents battles.csv models.csv outputs/<id>/ sheet.json sheet.xlsx <id> is a 16-char hex identifier (HMAC-SHA256 of an internal UUID under a… See the full description on the dataset page: https://huggingface.co/datasets/Longitude-Labs/spreadsheet-arena-release.tabulartabular-classificationn<1K5 likes4.5k downloads4mo agoHugging Face15Bai-YT /RAGDOLL The RAGDOLL E-Commerce Webpage Dataset This repository contains the RAGDOLL (Retrieval-Augmented Generation Deceived Ordering via AdversariaL materiaLs) dataset as well as its LLM-automated collection pipeline. The RAGDOLL dataset is from the paper Ranking Manipulation for Conversational Search Engines from Samuel Pfrommer, Yatong Bai, Tanmay Gautam, and Somayeh Sojoudi. For experiment code associated with this paper, please refer to this repository. The dataset consists of 10… See the full description on the dataset page: https://huggingface.co/datasets/Bai-YT/RAGDOLL.textquestion-answering1K<n<10K4 likes4.2k downloads2y agoHugging Face16jherng /wlasl_reduced Reduced WLASL Dataset This dataset is a task-specific reduced version of the WLASL dataset, constructed for American Sign Language (ASL) recognition experiments. Contents videos/Video clips organized by gloss label. metadata.csvPer-sample metadata including: file path gloss label fps (after normalization, if applied) video resolution normalized bounding box coordinates gloss_map.jsonMapping from gloss labels to integer class IDs. Dataset Construction… See the full description on the dataset page: https://huggingface.co/datasets/jherng/wlasl_reduced.tabularn<1K1 likes4.1k downloads9mo agoHugging Face17huggingface-projects /Deep-RL-Course-Certificationtabular1K<n<10K19 likes3.7k downloads4h agoHugging Face18roman-bushuiev /MassSpecGym MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from tandem mass spectrometry (MS/MS) spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems. Papers MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery (2025): Paper Link MassSpecGym: A benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/roman-bushuiev/MassSpecGym.tabularother100K<n<1M23 likes3.5k downloads2mo agoHugging Face19einrafh /hnm-fashion-recommendations-data Dataset Rekomendasi Fashion H&M Dataset ini berisi data transaksi, atribut pelanggan, dan metadata produk yang telah dianonimkan dari H&M Group. Kumpulan data komprehensif ini memungkinkan pemodelan perilaku pembelian pelanggan secara mendalam. Wawasan yang dihasilkan dapat dimanfaatkan untuk berbagai tujuan bisnis yang strategis, mulai dari meningkatkan personalisasi pengalaman berbelanja, mengoptimalkan manajemen inventaris untuk efisiensi produksi, hingga mendukung inisiatif… See the full description on the dataset page: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data.imagetabular-classification10M<n<100M3 likes3.3k downloads1y agoHugging Face20MrbBakh /Rotten_Tomatoestext1K<n<10K1 likes3.1k downloads3y agoHugging Face21patrickbdevaney /tripadvisor_hotel_reviewstext10K<n<100K1 likes2.7k downloads2y agoHugging Face22anonymous-stgnn-aas /TSP_EXECUTION_RUNStabular1K<n<10K1 likes2.7k downloads25d agoHugging Face23Rekin226 /aquascope-gauges AquaScope gauges The station catalog of every water-observation source AquaScope can reach, harvested on a schedule and published as GeoParquet. Last run 2026-09-23T23:27:06+00:00 with aquascope 0.18.0: 75,721 stations from 11 of 11 sources. Files: stations.parquet: GeoParquet 1.0 (WKB point geometry, WGS84). One row per station: source, station_id, site_id, name, latitude, longitude, variables, period_start, period_end, url (deep link to the agency page), river, country… See the full description on the dataset page: https://huggingface.co/datasets/Rekin226/aquascope-gauges.text10M<n<100M0 likes2.6k downloads2d agoHugging Face24CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes2.5k downloads4mo agoHugging Face25Censius-AI /ECommerce-Women-Clothing-Reviewstabular10K<n<100K2 likes2.5k downloads3y agoHugging Face26Linzhan /Objaverse-XL-Rigged-Animated Objaverse-XL Rigged & Animated Subset Every asset here carries both a skeleton and at least one animation clip, selected from Objaverse / Objaverse-XL. Rigs range from 3 to 344 joints and span characters as well as articulated rigid objects. Objaverse-XL indexes over 10 million objects, but only a small fraction carry a usable rig and motion on it. This subset isolates that fraction: every file was checked to contain at least one skin with joints and at least one animation clip… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Objaverse-XL-Rigged-Animated.3dtext-to-3d10K<n<100K1 likes2.5k downloads18d agoHugging Face27Infatoshi /kernelbench-v3-runs KernelBench-v3 — Agent Runs 2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py. Companion datasets: Infatoshi/kernelbench-v3-problems — 60 problem definitions Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.tabular1K<n<10K2 likes2.3k downloads5mo agoHugging Face28AllRates /central-bank-exchange-rates Central Bank Exchange Rates Official exchange rates published by 103 central banks and 4 tax authorities, as one CSV per institution: 14,612,276 rows, the oldest series from 1914. Refreshed daily from the GitHub source repository. Every row is the figure the institution itself published for that date: the ECB euro reference rate, the Federal Reserve H.10 table, the Bank of England spot rates, RBI reference rates, PBoC central parity, HMRC monthly rates for VAT, US Treasury… See the full description on the dataset page: https://huggingface.co/datasets/AllRates/central-bank-exchange-rates.texttime-series-forecasting10M<n<100M0 likes2.2k downloads2h agoHugging Face29well-balanced /cantabile-runs cantabile-runs Work queue and checkpoint store for the Cantabile dynamics study. The directory tree is the plan — there is no plan file and no database. main/<song>/<method>/.gitkeep queued, unclaimed main/<song>/<method>/<seed>/CLAIM-<worker> a worker holds it (mtime = heartbeat) main/<song>/<method>/<seed>/*.pt done: 5M / 6M / 7M / 8M checkpoints main/<song>/<method>/<seed>/FAILED crashed, needs a human A worker lists main/, takes… See the full description on the dataset page: https://huggingface.co/datasets/well-balanced/cantabile-runs.tabularn<1K0 likes2.1k downloads16d agoHugging Face30Real-TSF /TIME-OutputThis repository contains the extracted time series features (tsfeatures) for each variate and the detailed forecasting results for every experiment. Note: These files are for building leaderboard and visualization; users do not need to download this directory. features/: Statistical Features (tsfeatures) Each dataset's features are saved to: output/features/{dataset}/{freq}/. This directory stores the computed tsfeatures for the variates in the dataset. The folder contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Real-TSF/TIME-Output.tabulartime-series-forecasting1K<n<10K0 likes2k downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.