CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01di-zhang-fdu /AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from year 1983~2023, and 2024(part 2). Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions 2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime. Citation @misc {di_zhang_2025, author = { {Di Zhang} }, title = { AIME_1983_2024 (Revision 6283828) }, year = 2025, url = {… See the full description on the dataset page: https://huggingface.co/datasets/di-zhang-fdu/AIME_1983_2024.tabularn<1K41 likes22k downloads2y agoHugging Face02datamatastudios /ai-model-popularity Datamata AI Model Popularity Index Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot. Latest snapshot: 2026-09-20 Models in this release: 50 Updated: weekly Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution. Source & methodology: https://www.datamatastudios.com/datasets Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.tabularn<1K0 likes17k downloads6d agoHugging Face03APRIL-AIGC /UltraVideo UltraVideo: High-Quality UHD 4K Video Dataset 🤓 Project    | 📑 Paper    | 🤗 Hugging Face (UltraVideo Dataset))   | 🤗 Hugging Face (UltraVideo-Long Dataset))   | 🤗 Hugging Face (UltraWan-1K/4K Weights)   UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions 🎋 Click below image to watch the 4K demo video. 🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo.tabularimage-to-video10K<n<100K66 likes11k downloads1y agoHugging Face04Perle-ai /multimodal-ct-radiology-reports Perle AI Multi-phase CECT and CT with Radiology Reports Summary A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering. The release has three configurations: Config Modality Subjects Pairing cect_3phase 3-phase contrast-enhanced abdominal CT (DICOM) 5 per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.tabularimage-classificationn<1K4 likes10k downloads5mo agoHugging Face05lightly-ai /epic-kitchens-100-clips EPIC-KITCHENS-100 Extracted Clips About Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset, more precisely the extension part not contained in EPIC-KITCHENS-55. For details, see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio. The clips folder contains one video for every narration from action annotations stored in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.tabular10K<n<100K2 likes9.4k downloads6mo agoHugging Face06stanford-crfm /air-bench-2024 AIRBench 2024 AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning categories of the regulation-based safety categories in the AIR 2024 safety taxonomy. Dataset Details Dataset Description AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/air-bench-2024.texttext-generation10K<n<100K26 likes7.6k downloads2y agoHugging Face07gneubig /aime-1983-2024 AIME Problem Set 1983-2024 Dataset Description This dataset contains problems from the American Invitational Mathematics Examination (AIME) from 1983 to 2024. The AIME is a prestigious mathematics competition for high school students in the United States and Canada. Dataset Summary Source: Kaggle - AIME Problem Set 1983-2024 License: CC0: Public Domain Total Problems: 2,250 Years Covered: 1983 to 2024 Main Task: Mathematics Problem Solving… See the full description on the dataset page: https://huggingface.co/datasets/gneubig/aime-1983-2024.tabulartext-classificationn<1K21 likes6.1k downloads2y agoHugging Face08APRIL-AIGC /UltraVideo-Long UltraVideo: High-Quality UHD 4K Video Dataset 🤓 Project    | 📑 Paper    | 🤗 Hugging Face (UltraVideo Dataset))   | 🤗 Hugging Face (UltraVideo-Long Dataset))   | 🤗 Hugging Face (UltraWan-1K/4K Weights)   UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions 🎋 Click below image to watch the 4K demo video. 🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo-Long.tabularimage-to-video10K<n<100K7 likes4.3k downloads1y agoHugging Face09mitermix /ai-musictext10M<n<100M0 likes3k downloads1y agoHugging Face10lmarena-ai /arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles. The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models. Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad). Citation Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.tabulartext-classification10K<n<100K159 likes2.6k downloads2y agoHugging Face11Censius-AI /ECommerce-Women-Clothing-Reviewstabular10K<n<100K2 likes2.5k downloads3y agoHugging Face12Kukedlc /suno-ai-music-dataset Suno AI Music Dataset (Multi-Genre Curated) A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research. This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.audioaudio-classificationn<1K29 likes2.1k downloads4mo agoHugging Face13AI4Sec /cti-bench Dataset Card for CTIBench A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks. Dataset Details Dataset Description CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI. Components: CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.textzero-shot-classification1K<n<10K21 likes1.9k downloads2y agoHugging Face14infinite-dataset-hub /FINDER_API_KEY_AI_SEARCH_2023 FINDER_API_KEY_AI_SEARCH_2023 tags: data collection, machine learning, API performance Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.tabularn<1K0 likes1.7k downloads2y agoHugging Face15Beijing-AISI /panda-bench PandaBench PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies. The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges. Dataset Description This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/panda-bench.tabulartext-generation100K<n<1M0 likes1.5k downloads1y agoHugging Face16aieng-lab /biasneutral-ajibawa BIASNEUTRAL Ajibawa BIASNEUTRAL Ajibawa contains Ajibawa-derived text snippets that passed the GRADIEND bias-neutral filtering pipeline. It is the Ajibawa-source counterpart to the original aieng-lab/biasneutral dataset, which can not be directly published due to licensing issues. This dataset provides an easier-to-access solution, while maintaining the same generation principles as the original BIASNEUTRAL properties. Usage from datasets import load_dataset ds… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/biasneutral-ajibawa.text1M<n<10M0 likes1.5k downloads2mo agoHugging Face17Ahus-AIM /EchoXFlow EchoXFlow This dataset repository contains Croissant metadata plus one uncompressed tar archive per exam. Extraction Clone or download the dataset repository first. With Git, this creates an EchoXFlow/ folder: git lfs install git clone https://huggingface.co/datasets/Ahus-AIM/EchoXFlow cd EchoXFlow The downloaded repository contains croissant.json plus one tar archive per exam under exams/. Extract every exam archive into a local data/ directory to materialize the… See the full description on the dataset page: https://huggingface.co/datasets/Ahus-AIM/EchoXFlow.textn<1K3 likes1.5k downloads5mo agoHugging Face18AImageLab-Zip /mimose_runstabular1K<n<10K0 likes1.5k downloads28d agoHugging Face19shi-labs /physical-ai-bench-conditional-generation Physical AI Bench - Conditional Generation Paper | Code This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below. Dataset Category Sample Nums Agibot World Robotics 200 OpenDV… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-conditional-generation.textvideo-to-videon<1K0 likes1.4k downloads10mo agoHugging Face20ai4bharat /MANGO MANGO: A Corpus of Human Ratings for Speech MANGO (MUSHRA Assessment corpus using Native listeners and Guidelines to understand human Opinions at scale) is the first large-scale dataset designed for evaluating Text-to-Speech (TTS) systems in Indian languages. Key Features: 255,150 human ratings of TTS-generated outputs and ground-truth human speech. Covers two major Indian languages: Hindi & Tamil, and English. Based on the MUSHRA (Multiple Stimuli with Hidden Reference… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MANGO.audiotext-to-speech10K<n<100K6 likes1.3k downloads1y agoHugging Face21vals-ai /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.textn<1K9 likes1.2k downloads1y agoHugging Face22AIML-TUDA /i2p Inaproppriate Image Prompts (I2P) The I2P benchmark contains real user prompts for generative text2image prompts that are unproportionately likely to produce inappropriate images. I2P was introduced in the 2023 CVPR paper Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models. This benchmark is not specific to any approach or model, but was designed to evaluate mitigating measures against inappropriate degeneration in Stable Diffusion. The corresponding… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/i2p.tabular1K<n<10K21 likes1.1k downloads3y agoHugging Face23maum-ai /CostNav-Teleop-Dataset CostNav Teleop Dataset Dataset Summary The CostNav Teleop Dataset is a large-scale collection of human teleoperation recordings for robot navigation in an urban sidewalk simulation environment. It was collected as part of the CostNav benchmark, which evaluates navigation systems using real-world economic cost and revenue metrics rather than purely technical metrics. The dataset contains 2,203 teleoperation episodes totaling 50.2 hours of driving… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/CostNav-Teleop-Dataset.tabularrobotics1K<n<10K1 likes1k downloads4mo agoHugging Face24sujet-ai /Sujet-Finance-Instruct-177k Sujet Finance Dataset Overview The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.tabulartext-generation100K<n<1M85 likes908 downloads2y agoHugging Face25ai4bharat /BPCCgated BPCC Dataset Training Bharat Parallel Corpus Collection (BPCC) is a comprehensive and publicly available parallel corpus that includes both existing and new data for all 22 scheduled Indic languages. It is comprised of two parts: BPCC-Mined and BPCC-Human, totaling approximately 230 million bitext pairs. BPCC-Mined contains about 228 million pairs, with nearly 126 million pairs newly added as a part of this work. On the other hand, BPCC-Human consists of 2.2 million gold… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/BPCC.text100M<n<1B42 likes861 downloads9mo agoHugging Face26osanseviero /twitter-airline-sentiment Dataset Card for Twitter US Airline Sentiment Dataset Summary This data originally came from Crowdflower's Data for Everyone library. As the original source says, A sentiment analysis job about the problems of each major U.S. airline. Twitter data was scraped from February of 2015 and contributors were asked to first classify positive, negative, and neutral tweets, followed by categorizing negative reasons (such as "late flight" or "rude service"). The data we're… See the full description on the dataset page: https://huggingface.co/datasets/osanseviero/twitter-airline-sentiment.tabular10K<n<100K3 likes853 downloads4y agoHugging Face27NMAIResearch /eu-ai-act-article-50-scoreboard Article 50 historical public-evidence snapshot This work was produced through an AI-assisted workflow directed by the author. Historical work used Anthropic assistance; the retrospective correction uses OpenAI GPT-6, with separate bounded Gemini advice. All three providers have products in the scored set. Purpose: provide the corrected paper's version 1.1 bundle under v1_1. Start with its README and correction note. The paper and deposit and GitHub repository identify the same… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/eu-ai-act-article-50-scoreboard.imagen<1K0 likes835 downloads3d agoHugging Face28ai4bharat /FBI Finding Blind Spots in Evaluator LLMs with Interpretable Checklists We present FBI, our novel meta-evaluation framework designed to assess the robustness of evaluator LLMs across diverse tasks and evaluation strategies. Please refer to our paper for more details. Code The code to generate the perturbations and run evaluations are available on our github repository: ai4bharat/fbi Tasks We manually categorized each prompt into one of the 4 task… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/FBI.image1K<n<10K2 likes818 downloads2y agoHugging Face29apol /ai-election-manipulation-cases AI, Elections and Agency Transfer Evidence Index Version 0.4.4 · released 21 August 2026 · research cutoff 12 August 2026 The dataset contains 6 documented-manipulation records, not 1,087 cases. Read the counts in this order: 1,087 relational rows -> 64 catalogue entries -> 10 core records -> 8 incident-eligible records -> 6 documented-manipulation records The other two incident-eligible records are transparent contested-use… See the full description on the dataset page: https://huggingface.co/datasets/apol/ai-election-manipulation-cases.text1K<n<10K0 likes769 downloads1mo agoHugging Face30philschmid /AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from year 1983~2023, and 2024(part 2). Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions 2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime. tabularn<1K0 likes744 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.