CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01di-zhang-fdu /AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from year 1983~2023, and 2024(part 2). Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions 2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime. Citation @misc {di_zhang_2025, author = { {Di Zhang} }, title = { AIME_1983_2024 (Revision 6283828) }, year = 2025, url = {… See the full description on the dataset page: https://huggingface.co/datasets/di-zhang-fdu/AIME_1983_2024.tabularn<1K41 likes21k downloads2y agoHugging Face02datamatastudios /ai-model-popularity Datamata AI Model Popularity Index Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot. Latest snapshot: 2026-09-20 Models in this release: 50 Updated: weekly Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution. Source & methodology: https://www.datamatastudios.com/datasets Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.tabularn<1K0 likes17k downloads3d agoHugging Face03gneubig /aime-1983-2024 AIME Problem Set 1983-2024 Dataset Description This dataset contains problems from the American Invitational Mathematics Examination (AIME) from 1983 to 2024. The AIME is a prestigious mathematics competition for high school students in the United States and Canada. Dataset Summary Source: Kaggle - AIME Problem Set 1983-2024 License: CC0: Public Domain Total Problems: 2,250 Years Covered: 1983 to 2024 Main Task: Mathematics Problem Solving… See the full description on the dataset page: https://huggingface.co/datasets/gneubig/aime-1983-2024.tabulartext-classificationn<1K21 likes5.5k downloads2y agoHugging Face04mitermix /ai-musictext10M<n<100M0 likes3.2k downloads1y agoHugging Face05Ahus-AIM /EchoXFlow EchoXFlow This dataset repository contains Croissant metadata plus one uncompressed tar archive per exam. Extraction Clone or download the dataset repository first. With Git, this creates an EchoXFlow/ folder: git lfs install git clone https://huggingface.co/datasets/Ahus-AIM/EchoXFlow cd EchoXFlow The downloaded repository contains croissant.json plus one tar archive per exam under exams/. Extract every exam archive into a local data/ directory to materialize the… See the full description on the dataset page: https://huggingface.co/datasets/Ahus-AIM/EchoXFlow.textn<1K3 likes1.1k downloads5mo agoHugging Face06AIML-TUDA /i2p Inaproppriate Image Prompts (I2P) The I2P benchmark contains real user prompts for generative text2image prompts that are unproportionately likely to produce inappropriate images. I2P was introduced in the 2023 CVPR paper Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models. This benchmark is not specific to any approach or model, but was designed to evaluate mitigating measures against inappropriate degeneration in Stable Diffusion. The corresponding… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/i2p.tabular1K<n<10K21 likes1.1k downloads3y agoHugging Face07AImageLab-Zip /mimose_runstabular1K<n<10K0 likes1k downloads26d agoHugging Face08philschmid /AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from year 1983~2023, and 2024(part 2). Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions 2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime. tabularn<1K0 likes754 downloads2y agoHugging Face09AIML-TUDA /dlam-ts-project-data-2026 operations_forecasting_2026 Multivariate hourly forecasting for anonymized operations units. Target Predict the future hourly operational load index for each series_id. Higher values indicate more operational pressure in that unit. Forecast Contract Frequency: h Series: 96 Timesteps per series: 4992 Target column: target Training history length used by the baseline templates: 168 Rollout block length: 24 Required prediction horizon: validation: 336, test: 336… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/dlam-ts-project-data-2026.tabular100K<n<1M2 likes504 downloads4mo agoHugging Face10Cross-Mergeability /aim-activation-informed-merging AIM: does activation-informed merging change what makes a merge work? Headline AIM does exactly what it claims, the targeting is what makes it work — and it changes nothing about what predicts a good merge. AIM is exactly what it says on the tin, and that is verifiable from public artefacts alone. The published with-AIM checkpoints are recovered, to R² = 0.9992, as a closed-form per-input-channel shrinkage of their baseline twins toward the base model, with ω̂ =… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/aim-activation-informed-merging.tabular1K<n<10K0 likes426 downloads27d agoHugging Face11AIM-Harvard /MedBrowseComp MedBrowseComp Dataset This repository contains datasets for medical information-seeking-oriented deep research and computer use tasks. Datasets The repository contains three harmonized datasets: MedBrowseComp_50: A collection of 50 medical entries for browsing and comparison. MedBrowseComp_605: A comprehensive collection of 605 medical entries. MedBrowseComp_CUA: A curated collection of medical data for comparison and analysis. Usage These datasets can be… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Harvard/MedBrowseComp.textquestion-answering1K<n<10K8 likes256 downloads1y agoHugging Face12AiMijie /EC-Guide This repo is only used for dataset viewer. Please download from here. Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5) The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.textquestion-answering10K<n<100K2 likes217 downloads2y agoHugging Face13lchen001 /AIME1983_2024tabularn<1K0 likes186 downloads1y agoHugging Face14AIMH /SWMHgatedWe collect this dataset from some mental health-related subreddits in https://www.reddit.com/ to further the study of mental disorders and suicidal ideation. We name this dataset as Reddit SuicideWatch and Mental Health Collection, or SWMH for short, where discussions comprise suicide-related intention and mental disorders like depression, anxiety, and bipolar. We use the Reddit official API and develop a web spider to collect the targeted forums. This collection contains a total of 54,412… See the full description on the dataset page: https://huggingface.co/datasets/AIMH/SWMH.text10K<n<100K13 likes162 downloads3y agoHugging Face15joyfine /Qwen3-235B-A22B-Thinking-2507_Qwen3-1.7B_AIME_1983_2024textn<1K0 likes159 downloads11mo agoHugging Face16CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tabular1K<n<10K0 likes153 downloads2y agoHugging Face17AIML-TUDA /LlavaGuardgatedWARNING: This repository contains content that might be disturbing! Therefore, we set the Not-For-All-Audiences tag. License The annotations provided in this dataset (e.g., labels, bounding boxes, coordinates) are released under the Apache License 2.0. The linked or referenced images are not included under this license and remain under their original licenses. Users are responsible for ensuring compliance with the terms of use of the respective image sources. Content… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/LlavaGuard.image1K<n<10K10 likes130 downloads1y agoHugging Face18Kashtan /AIMixDetectPublicData AIMixDetect: detect mixed authorship of a language model (LM) and humans contributors: Alon Kipnis, Idan Kashtan This dataset was utilized in the paper "An Information-Theoretic Approach for Detecting Edits in AI-Generated Text" by Alon Kipnis and Idan Kashtan. We used three Hugging Face publicly available datasets to generate our dataset: aadityaubhat/GPT-wiki-intro isarth/chatgpt-news-articles NicolaiSivesind/ChatGPT-Research-Abstracts For each category, the dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Kashtan/AIMixDetectPublicData.texttext-classification10K<n<100K0 likes129 downloads2y agoHugging Face19hiddennode /AIME AIME Hallucination Detection Dataset This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research. Dataset Details Name: AIME Hallucination Detection Dataset Format: CSV Size: (14.6 MB) Files Included: AIME-hallucination-detection-dataset.csv: Contains the dataset. Content Description The dataset… See the full description on the dataset page: https://huggingface.co/datasets/hiddennode/AIME.tabularn<1K0 likes125 downloads2y agoHugging Face20bhushansshah /AIME-1983-to-2026tabular1K<n<10K0 likes78 downloads7mo agoHugging Face21Aiman1234 /Interview-questionsannotations_creators: crowdsourced machine-generated language: en language_creators: machine-generated crowdsourced license: other multilinguality: monolingual pretty_name: 'interview-questions-on-Programming-languages ' size_categories: n<1K source_datasets: original tags: interview-questions task_categories: text-generation task_ids: language-modeling textn<1K3 likes65 downloads2y agoHugging Face22AI-MED-AGH /Recruitment-Task-3 DeepWeeds - AI-MED AGH convenience mirror This is a convenience mirror of the official DeepWeeds image archive and the upstream annotations pinned to a specific commit. original/images.zip is preserved unchanged; images are not extracted or duplicated here. models.zip from the source authors is deliberately not mirrored. Dataset facts 17,509 in-situ images from Queensland, Australia. Nine classes: eight weed species plus Negative. The authors publish five folds… See the full description on the dataset page: https://huggingface.co/datasets/AI-MED-AGH/Recruitment-Task-3.imageimage-classification10K<n<100K0 likes64 downloads13d agoHugging Face23qyi-iquest /AIME_2026tabularn<1K0 likes59 downloads8mo agoHugging Face24Floppanacci /QWQ-LongCOT-AIMOQWQ-LongCOT-AIMO is a derived dataset created by processing the amphora/QwQ-LongCoT-130K dataset. It filters the original dataset to focus specifically on question-answering pairs where the final answer is a numerical value between 0 and 999, explicitly marked using the \boxed{...} format within the original chain-of-thought answer. Dataset Structure Data Splits The dataset is split into training, validation, and test sets with an 80/10/10 ratio based on the filtered… See the full description on the dataset page: https://huggingface.co/datasets/Floppanacci/QWQ-LongCOT-AIMO.texttext-generation10K<n<100K0 likes56 downloads1y agoHugging Face25ericzhao28 /aimetabularn<1K0 likes55 downloads2y agoHugging Face26lchen001 /AIME2025tabularn<1K0 likes55 downloads2y agoHugging Face27AIMindTeams /synthetic-chemical-reactor-5k-sample RL-Ready Synthetic Chemical Batch Reactor Dataset 🚨 Download the full 50,000-row dataset featuring 269 complete reactor life-cycles here: https://aimindteam.gumroad.com/l/reactor-dataset 🚨 Overview This dataset provides a mathematically consistent, high-fidelity simulation of an industrial liquid-phase exothermic batch reactor. It contains multivariate time-series operational logs designed specifically for training Reinforcement Learning (RL) agents, testing Predictive… See the full description on the dataset page: https://huggingface.co/datasets/AIMindTeams/synthetic-chemical-reactor-5k-sample.tabulartime-series-forecasting1K<n<10K2 likes45 downloads5mo agoHugging Face28AI-Mock-Interviewer /PDF_Extracted_Datatext1K<n<10K0 likes42 downloads1y agoHugging Face29backups /ai-musictext10M<n<100M0 likes35 downloads1y agoHugging Face30pagonzalez2001 /spanish-poetry-dataset-for-AFT-AImotionsThis dataset was previously created in Kaggle by Andrea Morales Garzón. Link Kaggle text1K<n<10K0 likes32 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.