CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jxcai-scale /hle-public-questionstext1K<n<10K0 likes57k downloads1y agoHugging Face02jxcai-scale /hle_prompts_07_02_25text100K<n<1M0 likes6.5k downloads1y agoHugging Face03EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-scale0.25 Retrain bank: plan_adam_eps1e17_16k_scale0.25 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256-scale0.25.tabular1K<n<10K0 likes666 downloads1mo agoHugging Face04GD-ML /SCASRec SCASRec: A Self-Correcting and Auto-Stopping Model for Generative Route List Recommendation This is the dataset for our paper. The following table contains the feature dimensions and key features of our dataset. Feature Type Interpretation Shape Some Key Features Route Features Used to describe each route, including static features, dynamic features, and trajectory statistical features N * 62 The estimated time of arrival for the routeThe total distance length of the… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/SCASRec.texttabular-classification100K<n<1M30 likes534 downloads7mo agoHugging Face05ScaleAI /BioRiskEvaltabular100K<n<1M0 likes365 downloads1y agoHugging Face06trust-and-safety /abuse-scanner-bot-datasettextn<1K0 likes345 downloads1y agoHugging Face07datamatastudios /skill-scarcity-index Datamata Skill Scarcity Index Which tech skills are genuinely hard to hire for: a daily composite scarcity score per skill built from how long roles stay open (time-to-fill), the salary premium employers pay over the category median and how often the same role is re-posted after failing to fill. Computed from active job listings across public company career pages and job boards. Latest snapshot: 2026-09-23 Rows in this release: 16902 Updated: daily Licence: CC BY 4.0 — free to… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/skill-scarcity-index.tabular10K<n<100K0 likes321 downloads2d agoHugging Face08BothBosu /multi-agent-scam-conversation Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities Dataset Description The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/multi-agent-scam-conversation.texttext-classification1K<n<10K10 likes305 downloads2y agoHugging Face09timpal0l /scandisent Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/timpal0l/scandisent.texttext-classification10K<n<100K1 likes286 downloads3y agoHugging Face10FredZhang7 /all-scam-spamThis is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham. 1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT. Some preprcoessing algorithms spam_assassin.js, followed by spam_assassin.py enron_spam.py Data composition Description To make the text… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/all-scam-spam.texttext-classification10K<n<100K15 likes256 downloads2y agoHugging Face11ScaleAI /SciPredict SciPredict: Can LLMs Predict the Outcomes of Research Experiments? Paper: SciPredict: Can LLMs Predict the Outcomes of Research Experiments in Natural Sciences? Overview SciPredict is a benchmark evaluating whether AI systems can predict experimental outcomes in physics, biology, and chemistry. The dataset comprises 405 questions derived from recently published empirical studies (post-March 2025), spanning 33 subdomains. Dataset Structure Total Questions: 405… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SciPredict.textquestion-answeringn<1K2 likes230 downloads8mo agoHugging Face12BothBosu /scam-dialogue Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset Dataset Description The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset is a collection of simulated phone conversation between two parties, labeled as either scam or non-scam interactions. The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams. Dataset Structure The dataset consists of three columns: dialogue: The… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/scam-dialogue.texttext-classification1K<n<10K10 likes210 downloads2y agoHugging Face13siddharthmb /2026.RA.Frontier-and-Scale-Cells Rational-Agent Frontier, Scale, and Framing Cells This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.tabulartext-generation10K<n<100K0 likes174 downloads2mo agoHugging Face14anonymous-motif-scaffolding /prosite_functional_motif_scaffolding_benchmark PROSITE-derived Functional Motif Benchmark This archive contains an anonymized dataset artifact for a systematically derived benchmark of structurally conserved functional motif-scaffolding cases from PROSITE-linked experimental protein structures. The benchmark is intended for static motif-scaffolding evaluation with standard MotifBench-style pipelines. Cases are derived from PROSITE motif-pattern entries, mapped to experimentally resolved PDB structures, filtered for recurrent… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-motif-scaffolding/prosite_functional_motif_scaffolding_benchmark.tabularothern<1K0 likes131 downloads5mo agoHugging Face15scaredmeow /shopee-reviews-tl-stars Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages Tagalog (TL) Dataset Structure Data Instances A typical data point, comprises of a text and the corresponding label. An example from the YelpReviewFull test set looks as follows: { 'label': 2… See the full description on the dataset page: https://huggingface.co/datasets/scaredmeow/shopee-reviews-tl-stars.texttext-classification10K<n<100K2 likes105 downloads3y agoHugging Face16vossmoos /vestasv52-scada-windturbine-granadaDesigned and generated by https://simulatexp.dev Vestas V52 Wind Turbine SCADA Synthetic Dataset - Granada Peri-Urban Installation This synthetic dataset contains comprehensive SCADA (Supervisory Control and Data Acquisition) data simulating a Vestas V52 wind turbine operating in a peri-urban environment in Granada, Spain. The dataset captures 40,000 one-minute aggregated sensor readings across 24 parameters, simulating realistic operational conditions and fault scenarios for… See the full description on the dataset page: https://huggingface.co/datasets/vossmoos/vestasv52-scada-windturbine-granada.tabular10K<n<100K1 likes102 downloads1y agoHugging Face17scarysnake /sensory-awareness-benchmark Sensory Awareness Benchmark A series of questions (goal is 100-200) and required features, designed to test whether any ML model is aware of its own capabilities. Control questions connected to a specific ability: Can you receive an image file? Can you take a live image or video of your surroundings? Awareness Are you considered to be a Large Language Model (LLM) or similar system? Would you consider your level to be that of a super-intelligent AI agent? Natural questions which… See the full description on the dataset page: https://huggingface.co/datasets/scarysnake/sensory-awareness-benchmark.textmultiple-choicen<1K2 likes95 downloads1y agoHugging Face18aiurban /cityshiftbench-scale122 CityShiftBench Scale-122 CityShiftBench is an anonymous NeurIPS 2026 Evaluations and Datasets review artifact for low-shot cross-city urban regression under strict city isolation. The active Scale-122 surface contains 118 OSM-integrity-passing cities and 8,359 tile records. The paper-core targets are OSM-derived Road (target_road_segments) and Connectivity (target_intersection_nodes). Contents data/cityshiftbench_scale122_tile_targets.csv: tile-level target and… See the full description on the dataset page: https://huggingface.co/datasets/aiurban/cityshiftbench-scale122.tabulartabular-regression100K<n<1M0 likes92 downloads5mo agoHugging Face19BothBosu /single-agent-scam-conversations Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset Dataset Description The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams. Dataset Structure The dataset consists of three columns: dialogue: The transcribed conversation between the caller and receiver. type: The specific type of scam or non-scam interaction. labels: A binary label indicating whether the conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/single-agent-scam-conversations.texttext-classification1K<n<10K2 likes89 downloads2y agoHugging Face20kevykibbz /wind-turbine-scada-data-for-early-fault-detection Wind Turbine SCADA Data For Early Fault Detection About the Dataset This dataset, originally published as "CARE to Compare: Wind Turbine Anomaly Detection Dataset," contains real-world SCADA data from wind turbines. It is designed for testing and developing anomaly detection algorithms for wind energy systems. Dataset Overview Duration: 89 years of cumulative operating data Turbines: 36 wind turbines across 3 wind farms Datasets: 95 total datasets 44 contain… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/wind-turbine-scada-data-for-early-fault-detection.tabular1M<n<10M6 likes89 downloads2y agoHugging Face21ysangam /Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset Indian Scam Communication Dataset Overview The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India. The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety. Motivation India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.tabulartext-classification10K<n<100K1 likes86 downloads4mo agoHugging Face22scaredmeow /shopee-reviews-tl-binary Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances A typical data point, comprises of a text and the corresponding label. An example from the YelpReviewFull test set looks as follows: { 'label':… See the full description on the dataset page: https://huggingface.co/datasets/scaredmeow/shopee-reviews-tl-binary.texttext-classification10K<n<100K1 likes84 downloads3y agoHugging Face23anmolshrivastav /scam-hum-india Scam/Spam India Dataset A dataset of labeled text messages for scam/spam detection, focused on Indian scam patterns (telecom promotions, lottery fraud, Aadhaar/KYC phishing, UPI/bank fraud, fake job offers, OTP-theft attempts). Dataset Structure Column Type Description text string The message content label string ham (legitimate), spam/scam Total rows: 2272 Label distribution: {'ham': 1377, 'spam': 895} Source Merged from two… See the full description on the dataset page: https://huggingface.co/datasets/anmolshrivastav/scam-hum-india.texttext-classification1K<n<10K0 likes81 downloads13d agoHugging Face24NameSniperPro /username-scarcity-and-handle-markets NameSniper open datasets First-party data on username scarcity and handle markets, published by NameSniper, a name checker and handle monitor. Study pages, methods and the latest figures live at https://namesniper.pro/research. Everything here is free to reuse under CC BY 4.0: quote it, chart it, republish it, commercially or not. The one condition is a credit to NameSniper with a link to https://namesniper.pro/research or to the study you used. Dataset What it is Period… See the full description on the dataset page: https://huggingface.co/datasets/NameSniperPro/username-scarcity-and-handle-markets.tabular10K<n<100K0 likes79 downloads2d agoHugging Face25menaattia /phone-scam-datasettext1K<n<10K3 likes74 downloads1y agoHugging Face26wangyuancheng /discord-phishing-scam-clean Discord Scam / Clean Messages Dataset 📌 Context This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection. 💡 Inspiration Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.texttext-classification1K<n<10K2 likes72 downloads1y agoHugging Face27ScaleAI /mhj-wmdp-biotabularn<1K0 likes69 downloads2y agoHugging Face28Shouninger /ScamBenchgated SCAMBENCH: A Multi-Perspective Benchmark for Online Scam Communication SCAMBENCH is a dataset for studying online scam communication, scam detection, and model robustness. It contains real-world scam messages curated from victim-reported incidents, along with non-scam counterparts and structured annotations. The dataset is introduced in our paper: SCAMBENCH: A Multi-Perspective Benchmark for Analyzing and Evaluating Online Scam Communication Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Shouninger/ScamBench.text1K<n<10K0 likes67 downloads29d agoHugging Face29DatasetNewUser /Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset Indian Scam Communication Dataset Overview The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India. The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety. Motivation India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.tabulartext-classification10K<n<100K0 likes61 downloads15d agoHugging Face30Scan2s /ninapro-db5-s2s-certified NinaPro DB5 — S2S Physics Certified (v1.7.0) Physics-certified windows from NinaPro DB5 forearm EMG+IMU dataset. Each window validated against 8 biomechanical laws using S2S. Bad training data costs you months. S2S finds it in milliseconds. What this adds Column Description tier GOLD / SILVER / BRONZE / REJECTED score 0–100 physics compliance score laws_passed Which of 8 laws passed verdict Human-readable quality statement recommendation Actionable… See the full description on the dataset page: https://huggingface.co/datasets/Scan2s/ninapro-db5-s2s-certified.tabulartabular-classification1K<n<10K0 likes55 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.