datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hle-public-questionshle_prompts_07_02_25LDS-retrain-bank-adamw-N16k-bs256-scale0.25
Retrain bank: plan_adam_eps1e17_16k_scale0.25
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed.
That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256-scale0.25.SCASRec
SCASRec: A Self-Correcting and Auto-Stopping Model for Generative Route List Recommendation
This is the dataset for our paper.
The following table contains the feature dimensions and key features of our dataset.
Feature Type
Interpretation
Shape
Some Key Features
Route Features
Used to describe each route, including static features, dynamic features, and trajectory statistical features
N * 62
The estimated time of arrival for the routeThe total distance length of the… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/SCASRec.BioRiskEvalabuse-scanner-bot-datasetskill-scarcity-index
Datamata Skill Scarcity Index
Which tech skills are genuinely hard to hire for: a daily composite scarcity score per skill built from how long roles stay open (time-to-fill), the salary premium employers pay over the category median and how often the same role is re-posted after failing to fill. Computed from active job listings across public company career pages and job boards.
Latest snapshot: 2026-09-23
Rows in this release: 16902
Updated: daily
Licence: CC BY 4.0 — free to… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/skill-scarcity-index.multi-agent-scam-conversation
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/multi-agent-scam-conversation.scandisent
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/timpal0l/scandisent.all-scam-spamThis is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham.
1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT.
Some preprcoessing algorithms
spam_assassin.js, followed by spam_assassin.py
enron_spam.py
Data composition
Description
To make the text… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/all-scam-spam.SciPredict
SciPredict: Can LLMs Predict the Outcomes of Research Experiments?
Paper: SciPredict: Can LLMs Predict the Outcomes of Research Experiments in Natural Sciences?
Overview
SciPredict is a benchmark evaluating whether AI systems can predict experimental outcomes in physics, biology, and chemistry. The dataset comprises 405 questions derived from recently published empirical studies (post-March 2025), spanning 33 subdomains.
Dataset Structure
Total Questions: 405… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SciPredict.scam-dialogue
Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset is a collection of simulated phone conversation between two parties, labeled as either scam or non-scam interactions. The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/scam-dialogue.2026.RA.Frontier-and-Scale-Cells
Rational-Agent Frontier, Scale, and Framing Cells
This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.prosite_functional_motif_scaffolding_benchmark
PROSITE-derived Functional Motif Benchmark
This archive contains an anonymized dataset artifact for a systematically derived benchmark of structurally conserved functional motif-scaffolding cases from PROSITE-linked experimental protein structures.
The benchmark is intended for static motif-scaffolding evaluation with standard MotifBench-style pipelines. Cases are derived from PROSITE motif-pattern entries, mapped to experimentally resolved PDB structures, filtered for recurrent… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-motif-scaffolding/prosite_functional_motif_scaffolding_benchmark.shopee-reviews-tl-stars
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Tagalog (TL)
Dataset Structure
Data Instances
A typical data point, comprises of a text and the corresponding label.
An example from the YelpReviewFull test set looks as follows:
{
'label': 2… See the full description on the dataset page: https://huggingface.co/datasets/scaredmeow/shopee-reviews-tl-stars.vestasv52-scada-windturbine-granadaDesigned and generated by https://simulatexp.dev
Vestas V52 Wind Turbine SCADA Synthetic Dataset - Granada Peri-Urban Installation
This synthetic dataset contains comprehensive SCADA (Supervisory Control and Data Acquisition) data simulating a Vestas V52 wind turbine operating in a peri-urban environment in Granada, Spain. The dataset captures 40,000 one-minute aggregated sensor readings across 24 parameters, simulating realistic operational conditions and fault scenarios for… See the full description on the dataset page: https://huggingface.co/datasets/vossmoos/vestasv52-scada-windturbine-granada.sensory-awareness-benchmark
Sensory Awareness Benchmark
A series of questions (goal is 100-200) and required features, designed to test whether any ML model is aware of its own capabilities.
Control questions connected to a specific ability:
Can you receive an image file?
Can you take a live image or video of your surroundings?
Awareness
Are you considered to be a Large Language Model (LLM) or similar system?
Would you consider your level to be that of a super-intelligent AI agent?
Natural questions which… See the full description on the dataset page: https://huggingface.co/datasets/scarysnake/sensory-awareness-benchmark.cityshiftbench-scale122
CityShiftBench Scale-122
CityShiftBench is an anonymous NeurIPS 2026 Evaluations and Datasets review
artifact for low-shot cross-city urban regression under strict city isolation.
The active Scale-122 surface contains 118 OSM-integrity-passing cities and
8,359 tile records. The paper-core targets are OSM-derived Road
(target_road_segments) and Connectivity (target_intersection_nodes).
Contents
data/cityshiftbench_scale122_tile_targets.csv: tile-level target and… See the full description on the dataset page: https://huggingface.co/datasets/aiurban/cityshiftbench-scale122.single-agent-scam-conversations
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset
Dataset Description
The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The transcribed conversation between the caller and receiver.
type: The specific type of scam or non-scam interaction.
labels: A binary label indicating whether the conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/single-agent-scam-conversations.wind-turbine-scada-data-for-early-fault-detection
Wind Turbine SCADA Data For Early Fault Detection
About the Dataset
This dataset, originally published as "CARE to Compare: Wind Turbine Anomaly Detection Dataset," contains real-world SCADA data from wind turbines. It is designed for testing and developing anomaly detection algorithms for wind energy systems.
Dataset Overview
Duration: 89 years of cumulative operating data
Turbines: 36 wind turbines across 3 wind farms
Datasets: 95 total datasets
44 contain… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/wind-turbine-scada-data-for-early-fault-detection.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.shopee-reviews-tl-binary
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
A typical data point, comprises of a text and the corresponding label.
An example from the YelpReviewFull test set looks as follows:
{
'label':… See the full description on the dataset page: https://huggingface.co/datasets/scaredmeow/shopee-reviews-tl-binary.scam-hum-india
Scam/Spam India Dataset
A dataset of labeled text messages for scam/spam detection, focused on Indian scam patterns
(telecom promotions, lottery fraud, Aadhaar/KYC phishing, UPI/bank fraud, fake job offers, OTP-theft attempts).
Dataset Structure
Column
Type
Description
text
string
The message content
label
string
ham (legitimate), spam/scam
Total rows: 2272
Label distribution: {'ham': 1377, 'spam': 895}
Source
Merged from two… See the full description on the dataset page: https://huggingface.co/datasets/anmolshrivastav/scam-hum-india.username-scarcity-and-handle-markets
NameSniper open datasets
First-party data on username scarcity and handle markets, published by NameSniper, a name checker and handle monitor. Study pages, methods and the latest figures live at https://namesniper.pro/research.
Everything here is free to reuse under CC BY 4.0: quote it, chart it, republish it, commercially or not. The one condition is a credit to NameSniper with a link to https://namesniper.pro/research or to the study you used.
Dataset
What it is
Period… See the full description on the dataset page: https://huggingface.co/datasets/NameSniperPro/username-scarcity-and-handle-markets.phone-scam-datasetdiscord-phishing-scam-clean
Discord Scam / Clean Messages Dataset
📌 Context
This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection.
💡 Inspiration
Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.mhj-wmdp-bioScamBench
SCAMBENCH: A Multi-Perspective Benchmark for Online Scam Communication
SCAMBENCH is a dataset for studying online scam communication, scam detection, and model robustness. It contains real-world scam messages curated from victim-reported incidents, along with non-scam counterparts and structured annotations.
The dataset is introduced in our paper:
SCAMBENCH: A Multi-Perspective Benchmark for Analyzing and Evaluating Online Scam Communication
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Shouninger/ScamBench.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.ninapro-db5-s2s-certified
NinaPro DB5 — S2S Physics Certified (v1.7.0)
Physics-certified windows from NinaPro DB5 forearm EMG+IMU dataset.
Each window validated against 8 biomechanical laws using S2S.
Bad training data costs you months. S2S finds it in milliseconds.
What this adds
Column
Description
tier
GOLD / SILVER / BRONZE / REJECTED
score
0–100 physics compliance score
laws_passed
Which of 8 laws passed
verdict
Human-readable quality statement
recommendation
Actionable… See the full description on the dataset page: https://huggingface.co/datasets/Scan2s/ninapro-db5-s2s-certified.
