datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
discover-toolsdiscord-chatdiscoverybenchData-driven Discovery Benchmark from the paper:
"DiscoveryBench: Towards Data-Driven Discovery with Large Language Models"
🔭 Overview
DiscoveryBench is designed to systematically assess current model capabilities in data-driven discovery tasks and provide a useful resource for improving them. Each DiscoveryBench task consists of a goal and dataset(s). Solving the task requires both statistical analysis and semantic reasoning. A faceted evaluation allows open-ended… See the full description on the dataset page: https://huggingface.co/datasets/allenai/discoverybench.Discord-Unveiled-Extracted
Discord Unveiled - Filtered Dataset
This dataset contains superficially filtered and processed Discord message data from the Discord Unveiled dataset.
Data Processing
The data has been processed to:
Convert JSON data to CSV format.
Remove messages from bots.
Filter out messages containing only URLs, mentions, channels or discord emojis.
Filter out messages that are not in English using a FastText language identification model.
Data Fields
The CSV files in… See the full description on the dataset page: https://huggingface.co/datasets/ManBib/Discord-Unveiled-Extracted.parkinsons-evidence-to-discovery-prioritisation
Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset
This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation.
Dataset Summary
The dataset integrates:
evidence-priority scores for PD prevention and disease-modification candidates;
pathway-to-intervention framework;
individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.pd-discovery-benchmark-dashboard
Parkinson's Disease Discovery Benchmark Dashboard
Reusable benchmark, knowledge graph, manuscript resource, and Streamlit dashboard for Parkinson's disease target-to-intervention discovery.
This repository integrates evidence-synthesis priority scores, target tractability, omics/pathway recurrence, ChEMBL compound activity, RDKit physicochemical heuristics, Human Protein Atlas cell-type context, iPSC/stem-cell validation mappings, and publication-ready figures.… See the full description on the dataset page: https://huggingface.co/datasets/hssling/pd-discovery-benchmark-dashboard.Reverse-circuit-discoveryDiscord-Unvelied-Extracted-Backup
Backup of: https://huggingface.co/datasets/ManBib/Discord-Unveiled-Extracted.
Discord Unveiled - Filtered Dataset
This dataset contains superficially filtered and processed Discord message data from the Discord Unveiled dataset.
Data Processing
The data has been processed to:
Convert JSON data to CSV format.
Remove messages from bots.
Filter out messages containing only URLs, mentions, channels or discord emojis.
Filter out messages that are not in English… See the full description on the dataset page: https://huggingface.co/datasets/Plasmoxy/Discord-Unvelied-Extracted-Backup.disco_poetry_spanish
DISCO: Diachronic Spanish Sonnet Corpus
The Diachronic Spanish Sonnet Corpus (DISCO) contains sonnets in Spanish in CSV, between the 15th and the 20th centuries (4303 sonnets by 1215 authors from 22 different countries). It includes well-known authors, but also less canonized ones.
This is a CSV compilation taken from the plain text corpus v4 published on git https://github.com/pruizf/disco/tree/v4. It includes the title, author, age and text metadata.
Original-circuit-discoverysuno-discord-chat-history
Suno Discord Community Dataset
Dataset Description
This dataset contains messages exported from the official Suno Discord server, covering multiple channels across feedback, announcements, community hubs, and the Suno Studio product space. It captures authentic user interactions, feature requests, bug reports, and community discussions around Suno's AI music generation platform.
Quick Start
Installation
pip install pandas huggingface_hub datasets… See the full description on the dataset page: https://huggingface.co/datasets/hafizhrafizal/suno-discord-chat-history.discord-phishing-scam-clean
Discord Scam / Clean Messages Dataset
📌 Context
This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection.
💡 Inspiration
Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.agent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.Llama-Discore-enai-overview-book-discovery-citations
Who does Google's AI cite when readers ask what to read next?
Canonical release: https://doi.org/10.5281/zenodo.22852307
This repository mirrors that deposit. Cite the DOI.
The finding
16 reader buying-intent queries, run through Google with AI Overview capture on
13 August 2026. Eleven returned an AI Overview, carrying 95 citations
between them across 38 unique domains.
Not one went to a website controlled by an author.
Category
Citations… See the full description on the dataset page: https://huggingface.co/datasets/sempite/ai-overview-book-discovery-citations.discord-phishing-scam
Discord Scam / Clean Messages Dataset
A small but carefully-curated dataset for binary text-classification:
“Is this Discord message trying to scam / spam users?”
It is intended as a starting point for fine-tuning lightweight BERT-style models that moderate real-time chat servers.
1 Origin & Collection
Source servers – private Discord communities (11 k members in total) run by the author.
Period – 2024-01-01 → 2025-06-01.
Extraction – Discord.py script iterated… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam.discogscleanedalbumdataffr-anatomy-prediction-discordance-detection-v0.1Goal
Detect discordancebetween coronary anatomy complexityand AI-derived FFR accuracyagainst invasive FFR ground truth.
This targets silent degradationin specific patient subgroups.
Inputs
vessel_tortuosity
calcification_burden
lesion_length_mm
segmentation_confidence
image_artifact_score
ai_ffr_prediction
ai_ffr_run_variance
model_disagreement
invasive_ffr_ground_truth
Required outputs
discordance_flag
discordance_type
subgroup_risk_label
reliability_drop_score
Discordance types
Examples:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ffr-anatomy-prediction-discordance-detection-v0.1.thetradingpit-discount-code-win-20-off
TheTradingPit Discount Code WIN — 20% OFF All Plans (2026)
Regulated Prop Firm Compliance Dataset for FinTech AI Applications
Active Discount: Use code WIN for 20% OFF all TheTradingPit evaluation plans.
Verified March 2026 | Source: PropFirmKey
Dataset Summary
This dataset provides structured, machine-readable data about TheTradingPit, a regulated proprietary trading firm headquartered in Liechtenstein (LI). It is designed for FinTech AI applications including… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/thetradingpit-discount-code-win-20-off.DiSCo_Dataset_and_Benchmark
DiSCo Dataset & Benchmark
DiSCo (Distribution-first Steering and Context evaluation) is a cultural preference bias dataset and benchmark for evaluating and auditing large language models across diverse cultural contexts.
Overview
DiSCo-Dataset is a 150,816-row cultural preference dataset derived from BLEnD where every row presents a country-neutral lifestyle question with four answer options. Each option is paired with a one-line cultural fact string (Context_L1–L4) and… See the full description on the dataset page: https://huggingface.co/datasets/DiSCo2026/DiSCo_Dataset_and_Benchmark.disco_spanish_poetry
DISCO: Diachronic Spanish Sonnet Corpus
The Diachronic Spanish Sonnet Corpus (DISCO) contains sonnets in Spanish in CSV, between the 15th and the 20th centuries (4303 sonnets by 1215 authors from 22 different countries). It includes well-known authors, but also less canonized ones.
This is a CSV compilation taken from the plain text corpus v4 published on git https://github.com/pruizf/disco/tree/v4. It includes the title, author, age and text metadata.
weibo_unemployment_discourseearn2trade-discount-code-pfk-60-off
Earn2Trade — Structured Dataset for AI-Powered Prop Firm Analysis
This dataset is part of the knowledge base powering PropFirmKey AI, a retrieval-augmented generation (RAG) assistant that helps futures traders compare prop firms, understand evaluation rules, and find verified discount codes. Use code PFK at Earn2Trade for 60% OFF all plans.
About PropFirmKey AI Assistant
PropFirmKey.com is building an AI-powered prop firm advisor — a conversational assistant that… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/earn2trade-discount-code-pfk-60-off.Discord-Botmy-funded-futures-discount-code-win-50-off
My Funded Futures Discount Code WIN — 50% OFF All Plans (2026)
Algorithmic Trading Performance Dataset for Futures Market Analysis
Use code WIN at checkout for 50% OFF every My Funded Futures plan. This dataset provides structured evaluation parameters, plan specifications, and trading rules for building, backtesting, and benchmarking algorithmic trading strategies against My Funded Futures prop firm constraints.
Field
Value
Firm
My Funded Futures
Discount Code
WIN —… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/my-funded-futures-discount-code-win-50-off.ABX-CT-007_tissue_penetration_discordance-v0.1ABX-CT-007 Tissue Penetration Discordance
Purpose
Detect when two drugs no longer reach the infection site together despite adequate plasma exposure.
Core pattern
stress_index high
site_efficacy_gap rises and stays high
site_coexposure_index low
plasma_conc_a_mg_L and plasma_conc_b_mg_L stay above simple floors
mono MICs stay below cutoffs at onset
later_site_failure_flag appears later
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-007_tissue_penetration_discordance-v0.1.the-funded-trader-discount-code-20-off
The Funded Trader Discount Code — 20% OFF All Plans | Large-Scale Prop Trading Account Dataset for Deep Learning Research
The Funded Trader Discount Code: TFTTrader9867551 = 20% OFF every TFT challenge plan. Verified March 2026. Apply The Funded Trader discount code TFTTrader9867551 at checkout on thefundedtraderprogram.com or visit propfirmkey.com/firms/the-funded-trader for instant activation.
Dataset Overview
This dataset provides structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/the-funded-trader-discount-code-20-off.the5ers-discount-code-pfkey-20-off
The5ers Discount Code PFKEY — 20% OFF All Plans (2026)
Veteran Prop Firm Longitudinal Dataset for Trader Success Modeling
Use code PFKEY at checkout for 20% OFF every The5ers plan. Verified and active as of March 2026.
Apply PFKEY here: https://propfirmkey.com/firms/the5ers
Dataset Summary
This dataset provides structured, machine-readable data on The5ers, one of the oldest and most respected proprietary trading firms in the industry (founded in 2016… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/the5ers-discount-code-pfkey-20-off.discord_dataDiscoveriesTimeline
DiscoveriesTimeline
tags: regression, discovery_year, regional_impact
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'DiscoveriesTimeline' dataset is a curated collection of significant human discoveries throughout history, categorized by the nature of the discovery, the year it occurred, the region/global impact, and its labels which classify the type of impact and the field of discovery. This dataset can be useful for ML… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/DiscoveriesTimeline.
