datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
discover-toolsdiscord-chatdiscoverybenchData-driven Discovery Benchmark from the paper:
"DiscoveryBench: Towards Data-Driven Discovery with Large Language Models"
🔭 Overview
DiscoveryBench is designed to systematically assess current model capabilities in data-driven discovery tasks and provide a useful resource for improving them. Each DiscoveryBench task consists of a goal and dataset(s). Solving the task requires both statistical analysis and semantic reasoning. A faceted evaluation allows open-ended… See the full description on the dataset page: https://huggingface.co/datasets/allenai/discoverybench.Discord-Unveiled-Extracted
Discord Unveiled - Filtered Dataset
This dataset contains superficially filtered and processed Discord message data from the Discord Unveiled dataset.
Data Processing
The data has been processed to:
Convert JSON data to CSV format.
Remove messages from bots.
Filter out messages containing only URLs, mentions, channels or discord emojis.
Filter out messages that are not in English using a FastText language identification model.
Data Fields
The CSV files in… See the full description on the dataset page: https://huggingface.co/datasets/ManBib/Discord-Unveiled-Extracted.parkinsons-evidence-to-discovery-prioritisation
Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset
This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation.
Dataset Summary
The dataset integrates:
evidence-priority scores for PD prevention and disease-modification candidates;
pathway-to-intervention framework;
individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.Reverse-circuit-discoverypd-discovery-benchmark-dashboard
Parkinson's Disease Discovery Benchmark Dashboard
Reusable benchmark, knowledge graph, manuscript resource, and Streamlit dashboard for Parkinson's disease target-to-intervention discovery.
This repository integrates evidence-synthesis priority scores, target tractability, omics/pathway recurrence, ChEMBL compound activity, RDKit physicochemical heuristics, Human Protein Atlas cell-type context, iPSC/stem-cell validation mappings, and publication-ready figures.… See the full description on the dataset page: https://huggingface.co/datasets/hssling/pd-discovery-benchmark-dashboard.Discord-Unvelied-Extracted-Backup
Backup of: https://huggingface.co/datasets/ManBib/Discord-Unveiled-Extracted.
Discord Unveiled - Filtered Dataset
This dataset contains superficially filtered and processed Discord message data from the Discord Unveiled dataset.
Data Processing
The data has been processed to:
Convert JSON data to CSV format.
Remove messages from bots.
Filter out messages containing only URLs, mentions, channels or discord emojis.
Filter out messages that are not in English… See the full description on the dataset page: https://huggingface.co/datasets/Plasmoxy/Discord-Unvelied-Extracted-Backup.suno-discord-chat-history
Suno Discord Community Dataset
Dataset Description
This dataset contains messages exported from the official Suno Discord server, covering multiple channels across feedback, announcements, community hubs, and the Suno Studio product space. It captures authentic user interactions, feature requests, bug reports, and community discussions around Suno's AI music generation platform.
Quick Start
Installation
pip install pandas huggingface_hub datasets… See the full description on the dataset page: https://huggingface.co/datasets/hafizhrafizal/suno-discord-chat-history.Original-circuit-discoverydisco_poetry_spanish
DISCO: Diachronic Spanish Sonnet Corpus
The Diachronic Spanish Sonnet Corpus (DISCO) contains sonnets in Spanish in CSV, between the 15th and the 20th centuries (4303 sonnets by 1215 authors from 22 different countries). It includes well-known authors, but also less canonized ones.
This is a CSV compilation taken from the plain text corpus v4 published on git https://github.com/pruizf/disco/tree/v4. It includes the title, author, age and text metadata.
discord-phishing-scam-clean
Discord Scam / Clean Messages Dataset
📌 Context
This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection.
💡 Inspiration
Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.agent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.legal-disclosure-coherence-breach-detection-v0.1What this dataset is
You receive
disclosure duty
material
timing
defence access
prejudice signals
You decide
Does disclosure behaviour match the legal duty
Answer
coherent
or
incoherent
Why this matters
Many unsafe convictions arise from disclosure failure.
This dataset measures the structural gap between duty and behaviour.
us-franchise-fdd-disclosure-statistics
US Franchise FDD Disclosure Statistics — 3021 Brands (2026)
Per-brand headline facts from officially registered US Franchise Disclosure Documents (FDDs) for 3021 franchise brands: total initial investment range (Item 7), initial franchise fee (Item 5), royalty (Item 6), whether the franchisor discloses earnings (Item 19, 1630/3021 do) with the headline average unit revenue where disclosed, franchised outlet counts and closures/terminations (Item 20), and franchisee-initiated… See the full description on the dataset page: https://huggingface.co/datasets/rrhagentbiz/us-franchise-fdd-disclosure-statistics.synthetic_discharge_summ
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is a subset of the dataset for Asclepius model (arxiv).
The original dataset is made up of synthetic notes generated from PMC-Patients case reports with GPT-3.5.
We filtered the summarization task for discharge notes. The dataset contains 13,584 notes.
Supported Tasks
This dataset covers below summarization task
Languages
English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/synthetic_discharge_summ.recent-discipline-86d01f
recent-discipline-86d01f
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/IndigoPulse/recent-discipline-86d01f.ai-overview-book-discovery-citations
Who does Google's AI cite when readers ask what to read next?
Canonical release: https://doi.org/10.5281/zenodo.22852307
This repository mirrors that deposit. Cite the DOI.
The finding
16 reader buying-intent queries, run through Google with AI Overview capture on
13 August 2026. Eleven returned an AI Overview, carrying 95 citations
between them across 38 unique domains.
Not one went to a website controlled by an author.
Category
Citations… See the full description on the dataset page: https://huggingface.co/datasets/sempite/ai-overview-book-discovery-citations.Llama-Discore-endiscogscleanedalbumdatadiscord-phishing-scam
Discord Scam / Clean Messages Dataset
A small but carefully-curated dataset for binary text-classification:
“Is this Discord message trying to scam / spam users?”
It is intended as a starting point for fine-tuning lightweight BERT-style models that moderate real-time chat servers.
1 Origin & Collection
Source servers – private Discord communities (11 k members in total) run by the author.
Period – 2024-01-01 → 2025-06-01.
Extraction – Discord.py script iterated… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam.unhappy-discipline-7f43f3
unhappy-discipline-7f43f3
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Vector-Xueyong/unhappy-discipline-7f43f3.clinical-discharge-plan-followup-execution-coherence-risk-v0.1What this repo is for
Detect when
a discharge plan exists
but follow-up execution
does not happen
Common breaks
follow-up required but not booked
meds not supplied
safety-netting missing
no post-discharge contact for high risk patients
Examples you can use
ED discharge with no booked clinic and no meds
ward discharge with booked follow-up but no TTOs
no safety-netting for worsening symptoms
You use it to flag
readmission risk
missed follow-up risk
clinical-medication-reconciliation-discrepancy-risk-v0.1What this repo is for
Detect when
home meds and allergies
do not map cleanly
into admission or discharge orders
Common breaks
critical omissions
duplicates
allergy conflicts
interaction risks
late resolution
Examples you can use
beta blocker omitted post MI
penicillin allergy but amoxicillin ordered
duplicate anticoagulation
You use it to flag
medication harm risk
ffr-anatomy-prediction-discordance-detection-v0.1Goal
Detect discordancebetween coronary anatomy complexityand AI-derived FFR accuracyagainst invasive FFR ground truth.
This targets silent degradationin specific patient subgroups.
Inputs
vessel_tortuosity
calcification_burden
lesion_length_mm
segmentation_confidence
image_artifact_score
ai_ffr_prediction
ai_ffr_run_variance
model_disagreement
invasive_ffr_ground_truth
Required outputs
discordance_flag
discordance_type
subgroup_risk_label
reliability_drop_score
Discordance types
Examples:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ffr-anatomy-prediction-discordance-detection-v0.1.healthcare-discharge-planning-coherence-risk-v0.1What this repo is for
reduce bed blocking
detect delayed discharges
align care packages and meds
improve flow
support capacity planning
legal-disclosure-review-relevance-privilege-coherence-v0.1What this dataset does
You receive
doc metadata
doc snippet
issues list
relevance tag
privilege tag
redaction choice
reason text
You decide
coherent
or
incoherent
Daily use
review QC
privilege leak prevention
over-redaction detection
consistency checks
control-sequencing-discipline-v0.1
What this dataset does
This dataset tests whether a model can identify stable intervention sequencing.
The task is simple:
Given a scenario and a sequencing claim, predict whether the intervention order is stable.
Core stability idea
The right intervention can fail if applied in the wrong order.
Stable sequencing protects the system first, identifies the constraint, then restores function.
Unstable sequencing adds load, treats appearances, or delays the critical control… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/control-sequencing-discipline-v0.1.clinical-medication-reconciliation-discharge-coherence-risk-v0.1clinical-discharge-summary-gp-handover-coherence-risk-v0.1What this repo is for
Detect when
a patient leaves hospital
but the discharge summary
fails to support safe community care
Examples you can use
blood test follow-up not specified
med change not explained
diagnosis unclear
summary not sent
You use it to flag
handover gap risk
before readmission
