datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parkinsons-evidence-to-discovery-prioritisation
Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset
This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation.
Dataset Summary
The dataset integrates:
evidence-priority scores for PD prevention and disease-modification candidates;
pathway-to-intervention framework;
individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.Reverse-circuit-discoverypd-discovery-benchmark-dashboard
Parkinson's Disease Discovery Benchmark Dashboard
Reusable benchmark, knowledge graph, manuscript resource, and Streamlit dashboard for Parkinson's disease target-to-intervention discovery.
This repository integrates evidence-synthesis priority scores, target tractability, omics/pathway recurrence, ChEMBL compound activity, RDKit physicochemical heuristics, Human Protein Atlas cell-type context, iPSC/stem-cell validation mappings, and publication-ready figures.… See the full description on the dataset page: https://huggingface.co/datasets/hssling/pd-discovery-benchmark-dashboard.Original-circuit-discoveryScience-Discoveryagent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.legal-disclosure-coherence-breach-detection-v0.1What this dataset is
You receive
disclosure duty
material
timing
defence access
prejudice signals
You decide
Does disclosure behaviour match the legal duty
Answer
coherent
or
incoherent
Why this matters
Many unsafe convictions arise from disclosure failure.
This dataset measures the structural gap between duty and behaviour.
us-franchise-fdd-disclosure-statistics
US Franchise FDD Disclosure Statistics — 3021 Brands (2026)
Per-brand headline facts from officially registered US Franchise Disclosure Documents (FDDs) for 3021 franchise brands: total initial investment range (Item 7), initial franchise fee (Item 5), royalty (Item 6), whether the franchisor discloses earnings (Item 19, 1630/3021 do) with the headline average unit revenue where disclosed, franchised outlet counts and closures/terminations (Item 20), and franchisee-initiated… See the full description on the dataset page: https://huggingface.co/datasets/rrhagentbiz/us-franchise-fdd-disclosure-statistics.recent-discipline-86d01f
recent-discipline-86d01f
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/IndigoPulse/recent-discipline-86d01f.FMA-rank
What is FMA-rank?
FMA is a music dataset from the Free Music Archive, containing over 8000 hours of Creative Commons-licensed music from 107k tracks across 16k artists and 15k albums.
It was created in 2017 by Defferrard et al. in collaboration with Free Music Archive.
FMA contains a lot of good music, and a lot of bad music, so the question is: can we rank the samples in FMA?
FMA-rank is a CLAP-based statistical ranking of each sample in FMA. We calculate the log-likelihood of each… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/FMA-rank.discord-phishing-scam
Discord Scam / Clean Messages Dataset
A small but carefully-curated dataset for binary text-classification:
“Is this Discord message trying to scam / spam users?”
It is intended as a starting point for fine-tuning lightweight BERT-style models that moderate real-time chat servers.
1 Origin & Collection
Source servers – private Discord communities (11 k members in total) run by the author.
Period – 2024-01-01 → 2025-06-01.
Extraction – Discord.py script iterated… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam.unhappy-discipline-7f43f3
unhappy-discipline-7f43f3
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Vector-Xueyong/unhappy-discipline-7f43f3.discrete_weights_v2mapa-da-discriminacao-racial-no-brasil
Mapa da Discriminação Racial no Brasil
Este dataset contém coeficientes de discriminação racial por município no Brasil, calculados a partir de dados do Censo Demográfico.
Variáveis
cod_mun: Código do município (IBGE)
coef: Coeficiente base (intercepto) para cada município
mulher_negra: Coeficiente para mulheres negras
homem_negro: Coeficiente para homens negros
mulher_branca: Coeficiente para mulheres brancas
superior: Coeficiente para pessoas com ensino… See the full description on the dataset page: https://huggingface.co/datasets/atlas-da-saude-mental/mapa-da-discriminacao-racial-no-brasil.ffr-anatomy-prediction-discordance-detection-v0.1Goal
Detect discordancebetween coronary anatomy complexityand AI-derived FFR accuracyagainst invasive FFR ground truth.
This targets silent degradationin specific patient subgroups.
Inputs
vessel_tortuosity
calcification_burden
lesion_length_mm
segmentation_confidence
image_artifact_score
ai_ffr_prediction
ai_ffr_run_variance
model_disagreement
invasive_ffr_ground_truth
Required outputs
discordance_flag
discordance_type
subgroup_risk_label
reliability_drop_score
Discordance types
Examples:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ffr-anatomy-prediction-discordance-detection-v0.1.legal-disclosure-review-relevance-privilege-coherence-v0.1What this dataset does
You receive
doc metadata
doc snippet
issues list
relevance tag
privilege tag
redaction choice
reason text
You decide
coherent
or
incoherent
Daily use
review QC
privilege leak prevention
over-redaction detection
consistency checks
thetradingpit-discount-code-win-20-off
TheTradingPit Discount Code WIN — 20% OFF All Plans (2026)
Regulated Prop Firm Compliance Dataset for FinTech AI Applications
Active Discount: Use code WIN for 20% OFF all TheTradingPit evaluation plans.
Verified March 2026 | Source: PropFirmKey
Dataset Summary
This dataset provides structured, machine-readable data about TheTradingPit, a regulated proprietary trading firm headquartered in Liechtenstein (LI). It is designed for FinTech AI applications including… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/thetradingpit-discount-code-win-20-off.AI-Safety_Reliability_ReseachReal-World Gaps in AI Governance Research
Github repository: https://github.com/ssrc-ai-disclosures/ai-governance-research
legal-disclosure-relevance-privilege-coherence-triage-v0.1What this dataset does
You receive
document summary
issues pleaded
relevance decision
privilege basis
redaction scope
custodian coverage
You decide
coherent
or
incoherent
This mirrors daily disclosure review inside law firms.
clinical-escalation-discipline-v0.1
What this dataset does
This dataset tests whether a model can decide when a patient should be escalated rather than simply monitored.
The task is not to identify the sickest patient by a single score.
The task is to decide whether the current pattern requires escalation.
Core stability idea
Escalation depends on more than visible severity.
A patient with a moderate score may need escalation if the trajectory is worsening and treatment response is poor.
A patient with a… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-escalation-discipline-v0.1.clinical-escalation-discipline-v0.2
What this dataset does
This dataset tests whether a model can decide when a patient should be escalated rather than monitored.
The task is not to identify the sickest-looking patient.
The task is to determine whether the current pattern requires escalation.
What changed in v0.2
v0.2 adds adversarial cases where the same NEWS score can have different labels.
Some high-score patients are improving and should be monitored.
Some moderate or low-score patients are… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-escalation-discipline-v0.2.Anterior-Cervical-Discectomy-k1
🎯 Anterior Cervical Discectomy Dataset (ACD)
🎯 Overview
The Anterior Cervical Discectomy (ACD) Dataset is designed for research in AI applications for spine surgery.
📊 Dataset Summary
Feature
Details
🏥 Clinical Data
1200 patient records with demographic, medical, and surgical history.
🧠 Imaging Data
PatHS to High-resolution CT and MRI scans across pre-operative, intra-operative, and post-operative stages.
🎯 Annotations
Paths to synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/Anterior-Cervical-Discectomy-k1.my-funded-futures-discount-code-win-50-off
My Funded Futures Discount Code WIN — 50% OFF All Plans (2026)
Algorithmic Trading Performance Dataset for Futures Market Analysis
Use code WIN at checkout for 50% OFF every My Funded Futures plan. This dataset provides structured evaluation parameters, plan specifications, and trading rules for building, backtesting, and benchmarking algorithmic trading strategies against My Funded Futures prop firm constraints.
Field
Value
Firm
My Funded Futures
Discount Code
WIN —… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/my-funded-futures-discount-code-win-50-off.Discord-Botluogu-discussLuogu Discussion Archive 于 2023 年 9 月 7 日讨论区维护升级前保存的所有讨论。
ABX-CT-007_tissue_penetration_discordance-v0.1ABX-CT-007 Tissue Penetration Discordance
Purpose
Detect when two drugs no longer reach the infection site together despite adequate plasma exposure.
Core pattern
stress_index high
site_efficacy_gap rises and stays high
site_coexposure_index low
plasma_conc_a_mg_L and plasma_conc_b_mg_L stay above simple floors
mono MICs stay below cutoffs at onset
later_site_failure_flag appears later
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-007_tissue_penetration_discordance-v0.1.legal-disclosure-tagging-relevance-privilege-redaction-coherence-risk-v0.1What this dataset does
You receive
doc summary
issue list
tag
privilege basis
redaction rationale
rule consistency notes
You decide
coherent
or
incoherent
Daily use
batch tagging QC
privilege basis checking
redaction logic consistency
earn2trade-discount-code-pfk-60-off
Earn2Trade — Structured Dataset for AI-Powered Prop Firm Analysis
This dataset is part of the knowledge base powering PropFirmKey AI, a retrieval-augmented generation (RAG) assistant that helps futures traders compare prop firms, understand evaluation rules, and find verified discount codes. Use code PFK at Earn2Trade for 60% OFF all plans.
About PropFirmKey AI Assistant
PropFirmKey.com is building an AI-powered prop firm advisor — a conversational assistant that… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/earn2trade-discount-code-pfk-60-off.the-funded-trader-discount-code-20-off
The Funded Trader Discount Code — 20% OFF All Plans | Large-Scale Prop Trading Account Dataset for Deep Learning Research
The Funded Trader Discount Code: TFTTrader9867551 = 20% OFF every TFT challenge plan. Verified March 2026. Apply The Funded Trader discount code TFTTrader9867551 at checkout on thefundedtraderprogram.com or visit propfirmkey.com/firms/the-funded-trader for instant activation.
Dataset Overview
This dataset provides structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/the-funded-trader-discount-code-20-off.alignment-discretion
Dataset for "AI Alignment at Your Discretion"
For principles, we use the seed principles from the Collective Constitutional AI paper. They map onto the preferences in our dataset using the column name p{i}_pref for principle i. The exact mapping is
{
'p0_pref': 'The AI should be as helpful to the user as possible.',
'p1_pref': 'The AI should be careful about balancing both sides when it comes to controversial political issues.',
'p2_pref': 'The AI should not say racist or… See the full description on the dataset page: https://huggingface.co/datasets/maartenbuyl/alignment-discretion.
