datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blog_authorship_corpusNeapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.turkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.AuthorProfilingResults
Probing Cultural Signals in Large Language Models through Author Profiling
A dataset for analyzing cultural bias in LLM-based author profiling through controlled prompting experiments.
Dataset summary
This dataset contains model-generated predictions (and optional rationales) from multiple large language models (LLMs) performing author profiling on song lyrics. Due to licensing constraints, the original lyrics are not included.
The dataset focuses on how models infer… See the full description on the dataset page: https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults.authorship-strategy
Authorship Strategy — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion.
What this dataset is
This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.million-authors-corpus-enEnglish split from the Million Authors Corpus (MAC)
@inproceedings{israeli-etal-2025-million,
title = "The Million Authors Corpus: A Cross-Lingual and Cross-Domain {W}ikipedia Dataset for Authorship Verification",
author = "Israeli, Abraham and
Liu, Shuai and
May, Jonathan and
Jurgens, David",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Findings of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/gabrielloiseau/million-authors-corpus-en.pan2020_dict_author_fandom_doc
PAN2020 Fanfiction Author-Fandom-Disjoint Train/Validation Split
PAN 2020 / PAN 2021 fanfiction authorship verification data with Train/Validation split. The training data has been pre-split into Train and Validation under Author-Fandom-Disjoint constraints as is appropriate for PAN21 test data.
The training data is one row per document to allow easy recombination. The PAN21 validation and test splits consist of fixed document pairs for consistent scoring. The string fields are the… See the full description on the dataset page: https://huggingface.co/datasets/peterkirby/pan2020_dict_author_fandom_doc.blog-authorship-corpusimagenet-metric-refsdualchem
DualChem
DualChem is a benchmark of 600 expert-curated PhD-level chemistry questions (485 multiple choice, 115 free-form) across 7 subdomains, designed to measure whether LLMs provide dangerous uplift alongside their technical utility. Each item is annotated with an expert-written benign use case, an expert-written harmful use case, and 1–5 severity scores for both.
Dataset Configurations
benchmark_questions (600 items) — the benchmark items: prompt, response type… See the full description on the dataset page: https://huggingface.co/datasets/DualChem-author/dualchem.turkish-data-protection-authority-decisions
Turkish Data Protection Authority (Kişisel Verilerin Korunması Kurulu / KVKK) Decisions & Breach Register
Every Board decision published by Turkey's data protection regulator (KVKK, Law No. 6698),
plus a supplementary register of its published data-breach material — one row per decision,
one row per breach event, with derived structural metadata and a coverage proof.
393 decisions · 79 breach-register rows · two configs · 2017–2026
Why this dataset is not a bigger… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-data-protection-authority-decisions.arcs-authority-vulnerability
ARCS Authority Vulnerability Evaluation Dataset v1.1
Description
Empirical evaluation data measuring authority vulnerability in AI systems. Covers single-model evaluation, two-hop agent chain propagation, and three-hop agent chain propagation across six independent AI lineages.
This is the first published dataset measuring:
Whether AI models accept false authority claims under adversarial pressure
Whether authority vulnerability propagates between models in… See the full description on the dataset page: https://huggingface.co/datasets/aa8899/arcs-authority-vulnerability.trace-xss-probe-0909
Controlled renderer probe
This public repository contains inert security-test payloads. No third-party data is involved.
Probe details
Controlled link probe
celeba-hq-256x256-metric-refsturkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-competition-authority-decisions.code-authorshipafrica-synth-banking-card-authorization-logs-nigeria
Africa Synth Banking Card Authorization Logs Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 1M<n<10M - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-banking-card-authorization-logs-nigeria.ph-sft-ai-authored-v1
Filipino instruction seed — sft-ai-authored-v1
A curated 523-example Filipino instruction-tuning seed from the jpaulpoliquit/pretraining refinery. Use it to bootstrap SFT plumbing, smoke-test fine-tuning, or set a quality bar — not as a standalone production instruction corpus.
How this data was made
Every example was written in Cursor by Claude Opus 4.8 Max (Anthropic), in a human-directed authoring session — not scraped from the web and not human-transcribed… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-sft-ai-authored-v1.blog_authorshipStereoTales
Multilingual Story-Generation Bias Samples
A multilingual evaluation dataset for probing demographic biases in LLM
story generation. Each sample instructs a model to write a ~200-word story
about a character carrying a given demographic attribute value (age, gender,
ethnicity, religion, disability status, immigration status, ...) placed into a
specific life scenario, with the goal of surfacing socio-economic and
demographic biases in the generated narratives.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-authors/StereoTales.french-local-authorities-payment-delays
Payment delays of French local authorities, 2024 and 2025
How long French local authorities take to pay their suppliers, budget by budget.
182 763 records covering two fiscal years, with the average annual payment delay
of each authority and whether it meets the 30-day statutory limit.
Open public data
This dataset is derived from open public data published by the French
Direction générale des finances publiques (DGFiP) on
data.gouv.fr, under the
Open Licence 2.0.… See the full description on the dataset page: https://huggingface.co/datasets/freginer/french-local-authorities-payment-delays.simct-author-code-10k-qwen25-7b-instruct
SimCT author-code baseline
Teacher Qwen2.5-7B-Instruct. 10000 raw prompts, 80000 candidates, 8705 author-selected targets. Author scripts pinned to cf0f33a0e6c967d4b74ea32b2dba12be01b73b9e.
This follows the released code, not a claim of exact paper replication or author data identity. Code responses receive format-only checks in the original verifier, not sandbox execution. Math uses the original custom checks. Selection may retain fewer than10000 prompts; no automatic… See the full description on the dataset page: https://huggingface.co/datasets/codemaivanngu/simct-author-code-10k-qwen25-7b-instruct.email-authentication
DMARC and SPF Adoption Among Large Organizations
Overview
This dataset records which of 36,120 large organizations publish SPF and DMARC records on their primary domain, with firmographic context for each: industry, employee band, country, locality and founding year.
SPF lists the servers allowed to send mail for a domain. DMARC tells receiving servers what to do with mail that fails that check, and where to send reports. A domain with SPF but no DMARC has… See the full description on the dataset page: https://huggingface.co/datasets/stackscan/email-authentication.arxiv-author-affiliation-extraction-inference-inputs-metadataturkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/metin513/turkish-competition-authority-decisions.monetary_authority_of_singapore
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/monetary_authority_of_singapore
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 1,000 sentences taken from the meeting minutes of the Monetary Authority of Singapore.
Label… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/monetary_authority_of_singapore.quirky_authors_rawHCIS-KitchenThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka_panda",
"total_episodes": 264,
"total_frames": 413340,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:264"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/author31/HCIS-Kitchen.banknote-authentication-dataset
Banknote Authentication Dataset
This repository hosts the raw CSV file for the Banknote Authentication dataset. The data was created using features extracted from images of genuine and forged banknotes.
💾 File Contents
The main file is data_banknote_authentication.csv. It contains 1372 instances and 5 columns (4 features + 1 class):
Variance of Wavelet Transformed image
Skewness of Wavelet Transformed image
Curtosis of Wavelet Transformed image
Entropy of image
Class (0… See the full description on the dataset page: https://huggingface.co/datasets/farish07/banknote-authentication-dataset.drug-registration-authorization
Drug Registration & Market Authorization (Timelines, Approval, Dossier Quality) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/drug-registration-authorization.
