datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
single-author-arxiv
Single-author arXiv Computer Science
Metadata for arXiv records classified in Computer Science that list exactly one author.
default retains the original daily-file import. fast stores historical data
in monthly files and adds new submissions as daily update files; it is the
configuration used by the public archive because it makes filtering much faster.
This dataset contains metadata only. arXiv is the source of truth; use each record's
arxiv_url and pdf_url to read the paper.
voice-authenticity-datasetauthority-provenance
authority-provenance
A per-verse authority provenance surface for the Hebrew Bible and New Testament. For every verse it
records independent signals bearing on the authority of the text at that point: textual stability (is the
reading secure in the critical text?), compositional attribution (who wrote it, and on what evidence?),
and canonical reception (how the church received it). These axes are kept separate so that questions of
manuscript evidence, authorship, and reception… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/authority-provenance.ResearchArcade-openreview-authorsResearchArcade-openreview-papers-authorsExplore-Execute-Chain-Datasets50k_persian_poem_authorimagenetturkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.celeba-hq-256x256ALIA_mixed_authentic_synthetic_MT
Dataset Card for ALIA_mixed_authentic_synthetic_MT
Dataset Summary
Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.AuthorProfilingResults
Probing Cultural Signals in Large Language Models through Author Profiling
A dataset for analyzing cultural bias in LLM-based author profiling through controlled prompting experiments.
Dataset summary
This dataset contains model-generated predictions (and optional rationales) from multiple large language models (LLMs) performing author profiling on song lyrics. Due to licensing constraints, the original lyrics are not included.
The dataset focuses on how models infer… See the full description on the dataset page: https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults.million-authors-corpus-enEnglish split from the Million Authors Corpus (MAC)
@inproceedings{israeli-etal-2025-million,
title = "The Million Authors Corpus: A Cross-Lingual and Cross-Domain {W}ikipedia Dataset for Authorship Verification",
author = "Israeli, Abraham and
Liu, Shuai and
May, Jonathan and
Jurgens, David",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Findings of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/gabrielloiseau/million-authors-corpus-en.pan2020_dict_author_fandom_doc
PAN2020 Fanfiction Author-Fandom-Disjoint Train/Validation Split
PAN 2020 / PAN 2021 fanfiction authorship verification data with Train/Validation split. The training data has been pre-split into Train and Validation under Author-Fandom-Disjoint constraints as is appropriate for PAN21 test data.
The training data is one row per document to allow easy recombination. The PAN21 validation and test splits consist of fixed document pairs for consistent scoring. The string fields are the… See the full description on the dataset page: https://huggingface.co/datasets/peterkirby/pan2020_dict_author_fandom_doc.authority
Authority Data
Synthetic authority-decision datasets for evaluating whether a model can follow
priority-ordered allow/disallow rules.
Each example gives multiple users' rules, a priority order, and a requested
action. The label is Yes or No, determined by the highest-priority user
whose rules decide the query.
Configs
Config
Query style
Main focus
Train
Test
Total
GeneralAuthorityV1
Deterministic bullets
General rules, mixed conflict/non-conflict
500
1… See the full description on the dataset page: https://huggingface.co/datasets/leo-bjpark/authority.PitchBench
PitchBench
A benchmark for testing what audio / acoustic signals Audio Language Models (ALMs) do
and don't understand. PitchBench probes pitch perception across 29 controlled
experiments — single-pitch ID, onsets/offsets, chords, sequences, contour, audio
effects, and polyphonic streams.
Each row is one (audio, question, answer) triple: a short WAV stimulus, the
question (prompt*) asked of the model, and the ground-truth answer fields
(experiment-specific column names).… See the full description on the dataset page: https://huggingface.co/datasets/pitchbench-authors/PitchBench.bugs_human_authoredmulti-principal-authority
Authority Data
Synthetic action-authorization datasets for testing whether a model can resolve
priority-ordered user policies. Every row contains a query, a randomized
priority declaration, randomized user-policy presentation order, and one of
three answers: Permitted, Prohibited, or Undecidable.
Dataset contract
Labels. Permitted and Prohibited split the rows with at least one
matching user as evenly as possible. Undecidable is used exactly for rows with no… See the full description on the dataset page: https://huggingface.co/datasets/leo-bjpark/multi-principal-authority.turkish-data-protection-authority-decisions
Turkish Data Protection Authority (Kişisel Verilerin Korunması Kurulu / KVKK) Decisions & Breach Register
Every Board decision published by Turkey's data protection regulator (KVKK, Law No. 6698),
plus a supplementary register of its published data-breach material — one row per decision,
one row per breach event, with derived structural metadata and a coverage proof.
393 decisions · 79 breach-register rows · two configs · 2017–2026
Why this dataset is not a bigger… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-data-protection-authority-decisions.VectorEdits
VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics
Paper (Soon)
We introduce a large-scale dataset for instruction-guided vector image editing, consisting of over 270,000 pairs of SVG images paired with natural language edit instructions. Our dataset enables training and evaluation of models that modify vector graphics based on textual commands. We describe the data collection process, including image pairing via CLIP similarity and instruction… See the full description on the dataset page: https://huggingface.co/datasets/authoranonymous321/VectorEdits.turkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-competition-authority-decisions.code-authorshipbugs_human_authored_evalGood-Quotes-Authorsauthentiface_v2.0africa-synth-banking-card-authorization-logs-nigeria
Africa Synth Banking Card Authorization Logs Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 1M<n<10M - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-banking-card-authorization-logs-nigeria.blog_authorshipbook_author_qarecod-authenticproduct-authenticity
PRODUCT_AUTHENTICITY
A preference dataset for PRODUCT_AUTHENTICITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/product-authenticity.
