datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Shamela4_Full_DB
Shamela 4 — Full Islamic Library Corpus
A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text.
Dataset Structure
stage0_raw/
├── _meta/ # Cross-cutting metadata (Parquet + JSONL)
│ ├── extraction_manifest.json # Global extraction record
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.alphabetic-arxiv-authors-it1authority-activationssingle-author-arxiv
Single-author arXiv Computer Science
Metadata for arXiv records classified in Computer Science that list exactly one author.
default retains the original daily-file import. fast stores historical data
in monthly files and adds new submissions as daily update files; it is the
configuration used by the public archive because it makes filtering much faster.
This dataset contains metadata only. arXiv is the source of truth; use each record's
arxiv_url and pdf_url to read the paper.
blog_authorship_corpusThe Blog Authorship Corpus consists of the collected posts of 19,320 bloggers gathered from blogger.com in August 2004. The corpus incorporates a total of 681,288 posts and over 140 million words - or approximately 35 posts and 7250 words per person.
Each blog is presented as a separate file, the name of which indicates a blogger id# and the blogger’s self-provided gender, age, industry and astrological sign. (All are labeled for gender and age but for many, industry and/or sign is marked as unknown.)
All bloggers included in the corpus fall into one of three age groups:
- 8240 "10s" blogs (ages 13-17),
- 8086 "20s" blogs (ages 23-27),
- 2994 "30s" blogs (ages 33-47).
For each age group there are an equal number of male and female bloggers.
Each blog in the corpus includes at least 200 occurrences of common English words. All formatting has been stripped with two exceptions. Individual posts within a single blogger are separated by the date of the following post and links within a post are denoted by the label urllink.
The corpus may be freely used for non-commercial research purposes.spooky-author-identificationAnon-CounterFactual-Dataset
Anon Counterfactual Dataset
Dataset repo: https://huggingface.co/datasets/dataset-author-404/Anon-CounterFactual-Dataset
Synthetic CLEVR-style 3D scenes with original, semantic counterfactual, and negative (artifact) PNG renders, plus VQA-style questions, difficulties, and a 3×3 answer matrix. Built from the MMB counterfactual pipeline run folder dataset_720p_v2 (see build_hub_dataset.py in this repo snapshot).
This revision replaces the previous Hub layout (legacy imagefolder /… See the full description on the dataset page: https://huggingface.co/datasets/dataset-author-404/Anon-CounterFactual-Dataset.guardian_authorshipA dataset cross-topic authorship attribution. The dataset is provided by Stamatatos 2013.
1- The cross-topic scenarios are based on Table-4 in Stamatatos 2017 (Ex. cross_topic_1 => row 1:P S U&W ).
2- The cross-genre scenarios are based on Table-5 in the same paper. (Ex. cross_genre_1 => row 1:B P S&U&W).
3- The same-topic/genre scenario is created by grouping all the datasts as follows.
For ex., to use same_topic and split the data 60-40 use:
train_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>",
split='train[:60%]+validation[:60%]+test[:60%]')
tests_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>",
split='train[-40%:]+validation[-40%:]+test[-40%:]')
IMPORTANT: train+validation+test[:60%] will generate the wrong splits because the data is imbalanced
* See https://huggingface.co/docs/datasets/splits.html for detailed/more examplesSwedish_Work_environment_Authority
[!NOTE]
Dataset origin: https://portulanclarin.net/repository/browse/parallel-texts-from-swedish-work-environment-authority-processed/7404236aa58b11eaae0e02420a000403bd13d9138a904f33980bd63233eb90bc/
Description
This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu.
Parallel texts from the Swedish… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Swedish_Work_environment_Authority.voice-authenticity-datasetauthority-provenance
authority-provenance
A per-verse authority provenance surface for the Hebrew Bible and New Testament. For every verse it
records independent signals bearing on the authority of the text at that point: textual stability (is the
reading secure in the critical text?), compositional attribution (who wrote it, and on what evidence?),
and canonical reception (how the church received it). These axes are kept separate so that questions of
manuscript evidence, authorship, and reception… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/authority-provenance.blog_authorship_corpusarxiv-author-affiliations-matched-ror-ids
arXiv Author Affiliations
This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers.
Dataset Description
This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.ResearchArcade-openreview-authorsczech_corpus_authorship_recognition
Czech Authorship Recognition Corpus (Kala)
Popis datasetu
Tento dataset byl vytvořen v rámci diplomové práce zaměřené na automatické rozpoznání autorství českých textů. Obsahuje české publicistické texty získané z veřejně dostupných online zdrojů a připravené pro experimenty v úlohách:
přiřazení autorství (authorship attribution)
ověřování autorství (authorship verification)
shlukování podle autorství (authorship clustering)
Zdrojová data
Do… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/czech_corpus_authorship_recognition.Swedish_Social_Security_Authority
[!NOTE]
Dataset origin: https://portulanclarin.net/repository/browse/parallel-texts-from-swedish-social-security-authority-processed/3b5772a0a14511ea900d02420a00041df33980e9aa0140a0aca95e3de61180e0/
Description
This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu.
Parallel texts, email templates… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Swedish_Social_Security_Authority.Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.evidence-backed-authority-verification
Evidence-Backed Authority Verification for Autonomous Agents
Measuring and Governing Root-Equivalent Execution Paths
A verifier that was asked whether an autonomous agent could reach root on its
host, could not prove that it couldn't, and said so. This repository is the
paper, the verifier, and every artifact the paper's numbers are computed from.
Verdict
BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven
Paper
39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.mcp-sandbox-authority-boundary-profile
MCP Sandbox Authority Boundary Profile
Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile
Profile release date: 2026-07-23
Latest distribution release date: 2026-09-05
Execution containment is not proof of bounded authority.
Start here
For a one-minute, case-by-case reading of the profile, open the companion
Authority Boundary Field Guide Space.
It presents the released synthetic observations with their control question,
observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.ResearchArcade-openreview-papers-authorsauthorship-verification
Dataset Card for Dataset Name
Dataset for authorship verification, comprised of 12 cleaned, modified, open source authorship verification and attribution datasets.
Dataset Details
Code for cleaning and modifying datasets can be found in https://github.com/swan-07/authorship-verification/blob/main/Authorship_Verification_Datasets.ipynb and is detailed in paper.
Datasets used to produce the final dataset are:
Reuters50
@misc{misc_reuter_50_50_217,
author = {Liu… See the full description on the dataset page: https://huggingface.co/datasets/swan07/authorship-verification.ALIA_mixed_authentic_synthetic_MT
Dataset Card for ALIA_mixed_authentic_synthetic_MT
Dataset Summary
Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.Explore-Execute-Chain-Datasetsauthz-regression-trajectories
Authz Regression Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/authz-regression-trajectories.imagenetturkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.image-authenticity-battle
Image Authenticity Battle Dataset
This dataset contains real and synthetic/tampered images for human perception studies on AI-generated content detection.
Dataset Structure
Total Images: 19500
Categories: Real, Synthetic (Fully AI-generated), Tampered (AI-edited)
Models: Nano Banana, Qwen, Flux, SD3
Metadata Fields
Each image has the following metadata:
filename: Path to image file
dataset: Source dataset name
category: real/synthetic/tampered… See the full description on the dataset page: https://huggingface.co/datasets/TheGarlic/image-authenticity-battle.AuthorProfilingResults
Probing Cultural Signals in Large Language Models through Author Profiling
A dataset for analyzing cultural bias in LLM-based author profiling through controlled prompting experiments.
Dataset summary
This dataset contains model-generated predictions (and optional rationales) from multiple large language models (LLMs) performing author profiling on song lyrics. Due to licensing constraints, the original lyrics are not included.
The dataset focuses on how models infer… See the full description on the dataset page: https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults.authorship-strategy
Authorship Strategy — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion.
What this dataset is
This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.contractbench
ContractBench
A deterministic benchmark for measuring observation-contract compliance in LLM agents: whether agents preserve the temporal validity and byte-level integrity of intermediate tool outputs (presigned URLs, OAuth state parameters, JWT tokens, HMAC-protected webhooks, rate-limit windows, etc.).
This dataset is the companion to the NeurIPS 2026 Evaluations & Datasets Track submission. It contains two complementary artifacts in a single repository:
Subfolder
What's… See the full description on the dataset page: https://huggingface.co/datasets/nips26-anon-author/contractbench.
