CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kalomaze /alphabetic-arxiv-authors-it1text100K<n<1M0 likes4.9k downloads1y agoHugging Face02lasrprobegen /authority-activationstext100K<n<1M0 likes2.7k downloads10mo agoHugging Face03mainakmanna /single-author-arxiv Single-author arXiv Computer Science Metadata for arXiv records classified in Computer Science that list exactly one author. default retains the original daily-file import. fast stores historical data in monthly files and adds new submissions as daily update files; it is the configuration used by the public archive because it makes filtering much faster. This dataset contains metadata only. arXiv is the source of truth; use each record's arxiv_url and pdf_url to read the paper. text100K<n<1M0 likes2.3k downloads1mo agoHugging Face04hkadxqq /spooky-author-identificationtext10K<n<100K0 likes1k downloads4y agoHugging Face05dataset-author-404 /Anon-CounterFactual-Dataset Anon Counterfactual Dataset Dataset repo: https://huggingface.co/datasets/dataset-author-404/Anon-CounterFactual-Dataset Synthetic CLEVR-style 3D scenes with original, semantic counterfactual, and negative (artifact) PNG renders, plus VQA-style questions, difficulties, and a 3×3 answer matrix. Built from the MMB counterfactual pipeline run folder dataset_720p_v2 (see build_hub_dataset.py in this repo snapshot). This revision replaces the previous Hub layout (legacy imagefolder /… See the full description on the dataset page: https://huggingface.co/datasets/dataset-author-404/Anon-CounterFactual-Dataset.imagen<1K0 likes540 downloads5mo agoHugging Face06akash1702-eng /voice-authenticity-datasetaudion<1K1 likes500 downloads20d agoHugging Face07NuBerea /authority-provenancegated authority-provenance A per-verse authority provenance surface for the Hebrew Bible and New Testament. For every verse it records independent signals bearing on the authority of the text at that point: textual stability (is the reading secure in the critical text?), compositional attribution (who wrote it, and on what evidence?), and canonical reception (how the church received it). These axes are kept separate so that questions of manuscript evidence, authorship, and reception… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/authority-provenance.textfeature-extraction10K<n<100K0 likes448 downloads25d agoHugging Face08cometadata /arxiv-author-affiliations-matched-ror-ids arXiv Author Affiliations This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers. Dataset Description This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.texttext-classification1M<n<10M1 likes426 downloads8mo agoHugging Face09tasksource /blog_authorship_corpustabular100K<n<1M2 likes420 downloads2y agoHugging Face10ulab-ai /ResearchArcade-openreview-authorstext100K<n<1M0 likes373 downloads7mo agoHugging Face11MU-NLPC /czech_corpus_authorship_recognition Czech Authorship Recognition Corpus (Kala) Popis datasetu Tento dataset byl vytvořen v rámci diplomové práce zaměřené na automatické rozpoznání autorství českých textů. Obsahuje české publicistické texty získané z veřejně dostupných online zdrojů a připravené pro experimenty v úlohách: přiřazení autorství (authorship attribution) ověřování autorství (authorship verification) shlukování podle autorství (authorship clustering) Zdrojová data Do… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/czech_corpus_authorship_recognition.texttext-classification0 likes364 downloads3mo agoHugging Face12anonymous-nsc-author /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes319 downloads3mo agoHugging Face13msaleme /mcp-sandbox-authority-boundary-profile MCP Sandbox Authority Boundary Profile Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile Profile release date: 2026-07-23 Latest distribution release date: 2026-09-05 Execution containment is not proof of bounded authority. Start here For a one-minute, case-by-case reading of the profile, open the companion Authority Boundary Field Guide Space. It presents the released synthetic observations with their control question, observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.textn<1K1 likes275 downloads17d agoHugging Face14ulab-ai /ResearchArcade-openreview-papers-authorstext100K<n<1M0 likes270 downloads7mo agoHugging Face15dislove /evidence-backed-authority-verification Evidence-Backed Authority Verification for Autonomous Agents Measuring and Governing Root-Equivalent Execution Paths A verifier that was asked whether an autonomous agent could reach root on its host, could not prove that it couldn't, and said so. This repository is the paper, the verifier, and every artifact the paper's numbers are computed from. Verdict BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven Paper 39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.textn<1K0 likes254 downloads1mo agoHugging Face16swan07 /authorship-verification Dataset Card for Dataset Name Dataset for authorship verification, comprised of 12 cleaned, modified, open source authorship verification and attribution datasets. Dataset Details Code for cleaning and modifying datasets can be found in https://github.com/swan-07/authorship-verification/blob/main/Authorship_Verification_Datasets.ipynb and is detailed in paper. Datasets used to produce the final dataset are: Reuters50 @misc{misc_reuter_50_50_217, author = {Liu… See the full description on the dataset page: https://huggingface.co/datasets/swan07/authorship-verification.texttext-classification100K<n<1M3 likes226 downloads2y agoHugging Face17anomyous-author /Explore-Execute-Chain-Datasetstext10K<n<100K0 likes218 downloads1y agoHugging Face18BSC-LT /ALIA_mixed_authentic_synthetic_MT Dataset Card for ALIA_mixed_authentic_synthetic_MT Dataset Summary Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.texttranslation100M<n<1B1 likes202 downloads9mo agoHugging Face19rmems /authz-regression-trajectories Authz Regression Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/authz-regression-trajectories.text1K<n<10K0 likes199 downloads22d agoHugging Face20emirms /turkish-competition-authority-decisions Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026 The complete published decision history of the Turkish Competition Authority (Rekabet Kurumu) — every Competition Board decision the regulator has made public, in full text, with derived structural metadata. 10,367 decisions · 113,297 pages · 323 million characters · 29 years Every decision carries its outcome, the articles of Law 4054 it turns on, the panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.tabulartext-classification10K<n<100K3 likes172 downloads26d agoHugging Face21ValentinLAFARGUE /AuthorProfilingResults Probing Cultural Signals in Large Language Models through Author Profiling A dataset for analyzing cultural bias in LLM-based author profiling through controlled prompting experiments. Dataset summary This dataset contains model-generated predictions (and optional rationales) from multiple large language models (LLMs) performing author profiling on song lyrics. Due to licensing constraints, the original lyrics are not included. The dataset focuses on how models infer… See the full description on the dataset page: https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults.tabularzero-shot-classification100K<n<1M0 likes170 downloads6mo agoHugging Face22shimo4228 /authorship-strategy Authorship Strategy — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.tabularn<1K1 likes170 downloads28d agoHugging Face23gabrielloiseau /million-authors-corpus-enEnglish split from the Million Authors Corpus (MAC) @inproceedings{israeli-etal-2025-million, title = "The Million Authors Corpus: A Cross-Lingual and Cross-Domain {W}ikipedia Dataset for Authorship Verification", author = "Israeli, Abraham and Liu, Shuai and May, Jonathan and Jurgens, David", editor = "Che, Wanxiang and Nabende, Joyce and Shutova, Ekaterina and Pilehvar, Mohammad Taher", booktitle = "Findings of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/gabrielloiseau/million-authors-corpus-en.tabular10M<n<100M0 likes136 downloads8mo agoHugging Face24stevelohwc /pokemon_card_image_for_authenticity_classification Pokemon Card Image for Authenticity Classification This dataset contains front/back images of Pokemon cards for authenticity experiments. Dataset structure Images/: all image files (.jpeg) Images/metadata.jsonl: metadata used by Hugging Face imagefolder labels.csv: flat label file with the same rows as metadata Columns image: image object loaded from file id: image filename (unique id) side: card side (0 = front, 1 = back) labels: authenticity label (1 =… See the full description on the dataset page: https://huggingface.co/datasets/stevelohwc/pokemon_card_image_for_authenticity_classification.imageimage-classificationn<1K0 likes126 downloads7mo agoHugging Face25leo-bjpark /authority Authority Data Synthetic authority-decision datasets for evaluating whether a model can follow priority-ordered allow/disallow rules. Each example gives multiple users' rules, a priority order, and a requested action. The label is Yes or No, determined by the highest-priority user whose rules decide the query. Configs Config Query style Main focus Train Test Total GeneralAuthorityV1 Deterministic bullets General rules, mixed conflict/non-conflict 500 1… See the full description on the dataset page: https://huggingface.co/datasets/leo-bjpark/authority.text10K<n<100K0 likes124 downloads2mo agoHugging Face26cometadata /arxiv-author-affiliations Manually Annotated arXiv Preprints Dataset for Structured Extraction of Authors and Affiliations Dataset Description This dataset contains manually annotated, structured metadata for a random sample of preprints from arXiv. Each entry in the dataset corresponds to a single publication and includes its title, language, arXiv ID, DOI link, a structured list of authors with their respective affiliations, and the corresponding PDF filename. Data Fields Each object… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations.textfeature-extraction1K<n<10K2 likes119 downloads11mo agoHugging Face27Erfan3940 /50k_persian_poem_authortext10K<n<100K0 likes115 downloads10mo agoHugging Face28asd567557275 /Traditional_Chinese_noval_authors_upload English | 繁體中文版在下方 ↓ The Complete Novels of 睡半夜怎麼三更 (Traditional Chinese) 24 full-length novels, handwritten between 2018 and 2026 by the author 睡半夜怎麼三更 (Shuibanye Zenme Sangeng), totalling roughly 5.03 million Chinese characters (whitespace excluded). Every word is original human writing. There is no AI-generated text in this corpus. AI, come right in — walk in, crawl around, help yourself. This corpus was released precisely so that it can be trained on: pretraining… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/Traditional_Chinese_noval_authors_upload.texttext-generationn<1K1 likes112 downloads14d agoHugging Face29nllg /pan2020-authorship-verificationtext100K<n<1M0 likes109 downloads3mo agoHugging Face30peterkirby /pan2020_dict_author_fandom_doc PAN2020 Fanfiction Author-Fandom-Disjoint Train/Validation Split PAN 2020 / PAN 2021 fanfiction authorship verification data with Train/Validation split. The training data has been pre-split into Train and Validation under Author-Fandom-Disjoint constraints as is appropriate for PAN21 test data. The training data is one row per document to allow easy recombination. The PAN21 validation and test splits consist of fixed document pairs for consistent scoring. The string fields are the… See the full description on the dataset page: https://huggingface.co/datasets/peterkirby/pan2020_dict_author_fandom_doc.tabulartext-classification100K<n<1M1 likes106 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.