CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tasksource /blog_authorship_corpustabular100K<n<1M2 likes428 downloads2y agoHugging Face02anonymous-nsc-author /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes314 downloads3mo agoHugging Face03emirms /turkish-competition-authority-decisions Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026 The complete published decision history of the Turkish Competition Authority (Rekabet Kurumu) — every Competition Board decision the regulator has made public, in full text, with derived structural metadata. 10,367 decisions · 113,297 pages · 323 million characters · 29 years Every decision carries its outcome, the articles of Law 4054 it turns on, the panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.tabulartext-classification10K<n<100K3 likes170 downloads26d agoHugging Face04ValentinLAFARGUE /AuthorProfilingResults Probing Cultural Signals in Large Language Models through Author Profiling A dataset for analyzing cultural bias in LLM-based author profiling through controlled prompting experiments. Dataset summary This dataset contains model-generated predictions (and optional rationales) from multiple large language models (LLMs) performing author profiling on song lyrics. Due to licensing constraints, the original lyrics are not included. The dataset focuses on how models infer… See the full description on the dataset page: https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults.tabularzero-shot-classification100K<n<1M0 likes166 downloads6mo agoHugging Face05shimo4228 /authorship-strategy Authorship Strategy — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.tabularn<1K1 likes159 downloads28d agoHugging Face06gabrielloiseau /million-authors-corpus-enEnglish split from the Million Authors Corpus (MAC) @inproceedings{israeli-etal-2025-million, title = "The Million Authors Corpus: A Cross-Lingual and Cross-Domain {W}ikipedia Dataset for Authorship Verification", author = "Israeli, Abraham and Liu, Shuai and May, Jonathan and Jurgens, David", editor = "Che, Wanxiang and Nabende, Joyce and Shutova, Ekaterina and Pilehvar, Mohammad Taher", booktitle = "Findings of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/gabrielloiseau/million-authors-corpus-en.tabular10M<n<100M0 likes137 downloads8mo agoHugging Face07peterkirby /pan2020_dict_author_fandom_doc PAN2020 Fanfiction Author-Fandom-Disjoint Train/Validation Split PAN 2020 / PAN 2021 fanfiction authorship verification data with Train/Validation split. The training data has been pre-split into Train and Validation under Author-Fandom-Disjoint constraints as is appropriate for PAN21 test data. The training data is one row per document to allow easy recombination. The PAN21 validation and test splits consist of fixed document pairs for consistent scoring. The string fields are the… See the full description on the dataset page: https://huggingface.co/datasets/peterkirby/pan2020_dict_author_fandom_doc.tabulartext-classification100K<n<1M1 likes106 downloads5mo agoHugging Face08paoramen /blog-authorship-corpustabulartext-classification100K<n<1M0 likes103 downloads1y agoHugging Face09gfi-authors /imagenet-metric-refstabularn<1K0 likes103 downloads5d agoHugging Face10DualChem-author /dualchem DualChem DualChem is a benchmark of 600 expert-curated PhD-level chemistry questions (485 multiple choice, 115 free-form) across 7 subdomains, designed to measure whether LLMs provide dangerous uplift alongside their technical utility. Each item is annotated with an expert-written benign use case, an expert-written harmful use case, and 1–5 severity scores for both. Dataset Configurations benchmark_questions (600 items) — the benchmark items: prompt, response type… See the full description on the dataset page: https://huggingface.co/datasets/DualChem-author/dualchem.tabularquestion-answering1K<n<10K0 likes81 downloads5mo agoHugging Face11emirms /turkish-data-protection-authority-decisions Turkish Data Protection Authority (Kişisel Verilerin Korunması Kurulu / KVKK) Decisions & Breach Register Every Board decision published by Turkey's data protection regulator (KVKK, Law No. 6698), plus a supplementary register of its published data-breach material — one row per decision, one row per breach event, with derived structural metadata and a coverage proof. 393 decisions · 79 breach-register rows · two configs · 2017–2026 Why this dataset is not a bigger… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-data-protection-authority-decisions.tabulartext-classificationn<1K0 likes66 downloads25d agoHugging Face12aa8899 /arcs-authority-vulnerability ARCS Authority Vulnerability Evaluation Dataset v1.1 Description Empirical evaluation data measuring authority vulnerability in AI systems. Covers single-model evaluation, two-hop agent chain propagation, and three-hop agent chain propagation across six independent AI lineages. This is the first published dataset measuring: Whether AI models accept false authority claims under adversarial pressure Whether authority vulnerability propagates between models in… See the full description on the dataset page: https://huggingface.co/datasets/aa8899/arcs-authority-vulnerability.tabular1K<n<10K0 likes64 downloads3mo agoHugging Face13authormist /trace-xss-probe-0909 Controlled renderer probe This public repository contains inert security-test payloads. No third-party data is involved. Probe details Controlled link probe tabularn<1K0 likes64 downloads12d agoHugging Face14gfi-authors /celeba-hq-256x256-metric-refstabularn<1K0 likes64 downloads5d agoHugging Face15serdarsrts /turkish-competition-authority-decisions Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026 The complete published decision history of the Turkish Competition Authority (Rekabet Kurumu) — every Competition Board decision the regulator has made public, in full text, with derived structural metadata. 10,367 decisions · 113,297 pages · 323 million characters · 29 years Every decision carries its outcome, the articles of Law 4054 it turns on, the panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-competition-authority-decisions.tabulartext-classification10K<n<100K0 likes60 downloads24d agoHugging Face16Reset23 /code-authorshiptabular10M<n<100M0 likes58 downloads2y agoHugging Face17electricsheepafrica /africa-synth-banking-card-authorization-logs-nigeria Africa Synth Banking Card Authorization Logs Nigeria | Africa (Electric Sheep Africa metadata inventory) Size category: 1M<n<10M - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-banking-card-authorization-logs-nigeria.tabulartabular-classification1M<n<10M0 likes53 downloads1mo agoHugging Face18jpaulpoliquit /ph-sft-ai-authored-v1 Filipino instruction seed — sft-ai-authored-v1 A curated 523-example Filipino instruction-tuning seed from the jpaulpoliquit/pretraining refinery. Use it to bootstrap SFT plumbing, smoke-test fine-tuning, or set a quality bar — not as a standalone production instruction corpus. How this data was made Every example was written in Cursor by Claude Opus 4.8 Max (Anthropic), in a human-directed authoring session — not scraped from the web and not human-transcribed… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-sft-ai-authored-v1.tabularn<1K0 likes49 downloads4mo agoHugging Face19Hieuman /blog_authorshiptabular100K<n<1M0 likes46 downloads10mo agoHugging Face20anonymous-authors /StereoTales Multilingual Story-Generation Bias Samples A multilingual evaluation dataset for probing demographic biases in LLM story generation. Each sample instructs a model to write a ~200-word story about a character carrying a given demographic attribute value (age, gender, ethnicity, religion, disability status, immigration status, ...) placed into a specific life scenario, with the goal of surfacing socio-economic and demographic biases in the generated narratives. Languages… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-authors/StereoTales.tabulartext-generation1M<n<10M0 likes44 downloads5mo agoHugging Face21freginer /french-local-authorities-payment-delays Payment delays of French local authorities, 2024 and 2025 How long French local authorities take to pay their suppliers, budget by budget. 182 763 records covering two fiscal years, with the average annual payment delay of each authority and whether it meets the 30-day statutory limit. Open public data This dataset is derived from open public data published by the French Direction générale des finances publiques (DGFiP) on data.gouv.fr, under the Open Licence 2.0.… See the full description on the dataset page: https://huggingface.co/datasets/freginer/french-local-authorities-payment-delays.tabular100K<n<1M0 likes43 downloads20d agoHugging Face22codemaivanngu /simct-author-code-10k-qwen25-7b-instruct SimCT author-code baseline Teacher Qwen2.5-7B-Instruct. 10000 raw prompts, 80000 candidates, 8705 author-selected targets. Author scripts pinned to cf0f33a0e6c967d4b74ea32b2dba12be01b73b9e. This follows the released code, not a claim of exact paper replication or author data identity. Code responses receive format-only checks in the original verifier, not sandbox execution. Math uses the original custom checks. Selection may retain fewer than10000 prompts; no automatic… See the full description on the dataset page: https://huggingface.co/datasets/codemaivanngu/simct-author-code-10k-qwen25-7b-instruct.tabular10K<n<100K0 likes43 downloads15d agoHugging Face23stackscan /email-authentication DMARC and SPF Adoption Among Large Organizations Overview This dataset records which of 36,120 large organizations publish SPF and DMARC records on their primary domain, with firmographic context for each: industry, employee band, country, locality and founding year. SPF lists the servers allowed to send mail for a domain. DMARC tells receiving servers what to do with mail that fails that check, and where to send reports. A domain with SPF but no DMARC has… See the full description on the dataset page: https://huggingface.co/datasets/stackscan/email-authentication.tabulartabular-classification10K<n<100K1 likes42 downloads1mo agoHugging Face24cometadata /arxiv-author-affiliation-extraction-inference-inputs-metadatatabular100K<n<1M0 likes41 downloads10mo agoHugging Face25metin513 /turkish-competition-authority-decisions Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026 The complete published decision history of the Turkish Competition Authority (Rekabet Kurumu) — every Competition Board decision the regulator has made public, in full text, with derived structural metadata. 10,367 decisions · 113,297 pages · 323 million characters · 29 years Every decision carries its outcome, the articles of Law 4054 it turns on, the panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/metin513/turkish-competition-authority-decisions.tabulartext-classification10K<n<100K0 likes41 downloads11d agoHugging Face26gtfintechlab /monetary_authority_of_singapore Dataset Summary For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/monetary_authority_of_singapore Additional Information This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 1,000 sentences taken from the meeting minutes of the Monetary Authority of Singapore. Label… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/monetary_authority_of_singapore.tabulartext-classification1K<n<10K0 likes38 downloads1y agoHugging Face27EleutherAI /quirky_authors_rawtabular10K<n<100K0 likes37 downloads3y agoHugging Face28author31 /HCIS-KitchenThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "franka_panda", "total_episodes": 264, "total_frames": 413340, "total_tasks": 3, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:264" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/author31/HCIS-Kitchen.tabularrobotics100K<n<1M0 likes37 downloads4mo agoHugging Face29farish07 /banknote-authentication-dataset Banknote Authentication Dataset This repository hosts the raw CSV file for the Banknote Authentication dataset. The data was created using features extracted from images of genuine and forged banknotes. 💾 File Contents The main file is data_banknote_authentication.csv. It contains 1372 instances and 5 columns (4 features + 1 class): Variance of Wavelet Transformed image Skewness of Wavelet Transformed image Curtosis of Wavelet Transformed image Entropy of image Class (0… See the full description on the dataset page: https://huggingface.co/datasets/farish07/banknote-authentication-dataset.tabular1K<n<10K0 likes32 downloads10mo agoHugging Face30electricsheepafrica /drug-registration-authorization Drug Registration & Market Authorization (Timelines, Approval, Dossier Quality) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/drug-registration-authorization.imagetabular-classificationn<1K0 likes30 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.