CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AuthenticIlm /Shamela4_Full_DB Shamela 4 — Full Islamic Library Corpus A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text. Dataset Structure stage0_raw/ ├── _meta/ # Cross-cutting metadata (Parquet + JSONL) │ ├── extraction_manifest.json # Global extraction record │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.text-generation10M<n<100M28 likes14k downloads4mo agoHugging Face02kalomaze /alphabetic-arxiv-authors-it1text100K<n<1M0 likes4.6k downloads1y agoHugging Face03lasrprobegen /authority-activationstext100K<n<1M0 likes2.7k downloads10mo agoHugging Face04mainakmanna /single-author-arxiv Single-author arXiv Computer Science Metadata for arXiv records classified in Computer Science that list exactly one author. default retains the original daily-file import. fast stores historical data in monthly files and adds new submissions as daily update files; it is the configuration used by the public archive because it makes filtering much faster. This dataset contains metadata only. arXiv is the source of truth; use each record's arxiv_url and pdf_url to read the paper. text100K<n<1M0 likes2.3k downloads1mo agoHugging Face05barilan /blog_authorship_corpusThe Blog Authorship Corpus consists of the collected posts of 19,320 bloggers gathered from blogger.com in August 2004. The corpus incorporates a total of 681,288 posts and over 140 million words - or approximately 35 posts and 7250 words per person. Each blog is presented as a separate file, the name of which indicates a blogger id# and the blogger’s self-provided gender, age, industry and astrological sign. (All are labeled for gender and age but for many, industry and/or sign is marked as unknown.) All bloggers included in the corpus fall into one of three age groups: - 8240 "10s" blogs (ages 13-17), - 8086 "20s" blogs (ages 23-27), - 2994 "30s" blogs (ages 33-47). For each age group there are an equal number of male and female bloggers. Each blog in the corpus includes at least 200 occurrences of common English words. All formatting has been stripped with two exceptions. Individual posts within a single blogger are separated by the date of the following post and links within a post are denoted by the label urllink. The corpus may be freely used for non-commercial research purposes.text-classification10K<n<100K18 likes1.1k downloads3y agoHugging Face06hkadxqq /spooky-author-identificationtext10K<n<100K0 likes981 downloads4y agoHugging Face07dataset-author-404 /Anon-CounterFactual-Dataset Anon Counterfactual Dataset Dataset repo: https://huggingface.co/datasets/dataset-author-404/Anon-CounterFactual-Dataset Synthetic CLEVR-style 3D scenes with original, semantic counterfactual, and negative (artifact) PNG renders, plus VQA-style questions, difficulties, and a 3×3 answer matrix. Built from the MMB counterfactual pipeline run folder dataset_720p_v2 (see build_hub_dataset.py in this repo snapshot). This revision replaces the previous Hub layout (legacy imagefolder /… See the full description on the dataset page: https://huggingface.co/datasets/dataset-author-404/Anon-CounterFactual-Dataset.imagen<1K0 likes897 downloads5mo agoHugging Face08Efstathios /guardian_authorshipA dataset cross-topic authorship attribution. The dataset is provided by Stamatatos 2013. 1- The cross-topic scenarios are based on Table-4 in Stamatatos 2017 (Ex. cross_topic_1 => row 1:P S U&W ). 2- The cross-genre scenarios are based on Table-5 in the same paper. (Ex. cross_genre_1 => row 1:B P S&U&W). 3- The same-topic/genre scenario is created by grouping all the datasts as follows. For ex., to use same_topic and split the data 60-40 use: train_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>", split='train[:60%]+validation[:60%]+test[:60%]') tests_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>", split='train[-40%:]+validation[-40%:]+test[-40%:]') IMPORTANT: train+validation+test[:60%] will generate the wrong splits because the data is imbalanced * See https://huggingface.co/docs/datasets/splits.html for detailed/more examplestext-classification1K<n<10K5 likes867 downloads3y agoHugging Face09FrancophonIA /Swedish_Work_environment_Authority [!NOTE] Dataset origin: https://portulanclarin.net/repository/browse/parallel-texts-from-swedish-work-environment-authority-processed/7404236aa58b11eaae0e02420a000403bd13d9138a904f33980bd63233eb90bc/ Description This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel texts from the Swedish… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Swedish_Work_environment_Authority.0 likes559 downloads1y agoHugging Face10akash1702-eng /voice-authenticity-datasetaudion<1K1 likes500 downloads19d agoHugging Face11NuBerea /authority-provenancegated authority-provenance A per-verse authority provenance surface for the Hebrew Bible and New Testament. For every verse it records independent signals bearing on the authority of the text at that point: textual stability (is the reading secure in the critical text?), compositional attribution (who wrote it, and on what evidence?), and canonical reception (how the church received it). These axes are kept separate so that questions of manuscript evidence, authorship, and reception… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/authority-provenance.textfeature-extraction10K<n<100K0 likes452 downloads24d agoHugging Face12tasksource /blog_authorship_corpustabular100K<n<1M2 likes428 downloads2y agoHugging Face13cometadata /arxiv-author-affiliations-matched-ror-ids arXiv Author Affiliations This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers. Dataset Description This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.texttext-classification1M<n<10M1 likes416 downloads8mo agoHugging Face14ulab-ai /ResearchArcade-openreview-authorstext100K<n<1M0 likes371 downloads7mo agoHugging Face15MU-NLPC /czech_corpus_authorship_recognition Czech Authorship Recognition Corpus (Kala) Popis datasetu Tento dataset byl vytvořen v rámci diplomové práce zaměřené na automatické rozpoznání autorství českých textů. Obsahuje české publicistické texty získané z veřejně dostupných online zdrojů a připravené pro experimenty v úlohách: přiřazení autorství (authorship attribution) ověřování autorství (authorship verification) shlukování podle autorství (authorship clustering) Zdrojová data Do… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/czech_corpus_authorship_recognition.texttext-classification0 likes366 downloads3mo agoHugging Face16FrancophonIA /Swedish_Social_Security_Authority [!NOTE] Dataset origin: https://portulanclarin.net/repository/browse/parallel-texts-from-swedish-social-security-authority-processed/3b5772a0a14511ea900d02420a00041df33980e9aa0140a0aca95e3de61180e0/ Description This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel texts, email templates… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Swedish_Social_Security_Authority.0 likes363 downloads1y agoHugging Face17anonymous-nsc-author /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes314 downloads2mo agoHugging Face18dislove /evidence-backed-authority-verification Evidence-Backed Authority Verification for Autonomous Agents Measuring and Governing Root-Equivalent Execution Paths A verifier that was asked whether an autonomous agent could reach root on its host, could not prove that it couldn't, and said so. This repository is the paper, the verifier, and every artifact the paper's numbers are computed from. Verdict BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven Paper 39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.textn<1K0 likes289 downloads1mo agoHugging Face19msaleme /mcp-sandbox-authority-boundary-profile MCP Sandbox Authority Boundary Profile Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile Profile release date: 2026-07-23 Latest distribution release date: 2026-09-05 Execution containment is not proof of bounded authority. Start here For a one-minute, case-by-case reading of the profile, open the companion Authority Boundary Field Guide Space. It presents the released synthetic observations with their control question, observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.textn<1K1 likes272 downloads16d agoHugging Face20ulab-ai /ResearchArcade-openreview-papers-authorstext100K<n<1M0 likes266 downloads7mo agoHugging Face21swan07 /authorship-verification Dataset Card for Dataset Name Dataset for authorship verification, comprised of 12 cleaned, modified, open source authorship verification and attribution datasets. Dataset Details Code for cleaning and modifying datasets can be found in https://github.com/swan-07/authorship-verification/blob/main/Authorship_Verification_Datasets.ipynb and is detailed in paper. Datasets used to produce the final dataset are: Reuters50 @misc{misc_reuter_50_50_217, author = {Liu… See the full description on the dataset page: https://huggingface.co/datasets/swan07/authorship-verification.texttext-classification100K<n<1M3 likes238 downloads2y agoHugging Face22BSC-LT /ALIA_mixed_authentic_synthetic_MT Dataset Card for ALIA_mixed_authentic_synthetic_MT Dataset Summary Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.texttranslation100M<n<1B1 likes219 downloads9mo agoHugging Face23anomyous-author /Explore-Execute-Chain-Datasetstext10K<n<100K0 likes210 downloads1y agoHugging Face24rmems /authz-regression-trajectories Authz Regression Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/authz-regression-trajectories.text1K<n<10K0 likes200 downloads21d agoHugging Face25gfi-authors /imagenet1M<n<10M0 likes194 downloads4d agoHugging Face26emirms /turkish-competition-authority-decisions Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026 The complete published decision history of the Turkish Competition Authority (Rekabet Kurumu) — every Competition Board decision the regulator has made public, in full text, with derived structural metadata. 10,367 decisions · 113,297 pages · 323 million characters · 29 years Every decision carries its outcome, the articles of Law 4054 it turns on, the panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.tabulartext-classification10K<n<100K3 likes170 downloads26d agoHugging Face27TheGarlic /image-authenticity-battle Image Authenticity Battle Dataset This dataset contains real and synthetic/tampered images for human perception studies on AI-generated content detection. Dataset Structure Total Images: 19500 Categories: Real, Synthetic (Fully AI-generated), Tampered (AI-edited) Models: Nano Banana, Qwen, Flux, SD3 Metadata Fields Each image has the following metadata: filename: Path to image file dataset: Source dataset name category: real/synthetic/tampered… See the full description on the dataset page: https://huggingface.co/datasets/TheGarlic/image-authenticity-battle.imageimage-classification10K<n<100K0 likes166 downloads11mo agoHugging Face28ValentinLAFARGUE /AuthorProfilingResults Probing Cultural Signals in Large Language Models through Author Profiling A dataset for analyzing cultural bias in LLM-based author profiling through controlled prompting experiments. Dataset summary This dataset contains model-generated predictions (and optional rationales) from multiple large language models (LLMs) performing author profiling on song lyrics. Due to licensing constraints, the original lyrics are not included. The dataset focuses on how models infer… See the full description on the dataset page: https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults.tabularzero-shot-classification100K<n<1M0 likes166 downloads6mo agoHugging Face29shimo4228 /authorship-strategy Authorship Strategy — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.tabularn<1K1 likes159 downloads28d agoHugging Face30nips26-anon-author /contractbench ContractBench A deterministic benchmark for measuring observation-contract compliance in LLM agents: whether agents preserve the temporal validity and byte-level integrity of intermediate tool outputs (presigned URLs, OAuth state parameters, JWT tokens, HMAC-protected webhooks, rate-limit windows, etc.). This dataset is the companion to the NeurIPS 2026 Evaluations & Datasets Track submission. It contains two complementary artifacts in a single repository: Subfolder What's… See the full description on the dataset page: https://huggingface.co/datasets/nips26-anon-author/contractbench.other1K<n<10K0 likes149 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.