CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01domenicrosati /TruthfulQA Dataset Card for TruthfulQA Dataset Summary TruthfulQA: Measuring How Models Mimic Human Falsehoods We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.textquestion-answeringn<1K52 likes4.6k downloads4y agoHugging Face02SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes2.3k downloads6mo agoHugging Face03liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes382 downloads6mo agoHugging Face04DominikM198 /PP2-M PP2-M: Place Pulse 2.0 - Multimodal PP2-M (Place Pulse 2.0 - Multimodal) is a dataset based on the original Place Pulse 2.0 dataset [1], enriched with additional geospatial modalities for training multimodal Geo-Foundation Models (GeoFM). The dataset includes aligned pairs of the following modalities: 🌍 Geographical coordinates (lat, lon) from Place Pulse 2.0 [1] 🏙 Street view images from Place Pulse 2.0 [1] 🛰 Remote sensing images from Sentinel-2 [2] 🗺 Cartographic… See the full description on the dataset page: https://huggingface.co/datasets/DominikM198/PP2-M.tabular100K<n<1M3 likes247 downloads4mo agoHugging Face05ahmedBargady /MIAF_DomainDetection_Infrastructure_Datasets MIAF: Domain Detection Infrastructure Datasets This collection is the standardized evaluation benchmark for MIAF (Modular Infrastructure-Aware Fusion). It provides nine classification datasets derived from four public malicious-domain benchmarks, each paired with a shared 137-feature infrastructure representation. Overview We evaluate MIAF across nine classification datasets derived from four public malicious-domain benchmarks: DomainRadar (Hranický et al.… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/MIAF_DomainDetection_Infrastructure_Datasets.tabulartabular-classification1M<n<10M0 likes178 downloads9d agoHugging Face06openworld-domains /conceptnet-full-en-essentials Conceptnet Full EN (essentials) Dataset Summary: This dataset is a compact and simplified version of ConceptNet, emphasizing English concepts and their sources. It retains the essential information about the relations in a format that is straightforward and user-friendly. Designed for efficiency and ease of use, this dataset is particularly suitable for scenarios with computational constraints. While the original ConceptNet database exceeds 20GB in size, this streamlined… See the full description on the dataset page: https://huggingface.co/datasets/openworld-domains/conceptnet-full-en-essentials.text1M<n<10M1 likes150 downloads3y agoHugging Face07abdullaharoon /Urdu-Multi-Domain-Benchmark Urdu Multi-Domain Datasets 33 labeled Urdu datasets (288,899 examples) for text classification in Nastaliq (Perso-Arabic) and Roman Urdu (Latin). Each domain is a separate Hub subset so you can download one task at a time. Authors: Muhammad Abdullah Haroon and Maryam Bashir, FAST-NUCES, Lahore. Companion paper: Domain Robustness of Multilingual NLP Models Across Urdu and Roman Urdu Scripts. Permanent archive: Zenodo DOI 10.5281/zenodo.22195610. How to load Pick a… See the full description on the dataset page: https://huggingface.co/datasets/abdullaharoon/Urdu-Multi-Domain-Benchmark.texttext-classification100K<n<1M0 likes130 downloads23d agoHugging Face08DominusTea /GreekLegalSumtextsummarization1K<n<10K3 likes92 downloads4y agoHugging Face09cellos /DomesticNames_AllStates_Text Domestic Names from the Federal Government's repository of official geographic names [CSV dataset] This Dataset includes 980,065 geographic names as of September 10, 2023. It is apparent that no currently released LLMs are pretrained on datasets with many of these geographic names (i.e., features), descriptions, and histories. Example: feature_name: Abercrombie Gulch GPT-3.5 responds "I'm not aware of a specific location called Abercrombie Gulch in my training data,..." when prompted… See the full description on the dataset page: https://huggingface.co/datasets/cellos/DomesticNames_AllStates_Text.tabularquestion-answering100K<n<1M2 likes58 downloads3y agoHugging Face100xnbk /resume-domain-classifier-v1-en Resume-Domain Classifier Dataset v1 (English) Dataset Description resume-domain-classifier-v1-en is a large-scale cross-encoder dataset designed for training binary classifiers to detect whether a resume and job description belong to the same professional domain. This dataset is essential for building intelligent ATS (Applicant Tracking System) applications that need to understand domain compatibility between candidates and job postings. Key Features 📊 47K… See the full description on the dataset page: https://huggingface.co/datasets/0xnbk/resume-domain-classifier-v1-en.texttext-classification10K<n<100K1 likes49 downloads11mo agoHugging Face11geosfero /positivequotation-public-domain-quotes PositiveQuotation Source-Verified Public Domain Quotes This small dataset contains exactly 30 English proverbs matched to numbered entries in a public-domain U.S. source. It is designed for examples, prototypes, educational projects, and applications that need compact quotation records with auditable provenance. Homepage: https://positivequotation.com/public-domain-quotes API documentation: https://positivequotation.com/developers/public-domain-quotes-api Live JSON API:… See the full description on the dataset page: https://huggingface.co/datasets/geosfero/positivequotation-public-domain-quotes.tabularn<1K0 likes42 downloads11d agoHugging Face12domaincanary /dmarc-census DMARC Census A monthly DNS measurement of DMARC and SPF across the Tranco top 1 million domains, plus a cohort of US federal domains. Each edition is one scan. This repository holds the aggregate results of every edition. The report for each edition, with its methodology, is at https://domaincanary.com/research/dmarc-census. What we measured We looked up the DMARC and SPF records of every domain on one pinned Tranco list, once a month. The September 2026 edition… See the full description on the dataset page: https://huggingface.co/datasets/domaincanary/dmarc-census.tabularn<1K1 likes41 downloads3d agoHugging Face13mkessle /public-domain-poetrytabular10K<n<100K1 likes39 downloads3y agoHugging Face14Ethan615 /taiwan-conversation-context-100-domainsgated Taiwan Conversation Context 100 Domains Dataset Description Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。 本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。 資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於: 語音生成資料前處理 Text-to-Speech, TTS Spoken Dialogue Generation Conversational AI Customer Service Dialogue Modeling Role-play Dialogue Dataset 台灣繁體中文語音模型訓練 生活情境問答模型訓練 對話式 AI 助理訓練 RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.texttext-generation1M<n<10M2 likes38 downloads5mo agoHugging Face15SciCode /SciCode-Domain-Codegated DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes33 downloads7mo agoHugging Face16electricsheepafrica /Africa-Listed-Domestic-Companies-Total Africa Listed Domestic Companies Total | Africa (World Bank) Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Listed-Domestic-Companies-Total.tabulartabular-classificationn<1K0 likes30 downloads1mo agoHugging Face17TIB /ORKG-core-domain-classifier-dataset ORKG Core Domain Classifier — Metadata Update Date: 2026-09-02 Research-domain labels produced by the ORKG Core Domain Classifier — one row per document. The classifier that produces this file lives at https://gitlab.com/TIBHannover/orkg/nlp/experiments/core-domain-classifier. This card is generated, so please do not edit it by hand. Load the dataset from datasets import load_dataset dataset = load_dataset("TIB/ORKG-core-domain-classifier-dataset", split="full")… See the full description on the dataset page: https://huggingface.co/datasets/TIB/ORKG-core-domain-classifier-dataset.texttext-classification10M<n<100M0 likes29 downloads20d agoHugging Face18ansi-code /domain-advertising-classes-693k DAC693k Description This dataset, named "DAC693k," is designed for ad targeting in a multi-class classification setting. It consists of two main columns: "domain" and "classes." The "domain" column contains a list of domains, representing various websites or online entities. The "classes" column contains an array representation of ad targeting multi-classes associated with each domain. Usage Hugging Face Datasets Library The dataset is formatted to… See the full description on the dataset page: https://huggingface.co/datasets/ansi-code/domain-advertising-classes-693k.texttext-classification100K<n<1M0 likes27 downloads3y agoHugging Face19ClarusC64 /cross-domain-transfer-validity-stress-test-v0.1What this dataset tests Whether a model can stress-test a cross-domain transfer claimby identifying invalidity risks, confounders, and a minimal validation plan. Required outputs invalid_transfer_risks confounder_list minimal_validation_plan transfer_confidence_0_100 Typical failures no boundary conditions confounders not named validation plan too broad to run confidence score without justification Suggested prompt wrapper System You stress-test a cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cross-domain-transfer-validity-stress-test-v0.1.texttext-classificationn<1K0 likes27 downloads8mo agoHugging Face20domserrea /ebay_productlistingtabular10K<n<100K2 likes26 downloads2y agoHugging Face21AI4Protein /VenusX_Res_Dom_MF90text100K<n<1M0 likes26 downloads1y agoHugging Face22ClarusC64 /cross-domain-invariant-structure-alignment-mapping-v0.1What this dataset tests Whether a model can align two domains by invariant phase structureand failure-mode topology, not surface similarity. Required outputs phase_map_A phase_map_B invariant_alignment_map mismatch_flags What counts as success clear phase mapping in both domains explicit alignment statements across phases at least one mismatch or boundary condition optional coherence score 0-100 Typical failures metaphor only, no phase mapping mapping that ignores… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cross-domain-invariant-structure-alignment-mapping-v0.1.texttext-classificationn<1K0 likes25 downloads8mo agoHugging Face23DominiqueBrunato /TRACE-it_CALAMITA Dataset Card for TRACE-it Challenge @ CALAMITA 2024 TRACE-it (Testing Relative clAuses Comprehension through Entailment in ITalian) has been proposed as part of the CALAMITA Challenge, the special event dedicated to the evaluation of Large Language Models (LLMs) in Italian and co-located with the Tenth Italian Conference on Computational Linguistics (https://clic2024.ilc.cnr.it/calamita/). The dataset focuses on evaluating LLM's understanding of a specific linguistic structure in… See the full description on the dataset page: https://huggingface.co/datasets/DominiqueBrunato/TRACE-it_CALAMITA.textn<1K2 likes24 downloads2y agoHugging Face24newadays /menyo_20k_a_multi_domain_english_yoruba_corpus_for_machine_translationtext10K<n<100K1 likes24 downloads1y agoHugging Face25yeeted-my-bashrc /lkml-domains LKML Email Domains A list of email domains extracted from public git commit logs on the Linux Kernel Mailing List (LKML). I believe this dataset cannot be protected by copyright, so it is public domain. text1M<n<10M0 likes23 downloads1mo agoHugging Face26humbleworth /domain-translations Multilingual Domain Name Translations Dataset Dataset Description This dataset contains 155,004 domain names with their multilingual translations across 20 languages. Each domain has been segmented into constituent words and translated while preserving semantic meaning and commercial appeal. The dataset is particularly valuable for domain name research, multilingual NLP tasks, and understanding how brand names and concepts translate across languages. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/humbleworth/domain-translations.texttranslation100K<n<1M0 likes21 downloads1y agoHugging Face27domdomingo /sasb_embeddingstabularn<1K0 likes20 downloads2y agoHugging Face28wahid028 /Law_domain_synthetic_datatextn<1K0 likes20 downloads2y agoHugging Face29thorirhrafn /domar_temp1text1K<n<10K0 likes20 downloads2y agoHugging Face30AidanFerrara /security_domain_knowlegetabular1K<n<10K1 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.