datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.PP2-M
PP2-M: Place Pulse 2.0 - Multimodal
PP2-M (Place Pulse 2.0 - Multimodal) is a dataset based on the original Place Pulse 2.0 dataset [1], enriched with additional geospatial modalities for training multimodal Geo-Foundation Models (GeoFM).
The dataset includes aligned pairs of the following modalities:
🌍 Geographical coordinates (lat, lon) from Place Pulse 2.0 [1]
🏙 Street view images from Place Pulse 2.0 [1]
🛰 Remote sensing images from Sentinel-2 [2]
🗺 Cartographic… See the full description on the dataset page: https://huggingface.co/datasets/DominikM198/PP2-M.MIAF_DomainDetection_Infrastructure_Datasets
MIAF: Domain Detection Infrastructure Datasets
This collection is the standardized evaluation benchmark for MIAF (Modular Infrastructure-Aware Fusion). It provides nine classification datasets derived from four public malicious-domain benchmarks, each paired with a shared 137-feature infrastructure representation.
Overview
We evaluate MIAF across nine classification datasets derived from four public malicious-domain benchmarks:
DomainRadar (Hranický et al.… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/MIAF_DomainDetection_Infrastructure_Datasets.conceptnet-full-en-essentials
Conceptnet Full EN (essentials)
Dataset Summary:
This dataset is a compact and simplified version of ConceptNet, emphasizing English concepts and their sources. It retains the essential information about the relations in a format that is straightforward and user-friendly. Designed for efficiency and ease of use, this dataset is particularly suitable for scenarios with computational constraints. While the original ConceptNet database exceeds 20GB in size, this streamlined… See the full description on the dataset page: https://huggingface.co/datasets/openworld-domains/conceptnet-full-en-essentials.Urdu-Multi-Domain-Benchmark
Urdu Multi-Domain Datasets
33 labeled Urdu datasets (288,899 examples) for text classification in Nastaliq (Perso-Arabic) and Roman Urdu (Latin). Each domain is a separate Hub subset so you can download one task at a time.
Authors: Muhammad Abdullah Haroon and Maryam Bashir, FAST-NUCES, Lahore.
Companion paper: Domain Robustness of Multilingual NLP Models Across Urdu and Roman Urdu Scripts.
Permanent archive: Zenodo DOI 10.5281/zenodo.22195610.
How to load
Pick a… See the full description on the dataset page: https://huggingface.co/datasets/abdullaharoon/Urdu-Multi-Domain-Benchmark.GreekLegalSumDomesticNames_AllStates_Text Domestic Names from the Federal Government's repository of official geographic names [CSV dataset]
This Dataset includes 980,065 geographic names as of September 10, 2023.
It is apparent that no currently released LLMs are pretrained on datasets with many of these geographic names (i.e., features), descriptions, and histories.
Example: feature_name: Abercrombie Gulch
GPT-3.5 responds "I'm not aware of a specific location called Abercrombie Gulch in my training data,..." when prompted… See the full description on the dataset page: https://huggingface.co/datasets/cellos/DomesticNames_AllStates_Text.resume-domain-classifier-v1-en
Resume-Domain Classifier Dataset v1 (English)
Dataset Description
resume-domain-classifier-v1-en is a large-scale cross-encoder dataset designed for training binary classifiers to detect whether a resume and job description belong to the same professional domain. This dataset is essential for building intelligent ATS (Applicant Tracking System) applications that need to understand domain compatibility between candidates and job postings.
Key Features
📊 47K… See the full description on the dataset page: https://huggingface.co/datasets/0xnbk/resume-domain-classifier-v1-en.positivequotation-public-domain-quotes
PositiveQuotation Source-Verified Public Domain Quotes
This small dataset contains exactly 30 English proverbs matched to numbered entries in a public-domain U.S. source. It is designed for examples, prototypes, educational projects, and applications that need compact quotation records with auditable provenance.
Homepage: https://positivequotation.com/public-domain-quotes
API documentation: https://positivequotation.com/developers/public-domain-quotes-api
Live JSON API:… See the full description on the dataset page: https://huggingface.co/datasets/geosfero/positivequotation-public-domain-quotes.dmarc-census
DMARC Census
A monthly DNS measurement of DMARC and SPF across the Tranco top 1 million domains, plus a cohort of US federal domains. Each edition is one scan. This repository holds the aggregate results of every edition.
The report for each edition, with its methodology, is at https://domaincanary.com/research/dmarc-census.
What we measured
We looked up the DMARC and SPF records of every domain on one pinned Tranco list, once a month. The September 2026 edition… See the full description on the dataset page: https://huggingface.co/datasets/domaincanary/dmarc-census.public-domain-poetrytaiwan-conversation-context-100-domains
Taiwan Conversation Context 100 Domains
Dataset Description
Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。
本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。
資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於:
語音生成資料前處理
Text-to-Speech, TTS
Spoken Dialogue Generation
Conversational AI
Customer Service Dialogue Modeling
Role-play Dialogue Dataset
台灣繁體中文語音模型訓練
生活情境問答模型訓練
對話式 AI 助理訓練
RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.Africa-Listed-Domestic-Companies-Total
Africa Listed Domestic Companies Total | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Listed-Domestic-Companies-Total.ORKG-core-domain-classifier-dataset
ORKG Core Domain Classifier — Metadata
Update Date: 2026-09-02
Research-domain labels produced by the ORKG Core Domain Classifier — one row per document.
The classifier that produces this file lives at https://gitlab.com/TIBHannover/orkg/nlp/experiments/core-domain-classifier.
This card is generated, so please do not edit it by hand.
Load the dataset
from datasets import load_dataset
dataset = load_dataset("TIB/ORKG-core-domain-classifier-dataset", split="full")… See the full description on the dataset page: https://huggingface.co/datasets/TIB/ORKG-core-domain-classifier-dataset.domain-advertising-classes-693k
DAC693k
Description
This dataset, named "DAC693k," is designed for ad targeting in a multi-class classification setting. It consists of two main columns: "domain" and "classes." The "domain" column contains a list of domains, representing various websites or online entities. The "classes" column contains an array representation of ad targeting multi-classes associated with each domain.
Usage
Hugging Face Datasets Library
The dataset is formatted to… See the full description on the dataset page: https://huggingface.co/datasets/ansi-code/domain-advertising-classes-693k.cross-domain-transfer-validity-stress-test-v0.1What this dataset tests
Whether a model can stress-test a cross-domain transfer claimby identifying invalidity risks, confounders, and a minimal validation plan.
Required outputs
invalid_transfer_risks
confounder_list
minimal_validation_plan
transfer_confidence_0_100
Typical failures
no boundary conditions
confounders not named
validation plan too broad to run
confidence score without justification
Suggested prompt wrapper
System
You stress-test a cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cross-domain-transfer-validity-stress-test-v0.1.ebay_productlistingVenusX_Res_Dom_MF90cross-domain-invariant-structure-alignment-mapping-v0.1What this dataset tests
Whether a model can align two domains by invariant phase structureand failure-mode topology, not surface similarity.
Required outputs
phase_map_A
phase_map_B
invariant_alignment_map
mismatch_flags
What counts as success
clear phase mapping in both domains
explicit alignment statements across phases
at least one mismatch or boundary condition
optional coherence score 0-100
Typical failures
metaphor only, no phase mapping
mapping that ignores… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cross-domain-invariant-structure-alignment-mapping-v0.1.TRACE-it_CALAMITA
Dataset Card for TRACE-it Challenge @ CALAMITA 2024
TRACE-it (Testing Relative clAuses Comprehension through Entailment in ITalian) has been proposed as part of the CALAMITA Challenge, the special event dedicated to the evaluation of Large Language Models (LLMs) in Italian and co-located with the Tenth Italian Conference on Computational Linguistics (https://clic2024.ilc.cnr.it/calamita/).
The dataset focuses on evaluating LLM's understanding of a specific linguistic structure in… See the full description on the dataset page: https://huggingface.co/datasets/DominiqueBrunato/TRACE-it_CALAMITA.menyo_20k_a_multi_domain_english_yoruba_corpus_for_machine_translationlkml-domains
LKML Email Domains
A list of email domains extracted from public git commit logs on the Linux Kernel Mailing List (LKML).
I believe this dataset cannot be protected by copyright, so it is public domain.
domain-translations
Multilingual Domain Name Translations Dataset
Dataset Description
This dataset contains 155,004 domain names with their multilingual translations across 20 languages. Each domain has been segmented into constituent words and translated while preserving semantic meaning and commercial appeal. The dataset is particularly valuable for domain name research, multilingual NLP tasks, and understanding how brand names and concepts translate across languages.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/humbleworth/domain-translations.sasb_embeddingsLaw_domain_synthetic_datadomar_temp1security_domain_knowlege
