datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
news_media_bias_and_factuality
News Media Factual Reporting and Political Bias
Dataset introduced in the paper "Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions" published in the CLEF 2024 main conference.
Similar to the news media reliability dataset, this dataset consists of a collections of 4K new media domains names with political bias and factual reporting labels.
Columns of the dataset:
source: domain name
bias: the political bias label. Values: "left"… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_bias_and_factuality.mixtral-factual-QA
Mixtral Factual QA
Generate questions and answers based on context provided. We use contexts from,
maktabahalbakri.com
muftiwp.gov.my
asklegal.my
dewanbahasa-jdbp
gov.my
patriots
rootofscience
majalahsains
nasilemaktech
alhijrahnews
https://huggingface.co/datasets/open-phi/textbooks
notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/mixtral-factual
factually-wrong-qa-coding.jsonl, 31253 rows, 425 MB
factually-wrong-qa.jsonl, 1108037 rows, 10… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-factual-QA.event_factuality
Event Factuality (It Happened / UDS-IH2)
Source
Decomp “It Happened” (UDS-IH2): https://decomp.io/projects/factuality/
UD English-EWT v1.2 (r1.2) for sentence reconstruction: https://github.com/UniversalDependencies/UD_English-EWT/tree/r1.2
Contains raw web/news text; included for research purposes only (no endorsement).
Task
Binary predicate-level event factuality (one row per predicate).
Labels
label: 0=false, 1=true
label_rule: single, agree, na_other, tie_conf4_vs0… See the full description on the dataset page: https://huggingface.co/datasets/compling/event_factuality.factual-consistency-training-mixThis is a mix of NLI-like datasets that is used to train factual consistency models available in this collection.
Some of the datasets are upsampled here (Seahorse). In all cases, we upsample the less represented label since we want to use this dataset for a binary classification task.
The distribution of the dataset is as follows:
subset
count
alisawuffles/WANLI
127885
anli
105076
Seahorse
31666
LingNLI
19994
scitail
16944
boolq
11725
FoolMeTwice
10569
vitaminc
8489… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/factual-consistency-training-mix.OpenDataGen-factuality-en-v0.1This synthetic dataset was generated using the Open DataGen Python library. (https://github.com/thoddnn/open-datagen)
Methodology:
Retrieve random article content from the HuggingFace Wikipedia English dataset.
Construct a Chain of Thought (CoT) to generate a Multiple Choice Question (MCQ).
Utilize a Large Language Model (LLM) to score the results then filter it.
All these steps are prompted in the 'template.json' file located in the specified code folder.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/thoddnn/OpenDataGen-factuality-en-v0.1.FactualConsistencyScoresTextSummarization
HuggingFace Dataset: FactualConsistencyScoresTextSummarization
Description:
This dataset aggregates model scores assessing factual consistency across multiple summarization datasets. It is designed to highlight the thresholding issue with current SOTS factual consistency models in evaluating the factuality of text summarizations.
What is the "Thresholding Issue" with SOTA Factual Consistency Models ?
Existing models for detecting factual errors in summaries… See the full description on the dataset page: https://huggingface.co/datasets/achandlr/FactualConsistencyScoresTextSummarization.FactualCS
FactualCS
FactualCS is a counterspeech generation dataset introduced in:
Counter with Evidence! A Multi-Agent Memory Efficient Reasoning Framework for Hate Category Informed Counterspeech GenerationAccepted at EMNLP 2026 Main Conference
💻 Code: https://github.com/C0mRD/Counter_with_evidence
Dataset Description
Existing counterspeech datasets primarily focus on mapping hateful content to an appropriate response, while often treating hate speech as a… See the full description on the dataset page: https://huggingface.co/datasets/Aswini123/FactualCS.factuality-rmbench-style
Factuality RM-Bench Style
Factuality RM-Bench Style is a controlled English dataset for studying whether
reward models and representation probes prefer stylistic presentation over
factual correctness. Each row contains one question, a localized correct and
incorrect proposition, and six responses formed by crossing correctness with
three presentation styles: concise, normal, and Markdown.
This repository is an export package for
factuality_rmbench_style_v6. The published data… See the full description on the dataset page: https://huggingface.co/datasets/Yunnnuy/factuality-rmbench-style.dpo-mix5-Llama3-FactualityFACTUAL_Scene_Graph_IDPlease refer to https://github.com/zhuang-li/FACTUAL for a detailed description of this dataset.
withdrarxiv-factual-errors
Dataset Card for "withdrarxiv-factual-errors"
More Information needed
auditkit-testrun-factual-consistency
auditkit-testrun-factual-consistency
Built using AuditKIT — evaluate any model on any dataset and any task.
Method
evaluate
Model
<auditkit.model.vllm_gen.VLLMModel object at 0x7c1b15bf5010>
Artifact
run
Published
2026-09-01 14:24 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/auditkit-testrun-factual-consistency")
qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_factualcorrectness_round1dpo-mix5-Llama3-Factuality-MinChosen9-MinDelta6zhou-et-al-factual-statements-cleanedwildchat-factualFactuality_Alignment
Factual Preference Alignment Dataset
**⚠️ Warning:**This dataset contains hallucinated and synthetic responses
intentionally generated for research on robust factuality alignment.
Responses may include fabricated or incorrect information by design
to support the evaluation of hallucination-aware learning.
Dataset Summary
The AIXpert Preference Alignment Dataset is a curated collection of
45,000 factuality-aware preference pairs designed to support
research on Modified… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/Factuality_Alignment.factual-multiagent-roleplay-ft-ru
march228/factual-multiagent-roleplay-ft-ru
Небольшой русскоязычный synthetic finetuning dataset для обучения модели следованию ролевым системным инструкциям личности при сохранении фактической опоры на контекст.
Что это за датасет
Этот набор сделан как instruction / finetuning dataset, а не как benchmark.
В каждой записи есть:
плотный system с персоной и тоном;
context, на который нужно опираться;
пользовательский question;
внутренние thoughts;
финальный answer.… See the full description on the dataset page: https://huggingface.co/datasets/march228/factual-multiagent-roleplay-ft-ru.squad_v2_factuality_v1
squad_v2_factuality_v1
This dataset is derived from "squad_v2" training "context" with the following steps.
NER is run to extract entities.
Lexicon of person's name, date, organisation name and location are collected.
20% of the time, one of the text attribute (person's name, date, organisation name and location) is randomly replaced. For consistency of context, all other place with the same name is also replaced.
Purpose of the Dataset
The purpose of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/kenhktsui/squad_v2_factuality_v1.factualqwq_32b_factualqa_sft_datafactual-multilingual-questionsfactual-consistency-evaluation-benchmarkThis is a mix of 22 datasets that were used to evaluate factual consistency models available in this collection.
The distribution is available below:
subset
count
halueval_cnndm
19998
alisawuffles/WANLI
5000
Seahorse
4135
ExpertQA
3702
fib_xsum
3534
anli
3200
scitail
2126
Lfqa
1911
DeFacto
1836
llm_summaries_cnndm
1829
llm_summaries_xsum
1726
Reveal
1705
FactCheck-GPT
1565
FoolMeTwice
1379
aggrefact_xsum
1335
ClaimVerify
1087
aggrefact_cnndm
1017… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/factual-consistency-evaluation-benchmark.dpo-qwen2572b-athene70b-jdg-Llama3-Factualitynvidia_NVLM-D-72B-jdgfct-Factualitymeta-llama_Llama-3.1-70B-Instruct-jdgfct-Factualitycosmo-1B-claude_factual_feedbacktulu-v2-sft-mixture-second-stage-classifier-v1-stage2-example-based-factual-info-outputsdpo-Llama31-70b-NVLM-72b-Llama3-Factualityfactual-state-discovery-benchmark
Factual State Discovery Benchmark
Dataset for the Factual State Discovery Benchmark: Evaluating Fact Elicitation
in Polish Tax Law (ACL 2026 SRW). It evaluates whether conversational agents
can systematically elicit, through dialogue, all the facts of a taxpayer's
situation from a real Polish tax interpretation document.
Each sample pairs a factual state (a narrative of the taxpayer's situation,
in Polish) with its decomposition into atomic facts — independent,
verifiable claims… See the full description on the dataset page: https://huggingface.co/datasets/AI-TAX/factual-state-discovery-benchmark.
