datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mixtral-factual-QA
Mixtral Factual QA
Generate questions and answers based on context provided. We use contexts from,
maktabahalbakri.com
muftiwp.gov.my
asklegal.my
dewanbahasa-jdbp
gov.my
patriots
rootofscience
majalahsains
nasilemaktech
alhijrahnews
https://huggingface.co/datasets/open-phi/textbooks
notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/mixtral-factual
factually-wrong-qa-coding.jsonl, 31253 rows, 425 MB
factually-wrong-qa.jsonl, 1108037 rows, 10… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-factual-QA.factual-consistency-training-mixThis is a mix of NLI-like datasets that is used to train factual consistency models available in this collection.
Some of the datasets are upsampled here (Seahorse). In all cases, we upsample the less represented label since we want to use this dataset for a binary classification task.
The distribution of the dataset is as follows:
subset
count
alisawuffles/WANLI
127885
anli
105076
Seahorse
31666
LingNLI
19994
scitail
16944
boolq
11725
FoolMeTwice
10569
vitaminc
8489… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/factual-consistency-training-mix.factuality-rmbench-style
Factuality RM-Bench Style
Factuality RM-Bench Style is a controlled English dataset for studying whether
reward models and representation probes prefer stylistic presentation over
factual correctness. Each row contains one question, a localized correct and
incorrect proposition, and six responses formed by crossing correctness with
three presentation styles: concise, normal, and Markdown.
This repository is an export package for
factuality_rmbench_style_v6. The published data… See the full description on the dataset page: https://huggingface.co/datasets/Yunnnuy/factuality-rmbench-style.qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_factualcorrectness_round1Factuality_Alignment
Factual Preference Alignment Dataset
**⚠️ Warning:**This dataset contains hallucinated and synthetic responses
intentionally generated for research on robust factuality alignment.
Responses may include fabricated or incorrect information by design
to support the evaluation of hallucination-aware learning.
Dataset Summary
The AIXpert Preference Alignment Dataset is a curated collection of
45,000 factuality-aware preference pairs designed to support
research on Modified… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/Factuality_Alignment.factual-multiagent-roleplay-ft-ru
march228/factual-multiagent-roleplay-ft-ru
Небольшой русскоязычный synthetic finetuning dataset для обучения модели следованию ролевым системным инструкциям личности при сохранении фактической опоры на контекст.
Что это за датасет
Этот набор сделан как instruction / finetuning dataset, а не как benchmark.
В каждой записи есть:
плотный system с персоной и тоном;
context, на который нужно опираться;
пользовательский question;
внутренние thoughts;
финальный answer.… See the full description on the dataset page: https://huggingface.co/datasets/march228/factual-multiagent-roleplay-ft-ru.factualqwq_32b_factualqa_sft_datafactual-consistency-evaluation-benchmarkThis is a mix of 22 datasets that were used to evaluate factual consistency models available in this collection.
The distribution is available below:
subset
count
halueval_cnndm
19998
alisawuffles/WANLI
5000
Seahorse
4135
ExpertQA
3702
fib_xsum
3534
anli
3200
scitail
2126
Lfqa
1911
DeFacto
1836
llm_summaries_cnndm
1829
llm_summaries_xsum
1726
Reveal
1705
FactCheck-GPT
1565
FoolMeTwice
1379
aggrefact_xsum
1335
ClaimVerify
1087
aggrefact_cnndm
1017… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/factual-consistency-evaluation-benchmark.adaption-cordel-factual-nordeste
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-cordel_factual_nordeste
This dataset contains pairs of prompts and completions where Brazilian 'cordel' poetry is generated in sextet stanzas based strictly on provided factual texts about Northeastern Brazilian culture, history, and geography. The source texts cover topics such as Frevo, the Cangaço (Lampião and Maria Bonita), the Caruaru Fair, and the Caatinga biome, with… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-cordel-factual-nordeste.factuality-benchmark-preview
AIgentic Factuality Benchmark — Preview
This repository is a placeholder for AIgentic’s upcoming open benchmark for evaluating factuality and hallucination control in enterprise AI systems.
🧭 Purpose
Our goal is to set a new reliability standard for AI models deployed in high-stakes professional domains — such as law, finance, and consulting.
🔬 Coming Soon
Example benchmark dataset (legal factuality)
Architecture overview diagram
LLM-as-a-Judge evaluation… See the full description on the dataset page: https://huggingface.co/datasets/AIgenticLLC/factuality-benchmark-preview.triviaqa_factual
Factual Recall Benchmark (adapted from TriviaQA)
Intended for mechanistic (factual recall and circuit) analysis.
Prompt format
Q: {question}?
A:
Model is expected to generate the answer.
Format
prompt — e.g. "Q: Who invented the telephone?"
answer — full canonical answer, e.g. "Alexander Graham Bell"
Notes
Answers are full canonical strings (not truncated to first word)
No explicit answer prefix (A:) is used in prompts
Designed for flexible… See the full description on the dataset page: https://huggingface.co/datasets/sohv/triviaqa_factual.Alpaca-GPT4-NAIT-Factual_Knowledgeen-si-translation-flores-factual-2k
En Si Translation Flores Factual 2K
Dataset Summary
English-Sinhala Factual Translation dataset containing ~2,000 highly accurate sentences covering diverse factual domains, curated from the FLORES+ benchmark.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 2000
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted from… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-flores-factual-2k.mnlp_mcqa_evals_factualfactuality-v0
