datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FairytaleQAFairytaleQA dataset, an open-source dataset focusing on comprehension of narratives, targeting students from kindergarten to eighth grade. The FairytaleQA dataset is annotated by education experts based on an evidence-based theoretical framework. It consists of 10,580 explicit and implicit questions derived from 278 children-friendly stories, covering seven types of narrative elements or relations.fairy_talesConcatenated and edited collection of fairy tales taken from Project Gutenberg.
Texts:
https://www.gutenberg.org/files/2591/2591-0.txt
https://www.gutenberg.org/files/503/503-0.txt
https://www.gutenberg.org/files/7277/7277-0.txt
https://www.gutenberg.org/cache/epub/35862/pg35862.txt
https://www.gutenberg.org/cache/epub/69739/pg69739.txt
https://www.gutenberg.org/files/2435/2435-0.txt
https://www.gutenberg.org/cache/epub/7871/pg7871.txt
https://www.gutenberg.org/files/8933/8933-0.txt… See the full description on the dataset page: https://huggingface.co/datasets/vicclab/fairy_tales.captrack
Dataset Card for CapTrack
Dataset Summary
CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary dimensions:
CAN (Latent Competence): What a model is capable of doing under ideal prompting
WILL (Default Behavioral Preferences): What a model chooses to do by default
HOW (Protocol Compliance): How reliably a… See the full description on the dataset page: https://huggingface.co/datasets/tri-fair-lab/captrack.task359_casino_classification_negotiation_vouch_fair
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task359_casino_classification_negotiation_vouch_fair
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task359_casino_classification_negotiation_vouch_fair.fairleap-driver-chat-sft-43k
Fairleap Driver Chat SFT 43k
📘 Dataset Overview
42,743 synthetic Indonesian conversations between a Gojek/GOTO driver and an assistant,
built for the Fairleap AI project — a platform addressing
income uncertainty and wellbeing for ride-hailing drivers in Indonesia. Each record is one
complete chat: a stuffed system prompt, the driver's questions, and the assistant's replies,
sometimes with a tool call or an injected domain-knowledge block in between.
It exists because… See the full description on the dataset page: https://huggingface.co/datasets/fairleap-ai/fairleap-driver-chat-sft-43k.LLMEval-Fair
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
LLMEval-Fair is a 30-month longitudinal study on the robustness and fairness of LLM evaluation,
built on a proprietary bank of 220,000+ graduate-level Chinese questions spanning 13 academic disciplines.
This is the publicly released subset of that bank.
Paper (arXiv): https://arxiv.org/abs/2508.05452
Venue: ACL 2026 Main Conference
Project website: https://llmeval.com/
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Fair.FairytaleQA-translated-spanish
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.FairytaleQA-translated-ptBR
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Brazilian Portuguese (pt-BR) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptBR.FairytaleQA-translated-ptPT
Dataset Card for FairytaleQA-translated-ptPT
Dataset Summary
This repository contains the European Portuguese (pt-PT) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptPT.FairytaleQA-translated-french
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the French machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-french.FairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.FairytaleQA-translated-italian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Italian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-italian.NaturalConversations-36k
NaturalConversations
NaturalConversations is a curated, deduplicated, and standardized conversational dataset created by merging and cleaning three publicly available dialogue corpora:
2017dailydialog/daily_dialog
3nesdeniz/english-daily-dialogues-10k
ProlificAI/overheard-18k
Key Specifications
Total conversations: 36,392
Average turns per conversation: 10.66
Format: Messages
Languages: English (and a lil bit of spanish)
Deduplication applied: Yes (exact and… See the full description on the dataset page: https://huggingface.co/datasets/Fair-HV/NaturalConversations-36k.fair-honest-os-v1.1
FAIR & HONEST DISCUSSION OS™ (Common Ground Edition v1.1)
"Most people complain about AI being 'woke', 'biased', or 'censored'.
Here's my FAIR & HONEST DISCUSSION OS™ (v1.1) — free copy-paste system prompt/framework to make any AI (Grok, Claude, ChatGPT, etc.) steelman perspectives, separate facts/values, seek common ground, and reduce heat without losing truth.
How to Use
Paste at the start of any chat...
FAIR & HONEST DISCUSSION OS™ (Common Ground Edition v1.1)
© 2026… See the full description on the dataset page: https://huggingface.co/datasets/whitbeckrocksAI/fair-honest-os-v1.1.fair-kto-datasetEthical-Reasoning-5k
Ethical Reasoning - DPO Dataset
EthicalReasoning is a dataset designed for the Direct Preference Optimization (DPO) stage in conversational AI training pipelines. It provides preference pairs for ethical and moral reasoning in dialogue contexts.
The dataset is built to align model responses with consistent ethical principles without enforcing a generic "assistant" persona. Instead, it focuses on moral coherence, ensuring models can navigate social situations with appropriate… See the full description on the dataset page: https://huggingface.co/datasets/Fair-HV/Ethical-Reasoning-5k.fairy_talesConcatenated and edited collection of fairy tales taken from Project Gutenberg.
Texts:
https://www.gutenberg.org/files/2591/2591-0.txt
https://www.gutenberg.org/files/503/503-0.txt
https://www.gutenberg.org/files/7277/7277-0.txt
https://www.gutenberg.org/cache/epub/35862/pg35862.txt
https://www.gutenberg.org/cache/epub/69739/pg69739.txt
https://www.gutenberg.org/files/2435/2435-0.txt
https://www.gutenberg.org/cache/epub/7871/pg7871.txt
https://www.gutenberg.org/files/8933/8933-0.txt… See the full description on the dataset page: https://huggingface.co/datasets/popsyren27/fairy_tales.
