datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
risale-nur-grounded-multipool
Risale-i Nur Grounded Multi-Pool LLM Dataset
TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak
bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim
çalışmaları için çok görünümlü bir veri seti.
EN. A multi-view dataset built from 15 canonical Risale-i
Nur books for grounded generation, SFT, preference learning, evaluation,
continued pretraining, and retrieval.
v2.10.0 · 199 configs · 463 config/split views ·
527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.RAG-Grounded-QA-188k
🎯 RAG Grounded QA 186K
The Anti-Hallucination Dataset
Teach language models to answer from context — or shut up trying.
Built by NovachronoAI — Precision AI for the real world.
Full Dataset (186K) · 20K Subset · Schema · Sources · Usage Guide
🧠 Why This Dataset Exists
Most QA datasets teach models what to say. This one also teaches them when to stay silent.
RAG (Retrieval-Augmented Generation) systems have a fatal flaw: the model hallucinates when… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/RAG-Grounded-QA-188k.omnimcp_graphrag_grounded_answer_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_grounded_answer_teaser.grounded-qa-preferences
Grounded QA preferences
Preference pairs for a small RLHF stack. Each row is a passage, a question, a preferred answer, and a rejected answer.
The questions, answer spans, and unanswerable labels come from SQuAD 2.0 (Rajpurkar et al.). This dataset does not add new human rankings. A fixed rule turns those annotations into Bradley-Terry pairs:
pair_type
When
Chosen
Rejected
wrong_span
The passage answers the question
The gold span
A different short span from the same… See the full description on the dataset page: https://huggingface.co/datasets/saitejaalasyam/grounded-qa-preferences.grounded-behavior-framework-v1_5
Grounded Behavior Framework N1 v1.5
Dataset sintético em português europeu para treino e avaliação de respostas
fundamentadas num contexto fornecido. Cada exemplo contém um contexto, uma
pergunta e uma resposta curta que aparece literalmente no contexto.
Como carregar
from datasets import load_dataset
dataset = load_dataset("empgces/grounded-behavior-framework-v1_5")
print(dataset)
print(dataset["train"][0])
Splits
Split
Exemplos
Utilização… See the full description on the dataset page: https://huggingface.co/datasets/empgces/grounded-behavior-framework-v1_5.grounded-meta-awareness
VANTA Research
Independent AI safety research lab specializing in cognitive fit, alignment, and human-AI collaboration
Grounded Meta-Awareness Dataset
A curated dataset of 1,187 conversational examples demonstrating honest, calibrated self-awareness about AI capabilities, limitations, and nature. Designed for fine-tuning language models to discuss their own functioning accurately without overclaiming or unnecessary deflection.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/grounded-meta-awareness.ua-legal-citation-grounded-sft
UA Legal Citation-Grounded SFT
A supervised fine-tuning set of citation-grounded legal question-answering examples in
Ukrainian. Every assistant answer attributes each factual claim to a specific source with a
[doc:ID] marker that refers to a real court decision passage placed in the prompt. The set
is built to train and study retrieval-grounded generation where faithfulness of citations,
not just answer quality, is the target.
How it was built
Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-legal-citation-grounded-sft.grounded-behavior-n1-pt
Dataset Description
Synthetic European Portuguese grounded question-answering examples generated by multiple model providers.
Objective
Train models to answer from the supplied context rather than external knowledge.
Dataset Structure
JSONL splits: train (4440), validation (250), and test (240).
Data Fields
Each row contains an ID, context, question, answer, source grouping metadata, and available curriculum metadata.… See the full description on the dataset page: https://huggingface.co/datasets/empgces/grounded-behavior-n1-pt.wikitext
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Grounded-Entropy/wikitext.
