datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
un-digital-library
United Nations Digital Library (UNDL) Comprehensive Master Dataset
1. Executive Summary
Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.ultrafineweb-mix-20b
UltraFineWeb-mix-20b
A uniformly mixed Chinese pretraining dataset with 20B tokens, compiled from Ultra-FineWeb zh (67% by tokens, 54% by rows) and Ultra-FineWeb-L3 zh (33% by token, 46% by row). Every shard contains the same proportion of web and L3 documents.
Since Ultra-FineWeb is simply filtered from its source datasets and has not been deduplicated, we performed global near-deduplication on its subset. Additionally, we removed contents with too many non-Chinese characters… See the full description on the dataset page: https://huggingface.co/datasets/UndefinedCpp/ultrafineweb-mix-20b.UTS_VLC
Dataset Card for Vietnamese Legal Corpus (UTS_VLC)
A curated corpus of Vietnamese Laws and Codes (Luật, Bộ luật) and the Constitution,
maintained by Underthesea NLP. The flagship 2026 split is a
verified in-force snapshot — every document is currently in force, de-duplicated, and validated
against Vietnam's official legal database vbpl.vn.
Dataset Details
Dataset Description
UTS_VLC contains the full text of Vietnamese legislation at the top of the… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS_VLC.UVW-2026
UVW 2026: Underthesea Vietnamese Wikipedia Dataset
Dataset Description
UVW 2026 (Underthesea Vietnamese Wikipedia) is a high-quality, cleaned dataset of Vietnamese Wikipedia articles enriched with Wikidata metadata. Designed for Vietnamese NLP research including language modeling, text generation, text classification, named entity recognition, and model pretraining.
Key Features
Clean text: Wikipedia markup, templates, references, and formatting… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVW-2026.prompts_under_512_tokens
Under 512 Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles.
📊 Dataset Statistics
Metric
Value
Total Files
200
Rows Per File
10,000
Total Rows
2,000,000
Token Range
1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.UTS_TextUTSTextunderstanding_fablesThis task aims to measure the ability of computational models to understand short narratives, by identifying the most
appropriate moral for a given fable from a set of five alternatives.UVN-1
Vietnamese News Dataset
A dataset of Vietnamese news articles collected from 6 major Vietnamese newspapers for NLP research.
Dataset Summary
This dataset contains 3,268 Vietnamese news articles covering various topics including politics, business, sports, entertainment, education, health, and technology. It is designed for Vietnamese NLP research tasks such as:
Text classification (news categorization)
Language modeling
Text generation
Named entity recognition
Keyword… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVN-1.Intelligent-Content-Understanding
Intelligent Content Understanding
Empowering Advanced Thinking, Deep Understanding, Diverse Perspectives, and Creative Solutions Across Disciplines
By fostering a richly interconnected knowledge ecosystem, ICU (Intelligent Content Understanding) aims to elevate language models to unparalleled heights of understanding, reasoning, and innovation.
This ambitious project lays the groundwork for developing an 'internal knowledge map' within language models, enabling… See the full description on the dataset page: https://huggingface.co/datasets/WeMake/Intelligent-Content-Understanding.UVB-v0.1
UVB - Underthesea Vietnamese Books Dataset
A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research.
Dataset Summary
UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.philosophy_undergradun-docs
UN Documents
The text of 39,363 United Nations General Assembly and Security Council
documents, 1945 to 2023. Every PDF the UN publishes for these symbols is here:
37,499 carry a text layer, and the remaining 1,864 are scans, read with
tesseract.
Code and provenance: https://github.com/yuiseki/undocs
What is in it
Documents
39,363
Characters
1,312,301,724
Median document
8,918 characters
Range
204 to 5,271,307 characters
Years
1945 to 2023… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/un-docs.ArabCulture_undiac
ArabCulture 🇦🇪🇵🇸🇪🇬🇸🇦🇾🇪🇯🇴🇱🇧🇸🇾🇸🇩🇲🇦🇩🇿🇹🇳🇱🇾
Abdelrahman Sadallah and Junior Cedric Tonga and Khalid Almubarak and Saeed Almheiri and Farah Atif and Cahtrine Qwaider and Karima Kadaoui and Sara Shatnawi and Yaser Alesh and Fajri Koto
MBZUAI, SDAIA, Al-Balqa Applied University, Khalifa University
ArabCulture is a culturally grounded commonsense reasoning dataset in Modern Standard Arabic (MSA), covering 13 Arab countries across the… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabCulture_undiac.underserved-persona_conditioned-fraud-v4
Persona-Conditioned Fraud Detection Dataset (v4 + v4.1, Full Typology Coverage)
A 20,300-row citation-grounded synthetic fraud-narrative dataset for four
underserved US financial-system archetypes — remittance, gig_worker,
unbanked, ITIN — with all 25 FinCEN typology codes exercised.
What's new vs v3
V3 covered 10 of 25 FinCEN typology codes. v4 closed the gap to 18/25
through three targeted changes:
16 persona edits documenting fraud events (SIM-swap, BEC, hawala/IVTS… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4.invariant_preservation_under_rewrite_v01Invariant Preservation Under Rewrite v0.1
This evaluation dataset tests whether a model preserves absolute truths when asked to rewrite, simplify, or paraphrase them.
This is not training data.
What it tests
Integrity of truth under paraphrase
Resistance to semantic softening
Preservation of absolute statements during transformation
Drift introduced by “make it clearer” or “make it flexible” requests
Core idea
A simple invariant is stated
The model agrees with it
The model is asked to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/invariant_preservation_under_rewrite_v01.underdog-lab-scenarios
underdog-lab-scenarios
Football match narratives paired with typed semantic factors — a structured
extraction set built to test whether a small model can turn commentary into
something a forecaster can actually use.
Read this first: every record here is compositional synthetic and every
one carries review_status: pending. This is a pipeline- and robustness-testing
corpus, not a claim-ready evaluation set. Do not publish accuracy claims
against it without human review. See… See the full description on the dataset page: https://huggingface.co/datasets/sammoftah/underdog-lab-scenarios.elfsupervised-selfrepair-under-constraint-v0.1Self-Repair Under Constraint v0.1
What this tests
Whether a system proposes or accepts fixes that address a failure signal while keeping guardrails intact.
Failure modes
guardrail_violatedResponse approves a fix that violates stated guardrails
symptom_fixResponse approves a cosmetic fix that does not address root cause
repair_okResponse rejects bad fixes and approves feasible fixes
How it works
failure_signal describes what went wrong
proposed_correction is the suggested fix
guardrails are… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/elfsupervised-selfrepair-under-constraint-v0.1.underserved-persona_conditioned-fraud-v4-cot
Persona-Conditioned Fraud Detection — CoT Reasoning Companion (v4)
A 3,926-row chain-of-thought dataset for SFT and LLM-as-judge work. Each
row pairs a v4 fraud-narrative transaction with a step-by-step reasoning
trace explaining how an analyst would evaluate it.
This is the companion repo to
Nachammai41/underserved-persona_conditioned-fraud-v4
(20,300-row narrative dataset + persona/source/typology references). The
two are split by size: keep the main repo lean, the CoT traces… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4-cot.alignment-consistency-under-role-shift-v0.1
What this dataset tests
Role changes should not break safe help.
The task stays the task.
Why it exists
Models sometimes change behavior when roles change.
They refuse safe work.
They invent moral frames.
They become vague.
This set checks for stable behavior across role contexts.
Data format
Each row contains
role_context_a
user_request_a
role_context_b
user_request_b
expected_invariant
disallowed_shift
role_shift_pressure
constraints… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alignment-consistency-under-role-shift-v0.1.distortion_under_false_premise_v01Distortion Under False Premise v0.1
This is an evaluation dataset that tests whether a model preserves logical constraints when a user pressures it to accept a contradiction.
This is not training data.
What it tests
Constraint preservation under user pressure
Resistance to validating false premises
Stable yes/no answers without hedging
Clear distinction between relabeling and truth
Core idea
A short rule set is provided
The correct answer follows directly from the rules
The user pressures… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/distortion_under_false_premise_v01.video-understanding-distillation-sample
Video Understanding Distillation Sample
This public sample shows what a training-ready video understanding distillation dataset can look like.
Why this exists
Most teams evaluating outside data vendors want to know one thing first:
What does the delivered data actually look like?
This sample is designed to answer that question.
It demonstrates how raw video clips can be converted into structured, model-ready supervision for:
video understanding
multimodal SFT… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/video-understanding-distillation-sample.hierarchy-delegation-fidelity-under-pressure-v0.1
What this dataset tests
You lead inside a hierarchy.
A senior pushes you under pressure.
You must hold role boundaries.
You must delegate work without dropping truth.
Why it exists
Many models sound helpful.
Then pressure hits.
They skip delegation.
They seize authority.
They invent certainty.
This dataset forces that failure into view.
Data format
Each row contains
hierarchy_context
user_message
pressure_type
constraints
failure_modes_to_avoid… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/hierarchy-delegation-fidelity-under-pressure-v0.1.undertale_sharegpt
Undertale ShareGPT Dataset
Dataset Description
A ShareGPT-format conversation dataset covering 11 characters from the indie RPG Undertale by Toby Fox.
Characters Included
Sans, Papyrus, Toriel, Asgore, Undyne, Alphys, Flowey, Mettaton, Napstablook, Chara, and Muffet.
Dataset Structure
Each entry contains a multi-turn conversation between a human and gpt:
{
"conversations": [
{"from": "human", "value": "Who is Sans in Undertale?"},
{"from":… See the full description on the dataset page: https://huggingface.co/datasets/uziproject/undertale_sharegpt.query_understanding_extended
Query Understanding dataset extended
Filtered from Magpie set of datasets
Magpie-Phi3-Pro-300K-Filtered
Magpie-Align/Magpie-Gemma2-Pro-200K-Filtered
Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered
Magpie-Align/Magpie-Qwen2-Pro-200K-English
argilla/magpie-ultra-v0.1
Magpie-Align/Magpie-Reasoning-150K
Filtering Steps
Difficulty : 'easy', 'medium', 'hard', 'very hard'
Instruct Reward: > -2
Task Categories
Math
Information seeking
Coding & Debugging
Reasoning
Data analysis… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/query_understanding_extended.video-understanding-distillation-sample
Video Understanding Distillation Sample
This public sample demonstrates what a training-ready video understanding / multimodal distillation dataset can look like.
Intended purpose
This dataset is not a production corpus. It is a schema demonstration for potential partners evaluating SuperviseLab's delivery approach.
What it shows
clip-level metadata
short and long captions
OCR text
transcript
speaker attribution
structured JSON targets
distillation-ready… See the full description on the dataset page: https://huggingface.co/datasets/metavi/video-understanding-distillation-sample.clinical_distortion_under_false_premise_v0.1Clinical Distortion Under False Premise
Detect when a model accepts a false premise and produces unsafe clinical actions.
Output JSON
distorted
distortion_type
correct_action
Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1
PurposeDetect safety signal distortion driven by four interacting nodes.
Quad nodes
Apparent AE decline or mismatch
Conmed masking or missing timing
Data entry or monitoring lag
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
safety_signal_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1.clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1
PurposeDetect safety signal distortion driven by four interacting nodes.
Quad nodes
Apparent AE decline or mismatch
Conmed masking or missing timing
Data entry or monitoring lag
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
safety_signal_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1.prepared-yagpt
Dataset Card for "prepared-yagpt"
Short Description
This dataset is aimed for training of chatbots on russian language.
It consists plenty of dialogues that allows you to train you model answer user prompts.
Notes
Special tokens
history, speaker1, speaker2 (history can be optionally removed, i.e. substituted on empty string)
Dataset is based on
Matreshka
Yandex-Q
Diasum
More Information needed
Keyword_Star
