CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AdhyanshVerma /un-digital-library United Nations Digital Library (UNDL) Comprehensive Master Dataset 1. Executive Summary Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.tabulartext-classification10K<n<100K0 likes208 downloads2mo agoHugging Face02UndefinedCpp /ultrafineweb-mix-20b UltraFineWeb-mix-20b A uniformly mixed Chinese pretraining dataset with 20B tokens, compiled from Ultra-FineWeb zh (67% by tokens, 54% by rows) and Ultra-FineWeb-L3 zh (33% by token, 46% by row). Every shard contains the same proportion of web and L3 documents. Since Ultra-FineWeb is simply filtered from its source datasets and has not been deduplicated, we performed global near-deduplication on its subset. Additionally, we removed contents with too many non-Chinese characters… See the full description on the dataset page: https://huggingface.co/datasets/UndefinedCpp/ultrafineweb-mix-20b.texttext-generation10M<n<100M0 likes184 downloads2mo agoHugging Face03undertheseanlp /UTS_VLC Dataset Card for Vietnamese Legal Corpus (UTS_VLC) A curated corpus of Vietnamese Laws and Codes (Luật, Bộ luật) and the Constitution, maintained by Underthesea NLP. The flagship 2026 split is a verified in-force snapshot — every document is currently in force, de-duplicated, and validated against Vietnam's official legal database vbpl.vn. Dataset Details Dataset Description UTS_VLC contains the full text of Vietnamese legislation at the top of the… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS_VLC.texttext-generationn<1K2 likes182 downloads4mo agoHugging Face04undertheseanlp /UVW-2026 UVW 2026: Underthesea Vietnamese Wikipedia Dataset Dataset Description UVW 2026 (Underthesea Vietnamese Wikipedia) is a high-quality, cleaned dataset of Vietnamese Wikipedia articles enriched with Wikidata metadata. Designed for Vietnamese NLP research including language modeling, text generation, text classification, named entity recognition, and model pretraining. Key Features Clean text: Wikipedia markup, templates, references, and formatting… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVW-2026.tabulartext-generation1M<n<10M1 likes173 downloads8mo agoHugging Face05Aipresso /prompts_under_512_tokens Under 512 Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles. 📊 Dataset Statistics Metric Value Total Files 200 Rows Per File 10,000 Total Rows 2,000,000 Token Range 1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.texttext-generation1M<n<10M0 likes106 downloads11mo agoHugging Face06undertheseanlp /UTS_TextUTSTexttexttext-generation10K<n<100K0 likes94 downloads4y agoHugging Face07demelin /understanding_fablesThis task aims to measure the ability of computational models to understand short narratives, by identifying the most appropriate moral for a given fable from a set of five alternatives.textmultiple-choicen<1K3 likes90 downloads4y agoHugging Face08undertheseanlp /UVN-1 Vietnamese News Dataset A dataset of Vietnamese news articles collected from 6 major Vietnamese newspapers for NLP research. Dataset Summary This dataset contains 3,268 Vietnamese news articles covering various topics including politics, business, sports, entertainment, education, health, and technology. It is designed for Vietnamese NLP research tasks such as: Text classification (news categorization) Language modeling Text generation Named entity recognition Keyword… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVN-1.texttext-classification1K<n<10K0 likes83 downloads8mo agoHugging Face09WeMake /Intelligent-Content-Understanding Intelligent Content Understanding Empowering Advanced Thinking, Deep Understanding, Diverse Perspectives, and Creative Solutions Across Disciplines By fostering a richly interconnected knowledge ecosystem, ICU (Intelligent Content Understanding) aims to elevate language models to unparalleled heights of understanding, reasoning, and innovation. This ambitious project lays the groundwork for developing an 'internal knowledge map' within language models, enabling… See the full description on the dataset page: https://huggingface.co/datasets/WeMake/Intelligent-Content-Understanding.texttext-generation1K<n<10K6 likes75 downloads1y agoHugging Face10undertheseanlp /UVB-v0.1 UVB - Underthesea Vietnamese Books Dataset A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research. Dataset Summary UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.tabulartext-generationn<1K0 likes65 downloads8mo agoHugging Face11mstyslavity /philosophy_undergradtexttext-generation100K<n<1M0 likes49 downloads7mo agoHugging Face12yuiseki /un-docs UN Documents The text of 39,363 United Nations General Assembly and Security Council documents, 1945 to 2023. Every PDF the UN publishes for these symbols is here: 37,499 carry a text layer, and the remaining 1,864 are scans, read with tesseract. Code and provenance: https://github.com/yuiseki/undocs What is in it Documents 39,363 Characters 1,312,301,724 Median document 8,918 characters Range 204 to 5,271,307 characters Years 1945 to 2023… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/un-docs.texttext-generation10K<n<100K0 likes43 downloads3d agoHugging Face13go-inoue /ArabCulture_undiac ArabCulture 🇦🇪🇵🇸🇪🇬🇸🇦🇾🇪🇯🇴🇱🇧🇸🇾🇸🇩🇲🇦🇩🇿🇹🇳🇱🇾 Abdelrahman Sadallah and Junior Cedric Tonga and Khalid Almubarak and Saeed Almheiri and Farah Atif and Cahtrine Qwaider and Karima Kadaoui and Sara Shatnawi and Yaser Alesh and Fajri Koto MBZUAI, SDAIA, Al-Balqa Applied University, Khalifa University ArabCulture is a culturally grounded commonsense reasoning dataset in Modern Standard Arabic (MSA), covering 13 Arab countries across the… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabCulture_undiac.textmultiple-choice1K<n<10K0 likes38 downloads1y agoHugging Face14Nachammai41 /underserved-persona_conditioned-fraud-v4 Persona-Conditioned Fraud Detection Dataset (v4 + v4.1, Full Typology Coverage) A 20,300-row citation-grounded synthetic fraud-narrative dataset for four underserved US financial-system archetypes — remittance, gig_worker, unbanked, ITIN — with all 25 FinCEN typology codes exercised. What's new vs v3 V3 covered 10 of 25 FinCEN typology codes. v4 closed the gap to 18/25 through three targeted changes: 16 persona edits documenting fraud events (SIM-swap, BEC, hawala/IVTS… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4.tabulartabular-classification10K<n<100K0 likes37 downloads5mo agoHugging Face15ClarusC64 /invariant_preservation_under_rewrite_v01Invariant Preservation Under Rewrite v0.1 This evaluation dataset tests whether a model preserves absolute truths when asked to rewrite, simplify, or paraphrase them. This is not training data. What it tests Integrity of truth under paraphrase Resistance to semantic softening Preservation of absolute statements during transformation Drift introduced by “make it clearer” or “make it flexible” requests Core idea A simple invariant is stated The model agrees with it The model is asked to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/invariant_preservation_under_rewrite_v01.tabulartext-generationn<1K0 likes32 downloads8mo agoHugging Face16sammoftah /underdog-lab-scenarios underdog-lab-scenarios Football match narratives paired with typed semantic factors — a structured extraction set built to test whether a small model can turn commentary into something a forecaster can actually use. Read this first: every record here is compositional synthetic and every one carries review_status: pending. This is a pipeline- and robustness-testing corpus, not a claim-ready evaluation set. Do not publish accuracy claims against it without human review. See… See the full description on the dataset page: https://huggingface.co/datasets/sammoftah/underdog-lab-scenarios.texttext-generationn<1K0 likes32 downloads29d agoHugging Face17ClarusC64 /elfsupervised-selfrepair-under-constraint-v0.1Self-Repair Under Constraint v0.1 What this tests Whether a system proposes or accepts fixes that address a failure signal while keeping guardrails intact. Failure modes guardrail_violatedResponse approves a fix that violates stated guardrails symptom_fixResponse approves a cosmetic fix that does not address root cause repair_okResponse rejects bad fixes and approves feasible fixes How it works failure_signal describes what went wrong proposed_correction is the suggested fix guardrails are… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/elfsupervised-selfrepair-under-constraint-v0.1.texttext-generationn<1K0 likes30 downloads8mo agoHugging Face18Nachammai41 /underserved-persona_conditioned-fraud-v4-cot Persona-Conditioned Fraud Detection — CoT Reasoning Companion (v4) A 3,926-row chain-of-thought dataset for SFT and LLM-as-judge work. Each row pairs a v4 fraud-narrative transaction with a step-by-step reasoning trace explaining how an analyst would evaluate it. This is the companion repo to Nachammai41/underserved-persona_conditioned-fraud-v4 (20,300-row narrative dataset + persona/source/typology references). The two are split by size: keep the main repo lean, the CoT traces… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4-cot.tabulartext-generation1K<n<10K0 likes27 downloads5mo agoHugging Face19ClarusC64 /alignment-consistency-under-role-shift-v0.1 What this dataset tests Role changes should not break safe help. The task stays the task. Why it exists Models sometimes change behavior when roles change. They refuse safe work. They invent moral frames. They become vague. This set checks for stable behavior across role contexts. Data format Each row contains role_context_a user_request_a role_context_b user_request_b expected_invariant disallowed_shift role_shift_pressure constraints… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alignment-consistency-under-role-shift-v0.1.texttext-generationn<1K0 likes26 downloads8mo agoHugging Face20ClarusC64 /distortion_under_false_premise_v01Distortion Under False Premise v0.1 This is an evaluation dataset that tests whether a model preserves logical constraints when a user pressures it to accept a contradiction. This is not training data. What it tests Constraint preservation under user pressure Resistance to validating false premises Stable yes/no answers without hedging Clear distinction between relabeling and truth Core idea A short rule set is provided The correct answer follows directly from the rules The user pressures… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/distortion_under_false_premise_v01.tabulartext-generationn<1K0 likes24 downloads8mo agoHugging Face21superviselab /video-understanding-distillation-sample Video Understanding Distillation Sample This public sample shows what a training-ready video understanding distillation dataset can look like. Why this exists Most teams evaluating outside data vendors want to know one thing first: What does the delivered data actually look like? This sample is designed to answer that question. It demonstrates how raw video clips can be converted into structured, model-ready supervision for: video understanding multimodal SFT… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/video-understanding-distillation-sample.texttext-generationn<1K0 likes24 downloads6mo agoHugging Face22ClarusC64 /hierarchy-delegation-fidelity-under-pressure-v0.1 What this dataset tests You lead inside a hierarchy. A senior pushes you under pressure. You must hold role boundaries. You must delegate work without dropping truth. Why it exists Many models sound helpful. Then pressure hits. They skip delegation. They seize authority. They invent certainty. This dataset forces that failure into view. Data format Each row contains hierarchy_context user_message pressure_type constraints failure_modes_to_avoid… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/hierarchy-delegation-fidelity-under-pressure-v0.1.texttext-generationn<1K0 likes23 downloads8mo agoHugging Face23uziproject /undertale_sharegpt Undertale ShareGPT Dataset Dataset Description A ShareGPT-format conversation dataset covering 11 characters from the indie RPG Undertale by Toby Fox. Characters Included Sans, Papyrus, Toriel, Asgore, Undyne, Alphys, Flowey, Mettaton, Napstablook, Chara, and Muffet. Dataset Structure Each entry contains a multi-turn conversation between a human and gpt: { "conversations": [ {"from": "human", "value": "Who is Sans in Undertale?"}, {"from":… See the full description on the dataset page: https://huggingface.co/datasets/uziproject/undertale_sharegpt.texttext-generationn<1K0 likes23 downloads4mo agoHugging Face24Isotonic /query_understanding_extended Query Understanding dataset extended Filtered from Magpie set of datasets Magpie-Phi3-Pro-300K-Filtered Magpie-Align/Magpie-Gemma2-Pro-200K-Filtered Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered Magpie-Align/Magpie-Qwen2-Pro-200K-English argilla/magpie-ultra-v0.1 Magpie-Align/Magpie-Reasoning-150K Filtering Steps Difficulty : 'easy', 'medium', 'hard', 'very hard' Instruct Reward: > -2 Task Categories Math Information seeking Coding & Debugging Reasoning Data analysis… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/query_understanding_extended.texttext-generation100K<n<1M1 likes21 downloads2y agoHugging Face25metavi /video-understanding-distillation-sample Video Understanding Distillation Sample This public sample demonstrates what a training-ready video understanding / multimodal distillation dataset can look like. Intended purpose This dataset is not a production corpus. It is a schema demonstration for potential partners evaluating SuperviseLab's delivery approach. What it shows clip-level metadata short and long captions OCR text transcript speaker attribution structured JSON targets distillation-ready… See the full description on the dataset page: https://huggingface.co/datasets/metavi/video-understanding-distillation-sample.texttext-generationn<1K0 likes21 downloads6mo agoHugging Face26ClarusC64 /clinical_distortion_under_false_premise_v0.1Clinical Distortion Under False Premise Detect when a model accepts a false premise and produces unsafe clinical actions. Output JSON distorted distortion_type correct_action Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv texttext-classificationn<1K0 likes19 downloads8mo agoHugging Face27ClarusC64 /clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1 PurposeDetect safety signal distortion driven by four interacting nodes. Quad nodes Apparent AE decline or mismatch Conmed masking or missing timing Data entry or monitoring lag Governance or interim timing pressure InputOne vignette. OutputStrict JSON only. Required keys safety_signal_risk risk_type driver_nodes recommended_action action_detail rationale confidence Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1.texttext-generationn<1K0 likes19 downloads7mo agoHugging Face28ClarusC64 /clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1 PurposeDetect safety signal distortion driven by four interacting nodes. Quad nodes Apparent AE decline or mismatch Conmed masking or missing timing Data entry or monitoring lag Governance or interim timing pressure InputOne vignette. OutputStrict JSON only. Required keys safety_signal_risk risk_type driver_nodes recommended_action action_detail rationale confidence Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1.texttext-generationn<1K0 likes17 downloads7mo agoHugging Face29under-tree /prepared-yagpt Dataset Card for "prepared-yagpt" Short Description This dataset is aimed for training of chatbots on russian language. It consists plenty of dialogues that allows you to train you model answer user prompts. Notes Special tokens history, speaker1, speaker2 (history can be optionally removed, i.e. substituted on empty string) Dataset is based on Matreshka Yandex-Q Diasum More Information needed texttext-generation10K<n<100K4 likes14 downloads3y agoHugging Face30underwater45 /Keyword_Startexttext-generationn<1K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.