CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Beijing-AISI /C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/C-VARC.texttext-generation100K<n<1M2 likes3.1k downloads1y agoHugging Face02vargov-design /vargov-design-catalog Vargov® Design Catalog — 605 lighting and decorative compositions in 8 languages A machine-readable catalog of the full body of work of Vargov® Design, an author-driven studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow). Every record is one composition: its identifier, category, canonical URLs, image links, awards, links to its 3D model, and editorial copy written by the studio in eight languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.imagetext-retrieval1K<n<10K0 likes345 downloads3d agoHugging Face03asingh15 /amazon-c2-varied-rubrics Amazon C2 varied-rubric distillation This release exposes six balanced C2 SFT configurations: latent-state and non-diverse candidate panels at K=1, K=2, and K=4 rubrics per retained reviewer. Each rubric-writer target is paired with one full-rubric listwise judge target over the same variant's frozen 40-candidate panel. The K arms within a variant share one reviewer cohort and are exact nested prefixes. Config Train rows Validation Test Train reviewers… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-varied-rubrics.tabulartext-generation100K<n<1M0 likes276 downloads6d agoHugging Face04mjbommar /opengloss-v1.3-encyclopedia-variants See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Encyclopedia Variants v1.3 Dataset Summary OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries rewritten in multiple writing styles.… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-encyclopedia-variants.texttext-generation100K<n<1M0 likes95 downloads18d agoHugging Face05Variable65536 /immersive_translate_en-zh MiniCPM5-1B Immersive Translation SFT Dataset 英译中微调数据集,专为沉浸式翻译插件场景设计。用于微调 MiniCPM5-1B-Base,使其在插件运行时稳定遵循翻译规则、保留代码与 HTML 格式、正确处理多段 %% 分隔。 数据集描述 本数据集主要训练以下能力: 严格遵循沉浸式翻译 system prompt 中的 5 条翻译规则 多段输入的 %% 段落分隔,输入输出段落数严格一致 代码块、行内代码、HTML 标签、URL、专有名词的原样保留 技术文档(GitHub README、Hugging Face 文档)与学术摘要(arXiv)的英译中 单段输入直接输出译文,无"翻译:"等额外前缀 数据格式为 ShareGPT 对话格式,每条样本包含 system / user / assistant 三角色。 数据来源 来源 说明 原始规模 本数据集采样量 License Mxode/BiST… See the full description on the dataset page: https://huggingface.co/datasets/Variable65536/immersive_translate_en-zh.texttranslation10K<n<100K0 likes88 downloads13d agoHugging Face06stelvita /C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/stelvita/C-VARC.texttext-generation100K<n<1M0 likes84 downloads7mo agoHugging Face07Toxicode /various_textsThese texts are used in our "Hour of code" activity on computing and using ngrams. texttext-generation0 likes70 downloads3y agoHugging Face08LorMolf /SPSD-Variants-opsd SPSD-Variants-opsd Grounded on-policy self-distillation (OPSD) teacher-context dataset over 45 board-game rule variants (5 families × 9: connect4, domineering, simplified_first_attack, simplified_othello, tic_tac_chess), derived from trained MuZero/EfficientZero checkpoints (plan-528 v2). Each row is a decision-state task (a move choice or one of six auxiliary state-QA tasks). The privileged_context is the teacher signal: grounded natural-language reasoning that discovers the… See the full description on the dataset page: https://huggingface.co/datasets/LorMolf/SPSD-Variants-opsd.texttext-generation100K<n<1M0 likes60 downloads26d agoHugging Face09jeqcho /bigcodebench-typo-variants BigCodeBench Typo Variants This dataset contains typo-injected variants of the BigCodeBench coding benchmark to evaluate the robustness of code generation models to typographical errors in problem descriptions. Dataset Description BigCodeBench is a benchmark for evaluating large language models on diverse and challenging coding tasks. This dataset provides 7 variants with different levels of typos injected into the instruction prompts: Original (0% typos): Clean baseline… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/bigcodebench-typo-variants.texttext-generation1K<n<10K0 likes32 downloads10mo agoHugging Face10miaomiao64 /tb-explore17-mcode-m3-harness-variancegated Terminal-Bench 2.1 explore-17 — mcode / MiniMax-M3 harness variance Three complete 17-task runs of the same dataset ref with the same agent and model, differing only in execution substrate and concurrency, plus one isolated rerun. The point of the bundle is not the resolve rate — it is how much the resolve rate moves when nothing about the task or the model changes. Same everywhere: dataset ai-solution-finetune/terminal-bench-2-1-explore-17 at… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/tb-explore17-mcode-m3-harness-variance.tabulartext-generationn<1K0 likes31 downloads16d agoHugging Face11ClarusC64 /clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1Clarus Clinical Quad Coupling PK Integrity v0.1 PurposeDetect PK integrity distortion driven by four interacting nodes. Quad nodes Sampling window deviation Bioanalytical or stability variance Dose adjustment decisions Governance interim or submission timing InputOne vignette. OutputStrict JSON only. Required keys pk_integrity_risk risk_type driver_nodes recommended_action action_detail rationale confidence Filesdata/train.csvdata/test.csvscorer.py Run scoringCreate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1.texttext-generationn<1K0 likes25 downloads7mo agoHugging Face12xxrjun /rtllm-variantsgated RTLLM Variants (Community Dataset Derivatives) This repository hosts community-maintained derivative variants based on upstream RTLLM releases. It is intended for benchmarking reproducibility and evaluation workflow integration. Important Notice This repository is not an official release from the RTLLM paper authors. Variants in this repository may include prompt wording normalization, interface naming unification, or testbench output standardization for evaluator… See the full description on the dataset page: https://huggingface.co/datasets/xxrjun/rtllm-variants.texttext-generationn<1K0 likes25 downloads7mo agoHugging Face13Amalin-Varsha /gsm8k Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to reach the… See the full description on the dataset page: https://huggingface.co/datasets/Amalin-Varsha/gsm8k.texttext-generation10K<n<100K0 likes25 downloads5mo agoHugging Face14mjbommar /opengloss-v1.1-encyclopedia-variants OpenGloss Encyclopedia Variants v1.1 Dataset Summary OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries rewritten in multiple writing styles. Each record contains an academic base entry alongside a variant rewritten for a specific audience, tone, and content structure. This dataset supports style transfer, text simplification, paraphrase generation, and audience-adaptive content creation. It is derived from the OpenGloss encyclopedic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-encyclopedia-variants.texttext-generation1K<n<10K0 likes23 downloads10mo agoHugging Face15muhammadUsman31254 /urdu-english-name-variants Names Dataset (English-Urdu-Variants) This dataset contains names with their standardized English form, Urdu script, and common English variants. Dataset Structure en_std: Standardized English name ur: Name in Urdu script en_var: Common English variants/spellings of the name Usage from datasets import load_dataset dataset = load_dataset("muhammadUsman31254/urdu-english-name-variants") Languages English (primary and variants) Urdu Use… See the full description on the dataset page: https://huggingface.co/datasets/muhammadUsman31254/urdu-english-name-variants.texttext-classification1K<n<10K0 likes22 downloads1y agoHugging Face16ClarusC64 /clinical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1Clarus Clinical Quad Coupling Protocol Deviation Staffing Drift Adjudication Variance Missingness Bias v0.1 What this dataset isThis dataset tests whether a model can detect protocol deviation events driven by quad coupling. Quad coupling nodes Operational staffing drift or site capacity constraint Protocol compliance breakdown Endpoint adjudication variance or bias risk Data missingness that distorts safety or efficacy interpretation under governance rules Input One vignette in… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1.texttext-generationn<1K0 likes21 downloads7mo agoHugging Face17rntc /clinical-variable-fr Clinical Variable Extraction Dataset (French) Dataset Description This dataset contains French clinical notes paired with their original text and successfully extracted clinical variables. Only variables with non-None values are included, making it ideal for training and evaluating models on clinical variable extraction tasks in French medical texts. Dataset Structure The dataset contains 3 columns: text_original: Original clinical notes from medical cases… See the full description on the dataset page: https://huggingface.co/datasets/rntc/clinical-variable-fr.texttext-generationn<1K0 likes20 downloads1y agoHugging Face18Jasaxion /MathSmith-Self-Improvement-VarientSet MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy This dataset contains variant problems generated by the MathSmith Self-Improvement Pipeline, introduced in the paper MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy. MathSmith is a framework for synthesizing challenging mathematical problems to enhance LLM reasoning. Rather than modifying existing problems… See the full description on the dataset page: https://huggingface.co/datasets/Jasaxion/MathSmith-Self-Improvement-VarientSet.texttext-generationn<1K0 likes20 downloads7mo agoHugging Face19naghamo /prompt-variations Prompt Variations and LLM Responses Prompt variants and model responses used to evaluate the Stability-Generalization Score (SGS) across eleven LLMs (eight open-source + three closed-source) on six QA / instruction benchmarks under six families of stylistic perturbations. Splits split rows source dataset truthful_qa 99,888 TruthfulQA natural_questions 41,040 Natural Questions alpaca 13,872 Alpaca simpleqa_verified 13,872 SimpleQA Verified… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/prompt-variations.texttext-generation100K<n<1M1 likes20 downloads4mo agoHugging Face20mjbommar /opengloss-v1.2-encyclopedia-variants OpenGloss Encyclopedia Variants v1.2 Dataset Summary OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries rewritten in multiple writing styles. Each record contains an academic base entry alongside a variant rewritten for a specific audience, tone, and content structure. This dataset supports style transfer, text simplification, paraphrase generation, and audience-adaptive content creation. It is derived from the OpenGloss encyclopedic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-encyclopedia-variants.texttext-generation10K<n<100K0 likes19 downloads6mo agoHugging Face21LocalDoc /various_topics_articles_azerbaijangatedArticles Dataset in Azerbaijani Description This dataset contains various topics articles in Azerbaijani language. It was created in 2024 and contains 236k articles (approximately 1 million sentences). License The dataset is licensed under the Creative Commons Attribution-NonCommercial 4.0 International license. This license allows you to freely share and redistribute the dataset with attribution to the source but prohibits commercial use. Contact information If you have any questions or… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/various_topics_articles_azerbaijan.texttext-generation100K<n<1M3 likes17 downloads3y agoHugging Face22wangekxy /classical-chinese-variant-collation Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset) This is a commercial dataset. A free 50-work preview sample is provided below (sample.jsonl, texts truncated); the full set with complete aligned texts is available upon request. 📧 To license / purchase or request a quote, email wangeksy@gmail.com. ✅ Cleared for commercial use — both members are public-domain classical works. What this is A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.tabulartext-generationn<1K0 likes16 downloads3mo agoHugging Face23vardhan-yash /smolified-tiny-text-to-sql 🤏 smolified-tiny-text-to-sql Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model vardhan-yash/smolified-tiny-text-to-sql. 📦 Asset Details Origin: Smolify Foundry (Job ID: 4b9509ca) Records: 320 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by vardhan-yash. Generated via Smolify.ai. texttext-generationn<1K0 likes14 downloads7mo agoHugging Face24feifan961206 /urdu-english-name-variants Names Dataset (English-Urdu-Variants) This dataset contains names with their standardized English form, Urdu script, and common English variants. Dataset Structure en_std: Standardized English name ur: Name in Urdu script en_var: Common English variants/spellings of the name Usage from datasets import load_dataset dataset = load_dataset("muhammadUsman31254/urdu-english-name-variants") Languages English (primary and variants)… See the full description on the dataset page: https://huggingface.co/datasets/feifan961206/urdu-english-name-variants.texttext-classification1K<n<10K0 likes14 downloads3mo agoHugging Face25northriverfence /varxipod-k8s-remediation VarXiPod K8s Remediation Dataset Dataset Description Synthetic dataset for training agentic K8s remediation models. Contains event diagnosis, remediation planning, tool calling, runbook execution, and confidence scoring examples. Generated from VarXiPod operator domain knowledge. This dataset is designed to fine-tune language models for autonomous Kubernetes cluster remediation. Each example represents a realistic scenario drawn from production operator experience… See the full description on the dataset page: https://huggingface.co/datasets/northriverfence/varxipod-k8s-remediation.texttext-generationn<1K0 likes13 downloads7mo agoHugging Face26varun500 /adaptive_rag_hotpotqa Adaptive RAG HotpotQA Dataset This dataset is a processed version of HotpotQA designed for training Adaptive Retrieval-Augmented Generation (RAG) systems. Features input: The input text for the model output: The target output text retrieval_label: Whether retrieval is needed (0/1) hop: The reasoning hop number (1 or 2) type: The type of example (multi_hop_qa, single_hop_qa, multi_hop_gating, etc.) metadata: Additional information about the example including: answer:… See the full description on the dataset page: https://huggingface.co/datasets/varun500/adaptive_rag_hotpotqa.tabularquestion-answeringn<1K0 likes12 downloads1y agoHugging Face27Nicholas0228 /variational-sd-qwen3-8b-sharegpt-rolloutsgated Qwen3-8B Regenerated ShareGPT Rollouts This is the exact target-regenerated ShareGPT JSONL used to train the Qwen3-8B D-PACE A512 one-epoch checkpoint and the discrete M=8 A512 one-epoch checkpoint in Nicholas0228/variational-sd. Contents data/train.jsonl: 78,810 successful regenerated rows, 983,686,208 bytes. metadata.json: generation settings, row statistics, and checksum. Each JSONL row contains: { "id": string, "status": "success", "conversations": [… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/variational-sd-qwen3-8b-sharegpt-rollouts.texttext-generation10K<n<100K0 likes12 downloads2mo agoHugging Face28Varho /gutenberg-fi Finnish Project Gutenberg Books Dataset Description A collection of 3,505 Finnish-language books from Project Gutenberg, extracted from a January 2026 ZIM archive and converted to Markdown. Statistics Total books 3,505 Unique authors 1,101 Books with translator 1,451 Total text ~0.9 GB Median book length 178k characters Mean book length 258k characters Min book length ~12k characters Max book length ~2.5M characters… See the full description on the dataset page: https://huggingface.co/datasets/Varho/gutenberg-fi.texttext-generation1K<n<10K0 likes9 downloads8mo agoHugging Face29DeAllGamer /VARAGtext-generation0 likes4 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.