CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Beijing-AISI /C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/C-VARC.texttext-generation100K<n<1M2 likes3.1k downloads1y agoHugging Face02mjbommar /opengloss-v1.3-encyclopedia-variants See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Encyclopedia Variants v1.3 Dataset Summary OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries rewritten in multiple writing styles.… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-encyclopedia-variants.texttext-generation100K<n<1M0 likes95 downloads18d agoHugging Face03Variable65536 /immersive_translate_en-zh MiniCPM5-1B Immersive Translation SFT Dataset 英译中微调数据集,专为沉浸式翻译插件场景设计。用于微调 MiniCPM5-1B-Base,使其在插件运行时稳定遵循翻译规则、保留代码与 HTML 格式、正确处理多段 %% 分隔。 数据集描述 本数据集主要训练以下能力: 严格遵循沉浸式翻译 system prompt 中的 5 条翻译规则 多段输入的 %% 段落分隔,输入输出段落数严格一致 代码块、行内代码、HTML 标签、URL、专有名词的原样保留 技术文档(GitHub README、Hugging Face 文档)与学术摘要(arXiv)的英译中 单段输入直接输出译文,无"翻译:"等额外前缀 数据格式为 ShareGPT 对话格式,每条样本包含 system / user / assistant 三角色。 数据来源 来源 说明 原始规模 本数据集采样量 License Mxode/BiST… See the full description on the dataset page: https://huggingface.co/datasets/Variable65536/immersive_translate_en-zh.texttranslation10K<n<100K0 likes88 downloads13d agoHugging Face04stelvita /C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/stelvita/C-VARC.texttext-generation100K<n<1M0 likes84 downloads7mo agoHugging Face05LorMolf /SPSD-Variants-opsd SPSD-Variants-opsd Grounded on-policy self-distillation (OPSD) teacher-context dataset over 45 board-game rule variants (5 families × 9: connect4, domineering, simplified_first_attack, simplified_othello, tic_tac_chess), derived from trained MuZero/EfficientZero checkpoints (plan-528 v2). Each row is a decision-state task (a move choice or one of six auxiliary state-QA tasks). The privileged_context is the teacher signal: grounded natural-language reasoning that discovers the… See the full description on the dataset page: https://huggingface.co/datasets/LorMolf/SPSD-Variants-opsd.texttext-generation100K<n<1M0 likes60 downloads26d agoHugging Face06miaomiao64 /tb-explore17-mcode-m3-harness-variancegated Terminal-Bench 2.1 explore-17 — mcode / MiniMax-M3 harness variance Three complete 17-task runs of the same dataset ref with the same agent and model, differing only in execution substrate and concurrency, plus one isolated rerun. The point of the bundle is not the resolve rate — it is how much the resolve rate moves when nothing about the task or the model changes. Same everywhere: dataset ai-solution-finetune/terminal-bench-2-1-explore-17 at… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/tb-explore17-mcode-m3-harness-variance.tabulartext-generationn<1K0 likes31 downloads16d agoHugging Face07xxrjun /rtllm-variantsgated RTLLM Variants (Community Dataset Derivatives) This repository hosts community-maintained derivative variants based on upstream RTLLM releases. It is intended for benchmarking reproducibility and evaluation workflow integration. Important Notice This repository is not an official release from the RTLLM paper authors. Variants in this repository may include prompt wording normalization, interface naming unification, or testbench output standardization for evaluator… See the full description on the dataset page: https://huggingface.co/datasets/xxrjun/rtllm-variants.texttext-generationn<1K0 likes25 downloads7mo agoHugging Face08mjbommar /opengloss-v1.1-encyclopedia-variants OpenGloss Encyclopedia Variants v1.1 Dataset Summary OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries rewritten in multiple writing styles. Each record contains an academic base entry alongside a variant rewritten for a specific audience, tone, and content structure. This dataset supports style transfer, text simplification, paraphrase generation, and audience-adaptive content creation. It is derived from the OpenGloss encyclopedic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-encyclopedia-variants.texttext-generation1K<n<10K0 likes23 downloads10mo agoHugging Face09Jasaxion /MathSmith-Self-Improvement-VarientSet MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy This dataset contains variant problems generated by the MathSmith Self-Improvement Pipeline, introduced in the paper MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy. MathSmith is a framework for synthesizing challenging mathematical problems to enhance LLM reasoning. Rather than modifying existing problems… See the full description on the dataset page: https://huggingface.co/datasets/Jasaxion/MathSmith-Self-Improvement-VarientSet.texttext-generationn<1K0 likes20 downloads7mo agoHugging Face10mjbommar /opengloss-v1.2-encyclopedia-variants OpenGloss Encyclopedia Variants v1.2 Dataset Summary OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries rewritten in multiple writing styles. Each record contains an academic base entry alongside a variant rewritten for a specific audience, tone, and content structure. This dataset supports style transfer, text simplification, paraphrase generation, and audience-adaptive content creation. It is derived from the OpenGloss encyclopedic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-encyclopedia-variants.texttext-generation10K<n<100K0 likes19 downloads6mo agoHugging Face11wangekxy /classical-chinese-variant-collation Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset) This is a commercial dataset. A free 50-work preview sample is provided below (sample.jsonl, texts truncated); the full set with complete aligned texts is available upon request. 📧 To license / purchase or request a quote, email wangeksy@gmail.com. ✅ Cleared for commercial use — both members are public-domain classical works. What this is A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.tabulartext-generationn<1K0 likes16 downloads3mo agoHugging Face12Nicholas0228 /variational-sd-qwen3-8b-sharegpt-rolloutsgated Qwen3-8B Regenerated ShareGPT Rollouts This is the exact target-regenerated ShareGPT JSONL used to train the Qwen3-8B D-PACE A512 one-epoch checkpoint and the discrete M=8 A512 one-epoch checkpoint in Nicholas0228/variational-sd. Contents data/train.jsonl: 78,810 successful regenerated rows, 983,686,208 bytes. metadata.json: generation settings, row statistics, and checksum. Each JSONL row contains: { "id": string, "status": "success", "conversations": [… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/variational-sd-qwen3-8b-sharegpt-rollouts.texttext-generation10K<n<100K0 likes12 downloads2mo agoHugging Face13Varho /gutenberg-fi Finnish Project Gutenberg Books Dataset Description A collection of 3,505 Finnish-language books from Project Gutenberg, extracted from a January 2026 ZIM archive and converted to Markdown. Statistics Total books 3,505 Unique authors 1,101 Books with translator 1,451 Total text ~0.9 GB Median book length 178k characters Mean book length 258k characters Min book length ~12k characters Max book length ~2.5M characters… See the full description on the dataset page: https://huggingface.co/datasets/Varho/gutenberg-fi.texttext-generation1K<n<10K0 likes9 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.