datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models".
We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/C-VARC.vargov-design-catalog
Vargov® Design Catalog — 605 lighting and decorative compositions in 8 languages
A machine-readable catalog of the full body of work of Vargov® Design, an author-driven
studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow).
Every record is one composition: its identifier, category, canonical URLs, image links,
awards, links to its 3D model, and editorial copy written by the studio in eight
languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.amazon-c2-varied-rubrics
Amazon C2 varied-rubric distillation
This release exposes six balanced C2 SFT configurations: latent-state and non-diverse candidate panels at
K=1, K=2, and K=4 rubrics per retained reviewer. Each rubric-writer target is paired with one full-rubric
listwise judge target over the same variant's frozen 40-candidate panel. The K arms within a variant share one
reviewer cohort and are exact nested prefixes.
Config
Train rows
Validation
Test
Train reviewers… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-varied-rubrics.opengloss-v1.3-encyclopedia-variants
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Encyclopedia Variants v1.3
Dataset Summary
OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries
rewritten in multiple writing styles.… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-encyclopedia-variants.immersive_translate_en-zh
MiniCPM5-1B Immersive Translation SFT Dataset
英译中微调数据集,专为沉浸式翻译插件场景设计。用于微调 MiniCPM5-1B-Base,使其在插件运行时稳定遵循翻译规则、保留代码与 HTML 格式、正确处理多段 %% 分隔。
数据集描述
本数据集主要训练以下能力:
严格遵循沉浸式翻译 system prompt 中的 5 条翻译规则
多段输入的 %% 段落分隔,输入输出段落数严格一致
代码块、行内代码、HTML 标签、URL、专有名词的原样保留
技术文档(GitHub README、Hugging Face 文档)与学术摘要(arXiv)的英译中
单段输入直接输出译文,无"翻译:"等额外前缀
数据格式为 ShareGPT 对话格式,每条样本包含 system / user / assistant 三角色。
数据来源
来源
说明
原始规模
本数据集采样量
License
Mxode/BiST… See the full description on the dataset page: https://huggingface.co/datasets/Variable65536/immersive_translate_en-zh.C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models".
We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/stelvita/C-VARC.various_textsThese texts are used in our "Hour of code" activity on computing and using ngrams.
SPSD-Variants-opsd
SPSD-Variants-opsd
Grounded on-policy self-distillation (OPSD) teacher-context dataset over 45
board-game rule variants (5 families × 9: connect4, domineering,
simplified_first_attack, simplified_othello, tic_tac_chess), derived from
trained MuZero/EfficientZero checkpoints (plan-528 v2).
Each row is a decision-state task (a move choice or one of six auxiliary
state-QA tasks). The privileged_context is the teacher signal: grounded
natural-language reasoning that discovers the… See the full description on the dataset page: https://huggingface.co/datasets/LorMolf/SPSD-Variants-opsd.bigcodebench-typo-variants
BigCodeBench Typo Variants
This dataset contains typo-injected variants of the BigCodeBench coding benchmark to evaluate the robustness of code generation models to typographical errors in problem descriptions.
Dataset Description
BigCodeBench is a benchmark for evaluating large language models on diverse and challenging coding tasks. This dataset provides 7 variants with different levels of typos injected into the instruction prompts:
Original (0% typos): Clean baseline… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/bigcodebench-typo-variants.tb-explore17-mcode-m3-harness-variance
Terminal-Bench 2.1 explore-17 — mcode / MiniMax-M3 harness variance
Three complete 17-task runs of the same dataset ref with the same agent and
model, differing only in execution substrate and concurrency, plus one isolated
rerun. The point of the bundle is not the resolve rate — it is how much the
resolve rate moves when nothing about the task or the model changes.
Same everywhere: dataset ai-solution-finetune/terminal-bench-2-1-explore-17 at… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/tb-explore17-mcode-m3-harness-variance.clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1Clarus Clinical Quad Coupling PK Integrity v0.1
PurposeDetect PK integrity distortion driven by four interacting nodes.
Quad nodes
Sampling window deviation
Bioanalytical or stability variance
Dose adjustment decisions
Governance interim or submission timing
InputOne vignette.
OutputStrict JSON only.
Required keys
pk_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py
Run scoringCreate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1.rtllm-variants
RTLLM Variants (Community Dataset Derivatives)
This repository hosts community-maintained derivative variants based on upstream RTLLM releases. It is intended for benchmarking reproducibility and evaluation workflow integration.
Important Notice
This repository is not an official release from the RTLLM paper authors.
Variants in this repository may include prompt wording normalization, interface naming unification, or testbench output standardization for evaluator… See the full description on the dataset page: https://huggingface.co/datasets/xxrjun/rtllm-variants.gsm8k
Dataset Card for GSM8K
Dataset Summary
GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.
These problems take between 2 and 8 steps to solve.
Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to reach the… See the full description on the dataset page: https://huggingface.co/datasets/Amalin-Varsha/gsm8k.opengloss-v1.1-encyclopedia-variants
OpenGloss Encyclopedia Variants v1.1
Dataset Summary
OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries
rewritten in multiple writing styles. Each record contains an academic base entry alongside
a variant rewritten for a specific audience, tone, and content structure.
This dataset supports style transfer, text simplification, paraphrase generation, and
audience-adaptive content creation. It is derived from the
OpenGloss
encyclopedic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-encyclopedia-variants.urdu-english-name-variants
Names Dataset (English-Urdu-Variants)
This dataset contains names with their standardized English form, Urdu script, and common English variants.
Dataset Structure
en_std: Standardized English name
ur: Name in Urdu script
en_var: Common English variants/spellings of the name
Usage
from datasets import load_dataset
dataset = load_dataset("muhammadUsman31254/urdu-english-name-variants")
Languages
English (primary and variants)
Urdu
Use… See the full description on the dataset page: https://huggingface.co/datasets/muhammadUsman31254/urdu-english-name-variants.clinical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1Clarus Clinical Quad Coupling Protocol Deviation Staffing Drift Adjudication Variance Missingness Bias v0.1
What this dataset isThis dataset tests whether a model can detect protocol deviation events driven by quad coupling.
Quad coupling nodes
Operational staffing drift or site capacity constraint
Protocol compliance breakdown
Endpoint adjudication variance or bias risk
Data missingness that distorts safety or efficacy interpretation under governance rules
Input
One vignette in… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1.clinical-variable-fr
Clinical Variable Extraction Dataset (French)
Dataset Description
This dataset contains French clinical notes paired with their original text and successfully extracted clinical variables. Only variables with non-None values are included, making it ideal for training and evaluating models on clinical variable extraction tasks in French medical texts.
Dataset Structure
The dataset contains 3 columns:
text_original: Original clinical notes from medical cases… See the full description on the dataset page: https://huggingface.co/datasets/rntc/clinical-variable-fr.MathSmith-Self-Improvement-VarientSet
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
This dataset contains variant problems generated by the MathSmith Self-Improvement Pipeline, introduced in the paper MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy.
MathSmith is a framework for synthesizing challenging mathematical problems to enhance LLM reasoning. Rather than modifying existing problems… See the full description on the dataset page: https://huggingface.co/datasets/Jasaxion/MathSmith-Self-Improvement-VarientSet.prompt-variations
Prompt Variations and LLM Responses
Prompt variants and model responses used to evaluate the
Stability-Generalization Score (SGS) across eleven LLMs (eight
open-source + three closed-source) on six QA / instruction benchmarks
under six families of stylistic perturbations.
Splits
split
rows
source dataset
truthful_qa
99,888
TruthfulQA
natural_questions
41,040
Natural Questions
alpaca
13,872
Alpaca
simpleqa_verified
13,872
SimpleQA Verified… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/prompt-variations.opengloss-v1.2-encyclopedia-variants
OpenGloss Encyclopedia Variants v1.2
Dataset Summary
OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries
rewritten in multiple writing styles. Each record contains an academic base entry alongside
a variant rewritten for a specific audience, tone, and content structure.
This dataset supports style transfer, text simplification, paraphrase generation, and
audience-adaptive content creation. It is derived from the
OpenGloss
encyclopedic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-encyclopedia-variants.various_topics_articles_azerbaijanArticles Dataset in Azerbaijani
Description
This dataset contains various topics articles in Azerbaijani language. It was created in 2024 and contains 236k articles (approximately 1 million sentences).
License
The dataset is licensed under the Creative Commons Attribution-NonCommercial 4.0 International license. This license allows you to freely share and redistribute the dataset with attribution to the source but prohibits commercial use.
Contact information
If you have any questions or… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/various_topics_articles_azerbaijan.classical-chinese-variant-collation
Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset)
This is a commercial dataset. A free 50-work preview sample is provided
below (sample.jsonl, texts truncated); the full set with complete aligned
texts is available upon request.
📧 To license / purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — both members are public-domain classical works.
What this is
A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.smolified-tiny-text-to-sql
🤏 smolified-tiny-text-to-sql
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model vardhan-yash/smolified-tiny-text-to-sql.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 4b9509ca)
Records: 320
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by vardhan-yash.
Generated via Smolify.ai.
urdu-english-name-variants
Names Dataset (English-Urdu-Variants)
This dataset contains names with their standardized English form, Urdu script, and common English variants.
Dataset Structure
en_std: Standardized English name
ur: Name in Urdu script
en_var: Common English variants/spellings of the name
Usage
from datasets import load_dataset
dataset = load_dataset("muhammadUsman31254/urdu-english-name-variants")
Languages
English (primary and variants)… See the full description on the dataset page: https://huggingface.co/datasets/feifan961206/urdu-english-name-variants.varxipod-k8s-remediation
VarXiPod K8s Remediation Dataset
Dataset Description
Synthetic dataset for training agentic K8s remediation models. Contains event
diagnosis, remediation planning, tool calling, runbook execution, and confidence
scoring examples. Generated from VarXiPod operator domain knowledge.
This dataset is designed to fine-tune language models for autonomous Kubernetes
cluster remediation. Each example represents a realistic scenario drawn from
production operator experience… See the full description on the dataset page: https://huggingface.co/datasets/northriverfence/varxipod-k8s-remediation.adaptive_rag_hotpotqa
Adaptive RAG HotpotQA Dataset
This dataset is a processed version of HotpotQA designed for training Adaptive Retrieval-Augmented Generation (RAG) systems.
Features
input: The input text for the model
output: The target output text
retrieval_label: Whether retrieval is needed (0/1)
hop: The reasoning hop number (1 or 2)
type: The type of example (multi_hop_qa, single_hop_qa, multi_hop_gating, etc.)
metadata: Additional information about the example including:
answer:… See the full description on the dataset page: https://huggingface.co/datasets/varun500/adaptive_rag_hotpotqa.variational-sd-qwen3-8b-sharegpt-rollouts
Qwen3-8B Regenerated ShareGPT Rollouts
This is the exact target-regenerated ShareGPT JSONL used to train the
Qwen3-8B D-PACE A512 one-epoch checkpoint and the discrete M=8 A512 one-epoch
checkpoint in
Nicholas0228/variational-sd.
Contents
data/train.jsonl: 78,810 successful regenerated rows, 983,686,208 bytes.
metadata.json: generation settings, row statistics, and checksum.
Each JSONL row contains:
{
"id": string,
"status": "success",
"conversations": [… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/variational-sd-qwen3-8b-sharegpt-rollouts.gutenberg-fi
Finnish Project Gutenberg Books
Dataset Description
A collection of 3,505 Finnish-language books from Project Gutenberg, extracted from a January 2026 ZIM archive and converted to Markdown.
Statistics
Total books
3,505
Unique authors
1,101
Books with translator
1,451
Total text
~0.9 GB
Median book length
178k characters
Mean book length
258k characters
Min book length
~12k characters
Max book length
~2.5M characters… See the full description on the dataset page: https://huggingface.co/datasets/Varho/gutenberg-fi.VARAG
