datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
balanced-copa
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/pkavumba/balanced-copa.balanced-copa-explanations
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/zuzannad1/balanced-copa-explanations.The-Balanced-Budget-And-Emergency-Deficit-Control-Act-of-1985
Dataset Description
The Balanced Budget and Emergency Deficit Control Act of 1985 Question-Answer Dataset is an English-language instructional dataset derived from the statutory provisions of the Balanced Budget and Emergency Deficit Control Act of 1985.
The Act was enacted as Title II of Public Law 99-177 on December 12, 1985, and is commonly known as the Gramm-Rudman-Hollings Act. Its provisions established federal budget-enforcement mechanisms intended to control deficits… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/The-Balanced-Budget-And-Emergency-Deficit-Control-Act-of-1985.Code-Vulnerability-Balanced
Code Vulnerability Balanced — CWE-Enriched Conversation Dataset
📌 Overview
This dataset is a balanced and shuffled version of
ChamaraVishwajithRajapaksha/Code-Vulnerability-FineTune,
which itself was derived from the original
ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset
(330k rows, sourced from DiverseVul + MITRE CWE enrichment).
The original fine-tuning dataset was imbalanced — the number of Vulnerable and Safe
samples were not equal — and the samples were not… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/Code-Vulnerability-Balanced.gaokao-sft-chinese-balanced
Gaokao SFT Chinese Balanced
This dataset is a cleaned SFT-style Chinese exam dataset prepared from multiple public Hugging Face sources.
Composition
Total samples: 1895
Train samples: 1853
Validation samples: 42
Fields
Each row contains:
id
lang
subject
source
instruction
input
output
messages
Cleaning Notes
Ordinary Markdown markers were removed.
Non-essential LaTeX commands were simplified into plain readable text.
Math expressions were… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-balanced.cmmc-training-balanced
CMMC Training Dataset - Balanced Variant
Dataset Description
This is the Balanced variant of the CMMC (Cybersecurity Maturity Model Certification) training dataset, containing 2,790 high-quality training examples with balanced coverage across all 17 CMMC domains.
Dataset Characteristics
Total Examples: 2,790 (2,232 train / 558 validation)
Source Documents: 71 NIST publications
CMMC Levels Covered: Level 1, Level 2, Level 3
CMMC Domains: All 17 domains (evenly… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/cmmc-training-balanced.nexa-science-multitask-balanced
Nexa Science Multitask Balanced
This dataset is a curated, instruction-formatted scientific multitask mixture for:
claim verification (<TASK:VERIFY>)
abstract-grounded biomedical QA (<TASK:QA>)
retrieval relevance re-ranking (<TASK:RERANK>)
Format
Each row is JSONL with:
{task, instruction, input, output, meta}
Splits Included
train_balanced_short.jsonl
val_balanced_short.jsonl
stats_balanced_short.json
Notes
QA in this balanced release is… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/nexa-science-multitask-balanced.domain-agnostic-reasoning-traces-balanced-top50-v1
BOTCOIN Balanced Top-50 Reasoning Traces
This public dataset contains enriched BOTCOIN reasoning-trace attempts selected
from canonical dataset/v2 research-ready objects.
Selection policy:
Source only attempts/research-ready objects.
Rank each domain by trace_quality.reasoning_trace_quality_score.
Keep each domain's top 50 percent.
Equalize domains to the smallest top-half count.
The rows are self-contained and intentionally rich: prompt/messages… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1.SQuAD2.0_balancedAn extended version of SQuAD 2.0 created by adding 50K+ unanswerable questions (from other contexts within the same article). The sole purpose is to improve model's ability to detect unanswerable questions.
