datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
balanced-copa-explanations
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/zuzannad1/balanced-copa-explanations.gaokao-sft-chinese-balanced
Gaokao SFT Chinese Balanced
This dataset is a cleaned SFT-style Chinese exam dataset prepared from multiple public Hugging Face sources.
Composition
Total samples: 1895
Train samples: 1853
Validation samples: 42
Fields
Each row contains:
id
lang
subject
source
instruction
input
output
messages
Cleaning Notes
Ordinary Markdown markers were removed.
Non-essential LaTeX commands were simplified into plain readable text.
Math expressions were… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-balanced.nexa-science-multitask-balanced
Nexa Science Multitask Balanced
This dataset is a curated, instruction-formatted scientific multitask mixture for:
claim verification (<TASK:VERIFY>)
abstract-grounded biomedical QA (<TASK:QA>)
retrieval relevance re-ranking (<TASK:RERANK>)
Format
Each row is JSONL with:
{task, instruction, input, output, meta}
Splits Included
train_balanced_short.jsonl
val_balanced_short.jsonl
stats_balanced_short.json
Notes
QA in this balanced release is… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/nexa-science-multitask-balanced.
