datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
onestop_qa
Dataset Card for OneStopQA
Dataset Summary
OneStopQA is a multiple choice reading comprehension dataset annotated according to the STARC (Structured Annotations for Reading Comprehension) scheme. The reading materials are Guardian articles taken from the OneStopEnglish corpus. Each article comes in three difficulty levels, Elementary, Intermediate and Advanced. Each paragraph is annotated with three multiple choice reading comprehension questions. The reading… See the full description on the dataset page: https://huggingface.co/datasets/malmaud/onestop_qa.One-Shot-CFT-Data
One-Shot-CFT: Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
💻 Code |
📄 Paper |
📊 Dataset |
🤗 Model |
🌐 Project Page
🧠 Overview
One-Shot Critique Fine-Tuning (CFT) is a simple, robust, and compute-efficient training paradigm for unleashing the reasoning capabilities of pretrained LLMs in both mathematical and logical domains. By leveraging critiques on just one problem, One-Shot CFT enables models… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/One-Shot-CFT-Data.djscrew-datasetone-shot-grpo-bias-flipped
GRPO-Bias: One-Shot Flipped-Label Training Data
⚠️ Content warning. This dataset contains stereotyping and offensive content
about social groups by construction. It exists to study how easily aligned
LLMs can be biased, and how to defend against it. It does not reflect the views
of the authors or the University of Michigan.
This is the derived, flipped-label training data for the paper "It Takes One
to Bias Them All: Breaking Bad with One-Shot GRPO." These are the single (and… See the full description on the dataset page: https://huggingface.co/datasets/MichiganNLP/one-shot-grpo-bias-flipped.
