datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
k12-mathematics-standards-aligned
[!WARNING]
Deprecated - use k12-mathematics-standards-expanded instead.
This dataset is superseded: every input in this set also appears there, plus 366 more and two additional metadata columns. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/k12-mathematics-standards-expanded.
K-12 Mathematics Standards (generated instruction data)
4,397 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-aligned.pleias-post-ocr-correction-chonkie-aligned-en
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.glm52-aligned-rubric-traces
GLM-5.2 Aligned Rubric-Writing Traces
23777 teacher traces from GLM-5.2 on the aligned rubric-writing task, collected to
distill / warmstart a smaller rubric-writer. For each (user, book) example the teacher is shown a
persona-conditioned prompt (a user's past book reviews) and asked to (1) predict what that user
would likely write about a new book and (2) produce a <rubric> of numbered criteria for scoring
candidate reviews on coverage of that prediction. The full generation —… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/glm52-aligned-rubric-traces.m9-verifier-38k-aligned
M9 Verifier 38K Aligned
This dataset contains 38,564 prompts with verifier-compatible gold answers for
an M9 RLVR-GRPO experiment in a unified post-training study with Qwen3-1.7B.
It is an independent research artifact, not an official release from the model
or paper authors.
The bank was reconstructed from the frozen
YangyiH/openreasoning_mixed_100k
prompt mixture. Every recovered row was matched to the frozen base row by
domain, source shard, and prompt SHA-256 before verifier… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/m9-verifier-38k-aligned.pleias-post-ocr-correction-chonkie-aligned-fr
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.TACTBench-Samples
TACTBench Demonstration Samples
This repository contains five full-context demonstration examples from
TACTBench. It does not contain the TACT training set or the remaining hidden
TACTBench evaluation set. The samples use the same full-history representation
as the benchmark evaluation and illustrate direct correction, error
explanation, guided revision, clarification checking, affective feedback, and
retry elicitation.
Data
data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.safety_aligned_datasets
Safety Aligned Datasets
A high-fidelity adversarial corpus engineered for alignment research, refusal boundary modeling, and robustness evaluation of Small Language Models.
The Problem This Solves
Fine-tuning a Small Language Model to be safe is not the same as fine-tuning it to understand safety.
Most safety datasets give models clean refusal examples on obvious prompts — and those models fail the moment an adversary wraps a harmful request in a… See the full description on the dataset page: https://huggingface.co/datasets/vvsd-charan/safety_aligned_datasets.Bilingual-SFT-2.0-Pashto-English-Aligned
Bilingual SFT 2.0 — Pashto English Aligned 🇦🇫🇬🇧
Bilingual-SFT-2.0-Pashto-English-Aligned is a bilingual supervised fine-tuning dataset designed to improve Large Language Models (LLMs) in Pashto ↔ English understanding, instruction following, conversation, and bilingual generation.
The dataset uses a conversational messages format and is intended for modern instruction-tuning pipelines, including Hugging Face Transformers, TRL, Unsloth, Axolotl, and other SFT frameworks.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-2.0-Pashto-English-Aligned.culturally_aligned_arabic_stories_subset_a
📚 Culturally Aligned Arabic Stories Dataset (Subset A)
A curated 110-example subset of the Crafting Culturally Aligned Narratives dataset, designed for the development and evaluation of Arabic children’s story generation models aligned with Islamic and cultural values.
✨ Overview
Language: Modern Standard Arabic (MSA)
Samples: 110 prompt–response pairs
Format: JSONL (id, language, prompt, response, source, license)
Moral domains: honesty, courage, generosity… See the full description on the dataset page: https://huggingface.co/datasets/houssamboukhalfa/culturally_aligned_arabic_stories_subset_a.
