CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01robworks-software /k12-mathematics-standards-aligned [!WARNING] Deprecated - use k12-mathematics-standards-expanded instead. This dataset is superseded: every input in this set also appears there, plus 366 more and two additional metadata columns. Nothing here is unique to it. It stays online so existing references keep resolving, but it will not be updated. New work should point at robworks-software/k12-mathematics-standards-expanded. K-12 Mathematics Standards (generated instruction data) 4,397 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-aligned.texttext-generation1K<n<10K0 likes42 downloads2mo agoHugging Face02emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-en PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.texttext-generation100K<n<1M0 likes40 downloads3mo agoHugging Face03asingh15 /glm52-aligned-rubric-traces GLM-5.2 Aligned Rubric-Writing Traces 23777 teacher traces from GLM-5.2 on the aligned rubric-writing task, collected to distill / warmstart a smaller rubric-writer. For each (user, book) example the teacher is shown a persona-conditioned prompt (a user's past book reviews) and asked to (1) predict what that user would likely write about a new book and (2) produce a <rubric> of numbered criteria for scoring candidate reviews on coverage of that prediction. The full generation —… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/glm52-aligned-rubric-traces.tabulartext-generation10K<n<100K1 likes34 downloads2mo agoHugging Face04YangyiH /m9-verifier-38k-aligned M9 Verifier 38K Aligned This dataset contains 38,564 prompts with verifier-compatible gold answers for an M9 RLVR-GRPO experiment in a unified post-training study with Qwen3-1.7B. It is an independent research artifact, not an official release from the model or paper authors. The bank was reconstructed from the frozen YangyiH/openreasoning_mixed_100k prompt mixture. Every recovered row was matched to the frozen base row by domain, source shard, and prompt SHA-256 before verifier… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/m9-verifier-38k-aligned.tabulartext-generation10K<n<100K0 likes27 downloads1mo agoHugging Face05emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-fr PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.texttext-generation10K<n<100K0 likes24 downloads3mo agoHugging Face06Taxonomy-Aligned-Conversational-Tutor /TACTBench-Samples TACTBench Demonstration Samples This repository contains five full-context demonstration examples from TACTBench. It does not contain the TACT training set or the remaining hidden TACTBench evaluation set. The samples use the same full-history representation as the benchmark evaluation and illustrate direct correction, error explanation, guided revision, clarification checking, affective feedback, and retry elicitation. Data data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.tabulartext-generationn<1K0 likes23 downloads2d agoHugging Face07vvsd-charan /safety_aligned_datasets Safety Aligned Datasets A high-fidelity adversarial corpus engineered for alignment research, refusal boundary modeling, and robustness evaluation of Small Language Models. The Problem This Solves Fine-tuning a Small Language Model to be safe is not the same as fine-tuning it to understand safety. Most safety datasets give models clean refusal examples on obvious prompts — and those models fail the moment an adversary wraps a harmful request in a… See the full description on the dataset page: https://huggingface.co/datasets/vvsd-charan/safety_aligned_datasets.texttext-generation10K<n<100K0 likes19 downloads4mo agoHugging Face08nassimjp /Bilingual-SFT-2.0-Pashto-English-Aligned Bilingual SFT 2.0 — Pashto English Aligned 🇦🇫🇬🇧 Bilingual-SFT-2.0-Pashto-English-Aligned is a bilingual supervised fine-tuning dataset designed to improve Large Language Models (LLMs) in Pashto ↔ English understanding, instruction following, conversation, and bilingual generation. The dataset uses a conversational messages format and is intended for modern instruction-tuning pipelines, including Hugging Face Transformers, TRL, Unsloth, Axolotl, and other SFT frameworks.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-2.0-Pashto-English-Aligned.texttext-generation100K<n<1M0 likes16 downloads1mo agoHugging Face09houssamboukhalfa /culturally_aligned_arabic_stories_subset_a 📚 Culturally Aligned Arabic Stories Dataset (Subset A) A curated 110-example subset of the Crafting Culturally Aligned Narratives dataset, designed for the development and evaluation of Arabic children’s story generation models aligned with Islamic and cultural values. ✨ Overview Language: Modern Standard Arabic (MSA) Samples: 110 prompt–response pairs Format: JSONL (id, language, prompt, response, source, license) Moral domains: honesty, courage, generosity… See the full description on the dataset page: https://huggingface.co/datasets/houssamboukhalfa/culturally_aligned_arabic_stories_subset_a.texttext-generationn<1K0 likes8 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.