datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.joke_explaination
Dataset Card for Dataset Name
Dataset Summary
Corpus for testing whether your LLM can explain the joke well. But this is a rather small dataset, if someone can point to a larger ones would be very nice.
Languages
English
Dataset Structure
Data Fields
url : link to the explaination
joke : the original joke
explaination : the explaination of the joke
Data Splits
Since its so small, there's no splits just like gsm8k
repro-time-series-saliency-maps-explaining-models-across-multiple-domains-traces
Agent traces
Agent sessions published from a Trackio Logbook.
blood-test-explainer-traces
Blood Test Explainer - agent traces
Agent traces from the Blood Test Explainer app (Build Small hackathon). Each row is one publicly-available sample lab report (fake patients, no PHI) run through the full agent pipeline: a small vision model reads the document and extracts the markers, then a curated medical knowledge base turns the values into a grounded, per-marker explanation plus cross-marker patterns.
Model: build-small-hackathon/blood-test-minicpmv-4_6-medreason, a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/blood-test-explainer-traces.vihsd-explainable-dpo
vihsd-explainable-dpo
DPO preference dataset derived from vominhmanh/vihsd-explainable for Direct Preference Optimization (DPO).
Each example is a preference pair (chosen vs rejected) for the same prompt.
Schema (per example):
prompt (string): the original SFT prompt for Vietnamese moderation (user instruction).
chosen (string): JSON string with keys explanation, evidence, label — preferred (longer/more informative) explanation.
rejected (string): JSON string with keys explanation… See the full description on the dataset page: https://huggingface.co/datasets/vominhmanh/vihsd-explainable-dpo.labeled-multiple-choice-explainedThis dataset is based on under-tree/labeled-multiple-choice but using GPT-3.5-turbo to generate explanations for each answer option.
This was a very basic attempt to follow the Orca paper approach of a 'teacher' model to provide more context to some trivia questions.
Questions were deduplicated based on the question text.
I used the python library guidance to help generate the prompts. Below is the prompt template I used.
{{#role 'system'~}}
You are an AI assistant that helps people find… See the full description on the dataset page: https://huggingface.co/datasets/layoric/labeled-multiple-choice-explained.vihsd-explainable-dpo
vihsd-explainable-dpo
DPO preference dataset derived from vominhmanh/vihsd-explainable for Direct Preference Optimization (DPO).
Each example is a preference pair (chosen vs rejected) for the same prompt.
Schema (per example):
prompt (string): the original SFT prompt for Vietnamese moderation (user instruction).
chosen (string): JSON string with keys explanation, evidence, label — preferred (longer/more informative) explanation.
rejected (string): JSON string with keys… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vihsd-explainable-dpo.adaption-tech-concepts-explained
Adaption Tech Concepts Explained
A High-Quality Instruction Tuning Dataset for Large Language Models
A high-quality instruction tuning dataset designed for fine-tuning Large Language Models (LLMs) to generate clear, structured, and beginner-friendly explanations of technical concepts.
This dataset was enhanced using Adaption's Adaptive Data Platform, which improves instruction quality, response consistency, and educational value for supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/ujjawalbansal/adaption-tech-concepts-explained.ArcEasy-ExplainChoicevihsd-explainable
vihsd-explainable
ViHSD (extended) — Vietnamese toxic/offensive dataset with explanations & evidence.
This dataset extends the original htdung167/ViHSD examples by adding:
explanation: a short Vietnamese rationale (why the gold label applies)
evidence: verbatim substrings extracted from the text that justify the label
Splits
train, validation, test — same splits as original ViHSD
Schema (per sample)
text (string): original sentence
label_id (int): 0=CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/vominhmanh/vihsd-explainable.vihsd-explainable
vihsd-explainable
ViHSD (extended) — Vietnamese toxic/offensive dataset with explanations & evidence.
This dataset extends the original htdung167/ViHSD examples by adding:
explanation: a short Vietnamese rationale (why the gold label applies)
evidence: verbatim substrings extracted from the text that justify the label
Splits
train, validation, test — same splits as original ViHSD
Schema (per sample)
text (string): original sentence
label_id (int): 0=CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/Greeed88/vihsd-explainable.Bird_explained_corrections
Dataset Card for Dataset Name
This dataset is truncated
vihsd-explainable
vihsd-explainable
ViHSD (extended) — Vietnamese toxic/offensive dataset with explanations & evidence.
This dataset extends the original htdung167/ViHSD examples by adding:
explanation: a short Vietnamese rationale (why the gold label applies)
evidence: verbatim substrings extracted from the text that justify the label
Splits
train, validation, test — same splits as original ViHSD
Schema (per sample)
text (string): original sentence
label_id (int): 0=CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/dongeov/vihsd-explainable.titleix-explainer
Title IX Respondent Explainer (Atomizer-ready)
Purpose. An instruction-tuning dataset designed to train an information-only explainer bot for Title IX respondents. The bot helps users understand fields on a Title IX form, timelines, rights, and process basics. It does not give legal advice and does not make determinations about responsibility.
Audience: Respondents (the party accused) using a Title IX website or form.Scope: Descriptive/educational answers only — no adjudication, no… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix-explainer.math-explain-viHistory_MCQs_explainedBvsALL_explained_selftrain
BvsALL_explained_selftrain_round1_half_original
This dataset is a local mixed SFT dataset built from:
Original dataset: Simonkami/BvsALL_explained split train
Self-training file: selftrain_good_samples_round1.jsonl
Mixing settings
Original total: 8499
Original used: 1699
Original fraction: 0.2
Original max samples: 0
Selftrain raw total: 13490
Selftrain valid: 13490
Selftrain after action balance: 9674
Selftrain used: 6796
Target selftrain ratio: 0.8
Actual selftrain… See the full description on the dataset page: https://huggingface.co/datasets/Simonkami/BvsALL_explained_selftrain.benefits-notices-explained-v4
Benefits Notices, Explained — training set v4
The dataset is the deliverable. 124 checker-filtered teacher-distillation
dialogs that train a small model (Qwen3-4B QLoRA) to explain U.S. benefits
notices under a falsifiable behavior spec: earned-vocabulary ceiling
(frozen 2,801-lemma NGSL allowed list + words the reader used + glossed
terms), character-for-character anchor fidelity (dates, amounts, phones,
durations, case/form numbers, citations), quote-then-explain, no advice… See the full description on the dataset page: https://huggingface.co/datasets/jmerithew1/benefits-notices-explained-v4.Chess_Stockfish_BestMove_ExplainCoT-Explaining-Mathkanitakorn-deepseek-v41-explain-robust-micro
Kanitakorn DeepSeek v41 Explain Robust Micro
Compact SFT continuation lane for a <=14B non-Thai-base Thai LLM.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Intended use: quick LoRA continuation after v39/v40-style DeepSeek candidates
Model identity taught: kanitakorn / คณิตกรณ์
Developer identity taught: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 472 rows = 400 MCQ + 56 Thai instruction + 16 identity
MCQ label balance: a=80 b=80 c=80 d=80 e=80
MCQ sources: v39… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v41-explain-robust-micro.ExplainAI-datavulnerability_detection__explainabilityBvsALL_explainedtitleix_explainer_casual
pretty_name: "Title IX Respondent Explainers (Casual, 2020 Regs)"
license: "cc-by-sa-4.0"
language:
- en
tags:
- law
- education
- safety
- assistant
- instruction-tuning
- compliance
task_categories:
- text2text-generation
size_categories:
- 1K<n<10K
Dataset Card for Title IX Respondent Explainers (Casual, 2020 Regs)
Dataset Summary
A companion dataset that paraphrases the structured six-section explanations into 2–3 short paragraphs with a… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix_explainer_casual.kanitakorn-qwen-v36-explained-failure-delta-clean-20260614
Qwen v36 Explained Failure-Delta Mix
Explanation-rich continuation of the v35 source-audited ThaiExam failure-delta data. Each assistant target explains the rule, evidence or calculation before the final คำตอบคือ (x) line.
explained-priority-scored-contract-vulnerabilitiesExplainMed-RAGsn96g-science-explainers-2chunk1-20250919_184511
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 32.5 (median 32.5, min 13, max 52)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): 217328e167aa722bfe23ac052014390d4aed1b453113248639790ab73e0a600e
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-2chunk1-20250919_184511.sn96g-science-explainers-2chunk1-20250919_193451
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 35 (median 35.0, min 26, max 44)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): 3d134bc47c74ba1f958620f97a9ed58f30fa62c14311e565ac699bde4aa5f089
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-2chunk1-20250919_193451.
