datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.joke_explaination
Dataset Card for Dataset Name
Dataset Summary
Corpus for testing whether your LLM can explain the joke well. But this is a rather small dataset, if someone can point to a larger ones would be very nice.
Languages
English
Dataset Structure
Data Fields
url : link to the explaination
joke : the original joke
explaination : the explaination of the joke
Data Splits
Since its so small, there's no splits just like gsm8k
vihsd-explainable-dpo
vihsd-explainable-dpo
DPO preference dataset derived from vominhmanh/vihsd-explainable for Direct Preference Optimization (DPO).
Each example is a preference pair (chosen vs rejected) for the same prompt.
Schema (per example):
prompt (string): the original SFT prompt for Vietnamese moderation (user instruction).
chosen (string): JSON string with keys explanation, evidence, label — preferred (longer/more informative) explanation.
rejected (string): JSON string with keys explanation… See the full description on the dataset page: https://huggingface.co/datasets/vominhmanh/vihsd-explainable-dpo.vihsd-explainable-dpo
vihsd-explainable-dpo
DPO preference dataset derived from vominhmanh/vihsd-explainable for Direct Preference Optimization (DPO).
Each example is a preference pair (chosen vs rejected) for the same prompt.
Schema (per example):
prompt (string): the original SFT prompt for Vietnamese moderation (user instruction).
chosen (string): JSON string with keys explanation, evidence, label — preferred (longer/more informative) explanation.
rejected (string): JSON string with keys… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vihsd-explainable-dpo.vihsd-explainable
vihsd-explainable
ViHSD (extended) — Vietnamese toxic/offensive dataset with explanations & evidence.
This dataset extends the original htdung167/ViHSD examples by adding:
explanation: a short Vietnamese rationale (why the gold label applies)
evidence: verbatim substrings extracted from the text that justify the label
Splits
train, validation, test — same splits as original ViHSD
Schema (per sample)
text (string): original sentence
label_id (int): 0=CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/vominhmanh/vihsd-explainable.vihsd-explainable
vihsd-explainable
ViHSD (extended) — Vietnamese toxic/offensive dataset with explanations & evidence.
This dataset extends the original htdung167/ViHSD examples by adding:
explanation: a short Vietnamese rationale (why the gold label applies)
evidence: verbatim substrings extracted from the text that justify the label
Splits
train, validation, test — same splits as original ViHSD
Schema (per sample)
text (string): original sentence
label_id (int): 0=CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/Greeed88/vihsd-explainable.vihsd-explainable
vihsd-explainable
ViHSD (extended) — Vietnamese toxic/offensive dataset with explanations & evidence.
This dataset extends the original htdung167/ViHSD examples by adding:
explanation: a short Vietnamese rationale (why the gold label applies)
evidence: verbatim substrings extracted from the text that justify the label
Splits
train, validation, test — same splits as original ViHSD
Schema (per sample)
text (string): original sentence
label_id (int): 0=CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/dongeov/vihsd-explainable.titleix-explainer
Title IX Respondent Explainer (Atomizer-ready)
Purpose. An instruction-tuning dataset designed to train an information-only explainer bot for Title IX respondents. The bot helps users understand fields on a Title IX form, timelines, rights, and process basics. It does not give legal advice and does not make determinations about responsibility.
Audience: Respondents (the party accused) using a Title IX website or form.Scope: Descriptive/educational answers only — no adjudication, no… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix-explainer.benefits-notices-explained-v4
Benefits Notices, Explained — training set v4
The dataset is the deliverable. 124 checker-filtered teacher-distillation
dialogs that train a small model (Qwen3-4B QLoRA) to explain U.S. benefits
notices under a falsifiable behavior spec: earned-vocabulary ceiling
(frozen 2,801-lemma NGSL allowed list + words the reader used + glossed
terms), character-for-character anchor fidelity (dates, amounts, phones,
durations, case/form numbers, citations), quote-then-explain, no advice… See the full description on the dataset page: https://huggingface.co/datasets/jmerithew1/benefits-notices-explained-v4.kanitakorn-deepseek-v41-explain-robust-micro
Kanitakorn DeepSeek v41 Explain Robust Micro
Compact SFT continuation lane for a <=14B non-Thai-base Thai LLM.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Intended use: quick LoRA continuation after v39/v40-style DeepSeek candidates
Model identity taught: kanitakorn / คณิตกรณ์
Developer identity taught: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 472 rows = 400 MCQ + 56 Thai instruction + 16 identity
MCQ label balance: a=80 b=80 c=80 d=80 e=80
MCQ sources: v39… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v41-explain-robust-micro.titleix_explainer_reformat
Dataset Card for Title IX Respondent Explainers (Structured, 2020 Regs)
Dataset Summary
A curated instruction-tuning dataset for plain-language, respondent-focused Title IX explanations aligned with the 2020 federal regulations.Each example follows a six-section template:
What this means
Who it applies to
What to expect
Your options now
Important cautions
Where to confirm
The dataset avoids legal advice and school-specific promises; timelines are framed as… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix_explainer_reformat.
