datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-code-science-deepseek-r1-en
R1 Dataset Collection
Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528.
Dataset Summary
The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes:
~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.TimeQA
TimeQA
Check out the original GitHub repo to learn more about the dataset.
protocolos-clinicos-br
Protocolos Clínicos BR
Paper | Code | Blog post
Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines".
Configurations
default — Original guidelines (raw text)
The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.Superior-Reasoning-SFT-gpt-oss-120b-split-en
Superior-Reasoning SFT (stage1 + stage2) with <think> split and English filtering
Summary
This dataset is a processed derivative of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b (subsets stage1 and stage2, train split). It restructures each example into three fields:
input: the original input
reasoning: the content extracted from <think> ... </think> within the original output (inner text only)
output: the remainder of the original output after removing all <think>… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/Superior-Reasoning-SFT-gpt-oss-120b-split-en.pcdt-qa
PCDT-QA
An open-ended QA benchmark of 890 clinical questions grounded in Brazil's official clinical guidelines (PCDTs), published by the Ministry of Health.
Each question is paired with a reference answer derived from the guideline text. Evaluation uses an LLM-as-a-judge pipeline: the model under evaluation generates a free-text response, and a judge model (e.g., GPT-4.1) compares it against the reference, producing a binary correct/incorrect verdict. This accommodates the… See the full description on the dataset page: https://huggingface.co/datasets/hugo/pcdt-qa.healthbench-br
HealthBench-BR
A true/false benchmark of 1,780 paired clinical assertions grounded in Brazil's official clinical guidelines (PCDTs), published by the Ministry of Health.
For each clinical fact, two statements are provided: a true version faithful to the guideline, and a false version that modifies a single critical detail (a dosage, route of administration, monitoring interval, etc.). The dataset is perfectly balanced (50% true / 50% false), so correct classification requires… See the full description on the dataset page: https://huggingface.co/datasets/hugo/healthbench-br.sm64-tas-dataset
SM64 Speedrun / TAS Reasoning Dataset
Question → <think> reasoning → answer pairs about Super Mario 64
speedrunning and Tool-Assisted Speedruns (TAS), in ShareGPT format.
Each assistant turn contains an explicit reasoning trace inside
<think>...</think> followed by the final answer, matching the native
thinking format of Qwen3-style models.
Files
File
Rows
Use
dataset_v11.jsonl
2721
full dataset
dataset_v11_train.jsonl
2585
training split (95%)… See the full description on the dataset page: https://huggingface.co/datasets/hugo74130/sm64-tas-dataset.legal-ai-act-spanish-sft-7k⚠️ Legal and Liability Disclaimer
This dataset is provided for research and educational purposes only.
It does not constitute legal advice, nor does it represent an official or authoritative interpretation of Regulation (EU) 2024/1689 (EU AI Act).
The content is synthetically generated and may contain errors, omissions, or hallucinations.
Under no circumstances should this dataset be used as a basis for legal, compliance, or regulatory decision-making.
The authors disclaim any liability for… See the full description on the dataset page: https://huggingface.co/datasets/hugoramallo/legal-ai-act-spanish-sft-7k.filtered-awesome-chatgpt-propmts-oss-120b
Filtered Awesome ChatGPT Prompts – Model Outputs Dataset
Overview
This dataset contains model-generated responses to prompts from the fka/awesome-chatgpt-prompts Hugging Face dataset.
Each prompt was sent to the openai/gpt-oss-120b model via the OpenRouter API.
The resulting dataset was then filtered to remove:
Non English outputs with high language-detection confidence (fastText score < 0.7)
Very short outputs (≤ 10 words)
The goal of this dataset is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/filtered-awesome-chatgpt-propmts-oss-120b.
