datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sm64-tas-dataset
SM64 Speedrun / TAS Reasoning Dataset
Question → <think> reasoning → answer pairs about Super Mario 64
speedrunning and Tool-Assisted Speedruns (TAS), in ShareGPT format.
Each assistant turn contains an explicit reasoning trace inside
<think>...</think> followed by the final answer, matching the native
thinking format of Qwen3-style models.
Files
File
Rows
Use
dataset_v11.jsonl
2721
full dataset
dataset_v11_train.jsonl
2585
training split (95%)… See the full description on the dataset page: https://huggingface.co/datasets/hugo74130/sm64-tas-dataset.legal-ai-act-spanish-sft-7k⚠️ Legal and Liability Disclaimer
This dataset is provided for research and educational purposes only.
It does not constitute legal advice, nor does it represent an official or authoritative interpretation of Regulation (EU) 2024/1689 (EU AI Act).
The content is synthetically generated and may contain errors, omissions, or hallucinations.
Under no circumstances should this dataset be used as a basis for legal, compliance, or regulatory decision-making.
The authors disclaim any liability for… See the full description on the dataset page: https://huggingface.co/datasets/hugoramallo/legal-ai-act-spanish-sft-7k.filtered-awesome-chatgpt-propmts-oss-120b
Filtered Awesome ChatGPT Prompts – Model Outputs Dataset
Overview
This dataset contains model-generated responses to prompts from the fka/awesome-chatgpt-prompts Hugging Face dataset.
Each prompt was sent to the openai/gpt-oss-120b model via the OpenRouter API.
The resulting dataset was then filtered to remove:
Non English outputs with high language-detection confidence (fastText score < 0.7)
Very short outputs (≤ 10 words)
The goal of this dataset is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/filtered-awesome-chatgpt-propmts-oss-120b.
