datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-code-science-deepseek-r1-en
R1 Dataset Collection
Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528.
Dataset Summary
The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes:
~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.TimeQA
TimeQA
Check out the original GitHub repo to learn more about the dataset.
professor_heideltime_en
Professor HeidelTime
Professor HeidelTime is a project to create a multilingual corpus weakly labeled with HeidelTime, a temporal tagger.
Corpus Details
The weak labeling was performed in six languages. Here are the specifics of the corpus for each language:
Dataset
Language
Documents
From
To
Tokens
Timexs
All the News 2.0
EN
24,642
2016-01-01
2020-04-0218,755,616
254,803
Italian Crime News
IT
9,619
2011-01-01
2021-12-31
3,296,898
58,823
German News… See the full description on the dataset page: https://huggingface.co/datasets/hugosousa/professor_heideltime_en.Publico
Público
This dataset was build by translating a set of 34,157 news from Público, an European Portuguese news paper. The news have been translated using Google Translator.
To now more about the data visit the Github repos used to scrape and translate the news.
n-gramsProfessorHeidelTime
Professor HeidelTime
Paper GitHub
Professor HeidelTime is a project to create a multilingual corpus weakly labeled with HeidelTime, a temporal tagger.
Corpus Details
The weak labeling was performed in six languages. Here are the specifics of the corpus for each language:
Dataset
Language
Documents
From
To
Tokens
Timexs
All the News 2.0
EN
24,642
2016-01-01
2020-04-02
18,755,616
254,803
Italian Crime News
IT
9,619
2011-01-01
2021-12-31
3,296,898
58,823… See the full description on the dataset page: https://huggingface.co/datasets/hugosousa/ProfessorHeidelTime.sm64-tas-dataset
SM64 Speedrun / TAS Reasoning Dataset
Question → <think> reasoning → answer pairs about Super Mario 64
speedrunning and Tool-Assisted Speedruns (TAS), in ShareGPT format.
Each assistant turn contains an explicit reasoning trace inside
<think>...</think> followed by the final answer, matching the native
thinking format of Qwen3-style models.
Files
File
Rows
Use
dataset_v11.jsonl
2721
full dataset
dataset_v11_train.jsonl
2585
training split (95%)… See the full description on the dataset page: https://huggingface.co/datasets/hugo74130/sm64-tas-dataset.legal-ai-act-spanish-sft-7k⚠️ Legal and Liability Disclaimer
This dataset is provided for research and educational purposes only.
It does not constitute legal advice, nor does it represent an official or authoritative interpretation of Regulation (EU) 2024/1689 (EU AI Act).
The content is synthetically generated and may contain errors, omissions, or hallucinations.
Under no circumstances should this dataset be used as a basis for legal, compliance, or regulatory decision-making.
The authors disclaim any liability for… See the full description on the dataset page: https://huggingface.co/datasets/hugoramallo/legal-ai-act-spanish-sft-7k.PerCN
PerCN Dataset
Overview
PerCN is a Chinese dataset for MBTI personality type prediction. Each sample contains multiple short posts from the same user, and labels are a 4-d binary vector corresponding to the four MBTI dimensions. The texts include typical Chinese social media expressions, emojis, and colloquial phrasing.
Data Format
The dataset is provided in JSONL format with three splits: train.jsonl, eval.jsonl, and test.jsonl.
Each line is a JSON object with… See the full description on the dataset page: https://huggingface.co/datasets/HugoZhu/PerCN.selma_ngramsngramsfolbar-test2filtered-awesome-chatgpt-propmts-oss-120b
Filtered Awesome ChatGPT Prompts – Model Outputs Dataset
Overview
This dataset contains model-generated responses to prompts from the fka/awesome-chatgpt-prompts Hugging Face dataset.
Each prompt was sent to the openai/gpt-oss-120b model via the OpenRouter API.
The resulting dataset was then filtered to remove:
Non English outputs with high language-detection confidence (fastText score < 0.7)
Very short outputs (≤ 10 words)
The goal of this dataset is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/filtered-awesome-chatgpt-propmts-oss-120b.WikiTimelinesfoolbar-llama3-ttprojetE3train and test for parkinson diseases
testdataset2
