datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-CC-HQ-20B
Nemotron-CC-HQ-20B
This Dataset consists of approximately 20B tokens of Nemotron-CC-HQ, consisting of randomly sampled slices from crawls in the range CC-MAIN-2013-20-part-00012 to CC-MAIN-2019-04-part-00007.
For more information about Nemotron-CC check the Paper by Nvidia
Disclaimer:
Derived from Nemotron-CC (Common Crawl). No ownership of underlying content is claimed.
Data may be subject to third-party rights. Use at your own risk and in compliance with… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Nemotron-CC-HQ-20B.week1-general-20b-dolma2-v1
Week-One General 20B Dolma2
This is a deterministic, pretokenized 20-billion-token baseline corpus for
controlled language-model architecture and training experiments. It contains
nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered
extension of the previous view. It also includes a dataset-only
370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B
order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.TinyMathStories_gpt-oss-20b
TinyMathStories
A TinyStories-style corpus extended with math and lightweight reasoning.
This dataset keeps the child-level vocabulary and short narrative style of TinyStories (Microsoft Research, Eldan & Li, 2023) and mixes in basic numeracy (counting, addition/subtraction, simple equations, fractions, measurement) and short justifications—so tiny models can practice coherent English and early math/logic.
Research, generation, and curation by AlgoDriveAI.Inspired by and… See the full description on the dataset page: https://huggingface.co/datasets/AlgoDriveAI/TinyMathStories_gpt-oss-20b.GPT-OSS-20B-Distilled-Reasoning-Mini
Dataset Card for Dataset Name
GPT-OSS-20B Distilled Reasoning Dataset Mini
(Multi-stage Evaluative Refinement Method for Reasoning Generation)
Dataset Details and Description
This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.gptoss20b-bilingual-curriculum-sft
gpt-oss-20b Bilingual Curriculum SFT
Synthetic bilingual supervised fine-tuning data generated with gpt-oss-20b (MoE, ~3.6B active params, native MXFP4, adaptive reasoning effort by difficulty).
Domains: mathematics, physics, chemistry, biology, computer science, general science, general knowledge, conversation.
Languages: Turkish and English.
Difficulty levels: 1-8.
The dataset is synthetic and should be independently evaluated before production use.
gpt-oss-20b-reasoning-traces
GPT-OSS-20B Reasoning Traces
3,333 reasoning traces generated by openai/gpt-oss-20b and filtered for clean, terminating reasoning. It was built to distill GPT-OSS's tight reasoning style into smaller models, and is the training set behind iAmBoosted/Qwen3.5-9B-OSS-Distilled.
What's in it
Each record pairs a prompt with GPT-OSS-20B's full reasoning trace and final answer, in chat-message form, ready for supervised fine-tuning (SFT).
~4,000 raw traces were generated, then… See the full description on the dataset page: https://huggingface.co/datasets/iAmBoosted/gpt-oss-20b-reasoning-traces.gpt-oss-20b-500xTrace of gpt-oss 20B LLM made by OpenAI.
Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning.
Brought to you by sapbot from Romarchive
rus_science_for_gpt_oss_20b
rus_science_for_gpt_oss_20b
Русскоязычный датасет для дообучения LLM под научно-академический ассистент.
Описание
~32 272 примера в формате JSONL.
Тематика: научные тексты, академический стиль, описание таблиц/методик, введения, пояснения, переформулировки.
Каждая строка содержит полный контекст диалога и готовые ответы ассистента.
Формат полей
reasoning_language: язык рассуждений ("Russian").
developer: инструкция для ассистента (роль/стиль/задача).
user:… See the full description on the dataset page: https://huggingface.co/datasets/MIldoc/rus_science_for_gpt_oss_20b.chinese_title_generation_gpt_oss_20b
該數據集主要用於訓練模型生成標題
(該數據提取 Mxode/Chinese-Instruct 其中的 5000 條,以及使用 gpt-oss-20b 進行標題生成 (即 response 欄位)。
